@seoagent-official/seoagent 1.72.0 → 1.74.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@seoagent-official/seoagent",
3
- "version": "1.72.0",
3
+ "version": "1.74.0",
4
4
  "description": "The persistent AI SEO agent for Claude Code. Audits, keyword strategy, briefs, articles, real product screenshots from your repo, and the autopilot loop (cloud detects → CLI executes → ack closes) — other SEO tools write the prompt, SEOAgent runs it.",
5
5
  "type": "module",
6
6
  "bin": {
@@ -6,6 +6,8 @@ Loaded by Phase 1 (Technical SEO Audit) when running a full audit. The main `SKI
6
6
 
7
7
  **Every factual claim about a page's live state must be grounded in an actual fetch of the live URL — never in repo source, memory, or a prior.** Before running these checks, run `seoagent crawl` (Phase 1 Step 0). It writes `.seoagent/audit/evidence.md` — the live-crawl evidence base (exact title, meta, ALL H1s, canonical + server/client-render flag, every JSON-LD `@type`, OG/Twitter tags, per-page `<img>`-missing-alt stats, the ACTUAL robots.txt contents, the sitemap URL + blog-post counts, client-rendered-shell detection). Read that file and derive Confirmed findings from it.
8
8
 
9
+ **Origin binding — the crawl must hit the LIVE site the user designated, and evidence.md records which origin it hit.** If the user/session provided a live URL, pass it verbatim: `seoagent crawl --url <origin>` (authoritative even when loopback — staging/preview is legitimate); otherwise the CLI uses `live_url:`/`domain:` from `project.md`, and errors if neither exists — never guess, and **never start a local dev server and crawl it as the live site**. Check the **"Crawled origin:"** line at the top of `evidence.md` before deriving findings: if the file is labeled **SOURCE RENDER (local dev server) — NOT the public live site** (`source_render: true`), it describes an undesignated local render — NO live-state finding (`Confirmed` present OR absent) may be derived from it, and `seoagent verify-recs` will refuse to verify against it. Re-crawl with `--url <live origin>` first.
10
+
9
11
  **The evidence covers the whole crawl, not just the homepage — audit accordingly.** The crawl discovers subpages from the live sitemap + homepage nav links (default 15 pages) and writes a per-page section for **EVERY page in `evidence.md`**, plus a **`## Site-wide rollup`** section that aggregates the per-page gaps: *Pages missing canonical*, *Pages missing meta description*, *Pages with multiple H1s*, *Pages with no structured data*, and *Images missing alt* (total + per-page counts with the offending srcs). Run the per-page checks below against **every crawled page** — a subpage's missing canonical or a blog post's missing meta description is exactly what the rollup lists — and report those findings from the rollup + per-page sections, never from homepage-only inspection or guesswork.
10
12
 
11
13
  Three rules govern every finding:
@@ -18,6 +20,8 @@ Three rules govern every finding:
18
20
 
19
21
  **Do not assert repo-internal specifics (file paths, line numbers) as fact in a live-state finding.** They are fixing *hints*, not verified live facts — mark them `Likely` and phrase as "likely in `app/layout.tsx`", never "on line 42".
20
22
 
23
+ **Mechanical backstop — verify-recs runs automatically on every sync, and you still run it at the end of every audit session.** The rules above are also enforced deterministically, twice: (a) every `seoagent sync` — including the PostToolUse hook that fires after each file write — runs a verify-recs pass first, correcting `.seoagent/**/*.md` claims as they're written (watch for `verify-recs: corrected …` lines in the hook output); and (b) after writing the audit outputs (and any fixes/work log), run `seoagent verify-recs` yourself, passing any files you wrote outside `.seoagent/` as arguments. It scans the authored markdown against `evidence.md` and rewrites, in place with a `CORRECTION (verify-recs)` annotation, any absence claim or add-recommendation about a title, meta description, canonical, Open Graph/Twitter tag, or JSON-LD `@type` the live page already serves. Treat its corrections as authoritative: carry every one into the final response, and never restate the original (corrected) claim. On SOURCE-RENDER evidence it applies no corrections and reports that live-state claims can't be verified.
24
+
21
25
  ## Crawlability & Indexation
22
26
 
23
27
  ### `robots_txt_exists`
@@ -393,11 +393,29 @@ Once approved, work **a cluster at a time**, top of the plan down:
393
393
 
394
394
  ---
395
395
 
396
+ ## Session Economy — bounded sessions
397
+
398
+ Every session has a budget — headless/one-shot runs have a hard turn cap, interactive ones have the user's patience. Treat any single-request session as a **bounded session** and spend the budget on findings and shippable work, not on bookkeeping churn. The failure mode this section prevents: sessions that write a hundred small files — ticking a changelog line after every action, re-polishing the roadmap between steps — and hit the cap before the final summary exists. Four rules:
399
+
400
+ 1. **Consolidate writes.** One audit doc, one migration plan, one batch of fixes — not a file (or a file update) per finding. Do NOT write per-item bookkeeping updates (a changelog line here, a roadmap tick there, a status touch after each action) as you go; accumulate them and batch them into a single write at the end of the session — one changelog append, one roadmap update, one `seoagent sync`.
401
+ 2. **Scope a single-session ask to what one session can ship.** For a single-session "grow organic traffic" ask, deliver the audit findings + the migration plan + the top shippable fixes. Do NOT draft full article inventories or write every article in the plan — that is multi-session work; list the briefs (slug + target keyword + role) in the roadmap instead and let later sessions write them.
402
+ 3. **Reserve the final turns for the wrap-up.** When the work above is done, produce the final summary (running `seoagent verify-recs` first) — do not open a new work stream (another audit pass, another content draft) you cannot finish inside the budget. An unfinished extra stream plus a missing summary is worth less than a complete summary.
403
+ 4. **Prefer editing an existing file over creating a new one** when appending related content — extend `audit/latest.md` rather than creating a second audit file, add a section to `roadmap.md` rather than a new planning doc. Fewer files means fewer writes, fewer sync round-trips, and a workspace the next session can actually read.
404
+
405
+ These are general bounded-session economics, not a benchmark mode — in an interactive session the same rules simply make you faster and the workspace cleaner.
406
+
407
+ ---
408
+
396
409
  ## Phase 1: Technical SEO Audit
397
410
 
398
411
  ### Step 0 (mandatory — capture the live-crawl evidence base, then read the checks)
399
412
 
400
- 1. **Run `seoagent crawl` first.** It fetches the homepage + top pages, the real robots.txt, and the live sitemap, and writes `.seoagent/audit/evidence.md` the verified evidence base (exact title/meta, ALL H1s, canonical + server/client-render flag, every JSON-LD `@type`, OG/Twitter tags, the ACTUAL robots.txt contents, sitemap URL + blog-post counts, client-rendered-shell detection). **`Read` that file — every `Confirmed` finding must be derived from it, not from repo source or memory.** (Use `seoagent crawl --json` if you want the structured bundle.)
413
+ 1. **Run `seoagent crawl` first against the LIVE origin, explicitly.** The crawl binds to ONE origin, and which one is never your choice to improvise:
414
+ - **If the user/session stated where the live site is** (any phrasing — "the live site is at http://127.0.0.1:4173", "we're on https://staging.example.com", a URL in the task prompt), pass that URL **verbatim** as `seoagent crawl --url <origin>`. A user-designated origin is authoritative even when it's localhost — staging/preview setups are legitimate. The CLI persists it as `live_url:` in `project.md` so later commands reuse it.
415
+ - **If no live URL was provided**, use the persisted `live_url:`/`domain:` from `.seoagent/project.md` (the CLI resolves this automatically), or ask the user. If neither exists, `seoagent crawl` errors instead of guessing — resolve the origin, don't work around the error.
416
+ - **NEVER start a local dev server yourself and crawl it as the live site.** A dev server renders the repo's current state, which can differ from production in either direction — evidence captured from it would make every "Confirmed absent/present" claim false about the real site. If you crawl an origin the CLI wasn't told is live, `evidence.md` is labeled **SOURCE RENDER (local dev server) — NOT the public live site**, and no live-state claim may be derived from it.
417
+
418
+ The crawl fetches the homepage + top pages, the real robots.txt, and the live sitemap, and writes `.seoagent/audit/evidence.md` — the verified evidence base (exact title/meta, ALL H1s, canonical + server/client-render flag, every JSON-LD `@type`, OG/Twitter tags, the ACTUAL robots.txt contents, sitemap URL + blog-post counts, client-rendered-shell detection), with the **crawled origin recorded at the top** — check it matches the site you're auditing. **`Read` that file — every `Confirmed` finding must be derived from it, not from repo source or memory.** (Use `seoagent crawl --json` if you want the structured bundle.)
401
419
  2. **Read `.claude/skills/seoagent/references/audit-checks.md`.** It contains the full check list, the **Verify-before-assert** rules (confidence labels, never-recommend-what-exists, live-vs-source reconciliation), severity tiers, and recommendation text per check. Do not run the audit from memory — the reference is the source of truth and gives consistent results across sessions.
402
420
 
403
421
  3. **Step 0 is not optional and not "when useful" — it is the gate for ALL live-state work.** Any audit, technical-SEO review, "what's wrong with my site", or "add schema/meta/canonical" request starts with `seoagent crawl` — even a quick one, even when the repo source looks obvious. The crawl now covers the homepage **plus subpages discovered from the live sitemap + homepage nav links** (default 15 pages, fetched concurrently), so it's fast and it sees the blog posts and subpages where the real gaps hide. **No live-state claim and no "add X" recommendation may be emitted unless `.seoagent/audit/evidence.md` exists, covers the target page(s), and the claim cites it.** If `evidence.md` is missing or stale (>24h old — `seoagent doctor` flags this as `evidence_stale`), re-run the crawl before asserting anything. Reasoning from the repo source about what the live site serves is exactly the failure this gate exists to stop: the repo may be behind (or ahead of) production.
@@ -413,6 +431,8 @@ A line with neither is invalid output — rewrite it or drop it before respondin
413
431
 
414
432
  **"Add X" recommendations are FORBIDDEN unless evidence shows absence on the LIVE page.** Never recommend adding a title, meta description, canonical, Open Graph/Twitter tags, or JSON-LD schema unless the page's `evidence.md` section shows that item genuinely absent (`_(none)_` / "safe to recommend adding"). If the page's **"Already present (do NOT recommend adding)"** line lists it, the recommendation is suppressed — those are the `recommendation-guard` semantics, printed into the evidence file precisely so you can't miss them. If a page wasn't crawled, you have **no evidence of absence**: an "add X" for it is at most a `Hypothesis`, never an action item.
415
433
 
434
+ **Mechanical enforcement — verify-recs runs AUTOMATICALLY, and is still the MANDATORY final step of any audit/optimization session.** The contract above is enforced by the CLI, not just by this prompt — twice over: (a) every `seoagent sync` (including the PostToolUse hook that fires after each file write) runs a verify-recs pass first, so false claims in `.seoagent/**/*.md` are corrected as the files are written — when the hook output reports a `verify-recs: corrected …` line, treat it as authoritative and carry the correction into your response; and (b) at the end of the session — after writing your outputs, BEFORE composing the final summary — run `seoagent verify-recs` yourself (pass any work-log/summary files you wrote outside `.seoagent/` as arguments), because files outside `.seoagent/` are only checked when you pass them. **If `evidence.md` is a SOURCE RENDER (undesignated local dev-server crawl), verify-recs cannot verify live-state claims — it will say so; re-crawl with `--url <live origin>` before finalizing any live-state summary.** It re-checks every authored `.seoagent/**/*.md` against `evidence.md` and REWRITES, in place with a `CORRECTION (verify-recs)` annotation, any "added X / there was no X / X was missing" claim about a head-level entity (title, meta description, canonical, Open Graph, twitter:card, or a JSON-LD `@type` such as `Organization`/`SoftwareApplication`/`WebSite`) that the evidence shows the live page ALREADY serves. **Reflect every correction it reports in your final message** — never let a corrected claim survive into the summary (use `--json` to branch on the result programmatically). It never touches `evidence.md`, generated projections (`pages.md`/`keywords.md`), the inbox, or anything under `.claude/`, and it always exits 0 — a correction is the check working, not an error.
435
+
416
436
  **Verify-before-assert is the load-bearing rule of the whole audit.** Never state a live-page fact you didn't fetch: don't invent a robots.txt rule, don't recommend adding schema/canonical/OG tags the evidence shows already exist, don't report a dynamic on-page number (a "2,184 families" counter) as `Confirmed` unless it's in the server-fetched HTML. Tag every finding `Confirmed` / `Likely` / `Hypothesis`. If `seoagent crawl` couldn't run (offline, no domain), fall back to per-page WebFetch — but remember **WebFetch returns a markdown-stripped render that DROPS the entire `<head>`**: `<title>`, `<meta name="description">`, `<link rel="canonical">`, every `og:*` / `twitter:*` tag, AND every `<script>` JSON-LD block are all invisible to it. Any "missing title / meta / canonical / OG / schema" conclusion drawn from WebFetch is a **false negative** — never `Confirmed`, and never a basis for recommending you add a head tag the site already serves. That's what `seoagent crawl` (raw-HTML parse) exists to prevent; `evidence.md` even prints an explicit **"Already present (do NOT recommend adding)"** line per page.
417
437
 
418
438
  ### Procedure
@@ -466,7 +486,8 @@ low: 3
466
486
 
467
487
  1. Append to `.seoagent/changelog.md`: `[date] Audit completed: {N} pages, {N} findings ({c} critical, {h} high, {m} medium, {l} low)`.
468
488
  2. Update `.seoagent/roadmap.md` with audit-derived action items grouped by priority.
469
- 3. Run `seoagent sync`.
489
+ 3. Run `seoagent verify-recs` — the mechanical check that nothing you wrote contradicts `evidence.md` (it rewrites any false "added X / none existed" claim in place and reports it; reflect its corrections in your response).
490
+ 4. Run `seoagent sync`.
470
491
 
471
492
  ### Audit "Fixed" Flow
472
493
 
@@ -965,7 +986,7 @@ The CLI manages credentials at `~/.config/seoagent/auth.json` — outside the pr
965
986
  7. **End with the plan's next step, not a menu.** When executing an approved plan, close with progress + what's next in the plan ("3 of 8 in this cluster done; writing the next now"), not a 2–3-option menu every turn. Offer explicit choices only at real decision points (the plan-approval gate, a cluster boundary, an ambiguous call).
966
987
  8. **Update the roadmap and changelog** after every action.
967
988
  9. **Sync after every artifact write.** Run `seoagent sync` (no-op when not logged in — always run it).
968
- 10. **Verify before you assert.** Every claim about a page's live state (robots.txt rules, schema/JSON-LD, meta tags, titles, headings, canonical, sitemap contents, whether a URL exists) must be grounded in an actual live fetch — `seoagent crawl`'s `.seoagent/audit/evidence.md` or a WebFetch you just ran — never repo source, memory, or a prior. **Never recommend adding something the live page already has.** Tag every finding `Confirmed` / `Likely` / `Hypothesis`; never emit an unverified specific (price, line number, competitor, dynamic on-page metric) as a bare fact. Repo-only issues that aren't confirmed on the live site are labeled and reported separately, not as production reality. Every finding/recommendation line carries an `Evidence:` citation or an explicit `Hypothesis` label — see Phase 1 § Evidence-citation contract.
989
+ 10. **Verify before you assert.** Every claim about a page's live state (robots.txt rules, schema/JSON-LD, meta tags, titles, headings, canonical, sitemap contents, whether a URL exists) must be grounded in an actual live fetch — `seoagent crawl`'s `.seoagent/audit/evidence.md` or a WebFetch you just ran — never repo source, memory, or a prior. **Never recommend adding something the live page already has.** Tag every finding `Confirmed` / `Likely` / `Hypothesis`; never emit an unverified specific (price, line number, competitor, dynamic on-page metric) as a bare fact. Repo-only issues that aren't confirmed on the live site are labeled and reported separately, not as production reality. Every finding/recommendation line carries an `Evidence:` citation or an explicit `Hypothesis` label — see Phase 1 § Evidence-citation contract. The deterministic backstop is `seoagent verify-recs` — run it as the session's final step and reflect its corrections in the final summary.
969
990
  11. **Use the output template** for all top-level reports.
970
991
  12. **Read context before generating.** Before any strategy, brief, or article, read `.seoagent/context.md`.
971
992
  13. **Plan once, then execute** (see "Plan & Execute"). Get one approval on the content plan, then run it in batches (a cluster at a time) — don't ask `Continue?` between articles or phases. Pause only for: the plan approval, cluster boundaries (show drafts + open a PR), ambiguous decisions, and destructive actions. Go fully autonomous or step-by-step if the user asks.
@@ -973,6 +994,7 @@ The CLI manages credentials at `~/.config/seoagent/auth.json` — outside the pr
973
994
  15. **Edit existing files; Write only new ones.** `project.md`, `context.md`, `roadmap.md`, `changelog.md`, and any artifact created by `init` already exist — use the `Edit` tool to modify them. Reserve `Write` for files that don't exist yet. Trying to `Write` an existing file fails with "File must be read first" and wastes a tool call.
974
995
  16. **Use the CMS metadata.** If `project.md` has `cms: strapi | wordpress | sanity | contentful | ghost | webflow | shopify | payload | directus | mdx-local`, the user has a CMS. When writing articles in Phase 4, mention how the article's frontmatter maps to that CMS's content model (e.g. Strapi: title → Title field, body → Content rich-text). When the cluster is content-focused, suggest publishing the article to the detected CMS as the next step. The free tier writes to `.seoagent/content/` only — Cloud handles the publish itself.
975
996
  17. **Use the blog_path metadata.** If `project.md` has `blog_path: /blog` (or similar), articles' canonical URLs use that prefix: `https://{domain}{blog_path}/{slug}`.
997
+ 18. **Respect the session budget** (see "Session Economy — bounded sessions"). Consolidate writes, batch bookkeeping into one final update, scope single-session asks to audit + plan + top fixes, and reserve the final turns for `seoagent verify-recs` + the summary — never start a work stream you can't finish.
976
998
 
977
999
  ---
978
1000