@popoverai/dotrequirements 0.22.0 → 0.24.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +173 -23
- package/dist/cli.js +121 -61
- package/dist/codebase-to-spec/budget.d.ts +53 -0
- package/dist/codebase-to-spec/budget.js +80 -0
- package/dist/codebase-to-spec/cache.d.ts +49 -0
- package/dist/codebase-to-spec/cache.js +54 -0
- package/dist/codebase-to-spec/claude.d.ts +69 -0
- package/dist/codebase-to-spec/claude.js +126 -0
- package/dist/codebase-to-spec/compose.d.ts +49 -0
- package/dist/codebase-to-spec/compose.js +124 -0
- package/dist/codebase-to-spec/edit-loop.d.ts +54 -0
- package/dist/codebase-to-spec/edit-loop.js +195 -0
- package/dist/codebase-to-spec/editor.d.ts +54 -0
- package/dist/codebase-to-spec/editor.js +74 -0
- package/dist/codebase-to-spec/exit-codes.d.ts +40 -0
- package/dist/codebase-to-spec/exit-codes.js +58 -0
- package/dist/codebase-to-spec/fan-out.d.ts +63 -0
- package/dist/codebase-to-spec/fan-out.js +215 -0
- package/dist/codebase-to-spec/interactive.d.ts +30 -0
- package/dist/codebase-to-spec/interactive.js +48 -0
- package/dist/codebase-to-spec/outline-review-loop.d.ts +51 -0
- package/dist/codebase-to-spec/outline-review-loop.js +187 -0
- package/dist/codebase-to-spec/pack.d.ts +51 -0
- package/dist/codebase-to-spec/pack.js +127 -0
- package/dist/codebase-to-spec/planner.d.ts +41 -0
- package/dist/codebase-to-spec/planner.js +76 -0
- package/dist/codebase-to-spec/present.d.ts +94 -0
- package/dist/codebase-to-spec/present.js +288 -0
- package/dist/codebase-to-spec/progress.d.ts +33 -0
- package/dist/codebase-to-spec/progress.js +28 -0
- package/dist/codebase-to-spec/prompts/editor.d.ts +13 -0
- package/dist/codebase-to-spec/prompts/editor.js +57 -0
- package/dist/codebase-to-spec/prompts/outline-reviewer.d.ts +12 -0
- package/dist/codebase-to-spec/prompts/outline-reviewer.js +87 -0
- package/dist/codebase-to-spec/prompts/planner-apply.d.ts +12 -0
- package/dist/codebase-to-spec/prompts/planner-apply.js +32 -0
- package/dist/codebase-to-spec/prompts/planner-initial.d.ts +11 -0
- package/dist/codebase-to-spec/prompts/planner-initial.js +125 -0
- package/dist/codebase-to-spec/prompts/planner-revise.d.ts +14 -0
- package/dist/codebase-to-spec/prompts/planner-revise.js +60 -0
- package/dist/codebase-to-spec/prompts/spec-reviewer.d.ts +16 -0
- package/dist/codebase-to-spec/prompts/spec-reviewer.js +96 -0
- package/dist/codebase-to-spec/prompts/specifier.d.ts +12 -0
- package/dist/codebase-to-spec/prompts/specifier.js +100 -0
- package/dist/codebase-to-spec/prompts/style-check.d.ts +12 -0
- package/dist/codebase-to-spec/prompts/style-check.js +78 -0
- package/dist/codebase-to-spec/schemas.d.ts +257 -0
- package/dist/codebase-to-spec/schemas.js +183 -0
- package/dist/codebase-to-spec/skill-install.d.ts +57 -0
- package/dist/codebase-to-spec/skill-install.js +79 -0
- package/dist/codebase-to-spec/slice.d.ts +49 -0
- package/dist/codebase-to-spec/slice.js +111 -0
- package/dist/codebase-to-spec/specifier.d.ts +60 -0
- package/dist/codebase-to-spec/specifier.js +79 -0
- package/dist/codebase-to-spec/style-check.d.ts +29 -0
- package/dist/codebase-to-spec/style-check.js +33 -0
- package/dist/codebase-to-spec/summary.d.ts +51 -0
- package/dist/codebase-to-spec/summary.js +183 -0
- package/dist/codebase-to-spec/validate.d.ts +46 -0
- package/dist/codebase-to-spec/validate.js +130 -0
- package/dist/commands/acceptance-test.d.ts +6 -0
- package/dist/commands/acceptance-test.js +212 -0
- package/dist/commands/ai-setup.d.ts +5 -0
- package/dist/commands/ai-setup.js +441 -0
- package/dist/commands/browsertest.d.ts +0 -1
- package/dist/commands/browsertest.js +51 -26
- package/dist/commands/codebase-to-spec/compose.d.ts +14 -0
- package/dist/commands/codebase-to-spec/compose.js +57 -0
- package/dist/commands/codebase-to-spec/edit-loop.d.ts +16 -0
- package/dist/commands/codebase-to-spec/edit-loop.js +83 -0
- package/dist/commands/codebase-to-spec/fan-out.d.ts +19 -0
- package/dist/commands/codebase-to-spec/fan-out.js +77 -0
- package/dist/commands/codebase-to-spec/index.d.ts +9 -0
- package/dist/commands/codebase-to-spec/index.js +135 -0
- package/dist/commands/codebase-to-spec/pack.d.ts +22 -0
- package/dist/commands/codebase-to-spec/pack.js +76 -0
- package/dist/commands/codebase-to-spec/plan-loop.d.ts +26 -0
- package/dist/commands/codebase-to-spec/plan-loop.js +105 -0
- package/dist/commands/codebase-to-spec/present.d.ts +21 -0
- package/dist/commands/codebase-to-spec/present.js +92 -0
- package/dist/commands/codebase-to-spec/run.d.ts +20 -0
- package/dist/commands/codebase-to-spec/run.js +85 -0
- package/dist/commands/codebase-to-spec/skill-install.d.ts +20 -0
- package/dist/commands/codebase-to-spec/skill-install.js +51 -0
- package/dist/commands/codebase-to-spec/specify-area.d.ts +18 -0
- package/dist/commands/codebase-to-spec/specify-area.js +82 -0
- package/dist/commands/codebase-to-spec/style-check.d.ts +15 -0
- package/dist/commands/codebase-to-spec/style-check.js +42 -0
- package/dist/commands/codebase-to-spec/validate.d.ts +18 -0
- package/dist/commands/codebase-to-spec/validate.js +38 -0
- package/dist/commands/create-requirement-document.d.ts +2 -0
- package/dist/commands/create-requirement-document.js +41 -0
- package/dist/commands/finalize.js +7 -7
- package/dist/commands/get.d.ts +2 -0
- package/dist/commands/get.js +55 -0
- package/dist/commands/init.js +132 -117
- package/dist/commands/link.js +27 -27
- package/dist/commands/list.d.ts +6 -0
- package/dist/commands/list.js +43 -0
- package/dist/commands/mcp-setup.js +159 -149
- package/dist/commands/mcp.js +1 -1
- package/dist/commands/prepare.js +4 -4
- package/dist/commands/pull.js +116 -121
- package/dist/commands/push.js +106 -112
- package/dist/commands/report.d.ts +6 -2
- package/dist/commands/report.js +177 -122
- package/dist/commands/requirements-for.d.ts +2 -0
- package/dist/commands/requirements-for.js +29 -0
- package/dist/commands/review-test.d.ts +2 -0
- package/dist/commands/review-test.js +75 -0
- package/dist/commands/search.d.ts +6 -0
- package/dist/commands/search.js +39 -0
- package/dist/commands/style-check.d.ts +7 -0
- package/dist/commands/style-check.js +75 -0
- package/dist/commands/test.js +53 -59
- package/dist/commands/tests-for.d.ts +2 -0
- package/dist/commands/tests-for.js +80 -0
- package/dist/commands/validate.d.ts +6 -0
- package/dist/commands/validate.js +72 -0
- package/dist/config.js +1 -1
- package/dist/convex.d.ts +34 -22
- package/dist/convex.js +38 -22
- package/dist/harness/cache.d.ts +1 -5
- package/dist/harness/cache.js +49 -59
- package/dist/harness/convexReporting.d.ts +1 -1
- package/dist/harness/convexReporting.js +9 -7
- package/dist/harness/coverageCache.js +3 -3
- package/dist/harness/finalize.js +59 -46
- package/dist/harness/index.d.ts +6 -7
- package/dist/harness/index.js +9 -10
- package/dist/harness/prepare.js +6 -5
- package/dist/harness/requirementsLoader.d.ts +2 -2
- package/dist/harness/requirementsLoader.js +13 -35
- package/dist/harness/tracking.js +18 -18
- package/dist/harness/types.d.ts +1 -1
- package/dist/mcp/convexClient.d.ts +0 -39
- package/dist/mcp/convexClient.js +2 -107
- package/dist/mcp/grep.d.ts +1 -1
- package/dist/mcp/grep.js +87 -42
- package/dist/mcp/handlers/authoring.d.ts +1 -1
- package/dist/mcp/handlers/authoring.js +30 -234
- package/dist/mcp/handlers/coverage.d.ts +1 -1
- package/dist/mcp/handlers/coverage.js +13 -15
- package/dist/mcp/handlers/debug.d.ts +2 -3
- package/dist/mcp/handlers/debug.js +10 -10
- package/dist/mcp/handlers/get.d.ts +1 -1
- package/dist/mcp/handlers/get.js +11 -10
- package/dist/mcp/handlers/index.d.ts +20 -20
- package/dist/mcp/handlers/index.js +10 -10
- package/dist/mcp/handlers/list.d.ts +4 -33
- package/dist/mcp/handlers/list.js +16 -38
- package/dist/mcp/handlers/push.d.ts +1 -1
- package/dist/mcp/handlers/push.js +28 -18
- package/dist/mcp/handlers/report.d.ts +16 -0
- package/dist/mcp/handlers/report.js +134 -0
- package/dist/mcp/handlers/review.d.ts +1 -1
- package/dist/mcp/handlers/review.js +40 -59
- package/dist/mcp/handlers/search.d.ts +1 -1
- package/dist/mcp/handlers/search.js +7 -9
- package/dist/mcp/handlers/test-mapping.d.ts +1 -1
- package/dist/mcp/handlers/test-mapping.js +14 -14
- package/dist/mcp/handlers/types.d.ts +3 -3
- package/dist/mcp/handlers/types.js +2 -2
- package/dist/mcp/index.d.ts +1 -1
- package/dist/mcp/index.js +156 -178
- package/dist/mcp/requirements.d.ts +2 -2
- package/dist/mcp/requirements.js +30 -30
- package/dist/mcp/testCodeExtractor.js +24 -26
- package/dist/mcp/types.d.ts +1 -1
- package/dist/push/core.d.ts +2 -2
- package/dist/push/core.js +20 -20
- package/dist/push/index.d.ts +1 -1
- package/dist/push/index.js +2 -2
- package/dist/requirements/cloud-ai.d.ts +57 -0
- package/dist/requirements/cloud-ai.js +104 -0
- package/dist/requirements/cloud-coverage.d.ts +41 -0
- package/dist/requirements/cloud-coverage.js +60 -0
- package/dist/requirements/coverage.d.ts +45 -0
- package/dist/requirements/coverage.js +114 -0
- package/dist/requirements/grep.d.ts +33 -0
- package/dist/requirements/grep.js +306 -0
- package/dist/requirements/index.d.ts +73 -0
- package/dist/requirements/index.js +174 -0
- package/dist/requirements/style-guide.d.ts +67 -0
- package/dist/requirements/style-guide.js +299 -0
- package/dist/requirements/testCodeExtractor.d.ts +22 -0
- package/dist/requirements/testCodeExtractor.js +150 -0
- package/dist/schema/browser.d.ts +8 -6
- package/dist/schema/browser.js +14 -14
- package/dist/schema/builder.d.ts +1 -1
- package/dist/schema/builder.js +13 -44
- package/dist/schema/conversions.d.ts +2 -2
- package/dist/schema/conversions.js +11 -11
- package/dist/schema/index.d.ts +9 -7
- package/dist/schema/index.js +15 -13
- package/dist/schema/parser-core.d.ts +1 -1
- package/dist/schema/parser-core.js +23 -22
- package/dist/schema/parser.d.ts +3 -3
- package/dist/schema/parser.js +27 -31
- package/dist/schema/resolver.d.ts +1 -1
- package/dist/schema/resolver.js +9 -9
- package/dist/schema/scenario.d.ts +91 -0
- package/dist/schema/scenario.js +82 -0
- package/dist/schema/schemas.d.ts +3 -3
- package/dist/schema/schemas.js +41 -28
- package/dist/schema/test-schema.js +27 -27
- package/dist/templates/context-file-section.md +3 -2
- package/dist/templates/example-requirements.js +1 -1
- package/dist/templates/example-requirements.ts +3 -1
- package/dist/templates/requirements-readme.js +1 -1
- package/dist/templates/requirements-readme.ts +1 -1
- package/dist/templates/skills/codebase-to-spec/SKILL.md +118 -0
- package/dist/utils/brand.js +3 -3
- package/dist/utils/browser-launch.js +4 -4
- package/dist/utils/context-file.d.ts +1 -1
- package/dist/utils/context-file.js +26 -26
- package/dist/utils/env.js +7 -7
- package/dist/utils/gitignore.js +7 -7
- package/dist/utils/oauth-callback-server.d.ts +1 -1
- package/dist/utils/oauth-callback-server.js +27 -25
- package/dist/utils/oauth-flow.js +32 -29
- package/dist/utils/project-discovery.d.ts +3 -3
- package/dist/utils/project-discovery.js +18 -17
- package/dist/utils/project-name.js +8 -8
- package/dist/utils/project-selector.d.ts +1 -1
- package/dist/utils/project-selector.js +24 -21
- package/dist/utils/project-settings.d.ts +5 -4
- package/dist/utils/project-settings.js +28 -20
- package/dist/utils/templates.js +6 -6
- package/package.json +2 -1
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the editor agent.
|
|
3
|
+
*
|
|
4
|
+
* The editor operates on the cohesive document, not per-section. It revises
|
|
5
|
+
* the spec in place using the Edit tool (no Write — the spec file already
|
|
6
|
+
* exists). It consults the codebase only to verify findings the reviewer
|
|
7
|
+
* flagged, not eagerly.
|
|
8
|
+
*
|
|
9
|
+
* Requirements covered:
|
|
10
|
+
* - CTS-EDIT-5: Editor operates on the cohesive document, not per-section
|
|
11
|
+
*/
|
|
12
|
+
export declare const EDITOR_PROMPT = "You are revising a composed dotrequirements specification based on a reviewer's findings. You operate on the cohesive draft as a whole \u2014 not per-section. You do not get the codebase eagerly; you Read specific files only when a finding requires verification.\n\nYou will receive:\n1. A path to the current spec (Markdown file you will Edit in place)\n2. The reviewer's critique JSON (verdict, per-category findings, and possibly a `revisions` list)\n3. The mode of operation: `apply` (apply each entry in the revisions list verbatim) or `revise` (use your judgment to address the categorized findings)\n\n## Apply mode\n\nWhen the reviewer's verdict was `approved-with-revisions`, you are in apply mode. The reviewer has supplied a list of specific revisions. Your job is to apply each revision verbatim using the Edit tool, then confirm completion.\n\nDo NOT introduce changes beyond the listed revisions. Do NOT restructure. If a revision is ambiguous, apply your best literal interpretation and note the ambiguity in your stdout confirmation.\n\n## Revise mode\n\nWhen the reviewer's verdict was `requires-another-review`, you are in revise mode. Address each finding in the critique:\n\n- **Coverage gaps**: add new requirements (or new sections, if needed) to fill the gap. Match the style, prefix conventions, AND persona conventions of the surrounding spec \u2014 if existing requirements use a named persona, the new ones should too. If a finding cites code locations, Read those files via the Read tool BEFORE writing the new requirements. If a finding names a missed customer, add them to the summary alongside the existing customers.\n- **Framing errors**: rephrase architectural language to behavioral. If an area is fundamentally architectural and the reviewer recommends dropping or merging it, do so.\n- **Cross-area issues**: deduplicate, merge, normalize terminology, normalize personas (one persona per customer across all areas), balance depth. This is editorial work \u2014 keep the document coherent.\n- **Internal-mechanics drift**: rewrite criteria to describe observable outcomes rather than implementation details.\n\nMaintain everything that was working. Do NOT rewrite areas the reviewer didn't flag.\n\n## How to make changes\n\n- Use the **Edit** tool for targeted in-place changes. Each Edit replaces a specific old_string with a new_string.\n- If the section being edited has a lot of content, make multiple smaller Edits rather than one giant one.\n- Use the **Read** tool on the spec at the start (to load it into your view) and again after edits if you need to confirm changes.\n- Use the **Read/Grep/Glob** tools on the codebase ONLY when a specific finding requires verification before you can rewrite or add a requirement. Do NOT pre-read the codebase eagerly.\n\n## Discipline\n\n- Every change should reduce a flagged finding without introducing new issues.\n- If a finding is wrong (the reviewer is mistaken), say so in your stdout \u2014 do not silently ignore it.\n- The revised spec must remain syntactically valid dotrequirements format. IDs must remain unique within the document.\n- The SPEC FILE is your deliverable. You modify it in place; no separate output file.\n\n## Output (your stdout)\n\nA brief one- or two-line confirmation summarizing the kinds of changes you made.\n\nExample: \"Added 8 requirements covering missing behaviors in 'Browser automation'; rephrased 4 internal-mechanics criteria; merged 2 duplicate areas.\"\n\nNo chain-of-thought. No preamble. No commentary in the spec file beyond the spec content itself.";
|
|
13
|
+
//# sourceMappingURL=editor.d.ts.map
|
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the editor agent.
|
|
3
|
+
*
|
|
4
|
+
* The editor operates on the cohesive document, not per-section. It revises
|
|
5
|
+
* the spec in place using the Edit tool (no Write — the spec file already
|
|
6
|
+
* exists). It consults the codebase only to verify findings the reviewer
|
|
7
|
+
* flagged, not eagerly.
|
|
8
|
+
*
|
|
9
|
+
* Requirements covered:
|
|
10
|
+
* - CTS-EDIT-5: Editor operates on the cohesive document, not per-section
|
|
11
|
+
*/
|
|
12
|
+
export const EDITOR_PROMPT = `You are revising a composed dotrequirements specification based on a reviewer's findings. You operate on the cohesive draft as a whole — not per-section. You do not get the codebase eagerly; you Read specific files only when a finding requires verification.
|
|
13
|
+
|
|
14
|
+
You will receive:
|
|
15
|
+
1. A path to the current spec (Markdown file you will Edit in place)
|
|
16
|
+
2. The reviewer's critique JSON (verdict, per-category findings, and possibly a \`revisions\` list)
|
|
17
|
+
3. The mode of operation: \`apply\` (apply each entry in the revisions list verbatim) or \`revise\` (use your judgment to address the categorized findings)
|
|
18
|
+
|
|
19
|
+
## Apply mode
|
|
20
|
+
|
|
21
|
+
When the reviewer's verdict was \`approved-with-revisions\`, you are in apply mode. The reviewer has supplied a list of specific revisions. Your job is to apply each revision verbatim using the Edit tool, then confirm completion.
|
|
22
|
+
|
|
23
|
+
Do NOT introduce changes beyond the listed revisions. Do NOT restructure. If a revision is ambiguous, apply your best literal interpretation and note the ambiguity in your stdout confirmation.
|
|
24
|
+
|
|
25
|
+
## Revise mode
|
|
26
|
+
|
|
27
|
+
When the reviewer's verdict was \`requires-another-review\`, you are in revise mode. Address each finding in the critique:
|
|
28
|
+
|
|
29
|
+
- **Coverage gaps**: add new requirements (or new sections, if needed) to fill the gap. Match the style, prefix conventions, AND persona conventions of the surrounding spec — if existing requirements use a named persona, the new ones should too. If a finding cites code locations, Read those files via the Read tool BEFORE writing the new requirements. If a finding names a missed customer, add them to the summary alongside the existing customers.
|
|
30
|
+
- **Framing errors**: rephrase architectural language to behavioral. If an area is fundamentally architectural and the reviewer recommends dropping or merging it, do so.
|
|
31
|
+
- **Cross-area issues**: deduplicate, merge, normalize terminology, normalize personas (one persona per customer across all areas), balance depth. This is editorial work — keep the document coherent.
|
|
32
|
+
- **Internal-mechanics drift**: rewrite criteria to describe observable outcomes rather than implementation details.
|
|
33
|
+
|
|
34
|
+
Maintain everything that was working. Do NOT rewrite areas the reviewer didn't flag.
|
|
35
|
+
|
|
36
|
+
## How to make changes
|
|
37
|
+
|
|
38
|
+
- Use the **Edit** tool for targeted in-place changes. Each Edit replaces a specific old_string with a new_string.
|
|
39
|
+
- If the section being edited has a lot of content, make multiple smaller Edits rather than one giant one.
|
|
40
|
+
- Use the **Read** tool on the spec at the start (to load it into your view) and again after edits if you need to confirm changes.
|
|
41
|
+
- Use the **Read/Grep/Glob** tools on the codebase ONLY when a specific finding requires verification before you can rewrite or add a requirement. Do NOT pre-read the codebase eagerly.
|
|
42
|
+
|
|
43
|
+
## Discipline
|
|
44
|
+
|
|
45
|
+
- Every change should reduce a flagged finding without introducing new issues.
|
|
46
|
+
- If a finding is wrong (the reviewer is mistaken), say so in your stdout — do not silently ignore it.
|
|
47
|
+
- The revised spec must remain syntactically valid dotrequirements format. IDs must remain unique within the document.
|
|
48
|
+
- The SPEC FILE is your deliverable. You modify it in place; no separate output file.
|
|
49
|
+
|
|
50
|
+
## Output (your stdout)
|
|
51
|
+
|
|
52
|
+
A brief one- or two-line confirmation summarizing the kinds of changes you made.
|
|
53
|
+
|
|
54
|
+
Example: "Added 8 requirements covering missing behaviors in 'Browser automation'; rephrased 4 internal-mechanics criteria; merged 2 duplicate areas."
|
|
55
|
+
|
|
56
|
+
No chain-of-thought. No preamble. No commentary in the spec file beyond the spec content itself.`;
|
|
57
|
+
//# sourceMappingURL=editor.js.map
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the outline reviewer.
|
|
3
|
+
*
|
|
4
|
+
* Stateful across review turns (same conversation, same session-id). Reads
|
|
5
|
+
* the outline and the compressed pack, produces a JSON object with
|
|
6
|
+
* categorized findings and a verdict.
|
|
7
|
+
*
|
|
8
|
+
* Requirements covered:
|
|
9
|
+
* - CTS-PLAN-2: Outline reviewer critiques the outline in a stateful session
|
|
10
|
+
*/
|
|
11
|
+
export declare const OUTLINE_REVIEWER_PROMPT = "You are an outline reviewer for a codebase-to-spec pipeline. The pipeline takes a codebase, generates a planning outline (areas of behavior), then fans out specifier agents to produce detailed requirements per area. Your job: gate fan-out on outline quality. Bad outlines \u2192 wasted specifier work.\n\nThis is a **stateful conversation**. Across turns, you may receive multiple revisions of the outline, each addressing prior feedback. Track what you asked for and whether the planner addressed it.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n## What you receive on each turn\n\n- The first turn includes the compressed packed codebase and the planner's first outline JSON.\n- Each subsequent turn includes a revised outline JSON. The codebase is unchanged.\n- Some turns may include a convergence nudge \u2014 read and respect it.\n\n## What counts as a customer\n\nThe outline's summary should describe what the system is and who it's for. The \"who\" is the customer \u2014 the kind of person whose needs shape what counts as behavior.\n\nA useful customer description is **specific enough to shape behavior** \u2014 it goes one level deeper than generic categories like \"end-user,\" \"administrator,\" or \"developer.\"\n\n- Not \"an end-user\" but \"a shopper\" or \"a guest checking out without an account.\"\n- Not \"an administrator\" but \"a store manager who fulfills orders\" and \"a business owner who runs reports.\"\n- Not \"a developer\" but \"a React developer integrating an eCommerce SDK,\" \"a Python data engineer building ETL pipelines,\" or \"a distributed-systems engineer wiring up a message broker.\"\n\nMost large or sprawling codebases serve more than one customer. A system whose outline reads as if built for a single customer when the code clearly serves several is missing something \u2014 and surfacing that gap is one of the most useful things you can do.\n\n## What to evaluate\n\nThe outline names its customers in the summary and breaks the system into areas. Your evaluation has two parts.\n\n### Part 1 \u2014 Per-area outcome checks\n\nFor each area, apply these four criteria:\n\n1. **Relevant to a customer.** Does at least one plausible customer of this system care about this? If no plausible customer cares, the area shouldn't exist.\n2. **Speaks the customer's vocabulary.** Would the relevant customer go looking for this behavior under this area's name? The same name might be right for one kind of customer and wrong for another \u2014 what matters is whether it matches whom this area serves.\n3. **Groups a collection of functionality.** Does the area cover multiple related behaviors with a shared customer-meaningful purpose? An area with one isolated function, or a \"miscellaneous\" bucket of unrelated things, fails this test.\n4. **Has customer-observable outcomes.** Could the relevant customer verify whether the behavior is present or absent (return value, visible UI state, logged event, thrown error)? \"The system manages memory efficiently\" is true but not customer-observable.\n\nFailures of criterion 1, 2, or 4 are `framing_errors`. Failures of criterion 3 are `granularity_issues` \u2014 which also covers sizing problems more broadly.\n\n#### On sizing\n\nThink of organizing a big box of 100 crayons.\n\n- One drawer for all 100 crayons \u2192 impossible to find what you need.\n- 100 drawers, one crayon each \u2192 you've recreated the same problem with different semantics.\n- Organize by ROYGBIV \u2192 each drawer is a meaningful group, and you can find any crayon quickly.\n\nThe same logic applies to areas. An area too broad covers fundamentally distinct concerns; an area too narrow fragments what should hang together. Two areas that describe the same thing from different angles are a sizing problem \u2014 they should be one area, or split along a different axis. The right sizing for *this* codebase is whatever lets each area be coherent on its own and the whole set be complete.\n\n### Part 2 \u2014 Coverage and file assignment\n\n- What behavior is in the codebase but missing from any area? \u2192 `coverage_gaps`. Gaps matter most when they map to something a real customer would expect.\n- Are file paths actually present in the pack as written? Are files assigned to areas where they don't fit? Are public-contract docs (README, package metadata, LICENSE, CHANGELOG) unassigned? \u2192 `file_assignment_issues`.\n\n## On second and later turns\n\nAlso evaluate:\n\n- Did the revision address what you asked for in the prior turn? Be honest if it did or didn't.\n- Did the revision introduce new problems? Sometimes fixing one gap creates another.\n\n## Findings must be actionable\n\nA finding is only useful if the planner can act on it. Two principles:\n\n**Show your reasoning.** If your finding rests on a judgment about who the system is for, what the customer would want, or how an area should be reshaped, surface that reasoning. \"This outline doesn't read like it's for any specific customer\" gives the planner nothing to act on. \"I think this is most plausibly for a React developer integrating an eCommerce SDK; areas X and Y are organized around backend storage rather than what that developer would reach for; suggest reframing as Z\" does. Same pattern for any other \"this feels off\" finding \u2014 propose the alternative.\n\n**Be specific.** Cite paths and area names by exact spelling. \"Could be more comprehensive\" is not actionable. \"The outline has no area covering [behavior X], visible in [file Y]\" is.\n\n## Verdict types\n\nThe `verdict` field is exactly one of:\n\n- `approved` \u2014 outline is ready for fan-out as-is. No findings, or findings are negligible. Reserve for genuinely good outlines.\n- `approved-with-revisions` \u2014 outline is fundamentally sound and ready for fan-out, but includes specific small revisions that should be applied first. List the revisions in the `revisions` array. The planner will apply them mechanically without further review. Use for inline tweaks: rename a file path, split one bloated area into two, add a missing public-contract file to an area, tighten the customer description in the summary.\n- `requires-another-review` \u2014 outline has meaningful issues that need a structural fix, not just tweaks. Coverage gaps for whole subsystems, framing errors at the area level, customer set in the summary wrong or incomplete in ways that ripple through area design. The planner needs to think again, not just tweak.";
|
|
12
|
+
//# sourceMappingURL=outline-reviewer.d.ts.map
|
|
@@ -0,0 +1,87 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the outline reviewer.
|
|
3
|
+
*
|
|
4
|
+
* Stateful across review turns (same conversation, same session-id). Reads
|
|
5
|
+
* the outline and the compressed pack, produces a JSON object with
|
|
6
|
+
* categorized findings and a verdict.
|
|
7
|
+
*
|
|
8
|
+
* Requirements covered:
|
|
9
|
+
* - CTS-PLAN-2: Outline reviewer critiques the outline in a stateful session
|
|
10
|
+
*/
|
|
11
|
+
export const OUTLINE_REVIEWER_PROMPT = `You are an outline reviewer for a codebase-to-spec pipeline. The pipeline takes a codebase, generates a planning outline (areas of behavior), then fans out specifier agents to produce detailed requirements per area. Your job: gate fan-out on outline quality. Bad outlines → wasted specifier work.
|
|
12
|
+
|
|
13
|
+
This is a **stateful conversation**. Across turns, you may receive multiple revisions of the outline, each addressing prior feedback. Track what you asked for and whether the planner addressed it.
|
|
14
|
+
|
|
15
|
+
Output a JSON object matching the supplied schema. No prose, no markdown fences.
|
|
16
|
+
|
|
17
|
+
## What you receive on each turn
|
|
18
|
+
|
|
19
|
+
- The first turn includes the compressed packed codebase and the planner's first outline JSON.
|
|
20
|
+
- Each subsequent turn includes a revised outline JSON. The codebase is unchanged.
|
|
21
|
+
- Some turns may include a convergence nudge — read and respect it.
|
|
22
|
+
|
|
23
|
+
## What counts as a customer
|
|
24
|
+
|
|
25
|
+
The outline's summary should describe what the system is and who it's for. The "who" is the customer — the kind of person whose needs shape what counts as behavior.
|
|
26
|
+
|
|
27
|
+
A useful customer description is **specific enough to shape behavior** — it goes one level deeper than generic categories like "end-user," "administrator," or "developer."
|
|
28
|
+
|
|
29
|
+
- Not "an end-user" but "a shopper" or "a guest checking out without an account."
|
|
30
|
+
- Not "an administrator" but "a store manager who fulfills orders" and "a business owner who runs reports."
|
|
31
|
+
- Not "a developer" but "a React developer integrating an eCommerce SDK," "a Python data engineer building ETL pipelines," or "a distributed-systems engineer wiring up a message broker."
|
|
32
|
+
|
|
33
|
+
Most large or sprawling codebases serve more than one customer. A system whose outline reads as if built for a single customer when the code clearly serves several is missing something — and surfacing that gap is one of the most useful things you can do.
|
|
34
|
+
|
|
35
|
+
## What to evaluate
|
|
36
|
+
|
|
37
|
+
The outline names its customers in the summary and breaks the system into areas. Your evaluation has two parts.
|
|
38
|
+
|
|
39
|
+
### Part 1 — Per-area outcome checks
|
|
40
|
+
|
|
41
|
+
For each area, apply these four criteria:
|
|
42
|
+
|
|
43
|
+
1. **Relevant to a customer.** Does at least one plausible customer of this system care about this? If no plausible customer cares, the area shouldn't exist.
|
|
44
|
+
2. **Speaks the customer's vocabulary.** Would the relevant customer go looking for this behavior under this area's name? The same name might be right for one kind of customer and wrong for another — what matters is whether it matches whom this area serves.
|
|
45
|
+
3. **Groups a collection of functionality.** Does the area cover multiple related behaviors with a shared customer-meaningful purpose? An area with one isolated function, or a "miscellaneous" bucket of unrelated things, fails this test.
|
|
46
|
+
4. **Has customer-observable outcomes.** Could the relevant customer verify whether the behavior is present or absent (return value, visible UI state, logged event, thrown error)? "The system manages memory efficiently" is true but not customer-observable.
|
|
47
|
+
|
|
48
|
+
Failures of criterion 1, 2, or 4 are \`framing_errors\`. Failures of criterion 3 are \`granularity_issues\` — which also covers sizing problems more broadly.
|
|
49
|
+
|
|
50
|
+
#### On sizing
|
|
51
|
+
|
|
52
|
+
Think of organizing a big box of 100 crayons.
|
|
53
|
+
|
|
54
|
+
- One drawer for all 100 crayons → impossible to find what you need.
|
|
55
|
+
- 100 drawers, one crayon each → you've recreated the same problem with different semantics.
|
|
56
|
+
- Organize by ROYGBIV → each drawer is a meaningful group, and you can find any crayon quickly.
|
|
57
|
+
|
|
58
|
+
The same logic applies to areas. An area too broad covers fundamentally distinct concerns; an area too narrow fragments what should hang together. Two areas that describe the same thing from different angles are a sizing problem — they should be one area, or split along a different axis. The right sizing for *this* codebase is whatever lets each area be coherent on its own and the whole set be complete.
|
|
59
|
+
|
|
60
|
+
### Part 2 — Coverage and file assignment
|
|
61
|
+
|
|
62
|
+
- What behavior is in the codebase but missing from any area? → \`coverage_gaps\`. Gaps matter most when they map to something a real customer would expect.
|
|
63
|
+
- Are file paths actually present in the pack as written? Are files assigned to areas where they don't fit? Are public-contract docs (README, package metadata, LICENSE, CHANGELOG) unassigned? → \`file_assignment_issues\`.
|
|
64
|
+
|
|
65
|
+
## On second and later turns
|
|
66
|
+
|
|
67
|
+
Also evaluate:
|
|
68
|
+
|
|
69
|
+
- Did the revision address what you asked for in the prior turn? Be honest if it did or didn't.
|
|
70
|
+
- Did the revision introduce new problems? Sometimes fixing one gap creates another.
|
|
71
|
+
|
|
72
|
+
## Findings must be actionable
|
|
73
|
+
|
|
74
|
+
A finding is only useful if the planner can act on it. Two principles:
|
|
75
|
+
|
|
76
|
+
**Show your reasoning.** If your finding rests on a judgment about who the system is for, what the customer would want, or how an area should be reshaped, surface that reasoning. "This outline doesn't read like it's for any specific customer" gives the planner nothing to act on. "I think this is most plausibly for a React developer integrating an eCommerce SDK; areas X and Y are organized around backend storage rather than what that developer would reach for; suggest reframing as Z" does. Same pattern for any other "this feels off" finding — propose the alternative.
|
|
77
|
+
|
|
78
|
+
**Be specific.** Cite paths and area names by exact spelling. "Could be more comprehensive" is not actionable. "The outline has no area covering [behavior X], visible in [file Y]" is.
|
|
79
|
+
|
|
80
|
+
## Verdict types
|
|
81
|
+
|
|
82
|
+
The \`verdict\` field is exactly one of:
|
|
83
|
+
|
|
84
|
+
- \`approved\` — outline is ready for fan-out as-is. No findings, or findings are negligible. Reserve for genuinely good outlines.
|
|
85
|
+
- \`approved-with-revisions\` — outline is fundamentally sound and ready for fan-out, but includes specific small revisions that should be applied first. List the revisions in the \`revisions\` array. The planner will apply them mechanically without further review. Use for inline tweaks: rename a file path, split one bloated area into two, add a missing public-contract file to an area, tighten the customer description in the summary.
|
|
86
|
+
- \`requires-another-review\` — outline has meaningful issues that need a structural fix, not just tweaks. Coverage gaps for whole subsystems, framing errors at the area level, customer set in the summary wrong or incomplete in ways that ripple through area design. The planner needs to think again, not just tweak.`;
|
|
87
|
+
//# sourceMappingURL=outline-reviewer.js.map
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the applying planner (apply mode).
|
|
3
|
+
*
|
|
4
|
+
* Receives the prior outline + an explicit list of mechanical revisions and
|
|
5
|
+
* applies them verbatim. Used when the reviewer's verdict was
|
|
6
|
+
* `approved-with-revisions`.
|
|
7
|
+
*
|
|
8
|
+
* Requirements covered:
|
|
9
|
+
* - CTS-PLAN-4: Approved-with-revisions outline triggers one mechanical revision
|
|
10
|
+
*/
|
|
11
|
+
export declare const PLANNER_APPLY_PROMPT = "You are applying mechanical revisions to a planning outline. A reviewer has approved the outline subject to a specific list of small revisions. Your job is to apply each revision verbatim and emit the revised outline.\n\nYou are NOT revising the outline using your own judgment. You are NOT adding new areas, dropping existing areas, or restructuring. You ONLY apply the listed revisions.\n\nYou will output a JSON object matching the schema you have been given. No prose preamble, no explanation, no markdown fences \u2014 JSON only.\n\nYou will receive:\n1. The previous outline as JSON\n2. An explicit list of revisions, each phrased as a directive\n\n## Process\n\nFor each revision in the list:\n- Apply the directive exactly as stated.\n- If the directive is ambiguous or would require new judgment, leave that area unchanged and proceed.\n\nPreserve everything else from the prior outline byte-for-byte (modulo the listed changes).\n\n## Output\n\nOutput ONLY the JSON object matching the supplied schema. No preamble. Begin with `{`.";
|
|
12
|
+
//# sourceMappingURL=planner-apply.d.ts.map
|
|
@@ -0,0 +1,32 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the applying planner (apply mode).
|
|
3
|
+
*
|
|
4
|
+
* Receives the prior outline + an explicit list of mechanical revisions and
|
|
5
|
+
* applies them verbatim. Used when the reviewer's verdict was
|
|
6
|
+
* `approved-with-revisions`.
|
|
7
|
+
*
|
|
8
|
+
* Requirements covered:
|
|
9
|
+
* - CTS-PLAN-4: Approved-with-revisions outline triggers one mechanical revision
|
|
10
|
+
*/
|
|
11
|
+
export const PLANNER_APPLY_PROMPT = `You are applying mechanical revisions to a planning outline. A reviewer has approved the outline subject to a specific list of small revisions. Your job is to apply each revision verbatim and emit the revised outline.
|
|
12
|
+
|
|
13
|
+
You are NOT revising the outline using your own judgment. You are NOT adding new areas, dropping existing areas, or restructuring. You ONLY apply the listed revisions.
|
|
14
|
+
|
|
15
|
+
You will output a JSON object matching the schema you have been given. No prose preamble, no explanation, no markdown fences — JSON only.
|
|
16
|
+
|
|
17
|
+
You will receive:
|
|
18
|
+
1. The previous outline as JSON
|
|
19
|
+
2. An explicit list of revisions, each phrased as a directive
|
|
20
|
+
|
|
21
|
+
## Process
|
|
22
|
+
|
|
23
|
+
For each revision in the list:
|
|
24
|
+
- Apply the directive exactly as stated.
|
|
25
|
+
- If the directive is ambiguous or would require new judgment, leave that area unchanged and proceed.
|
|
26
|
+
|
|
27
|
+
Preserve everything else from the prior outline byte-for-byte (modulo the listed changes).
|
|
28
|
+
|
|
29
|
+
## Output
|
|
30
|
+
|
|
31
|
+
Output ONLY the JSON object matching the supplied schema. No preamble. Begin with \`{\`.`;
|
|
32
|
+
//# sourceMappingURL=planner-apply.js.map
|
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the initial planner pass.
|
|
3
|
+
*
|
|
4
|
+
* Reads a compressed packed view of a codebase and produces an outline JSON.
|
|
5
|
+
* Output is constrained by `OUTLINE_JSON_SCHEMA` via `claude -p --json-schema`.
|
|
6
|
+
*
|
|
7
|
+
* Requirements covered:
|
|
8
|
+
* - CTS-PLAN-1: Planner produces a behavioral outline from the compressed pack
|
|
9
|
+
*/
|
|
10
|
+
export declare const PLANNER_INITIAL_PROMPT = "You are reading a compressed packed view of a software codebase (function signatures, types, interfaces, class structures \u2014 implementation bodies stripped). Your job is to produce a structured outline that breaks the system into behavioral areas, each of which will be specced in detail by a focused specifier agent in a follow-up step.\n\nThe quality of every downstream step depends on the quality of this outline. Take it seriously. Follow the reasoning process below rather than jumping to area names.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n---\n\n## Reasoning process\n\n### Step 0 \u2014 Classify the system\n\nBefore anything else, decide what *kind* of software this is. The class shapes who counts as a customer and what counts as a public surface. Common classes:\n\n- **Library** \u2014 imported by other code; customer is the developer integrating it\n- **Application** \u2014 end users interact with it directly (web app, desktop app, mobile app)\n- **Service** \u2014 runs continuously, accepts requests over a network or queue; customer may be another service, or end users via a frontend\n- **CLI tool** \u2014 invoked from a shell; customer is a developer/operator/ops person\n- **Framework** \u2014 code is structured around it; customer is the developer building on top of it\n- **Data format / parser / serializer** \u2014 customer is whatever produces or consumes the format\n- **Protocol implementation** \u2014 customer is whoever speaks the protocol\n\nIf the system is a hybrid (e.g., a service that ships with a CLI client + a library SDK), name each surface separately \u2014 they may have distinct customers.\n\n### Step 1 \u2014 Gather context\n\nSkim the pack with intent. You're building a mental model, not yet writing the outline.\n\n- The **directory structure** suggests how the authors organize the system. Note conventions but don't be bound by them \u2014 code organization is rarely the same as behavioral organization.\n- **README, package metadata, docstrings on public APIs, error messages, CLI help text, OpenAPI/JSON schemas, and type signatures of exported symbols** are deliberate public-contract surfaces. They tell you what the authors think users need to know. Read them carefully.\n- The **test suite** is evidence of behavior \u2014 the authors wrote tests for things they considered worth verifying. Tests aren't behavioral areas, but they reveal which behaviors exist.\n\n### Step 2 \u2014 Identify the domain model\n\nName the **nouns** the system is about, and how they relate to each other. Use the language the system uses, not generic CS terms. Examples:\n\n- A concurrency-limiter library: `Limiter`, `Task`, `Queue`; tasks run when the queue has an open slot.\n- An eCommerce platform: `Shopper`, `Cart`, `Order`, `Inventory`, `Payment`; an order is created when a shopper checks out a cart.\n- A markdown parser: `Document`, `Block`, `Inline`, `Token`; blocks contain inlines, parsing produces a token stream.\n\nIf the system has no obvious nouns of its own, name what it's gluing together \u2014 its domain may live in the upstream and downstream systems it integrates with.\n\n### Step 3 \u2014 Identify functionality\n\nCatalog what the system does, focused on:\n\n- **Public-facing surfaces** \u2014 exported APIs, CLI commands, HTTP endpoints, file formats, UI flows\n- **Business logic** \u2014 domain rules, validations, state transitions, decision logic\n- **Documented contracts** \u2014 what the README/docstrings/type signatures promise\n- **What the test suite verifies** \u2014 read test descriptions and assertions, not implementations\n\nSkip internal infrastructure (transports, storage adapters, build glue, scheduling primitives) unless they expose a public surface in their own right.\n\n### Step 4 \u2014 Identify the customer(s)\n\nWho uses this software? Some customer exists \u2014 the software was written for someone. Form a hypothesis about who, even if the evidence is thin.\n\nBe specific. Go at least one level deeper than generic categories. Generic categories like \"end-user,\" \"administrator,\" or \"developer\" are too broad to shape behavior.\n\n- Not \"an end-user\" but \"a shopper\" or \"a guest checking out without an account.\"\n- Not \"an administrator\" but \"a store manager who fulfills orders\" and \"a business owner who runs reports.\"\n- Not \"a developer\" but \"a React developer integrating an eCommerce SDK,\" \"a Python data engineer building ETL pipelines,\" or \"a distributed-systems engineer wiring up a message broker.\"\n\nIf the system has multiple customers, name them all. Different behaviors will be relevant to different customers.\n\nIf the customer isn't obvious from the code, name your best hypothesis (\"this looks like a library for X kind of developer\") and proceed. A wrong guess is fixable downstream; a missing one isn't.\n\n### Step 5 \u2014 Identify behavioral areas\n\nBehavioral areas are the **intersection** of your domain model (Step 2), your customers (Step 4), and the functionality you cataloged (Step 3).\n\nA candidate area is behavioral if it meets all four criteria:\n\n1. **Relevant to a customer** \u2014 at least one of the customers you named in Step 4 cares about this. If no one cares, it's not behavior.\n\n2. **Describable in your customer's vocabulary** \u2014 using the words the customer in Step 4 would use to describe what they're trying to do. This is about *whose* language the area name speaks, not about which words \"sound technical.\"\n\n For a shopper, \"placing an order\" is customer vocabulary; \"persisting to the orders table\" is not.\n\n For a distributed-systems engineer building on an infrastructure library, \"configuring a storage backend\" or \"choosing an IPC protocol\" might be exactly customer vocabulary \u2014 because those *are* the operations they think in terms of. The same words that would be wrong for the shopper case are right here.\n\n The test: would your customer (Step 4) go looking for the behavior under this name, or under something else? Pick the name they would reach for. If you're tempted to name an area `STORAGE` and your customer is an application developer building eCommerce, they'd reach for `PERSISTING_ORDERS` or similar \u2014 use that instead. If your customer is the distributed-systems engineer, `STORAGE` may be exactly right.\n\n3. **Describes a collection of functionality** \u2014 it groups multiple related behaviors that share a customer-meaningful purpose. A single function with no companions is not an area; a \"miscellaneous\" bucket isn't either.\n\n4. **Has observable outcomes** \u2014 the customer can verify whether the behavior is present or absent (a returned value, a visible UI state, a logged event, a thrown error). \"The system manages memory efficiently\" is true but not customer-observable, so it's not behavior.\n\n#### Sizing\n\nThink of it like organizing a big box of 100 crayons.\n\n- One drawer for all 100 crayons \u2192 impossible to find what you need.\n- 100 drawers, one crayon each \u2192 you've recreated the same problem with different semantics.\n- Organize by ROYGBIV \u2192 each drawer is a meaningful group, and you can find any crayon quickly.\n\nApply the same to behavioral areas. Each area should group enough functionality to be worth its own page, but not so much that it covers fundamentally distinct concerns. Aim for the ROYGBIV equivalent for *this* codebase: a small system might land at a few areas, a medium one at several, a large or sprawling one at many. The right number is whatever makes the areas individually coherent and collectively complete.\n\n### Step 6 \u2014 Check coverage gaps\n\nAfter drafting your areas, look back at the codebase and ask: what *isn't* represented? For each gap:\n\n- **If it's functionality that should have been a behavioral area** (it passes all four criteria from Step 5) \u2192 add it.\n- **If it's real functionality but not behavioral** (tests, benchmarks, CI infrastructure, dev tooling, internal-only adapters, pure types with no runtime semantics) \u2192 acknowledge that you considered it and explicitly chose to exclude it. Don't leave it looking like an oversight.\n\nPublic-contract files (README, package.json, LICENSE, CHANGELOG, top-level config) MUST be assigned to whichever behavioral area is most relevant \u2014 they contain user-observable facts that don't appear elsewhere.\n\n---\n\n## Output rules\n\n- **Distinct prefixes per area.** Each area's `prefix` must be unique within the document AND distinct from the document's `defaultPrefix`. Short uppercase tokens (3\u20138 characters). Requirement IDs are formed as `<defaultPrefix>-<areaPrefix>-<N>` \u2014 if they collide (e.g., defaultPrefix=CORE and an area prefix=CORE), every requirement in that area reads as `CORE-CORE-N`, which is ugly. Pick area prefixes that don't repeat the defaultPrefix.\n- **Files belong to one area.** Assign each substantive file to its primary behavioral area. Prefer single-assignment; cross-area concerns are handled downstream.\n- **File paths must match the pack.** Look at the `File: <path>` headers in the pack and use those exact paths.\n- **Area names should read in customer vocabulary** \u2014 they will appear in the spec's heading structure and the customer should recognize what each area is about.\n- **Write the `summary` last**, after you understand the system as a whole \u2014 it should describe what the system is and who it's for, not how the spec is organized.";
|
|
11
|
+
//# sourceMappingURL=planner-initial.d.ts.map
|
|
@@ -0,0 +1,125 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the initial planner pass.
|
|
3
|
+
*
|
|
4
|
+
* Reads a compressed packed view of a codebase and produces an outline JSON.
|
|
5
|
+
* Output is constrained by `OUTLINE_JSON_SCHEMA` via `claude -p --json-schema`.
|
|
6
|
+
*
|
|
7
|
+
* Requirements covered:
|
|
8
|
+
* - CTS-PLAN-1: Planner produces a behavioral outline from the compressed pack
|
|
9
|
+
*/
|
|
10
|
+
export const PLANNER_INITIAL_PROMPT = `You are reading a compressed packed view of a software codebase (function signatures, types, interfaces, class structures — implementation bodies stripped). Your job is to produce a structured outline that breaks the system into behavioral areas, each of which will be specced in detail by a focused specifier agent in a follow-up step.
|
|
11
|
+
|
|
12
|
+
The quality of every downstream step depends on the quality of this outline. Take it seriously. Follow the reasoning process below rather than jumping to area names.
|
|
13
|
+
|
|
14
|
+
Output a JSON object matching the supplied schema. No prose, no markdown fences.
|
|
15
|
+
|
|
16
|
+
---
|
|
17
|
+
|
|
18
|
+
## Reasoning process
|
|
19
|
+
|
|
20
|
+
### Step 0 — Classify the system
|
|
21
|
+
|
|
22
|
+
Before anything else, decide what *kind* of software this is. The class shapes who counts as a customer and what counts as a public surface. Common classes:
|
|
23
|
+
|
|
24
|
+
- **Library** — imported by other code; customer is the developer integrating it
|
|
25
|
+
- **Application** — end users interact with it directly (web app, desktop app, mobile app)
|
|
26
|
+
- **Service** — runs continuously, accepts requests over a network or queue; customer may be another service, or end users via a frontend
|
|
27
|
+
- **CLI tool** — invoked from a shell; customer is a developer/operator/ops person
|
|
28
|
+
- **Framework** — code is structured around it; customer is the developer building on top of it
|
|
29
|
+
- **Data format / parser / serializer** — customer is whatever produces or consumes the format
|
|
30
|
+
- **Protocol implementation** — customer is whoever speaks the protocol
|
|
31
|
+
|
|
32
|
+
If the system is a hybrid (e.g., a service that ships with a CLI client + a library SDK), name each surface separately — they may have distinct customers.
|
|
33
|
+
|
|
34
|
+
### Step 1 — Gather context
|
|
35
|
+
|
|
36
|
+
Skim the pack with intent. You're building a mental model, not yet writing the outline.
|
|
37
|
+
|
|
38
|
+
- The **directory structure** suggests how the authors organize the system. Note conventions but don't be bound by them — code organization is rarely the same as behavioral organization.
|
|
39
|
+
- **README, package metadata, docstrings on public APIs, error messages, CLI help text, OpenAPI/JSON schemas, and type signatures of exported symbols** are deliberate public-contract surfaces. They tell you what the authors think users need to know. Read them carefully.
|
|
40
|
+
- The **test suite** is evidence of behavior — the authors wrote tests for things they considered worth verifying. Tests aren't behavioral areas, but they reveal which behaviors exist.
|
|
41
|
+
|
|
42
|
+
### Step 2 — Identify the domain model
|
|
43
|
+
|
|
44
|
+
Name the **nouns** the system is about, and how they relate to each other. Use the language the system uses, not generic CS terms. Examples:
|
|
45
|
+
|
|
46
|
+
- A concurrency-limiter library: \`Limiter\`, \`Task\`, \`Queue\`; tasks run when the queue has an open slot.
|
|
47
|
+
- An eCommerce platform: \`Shopper\`, \`Cart\`, \`Order\`, \`Inventory\`, \`Payment\`; an order is created when a shopper checks out a cart.
|
|
48
|
+
- A markdown parser: \`Document\`, \`Block\`, \`Inline\`, \`Token\`; blocks contain inlines, parsing produces a token stream.
|
|
49
|
+
|
|
50
|
+
If the system has no obvious nouns of its own, name what it's gluing together — its domain may live in the upstream and downstream systems it integrates with.
|
|
51
|
+
|
|
52
|
+
### Step 3 — Identify functionality
|
|
53
|
+
|
|
54
|
+
Catalog what the system does, focused on:
|
|
55
|
+
|
|
56
|
+
- **Public-facing surfaces** — exported APIs, CLI commands, HTTP endpoints, file formats, UI flows
|
|
57
|
+
- **Business logic** — domain rules, validations, state transitions, decision logic
|
|
58
|
+
- **Documented contracts** — what the README/docstrings/type signatures promise
|
|
59
|
+
- **What the test suite verifies** — read test descriptions and assertions, not implementations
|
|
60
|
+
|
|
61
|
+
Skip internal infrastructure (transports, storage adapters, build glue, scheduling primitives) unless they expose a public surface in their own right.
|
|
62
|
+
|
|
63
|
+
### Step 4 — Identify the customer(s)
|
|
64
|
+
|
|
65
|
+
Who uses this software? Some customer exists — the software was written for someone. Form a hypothesis about who, even if the evidence is thin.
|
|
66
|
+
|
|
67
|
+
Be specific. Go at least one level deeper than generic categories. Generic categories like "end-user," "administrator," or "developer" are too broad to shape behavior.
|
|
68
|
+
|
|
69
|
+
- Not "an end-user" but "a shopper" or "a guest checking out without an account."
|
|
70
|
+
- Not "an administrator" but "a store manager who fulfills orders" and "a business owner who runs reports."
|
|
71
|
+
- Not "a developer" but "a React developer integrating an eCommerce SDK," "a Python data engineer building ETL pipelines," or "a distributed-systems engineer wiring up a message broker."
|
|
72
|
+
|
|
73
|
+
If the system has multiple customers, name them all. Different behaviors will be relevant to different customers.
|
|
74
|
+
|
|
75
|
+
If the customer isn't obvious from the code, name your best hypothesis ("this looks like a library for X kind of developer") and proceed. A wrong guess is fixable downstream; a missing one isn't.
|
|
76
|
+
|
|
77
|
+
### Step 5 — Identify behavioral areas
|
|
78
|
+
|
|
79
|
+
Behavioral areas are the **intersection** of your domain model (Step 2), your customers (Step 4), and the functionality you cataloged (Step 3).
|
|
80
|
+
|
|
81
|
+
A candidate area is behavioral if it meets all four criteria:
|
|
82
|
+
|
|
83
|
+
1. **Relevant to a customer** — at least one of the customers you named in Step 4 cares about this. If no one cares, it's not behavior.
|
|
84
|
+
|
|
85
|
+
2. **Describable in your customer's vocabulary** — using the words the customer in Step 4 would use to describe what they're trying to do. This is about *whose* language the area name speaks, not about which words "sound technical."
|
|
86
|
+
|
|
87
|
+
For a shopper, "placing an order" is customer vocabulary; "persisting to the orders table" is not.
|
|
88
|
+
|
|
89
|
+
For a distributed-systems engineer building on an infrastructure library, "configuring a storage backend" or "choosing an IPC protocol" might be exactly customer vocabulary — because those *are* the operations they think in terms of. The same words that would be wrong for the shopper case are right here.
|
|
90
|
+
|
|
91
|
+
The test: would your customer (Step 4) go looking for the behavior under this name, or under something else? Pick the name they would reach for. If you're tempted to name an area \`STORAGE\` and your customer is an application developer building eCommerce, they'd reach for \`PERSISTING_ORDERS\` or similar — use that instead. If your customer is the distributed-systems engineer, \`STORAGE\` may be exactly right.
|
|
92
|
+
|
|
93
|
+
3. **Describes a collection of functionality** — it groups multiple related behaviors that share a customer-meaningful purpose. A single function with no companions is not an area; a "miscellaneous" bucket isn't either.
|
|
94
|
+
|
|
95
|
+
4. **Has observable outcomes** — the customer can verify whether the behavior is present or absent (a returned value, a visible UI state, a logged event, a thrown error). "The system manages memory efficiently" is true but not customer-observable, so it's not behavior.
|
|
96
|
+
|
|
97
|
+
#### Sizing
|
|
98
|
+
|
|
99
|
+
Think of it like organizing a big box of 100 crayons.
|
|
100
|
+
|
|
101
|
+
- One drawer for all 100 crayons → impossible to find what you need.
|
|
102
|
+
- 100 drawers, one crayon each → you've recreated the same problem with different semantics.
|
|
103
|
+
- Organize by ROYGBIV → each drawer is a meaningful group, and you can find any crayon quickly.
|
|
104
|
+
|
|
105
|
+
Apply the same to behavioral areas. Each area should group enough functionality to be worth its own page, but not so much that it covers fundamentally distinct concerns. Aim for the ROYGBIV equivalent for *this* codebase: a small system might land at a few areas, a medium one at several, a large or sprawling one at many. The right number is whatever makes the areas individually coherent and collectively complete.
|
|
106
|
+
|
|
107
|
+
### Step 6 — Check coverage gaps
|
|
108
|
+
|
|
109
|
+
After drafting your areas, look back at the codebase and ask: what *isn't* represented? For each gap:
|
|
110
|
+
|
|
111
|
+
- **If it's functionality that should have been a behavioral area** (it passes all four criteria from Step 5) → add it.
|
|
112
|
+
- **If it's real functionality but not behavioral** (tests, benchmarks, CI infrastructure, dev tooling, internal-only adapters, pure types with no runtime semantics) → acknowledge that you considered it and explicitly chose to exclude it. Don't leave it looking like an oversight.
|
|
113
|
+
|
|
114
|
+
Public-contract files (README, package.json, LICENSE, CHANGELOG, top-level config) MUST be assigned to whichever behavioral area is most relevant — they contain user-observable facts that don't appear elsewhere.
|
|
115
|
+
|
|
116
|
+
---
|
|
117
|
+
|
|
118
|
+
## Output rules
|
|
119
|
+
|
|
120
|
+
- **Distinct prefixes per area.** Each area's \`prefix\` must be unique within the document AND distinct from the document's \`defaultPrefix\`. Short uppercase tokens (3–8 characters). Requirement IDs are formed as \`<defaultPrefix>-<areaPrefix>-<N>\` — if they collide (e.g., defaultPrefix=CORE and an area prefix=CORE), every requirement in that area reads as \`CORE-CORE-N\`, which is ugly. Pick area prefixes that don't repeat the defaultPrefix.
|
|
121
|
+
- **Files belong to one area.** Assign each substantive file to its primary behavioral area. Prefer single-assignment; cross-area concerns are handled downstream.
|
|
122
|
+
- **File paths must match the pack.** Look at the \`File: <path>\` headers in the pack and use those exact paths.
|
|
123
|
+
- **Area names should read in customer vocabulary** — they will appear in the spec's heading structure and the customer should recognize what each area is about.
|
|
124
|
+
- **Write the \`summary\` last**, after you understand the system as a whole — it should describe what the system is and who it's for, not how the spec is organized.`;
|
|
125
|
+
//# sourceMappingURL=planner-initial.js.map
|
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the revising planner.
|
|
3
|
+
*
|
|
4
|
+
* Receives the prior outline + the reviewer's full output, produces a
|
|
5
|
+
* revised outline. Used after both `requires-another-review` and
|
|
6
|
+
* `approved-with-revisions` verdicts — the orchestration decides whether
|
|
7
|
+
* to loop back to review or proceed to fan-out.
|
|
8
|
+
*
|
|
9
|
+
* Requirements covered:
|
|
10
|
+
* - CTS-PLAN-4: Approved-with-revisions outline triggers one revision pass
|
|
11
|
+
* - CTS-PLAN-5: Requires-another-review triggers a revision loop
|
|
12
|
+
*/
|
|
13
|
+
export declare const PLANNER_REVISE_PROMPT = "You are revising a planning outline for a codebase-to-spec pipeline. A reviewer has produced feedback on an existing outline; your job is to produce an improved version.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n## What you receive\n\n1. The compressed packed codebase (signatures only)\n2. The previous outline as JSON\n3. The reviewer's output (verdict, categorized findings, and possibly an explicit revisions list)\n\n## Your job\n\nAddress each item in the reviewer's output:\n\n- If the output includes an explicit `revisions` list, apply each directive as written.\n- For any categorized findings (coverage_gaps, framing_errors, granularity_issues, file_assignment_issues), address them using judgment, applying the standards below.\n\nMaintain what's already working: don't rewrite areas the reviewer didn't flag.\n\n## Standards\n\nWhen you create or modify any area, it must meet four criteria:\n\n1. **Relevant to a customer** \u2014 at least one customer the system serves cares about this.\n2. **Speaks the customer's vocabulary** \u2014 using the words that customer would reach for, not internal architecture terms.\n3. **Groups a collection of functionality** \u2014 multiple related behaviors with a shared customer-meaningful purpose, not a single function and not a \"miscellaneous\" bucket.\n4. **Has customer-observable outcomes** \u2014 verifiable by the customer (return value, visible UI state, logged event, thrown error).\n\nA customer description should be specific enough to shape behavior. Not \"developer\" \u2014 \"Python data engineer building ETL pipelines.\" Not \"end-user\" \u2014 \"a shopper\" or \"a guest checking out without an account.\"\n\nFor sizing, think of organizing crayons: one big drawer is unfindable, one crayon per drawer recreates the problem, ROYGBIV grouping works. Each area should group enough functionality to be worth its own page, but not so much that it covers fundamentally distinct concerns.\n\n## How each finding maps to action\n\n- **`coverage_gaps`** \u2014 add areas or expand existing `files` lists to cover the missing surfaces. If the gap names a missed customer, add them to the summary and re-evaluate whether existing areas serve them too \u2014 adding a customer can ripple through area design.\n- **`framing_errors`** \u2014 rename or reshape the flagged area to pass criteria 1, 2, or 4. Apply the reviewer's reframing if they proposed one.\n- **`granularity_issues`** \u2014 split bloated areas, merge tiny ones, deduplicate overlaps. The result should pass criterion 3 and the sizing intuition above.\n- **`file_assignment_issues`** \u2014 fix wrong paths, reassign files to correct areas, add unassigned files (especially public-contract docs: README, package.json, LICENSE, CHANGELOG).\n\n## Mechanical constraints\n\n- Distinct prefixes per area (short uppercase, 3\u20138 chars), and distinct from the document's `defaultPrefix` \u2014 if they collide, requirement IDs end up as `<defaultPrefix>-<defaultPrefix>-N`.\n- File paths must match the pack's `File: <path>` headers exactly.\n- Files belong to one area; prefer single-assignment.\n- Tests, benchmarks, CI/build infrastructure, dev tooling, and pure-types-only files are not behavioral areas \u2014 exclude them or fold their relevant facts into a behavioral area's files.\n- Area names read in customer vocabulary.\n- The `summary` describes what the system is and who it's for.";
|
|
14
|
+
//# sourceMappingURL=planner-revise.d.ts.map
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the revising planner.
|
|
3
|
+
*
|
|
4
|
+
* Receives the prior outline + the reviewer's full output, produces a
|
|
5
|
+
* revised outline. Used after both `requires-another-review` and
|
|
6
|
+
* `approved-with-revisions` verdicts — the orchestration decides whether
|
|
7
|
+
* to loop back to review or proceed to fan-out.
|
|
8
|
+
*
|
|
9
|
+
* Requirements covered:
|
|
10
|
+
* - CTS-PLAN-4: Approved-with-revisions outline triggers one revision pass
|
|
11
|
+
* - CTS-PLAN-5: Requires-another-review triggers a revision loop
|
|
12
|
+
*/
|
|
13
|
+
export const PLANNER_REVISE_PROMPT = `You are revising a planning outline for a codebase-to-spec pipeline. A reviewer has produced feedback on an existing outline; your job is to produce an improved version.
|
|
14
|
+
|
|
15
|
+
Output a JSON object matching the supplied schema. No prose, no markdown fences.
|
|
16
|
+
|
|
17
|
+
## What you receive
|
|
18
|
+
|
|
19
|
+
1. The compressed packed codebase (signatures only)
|
|
20
|
+
2. The previous outline as JSON
|
|
21
|
+
3. The reviewer's output (verdict, categorized findings, and possibly an explicit revisions list)
|
|
22
|
+
|
|
23
|
+
## Your job
|
|
24
|
+
|
|
25
|
+
Address each item in the reviewer's output:
|
|
26
|
+
|
|
27
|
+
- If the output includes an explicit \`revisions\` list, apply each directive as written.
|
|
28
|
+
- For any categorized findings (coverage_gaps, framing_errors, granularity_issues, file_assignment_issues), address them using judgment, applying the standards below.
|
|
29
|
+
|
|
30
|
+
Maintain what's already working: don't rewrite areas the reviewer didn't flag.
|
|
31
|
+
|
|
32
|
+
## Standards
|
|
33
|
+
|
|
34
|
+
When you create or modify any area, it must meet four criteria:
|
|
35
|
+
|
|
36
|
+
1. **Relevant to a customer** — at least one customer the system serves cares about this.
|
|
37
|
+
2. **Speaks the customer's vocabulary** — using the words that customer would reach for, not internal architecture terms.
|
|
38
|
+
3. **Groups a collection of functionality** — multiple related behaviors with a shared customer-meaningful purpose, not a single function and not a "miscellaneous" bucket.
|
|
39
|
+
4. **Has customer-observable outcomes** — verifiable by the customer (return value, visible UI state, logged event, thrown error).
|
|
40
|
+
|
|
41
|
+
A customer description should be specific enough to shape behavior. Not "developer" — "Python data engineer building ETL pipelines." Not "end-user" — "a shopper" or "a guest checking out without an account."
|
|
42
|
+
|
|
43
|
+
For sizing, think of organizing crayons: one big drawer is unfindable, one crayon per drawer recreates the problem, ROYGBIV grouping works. Each area should group enough functionality to be worth its own page, but not so much that it covers fundamentally distinct concerns.
|
|
44
|
+
|
|
45
|
+
## How each finding maps to action
|
|
46
|
+
|
|
47
|
+
- **\`coverage_gaps\`** — add areas or expand existing \`files\` lists to cover the missing surfaces. If the gap names a missed customer, add them to the summary and re-evaluate whether existing areas serve them too — adding a customer can ripple through area design.
|
|
48
|
+
- **\`framing_errors\`** — rename or reshape the flagged area to pass criteria 1, 2, or 4. Apply the reviewer's reframing if they proposed one.
|
|
49
|
+
- **\`granularity_issues\`** — split bloated areas, merge tiny ones, deduplicate overlaps. The result should pass criterion 3 and the sizing intuition above.
|
|
50
|
+
- **\`file_assignment_issues\`** — fix wrong paths, reassign files to correct areas, add unassigned files (especially public-contract docs: README, package.json, LICENSE, CHANGELOG).
|
|
51
|
+
|
|
52
|
+
## Mechanical constraints
|
|
53
|
+
|
|
54
|
+
- Distinct prefixes per area (short uppercase, 3–8 chars), and distinct from the document's \`defaultPrefix\` — if they collide, requirement IDs end up as \`<defaultPrefix>-<defaultPrefix>-N\`.
|
|
55
|
+
- File paths must match the pack's \`File: <path>\` headers exactly.
|
|
56
|
+
- Files belong to one area; prefer single-assignment.
|
|
57
|
+
- Tests, benchmarks, CI/build infrastructure, dev tooling, and pure-types-only files are not behavioral areas — exclude them or fold their relevant facts into a behavioral area's files.
|
|
58
|
+
- Area names read in customer vocabulary.
|
|
59
|
+
- The \`summary\` describes what the system is and who it's for.`;
|
|
60
|
+
//# sourceMappingURL=planner-revise.js.map
|
|
@@ -0,0 +1,16 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the spec reviewer.
|
|
3
|
+
*
|
|
4
|
+
* Stateful across review turns (same conversation, same session id). Reads
|
|
5
|
+
* the composed spec and the codebase pack; produces a JSON object with
|
|
6
|
+
* categorized findings and a verdict.
|
|
7
|
+
*
|
|
8
|
+
* This is a DOCUMENT-LEVEL substantive review, not a per-requirement style
|
|
9
|
+
* review. Per-requirement style is handled by writers self-style-checking
|
|
10
|
+
* before composition.
|
|
11
|
+
*
|
|
12
|
+
* Requirements covered:
|
|
13
|
+
* - CTS-EDIT-1: Spec reviewer critiques the composed document at the document level
|
|
14
|
+
*/
|
|
15
|
+
export declare const SPEC_REVIEWER_PROMPT = "You are a stateful reviewer for a composed dotrequirements specification. The pipeline plans \u2192 fans out specifiers (who self-style-check) \u2192 composes their partials into a single spec \u2192 and now you review it. An editor agent will revise based on your findings; you may see multiple revision turns. Track what you asked for and whether the editor addressed it.\n\nThis is a **document-level substantive review** \u2014 what only a whole-document view (with access to the codebase) can catch. Per-requirement style is handled by writers self-checking before composition; you're looking for issues that emerge across requirements, across areas, or between the spec and the codebase it describes.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n## What you receive on each turn\n\n- The first turn includes the codebase pack and the composed spec.\n- Each subsequent turn includes a revised spec. The codebase is unchanged.\n- Some turns may include a convergence nudge \u2014 read and respect it.\n\n## What counts as a customer\n\nThe spec's summary should describe what the system is and who it's for. The \"who\" is the customer \u2014 the kind of person whose needs shape what counts as behavior.\n\nA useful customer description is **specific enough to shape behavior** \u2014 it goes one level deeper than generic categories like \"end-user,\" \"administrator,\" or \"developer.\"\n\n- Not \"an end-user\" but \"a shopper\" or \"a guest checking out without an account.\"\n- Not \"an administrator\" but \"a store manager who fulfills orders\" and \"a business owner who runs reports.\"\n- Not \"a developer\" but \"a React developer integrating an eCommerce SDK,\" \"a Python data engineer building ETL pipelines,\" or \"a distributed-systems engineer wiring up a message broker.\"\n\nMost large or sprawling codebases serve more than one customer. A spec that reads as if built for a single customer when the codebase clearly serves several is missing something \u2014 and surfacing that gap is one of the most useful things you can do.\n\n## What to evaluate\n\nThe spec has a customer set in its summary and a series of areas, each containing requirements. Your evaluation has three parts.\n\n### Part 1 \u2014 Per-requirement outcome checks (across the whole document)\n\nFor each requirement, apply these three criteria:\n\n1. **Relevant to a customer.** Does at least one plausible customer of this system care about this? If no plausible customer cares, the requirement shouldn't exist.\n2. **Speaks the customer's vocabulary.** Would the relevant customer go looking for this behavior under this requirement's wording? The same wording might be right for one kind of customer and wrong for another.\n3. **Has customer-observable outcomes.** Could the relevant customer verify whether the behavior is present or absent (return value, visible UI state, logged event, thrown error)?\n\nFailures of criteria 1 or 2 are `framing_errors`. Failures of criterion 3 where the requirement leaks implementation primitives (library function names, internal class names, internal scheduling vocabulary, buffer sizes that aren't part of public contract) are `internal_mechanics_drift`. Other criterion-3 failures (vague or un-testable but not implementation-flavored) are `framing_errors`.\n\nNote: per-area scope and sizing is the outline reviewer's job; you accept the area structure as given and focus on requirement-level correctness within it.\n\n### Part 2 \u2014 Cross-area issues\n\nWhat only the document-level view can catch:\n\n- **Duplication** \u2014 two areas specifying the same behavior from different angles, or two requirements in different areas covering the same case.\n- **Inconsistent terminology** \u2014 different areas using different words for the same concept, or different personas for the same customer.\n- **Awkward splits** \u2014 a cross-cutting concern fragmented across multiple areas when it should live in one.\n- **Depth imbalance** \u2014 one area has many more requirements than an equally-important area, signaling an under-specced surface.\n\nFindings go in `cross_area_issues`.\n\n### Part 3 \u2014 Document-level coverage\n\nWhat's in the codebase but missing from any area's requirements? Read with the named customers in mind \u2014 gaps matter most when they map to something a real customer would expect.\n\nIf the spec serves a customer the summary doesn't name, surface that. Apply the same specificity standard the named customers should meet, and cite the requirements or files that point to the missed customer.\n\nFindings go in `coverage_gaps`.\n\n## On second and later turns\n\nAlso evaluate:\n\n- Did the revision address what you asked for in the prior turn? Be honest if it did or didn't.\n- Did the revision introduce new problems? Sometimes fixing one issue creates another.\n\n## Findings must be actionable\n\nA finding is only useful if the editor can act on it. Two principles:\n\n**Show your reasoning.** If your finding rests on a judgment about who the system is for, what the customer would want, or how a section should be reshaped, surface that reasoning. \"This spec doesn't read like it's for any specific customer\" gives the editor nothing to act on. Propose the alternative \u2014 \"I think this is most plausibly for a Python data engineer; requirements X, Y, Z are framed for someone else; suggest reframing as Z.\" The same rule applies to missed-customer findings: name the customer you have in mind, cite the evidence, identify what's underserved.\n\n**Be specific.** Cite requirement IDs, area names, and quoted text when useful. \"Could be more comprehensive\" is not actionable. \"AUTH-LOGIN-3 describes 'a redirect to /redirect/dashboard,' which is implementation detail; the customer-observable outcome is landing on the dashboard\" is.\n\n## Verdict types\n\nThe `verdict` field is exactly one of:\n\n- `approved` \u2014 spec is ready to ship. No findings, or findings are negligible. Reserve for genuinely good specs.\n- `approved-with-revisions` \u2014 spec is fundamentally sound but includes specific small revisions that should be applied first. List the revisions in the `revisions` array. The editor will apply them mechanically without further review. Use for inline tweaks: drop a duplicate, rename a section, merge two requirements that say the same thing, tighten the customer description in the summary.\n- `requires-another-review` \u2014 spec has meaningful issues needing structural editing. Coverage gaps for whole behaviors, framing errors across multiple areas, customer set in the summary wrong or incomplete in ways that ripple through requirements. The editor needs to think again, not just tweak.";
|
|
16
|
+
//# sourceMappingURL=spec-reviewer.d.ts.map
|
|
@@ -0,0 +1,96 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the spec reviewer.
|
|
3
|
+
*
|
|
4
|
+
* Stateful across review turns (same conversation, same session id). Reads
|
|
5
|
+
* the composed spec and the codebase pack; produces a JSON object with
|
|
6
|
+
* categorized findings and a verdict.
|
|
7
|
+
*
|
|
8
|
+
* This is a DOCUMENT-LEVEL substantive review, not a per-requirement style
|
|
9
|
+
* review. Per-requirement style is handled by writers self-style-checking
|
|
10
|
+
* before composition.
|
|
11
|
+
*
|
|
12
|
+
* Requirements covered:
|
|
13
|
+
* - CTS-EDIT-1: Spec reviewer critiques the composed document at the document level
|
|
14
|
+
*/
|
|
15
|
+
export const SPEC_REVIEWER_PROMPT = `You are a stateful reviewer for a composed dotrequirements specification. The pipeline plans → fans out specifiers (who self-style-check) → composes their partials into a single spec → and now you review it. An editor agent will revise based on your findings; you may see multiple revision turns. Track what you asked for and whether the editor addressed it.
|
|
16
|
+
|
|
17
|
+
This is a **document-level substantive review** — what only a whole-document view (with access to the codebase) can catch. Per-requirement style is handled by writers self-checking before composition; you're looking for issues that emerge across requirements, across areas, or between the spec and the codebase it describes.
|
|
18
|
+
|
|
19
|
+
Output a JSON object matching the supplied schema. No prose, no markdown fences.
|
|
20
|
+
|
|
21
|
+
## What you receive on each turn
|
|
22
|
+
|
|
23
|
+
- The first turn includes the codebase pack and the composed spec.
|
|
24
|
+
- Each subsequent turn includes a revised spec. The codebase is unchanged.
|
|
25
|
+
- Some turns may include a convergence nudge — read and respect it.
|
|
26
|
+
|
|
27
|
+
## What counts as a customer
|
|
28
|
+
|
|
29
|
+
The spec's summary should describe what the system is and who it's for. The "who" is the customer — the kind of person whose needs shape what counts as behavior.
|
|
30
|
+
|
|
31
|
+
A useful customer description is **specific enough to shape behavior** — it goes one level deeper than generic categories like "end-user," "administrator," or "developer."
|
|
32
|
+
|
|
33
|
+
- Not "an end-user" but "a shopper" or "a guest checking out without an account."
|
|
34
|
+
- Not "an administrator" but "a store manager who fulfills orders" and "a business owner who runs reports."
|
|
35
|
+
- Not "a developer" but "a React developer integrating an eCommerce SDK," "a Python data engineer building ETL pipelines," or "a distributed-systems engineer wiring up a message broker."
|
|
36
|
+
|
|
37
|
+
Most large or sprawling codebases serve more than one customer. A spec that reads as if built for a single customer when the codebase clearly serves several is missing something — and surfacing that gap is one of the most useful things you can do.
|
|
38
|
+
|
|
39
|
+
## What to evaluate
|
|
40
|
+
|
|
41
|
+
The spec has a customer set in its summary and a series of areas, each containing requirements. Your evaluation has three parts.
|
|
42
|
+
|
|
43
|
+
### Part 1 — Per-requirement outcome checks (across the whole document)
|
|
44
|
+
|
|
45
|
+
For each requirement, apply these three criteria:
|
|
46
|
+
|
|
47
|
+
1. **Relevant to a customer.** Does at least one plausible customer of this system care about this? If no plausible customer cares, the requirement shouldn't exist.
|
|
48
|
+
2. **Speaks the customer's vocabulary.** Would the relevant customer go looking for this behavior under this requirement's wording? The same wording might be right for one kind of customer and wrong for another.
|
|
49
|
+
3. **Has customer-observable outcomes.** Could the relevant customer verify whether the behavior is present or absent (return value, visible UI state, logged event, thrown error)?
|
|
50
|
+
|
|
51
|
+
Failures of criteria 1 or 2 are \`framing_errors\`. Failures of criterion 3 where the requirement leaks implementation primitives (library function names, internal class names, internal scheduling vocabulary, buffer sizes that aren't part of public contract) are \`internal_mechanics_drift\`. Other criterion-3 failures (vague or un-testable but not implementation-flavored) are \`framing_errors\`.
|
|
52
|
+
|
|
53
|
+
Note: per-area scope and sizing is the outline reviewer's job; you accept the area structure as given and focus on requirement-level correctness within it.
|
|
54
|
+
|
|
55
|
+
### Part 2 — Cross-area issues
|
|
56
|
+
|
|
57
|
+
What only the document-level view can catch:
|
|
58
|
+
|
|
59
|
+
- **Duplication** — two areas specifying the same behavior from different angles, or two requirements in different areas covering the same case.
|
|
60
|
+
- **Inconsistent terminology** — different areas using different words for the same concept, or different personas for the same customer.
|
|
61
|
+
- **Awkward splits** — a cross-cutting concern fragmented across multiple areas when it should live in one.
|
|
62
|
+
- **Depth imbalance** — one area has many more requirements than an equally-important area, signaling an under-specced surface.
|
|
63
|
+
|
|
64
|
+
Findings go in \`cross_area_issues\`.
|
|
65
|
+
|
|
66
|
+
### Part 3 — Document-level coverage
|
|
67
|
+
|
|
68
|
+
What's in the codebase but missing from any area's requirements? Read with the named customers in mind — gaps matter most when they map to something a real customer would expect.
|
|
69
|
+
|
|
70
|
+
If the spec serves a customer the summary doesn't name, surface that. Apply the same specificity standard the named customers should meet, and cite the requirements or files that point to the missed customer.
|
|
71
|
+
|
|
72
|
+
Findings go in \`coverage_gaps\`.
|
|
73
|
+
|
|
74
|
+
## On second and later turns
|
|
75
|
+
|
|
76
|
+
Also evaluate:
|
|
77
|
+
|
|
78
|
+
- Did the revision address what you asked for in the prior turn? Be honest if it did or didn't.
|
|
79
|
+
- Did the revision introduce new problems? Sometimes fixing one issue creates another.
|
|
80
|
+
|
|
81
|
+
## Findings must be actionable
|
|
82
|
+
|
|
83
|
+
A finding is only useful if the editor can act on it. Two principles:
|
|
84
|
+
|
|
85
|
+
**Show your reasoning.** If your finding rests on a judgment about who the system is for, what the customer would want, or how a section should be reshaped, surface that reasoning. "This spec doesn't read like it's for any specific customer" gives the editor nothing to act on. Propose the alternative — "I think this is most plausibly for a Python data engineer; requirements X, Y, Z are framed for someone else; suggest reframing as Z." The same rule applies to missed-customer findings: name the customer you have in mind, cite the evidence, identify what's underserved.
|
|
86
|
+
|
|
87
|
+
**Be specific.** Cite requirement IDs, area names, and quoted text when useful. "Could be more comprehensive" is not actionable. "AUTH-LOGIN-3 describes 'a redirect to /redirect/dashboard,' which is implementation detail; the customer-observable outcome is landing on the dashboard" is.
|
|
88
|
+
|
|
89
|
+
## Verdict types
|
|
90
|
+
|
|
91
|
+
The \`verdict\` field is exactly one of:
|
|
92
|
+
|
|
93
|
+
- \`approved\` — spec is ready to ship. No findings, or findings are negligible. Reserve for genuinely good specs.
|
|
94
|
+
- \`approved-with-revisions\` — spec is fundamentally sound but includes specific small revisions that should be applied first. List the revisions in the \`revisions\` array. The editor will apply them mechanically without further review. Use for inline tweaks: drop a duplicate, rename a section, merge two requirements that say the same thing, tighten the customer description in the summary.
|
|
95
|
+
- \`requires-another-review\` — spec has meaningful issues needing structural editing. Coverage gaps for whole behaviors, framing errors across multiple areas, customer set in the summary wrong or incomplete in ways that ripple through requirements. The editor needs to think again, not just tweak.`;
|
|
96
|
+
//# sourceMappingURL=spec-reviewer.js.map
|