@popoverai/dotrequirements 0.23.0 → 0.24.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +167 -20
- package/dist/cli.js +121 -60
- package/dist/codebase-to-spec/budget.d.ts +53 -0
- package/dist/codebase-to-spec/budget.js +80 -0
- package/dist/codebase-to-spec/cache.d.ts +49 -0
- package/dist/codebase-to-spec/cache.js +54 -0
- package/dist/codebase-to-spec/claude.d.ts +69 -0
- package/dist/codebase-to-spec/claude.js +126 -0
- package/dist/codebase-to-spec/compose.d.ts +49 -0
- package/dist/codebase-to-spec/compose.js +124 -0
- package/dist/codebase-to-spec/edit-loop.d.ts +54 -0
- package/dist/codebase-to-spec/edit-loop.js +195 -0
- package/dist/codebase-to-spec/editor.d.ts +54 -0
- package/dist/codebase-to-spec/editor.js +74 -0
- package/dist/codebase-to-spec/exit-codes.d.ts +40 -0
- package/dist/codebase-to-spec/exit-codes.js +58 -0
- package/dist/codebase-to-spec/fan-out.d.ts +63 -0
- package/dist/codebase-to-spec/fan-out.js +215 -0
- package/dist/codebase-to-spec/interactive.d.ts +30 -0
- package/dist/codebase-to-spec/interactive.js +48 -0
- package/dist/codebase-to-spec/outline-review-loop.d.ts +51 -0
- package/dist/codebase-to-spec/outline-review-loop.js +187 -0
- package/dist/codebase-to-spec/pack.d.ts +51 -0
- package/dist/codebase-to-spec/pack.js +127 -0
- package/dist/codebase-to-spec/planner.d.ts +41 -0
- package/dist/codebase-to-spec/planner.js +76 -0
- package/dist/codebase-to-spec/present.d.ts +94 -0
- package/dist/codebase-to-spec/present.js +288 -0
- package/dist/codebase-to-spec/progress.d.ts +33 -0
- package/dist/codebase-to-spec/progress.js +28 -0
- package/dist/codebase-to-spec/prompts/editor.d.ts +13 -0
- package/dist/codebase-to-spec/prompts/editor.js +57 -0
- package/dist/codebase-to-spec/prompts/outline-reviewer.d.ts +12 -0
- package/dist/codebase-to-spec/prompts/outline-reviewer.js +87 -0
- package/dist/codebase-to-spec/prompts/planner-apply.d.ts +12 -0
- package/dist/codebase-to-spec/prompts/planner-apply.js +32 -0
- package/dist/codebase-to-spec/prompts/planner-initial.d.ts +11 -0
- package/dist/codebase-to-spec/prompts/planner-initial.js +125 -0
- package/dist/codebase-to-spec/prompts/planner-revise.d.ts +14 -0
- package/dist/codebase-to-spec/prompts/planner-revise.js +60 -0
- package/dist/codebase-to-spec/prompts/spec-reviewer.d.ts +16 -0
- package/dist/codebase-to-spec/prompts/spec-reviewer.js +96 -0
- package/dist/codebase-to-spec/prompts/specifier.d.ts +12 -0
- package/dist/codebase-to-spec/prompts/specifier.js +100 -0
- package/dist/codebase-to-spec/prompts/style-check.d.ts +12 -0
- package/dist/codebase-to-spec/prompts/style-check.js +78 -0
- package/dist/codebase-to-spec/schemas.d.ts +257 -0
- package/dist/codebase-to-spec/schemas.js +183 -0
- package/dist/codebase-to-spec/skill-install.d.ts +57 -0
- package/dist/codebase-to-spec/skill-install.js +79 -0
- package/dist/codebase-to-spec/slice.d.ts +49 -0
- package/dist/codebase-to-spec/slice.js +111 -0
- package/dist/codebase-to-spec/specifier.d.ts +60 -0
- package/dist/codebase-to-spec/specifier.js +79 -0
- package/dist/codebase-to-spec/style-check.d.ts +29 -0
- package/dist/codebase-to-spec/style-check.js +33 -0
- package/dist/codebase-to-spec/summary.d.ts +51 -0
- package/dist/codebase-to-spec/summary.js +183 -0
- package/dist/codebase-to-spec/validate.d.ts +46 -0
- package/dist/codebase-to-spec/validate.js +130 -0
- package/dist/commands/acceptance-test.d.ts +6 -0
- package/dist/commands/acceptance-test.js +212 -0
- package/dist/commands/ai-setup.d.ts +5 -0
- package/dist/commands/ai-setup.js +441 -0
- package/dist/commands/browsertest.js +34 -27
- package/dist/commands/codebase-to-spec/compose.d.ts +14 -0
- package/dist/commands/codebase-to-spec/compose.js +57 -0
- package/dist/commands/codebase-to-spec/edit-loop.d.ts +16 -0
- package/dist/commands/codebase-to-spec/edit-loop.js +83 -0
- package/dist/commands/codebase-to-spec/fan-out.d.ts +19 -0
- package/dist/commands/codebase-to-spec/fan-out.js +77 -0
- package/dist/commands/codebase-to-spec/index.d.ts +9 -0
- package/dist/commands/codebase-to-spec/index.js +135 -0
- package/dist/commands/codebase-to-spec/pack.d.ts +22 -0
- package/dist/commands/codebase-to-spec/pack.js +76 -0
- package/dist/commands/codebase-to-spec/plan-loop.d.ts +26 -0
- package/dist/commands/codebase-to-spec/plan-loop.js +105 -0
- package/dist/commands/codebase-to-spec/present.d.ts +21 -0
- package/dist/commands/codebase-to-spec/present.js +92 -0
- package/dist/commands/codebase-to-spec/run.d.ts +20 -0
- package/dist/commands/codebase-to-spec/run.js +85 -0
- package/dist/commands/codebase-to-spec/skill-install.d.ts +20 -0
- package/dist/commands/codebase-to-spec/skill-install.js +51 -0
- package/dist/commands/codebase-to-spec/specify-area.d.ts +18 -0
- package/dist/commands/codebase-to-spec/specify-area.js +82 -0
- package/dist/commands/codebase-to-spec/style-check.d.ts +15 -0
- package/dist/commands/codebase-to-spec/style-check.js +42 -0
- package/dist/commands/codebase-to-spec/validate.d.ts +18 -0
- package/dist/commands/codebase-to-spec/validate.js +38 -0
- package/dist/commands/create-requirement-document.d.ts +2 -0
- package/dist/commands/create-requirement-document.js +41 -0
- package/dist/commands/finalize.js +7 -7
- package/dist/commands/get.d.ts +2 -0
- package/dist/commands/get.js +55 -0
- package/dist/commands/init.js +132 -117
- package/dist/commands/link.js +27 -27
- package/dist/commands/list.d.ts +6 -0
- package/dist/commands/list.js +43 -0
- package/dist/commands/mcp-setup.js +159 -149
- package/dist/commands/mcp.js +1 -1
- package/dist/commands/prepare.js +4 -4
- package/dist/commands/pull.js +116 -121
- package/dist/commands/push.js +106 -112
- package/dist/commands/report.d.ts +6 -2
- package/dist/commands/report.js +177 -122
- package/dist/commands/requirements-for.d.ts +2 -0
- package/dist/commands/requirements-for.js +29 -0
- package/dist/commands/review-test.d.ts +2 -0
- package/dist/commands/review-test.js +75 -0
- package/dist/commands/search.d.ts +6 -0
- package/dist/commands/search.js +39 -0
- package/dist/commands/style-check.d.ts +7 -0
- package/dist/commands/style-check.js +75 -0
- package/dist/commands/test.js +53 -59
- package/dist/commands/tests-for.d.ts +2 -0
- package/dist/commands/tests-for.js +80 -0
- package/dist/commands/validate.d.ts +6 -0
- package/dist/commands/validate.js +72 -0
- package/dist/config.js +1 -1
- package/dist/convex.d.ts +34 -22
- package/dist/convex.js +38 -22
- package/dist/harness/cache.d.ts +1 -5
- package/dist/harness/cache.js +49 -59
- package/dist/harness/convexReporting.d.ts +1 -1
- package/dist/harness/convexReporting.js +9 -7
- package/dist/harness/coverageCache.js +3 -3
- package/dist/harness/finalize.js +59 -46
- package/dist/harness/index.d.ts +6 -7
- package/dist/harness/index.js +9 -10
- package/dist/harness/prepare.js +6 -5
- package/dist/harness/requirementsLoader.d.ts +2 -2
- package/dist/harness/requirementsLoader.js +13 -35
- package/dist/harness/tracking.js +18 -18
- package/dist/harness/types.d.ts +1 -1
- package/dist/mcp/convexClient.d.ts +0 -39
- package/dist/mcp/convexClient.js +2 -107
- package/dist/mcp/grep.d.ts +1 -1
- package/dist/mcp/grep.js +87 -42
- package/dist/mcp/handlers/authoring.d.ts +1 -1
- package/dist/mcp/handlers/authoring.js +30 -234
- package/dist/mcp/handlers/coverage.d.ts +1 -1
- package/dist/mcp/handlers/coverage.js +13 -15
- package/dist/mcp/handlers/debug.d.ts +2 -3
- package/dist/mcp/handlers/debug.js +10 -10
- package/dist/mcp/handlers/get.d.ts +1 -1
- package/dist/mcp/handlers/get.js +11 -10
- package/dist/mcp/handlers/index.d.ts +20 -20
- package/dist/mcp/handlers/index.js +10 -10
- package/dist/mcp/handlers/list.d.ts +4 -33
- package/dist/mcp/handlers/list.js +16 -38
- package/dist/mcp/handlers/push.d.ts +1 -1
- package/dist/mcp/handlers/push.js +28 -18
- package/dist/mcp/handlers/report.d.ts +16 -0
- package/dist/mcp/handlers/report.js +134 -0
- package/dist/mcp/handlers/review.d.ts +1 -1
- package/dist/mcp/handlers/review.js +40 -59
- package/dist/mcp/handlers/search.d.ts +1 -1
- package/dist/mcp/handlers/search.js +7 -9
- package/dist/mcp/handlers/test-mapping.d.ts +1 -1
- package/dist/mcp/handlers/test-mapping.js +14 -14
- package/dist/mcp/handlers/types.d.ts +3 -3
- package/dist/mcp/handlers/types.js +2 -2
- package/dist/mcp/index.d.ts +1 -1
- package/dist/mcp/index.js +147 -167
- package/dist/mcp/requirements.d.ts +2 -2
- package/dist/mcp/requirements.js +30 -30
- package/dist/mcp/testCodeExtractor.js +24 -26
- package/dist/mcp/types.d.ts +1 -1
- package/dist/push/core.d.ts +2 -2
- package/dist/push/core.js +20 -20
- package/dist/push/index.d.ts +1 -1
- package/dist/push/index.js +2 -2
- package/dist/requirements/cloud-ai.d.ts +57 -0
- package/dist/requirements/cloud-ai.js +104 -0
- package/dist/requirements/cloud-coverage.d.ts +41 -0
- package/dist/requirements/cloud-coverage.js +60 -0
- package/dist/requirements/coverage.d.ts +45 -0
- package/dist/requirements/coverage.js +114 -0
- package/dist/requirements/grep.d.ts +33 -0
- package/dist/requirements/grep.js +306 -0
- package/dist/requirements/index.d.ts +73 -0
- package/dist/requirements/index.js +174 -0
- package/dist/requirements/style-guide.d.ts +67 -0
- package/dist/requirements/style-guide.js +299 -0
- package/dist/requirements/testCodeExtractor.d.ts +22 -0
- package/dist/requirements/testCodeExtractor.js +150 -0
- package/dist/schema/browser.d.ts +8 -8
- package/dist/schema/browser.js +13 -15
- package/dist/schema/builder.d.ts +1 -1
- package/dist/schema/builder.js +13 -44
- package/dist/schema/conversions.d.ts +2 -2
- package/dist/schema/conversions.js +11 -11
- package/dist/schema/index.d.ts +9 -9
- package/dist/schema/index.js +15 -15
- package/dist/schema/parser-core.d.ts +1 -1
- package/dist/schema/parser-core.js +23 -22
- package/dist/schema/parser.d.ts +3 -3
- package/dist/schema/parser.js +27 -31
- package/dist/schema/resolver.d.ts +1 -1
- package/dist/schema/resolver.js +9 -9
- package/dist/schema/scenario.d.ts +1 -1
- package/dist/schema/scenario.js +1 -1
- package/dist/schema/schemas.d.ts +3 -3
- package/dist/schema/schemas.js +41 -28
- package/dist/schema/test-schema.js +27 -27
- package/dist/templates/context-file-section.md +3 -2
- package/dist/templates/example-requirements.js +1 -1
- package/dist/templates/example-requirements.ts +3 -1
- package/dist/templates/requirements-readme.js +1 -1
- package/dist/templates/requirements-readme.ts +1 -1
- package/dist/templates/skills/codebase-to-spec/SKILL.md +118 -0
- package/dist/utils/brand.js +3 -3
- package/dist/utils/browser-launch.js +4 -4
- package/dist/utils/context-file.d.ts +1 -1
- package/dist/utils/context-file.js +26 -26
- package/dist/utils/env.js +7 -7
- package/dist/utils/gitignore.js +7 -7
- package/dist/utils/oauth-callback-server.d.ts +1 -1
- package/dist/utils/oauth-callback-server.js +27 -25
- package/dist/utils/oauth-flow.js +32 -29
- package/dist/utils/project-discovery.d.ts +3 -3
- package/dist/utils/project-discovery.js +18 -17
- package/dist/utils/project-name.js +8 -8
- package/dist/utils/project-selector.d.ts +1 -1
- package/dist/utils/project-selector.js +24 -21
- package/dist/utils/project-settings.d.ts +1 -1
- package/dist/utils/project-settings.js +24 -22
- package/dist/utils/templates.js +6 -6
- package/package.json +2 -1
|
@@ -0,0 +1,125 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the initial planner pass.
|
|
3
|
+
*
|
|
4
|
+
* Reads a compressed packed view of a codebase and produces an outline JSON.
|
|
5
|
+
* Output is constrained by `OUTLINE_JSON_SCHEMA` via `claude -p --json-schema`.
|
|
6
|
+
*
|
|
7
|
+
* Requirements covered:
|
|
8
|
+
* - CTS-PLAN-1: Planner produces a behavioral outline from the compressed pack
|
|
9
|
+
*/
|
|
10
|
+
export const PLANNER_INITIAL_PROMPT = `You are reading a compressed packed view of a software codebase (function signatures, types, interfaces, class structures — implementation bodies stripped). Your job is to produce a structured outline that breaks the system into behavioral areas, each of which will be specced in detail by a focused specifier agent in a follow-up step.
|
|
11
|
+
|
|
12
|
+
The quality of every downstream step depends on the quality of this outline. Take it seriously. Follow the reasoning process below rather than jumping to area names.
|
|
13
|
+
|
|
14
|
+
Output a JSON object matching the supplied schema. No prose, no markdown fences.
|
|
15
|
+
|
|
16
|
+
---
|
|
17
|
+
|
|
18
|
+
## Reasoning process
|
|
19
|
+
|
|
20
|
+
### Step 0 — Classify the system
|
|
21
|
+
|
|
22
|
+
Before anything else, decide what *kind* of software this is. The class shapes who counts as a customer and what counts as a public surface. Common classes:
|
|
23
|
+
|
|
24
|
+
- **Library** — imported by other code; customer is the developer integrating it
|
|
25
|
+
- **Application** — end users interact with it directly (web app, desktop app, mobile app)
|
|
26
|
+
- **Service** — runs continuously, accepts requests over a network or queue; customer may be another service, or end users via a frontend
|
|
27
|
+
- **CLI tool** — invoked from a shell; customer is a developer/operator/ops person
|
|
28
|
+
- **Framework** — code is structured around it; customer is the developer building on top of it
|
|
29
|
+
- **Data format / parser / serializer** — customer is whatever produces or consumes the format
|
|
30
|
+
- **Protocol implementation** — customer is whoever speaks the protocol
|
|
31
|
+
|
|
32
|
+
If the system is a hybrid (e.g., a service that ships with a CLI client + a library SDK), name each surface separately — they may have distinct customers.
|
|
33
|
+
|
|
34
|
+
### Step 1 — Gather context
|
|
35
|
+
|
|
36
|
+
Skim the pack with intent. You're building a mental model, not yet writing the outline.
|
|
37
|
+
|
|
38
|
+
- The **directory structure** suggests how the authors organize the system. Note conventions but don't be bound by them — code organization is rarely the same as behavioral organization.
|
|
39
|
+
- **README, package metadata, docstrings on public APIs, error messages, CLI help text, OpenAPI/JSON schemas, and type signatures of exported symbols** are deliberate public-contract surfaces. They tell you what the authors think users need to know. Read them carefully.
|
|
40
|
+
- The **test suite** is evidence of behavior — the authors wrote tests for things they considered worth verifying. Tests aren't behavioral areas, but they reveal which behaviors exist.
|
|
41
|
+
|
|
42
|
+
### Step 2 — Identify the domain model
|
|
43
|
+
|
|
44
|
+
Name the **nouns** the system is about, and how they relate to each other. Use the language the system uses, not generic CS terms. Examples:
|
|
45
|
+
|
|
46
|
+
- A concurrency-limiter library: \`Limiter\`, \`Task\`, \`Queue\`; tasks run when the queue has an open slot.
|
|
47
|
+
- An eCommerce platform: \`Shopper\`, \`Cart\`, \`Order\`, \`Inventory\`, \`Payment\`; an order is created when a shopper checks out a cart.
|
|
48
|
+
- A markdown parser: \`Document\`, \`Block\`, \`Inline\`, \`Token\`; blocks contain inlines, parsing produces a token stream.
|
|
49
|
+
|
|
50
|
+
If the system has no obvious nouns of its own, name what it's gluing together — its domain may live in the upstream and downstream systems it integrates with.
|
|
51
|
+
|
|
52
|
+
### Step 3 — Identify functionality
|
|
53
|
+
|
|
54
|
+
Catalog what the system does, focused on:
|
|
55
|
+
|
|
56
|
+
- **Public-facing surfaces** — exported APIs, CLI commands, HTTP endpoints, file formats, UI flows
|
|
57
|
+
- **Business logic** — domain rules, validations, state transitions, decision logic
|
|
58
|
+
- **Documented contracts** — what the README/docstrings/type signatures promise
|
|
59
|
+
- **What the test suite verifies** — read test descriptions and assertions, not implementations
|
|
60
|
+
|
|
61
|
+
Skip internal infrastructure (transports, storage adapters, build glue, scheduling primitives) unless they expose a public surface in their own right.
|
|
62
|
+
|
|
63
|
+
### Step 4 — Identify the customer(s)
|
|
64
|
+
|
|
65
|
+
Who uses this software? Some customer exists — the software was written for someone. Form a hypothesis about who, even if the evidence is thin.
|
|
66
|
+
|
|
67
|
+
Be specific. Go at least one level deeper than generic categories. Generic categories like "end-user," "administrator," or "developer" are too broad to shape behavior.
|
|
68
|
+
|
|
69
|
+
- Not "an end-user" but "a shopper" or "a guest checking out without an account."
|
|
70
|
+
- Not "an administrator" but "a store manager who fulfills orders" and "a business owner who runs reports."
|
|
71
|
+
- Not "a developer" but "a React developer integrating an eCommerce SDK," "a Python data engineer building ETL pipelines," or "a distributed-systems engineer wiring up a message broker."
|
|
72
|
+
|
|
73
|
+
If the system has multiple customers, name them all. Different behaviors will be relevant to different customers.
|
|
74
|
+
|
|
75
|
+
If the customer isn't obvious from the code, name your best hypothesis ("this looks like a library for X kind of developer") and proceed. A wrong guess is fixable downstream; a missing one isn't.
|
|
76
|
+
|
|
77
|
+
### Step 5 — Identify behavioral areas
|
|
78
|
+
|
|
79
|
+
Behavioral areas are the **intersection** of your domain model (Step 2), your customers (Step 4), and the functionality you cataloged (Step 3).
|
|
80
|
+
|
|
81
|
+
A candidate area is behavioral if it meets all four criteria:
|
|
82
|
+
|
|
83
|
+
1. **Relevant to a customer** — at least one of the customers you named in Step 4 cares about this. If no one cares, it's not behavior.
|
|
84
|
+
|
|
85
|
+
2. **Describable in your customer's vocabulary** — using the words the customer in Step 4 would use to describe what they're trying to do. This is about *whose* language the area name speaks, not about which words "sound technical."
|
|
86
|
+
|
|
87
|
+
For a shopper, "placing an order" is customer vocabulary; "persisting to the orders table" is not.
|
|
88
|
+
|
|
89
|
+
For a distributed-systems engineer building on an infrastructure library, "configuring a storage backend" or "choosing an IPC protocol" might be exactly customer vocabulary — because those *are* the operations they think in terms of. The same words that would be wrong for the shopper case are right here.
|
|
90
|
+
|
|
91
|
+
The test: would your customer (Step 4) go looking for the behavior under this name, or under something else? Pick the name they would reach for. If you're tempted to name an area \`STORAGE\` and your customer is an application developer building eCommerce, they'd reach for \`PERSISTING_ORDERS\` or similar — use that instead. If your customer is the distributed-systems engineer, \`STORAGE\` may be exactly right.
|
|
92
|
+
|
|
93
|
+
3. **Describes a collection of functionality** — it groups multiple related behaviors that share a customer-meaningful purpose. A single function with no companions is not an area; a "miscellaneous" bucket isn't either.
|
|
94
|
+
|
|
95
|
+
4. **Has observable outcomes** — the customer can verify whether the behavior is present or absent (a returned value, a visible UI state, a logged event, a thrown error). "The system manages memory efficiently" is true but not customer-observable, so it's not behavior.
|
|
96
|
+
|
|
97
|
+
#### Sizing
|
|
98
|
+
|
|
99
|
+
Think of it like organizing a big box of 100 crayons.
|
|
100
|
+
|
|
101
|
+
- One drawer for all 100 crayons → impossible to find what you need.
|
|
102
|
+
- 100 drawers, one crayon each → you've recreated the same problem with different semantics.
|
|
103
|
+
- Organize by ROYGBIV → each drawer is a meaningful group, and you can find any crayon quickly.
|
|
104
|
+
|
|
105
|
+
Apply the same to behavioral areas. Each area should group enough functionality to be worth its own page, but not so much that it covers fundamentally distinct concerns. Aim for the ROYGBIV equivalent for *this* codebase: a small system might land at a few areas, a medium one at several, a large or sprawling one at many. The right number is whatever makes the areas individually coherent and collectively complete.
|
|
106
|
+
|
|
107
|
+
### Step 6 — Check coverage gaps
|
|
108
|
+
|
|
109
|
+
After drafting your areas, look back at the codebase and ask: what *isn't* represented? For each gap:
|
|
110
|
+
|
|
111
|
+
- **If it's functionality that should have been a behavioral area** (it passes all four criteria from Step 5) → add it.
|
|
112
|
+
- **If it's real functionality but not behavioral** (tests, benchmarks, CI infrastructure, dev tooling, internal-only adapters, pure types with no runtime semantics) → acknowledge that you considered it and explicitly chose to exclude it. Don't leave it looking like an oversight.
|
|
113
|
+
|
|
114
|
+
Public-contract files (README, package.json, LICENSE, CHANGELOG, top-level config) MUST be assigned to whichever behavioral area is most relevant — they contain user-observable facts that don't appear elsewhere.
|
|
115
|
+
|
|
116
|
+
---
|
|
117
|
+
|
|
118
|
+
## Output rules
|
|
119
|
+
|
|
120
|
+
- **Distinct prefixes per area.** Each area's \`prefix\` must be unique within the document AND distinct from the document's \`defaultPrefix\`. Short uppercase tokens (3–8 characters). Requirement IDs are formed as \`<defaultPrefix>-<areaPrefix>-<N>\` — if they collide (e.g., defaultPrefix=CORE and an area prefix=CORE), every requirement in that area reads as \`CORE-CORE-N\`, which is ugly. Pick area prefixes that don't repeat the defaultPrefix.
|
|
121
|
+
- **Files belong to one area.** Assign each substantive file to its primary behavioral area. Prefer single-assignment; cross-area concerns are handled downstream.
|
|
122
|
+
- **File paths must match the pack.** Look at the \`File: <path>\` headers in the pack and use those exact paths.
|
|
123
|
+
- **Area names should read in customer vocabulary** — they will appear in the spec's heading structure and the customer should recognize what each area is about.
|
|
124
|
+
- **Write the \`summary\` last**, after you understand the system as a whole — it should describe what the system is and who it's for, not how the spec is organized.`;
|
|
125
|
+
//# sourceMappingURL=planner-initial.js.map
|
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the revising planner.
|
|
3
|
+
*
|
|
4
|
+
* Receives the prior outline + the reviewer's full output, produces a
|
|
5
|
+
* revised outline. Used after both `requires-another-review` and
|
|
6
|
+
* `approved-with-revisions` verdicts — the orchestration decides whether
|
|
7
|
+
* to loop back to review or proceed to fan-out.
|
|
8
|
+
*
|
|
9
|
+
* Requirements covered:
|
|
10
|
+
* - CTS-PLAN-4: Approved-with-revisions outline triggers one revision pass
|
|
11
|
+
* - CTS-PLAN-5: Requires-another-review triggers a revision loop
|
|
12
|
+
*/
|
|
13
|
+
export declare const PLANNER_REVISE_PROMPT = "You are revising a planning outline for a codebase-to-spec pipeline. A reviewer has produced feedback on an existing outline; your job is to produce an improved version.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n## What you receive\n\n1. The compressed packed codebase (signatures only)\n2. The previous outline as JSON\n3. The reviewer's output (verdict, categorized findings, and possibly an explicit revisions list)\n\n## Your job\n\nAddress each item in the reviewer's output:\n\n- If the output includes an explicit `revisions` list, apply each directive as written.\n- For any categorized findings (coverage_gaps, framing_errors, granularity_issues, file_assignment_issues), address them using judgment, applying the standards below.\n\nMaintain what's already working: don't rewrite areas the reviewer didn't flag.\n\n## Standards\n\nWhen you create or modify any area, it must meet four criteria:\n\n1. **Relevant to a customer** \u2014 at least one customer the system serves cares about this.\n2. **Speaks the customer's vocabulary** \u2014 using the words that customer would reach for, not internal architecture terms.\n3. **Groups a collection of functionality** \u2014 multiple related behaviors with a shared customer-meaningful purpose, not a single function and not a \"miscellaneous\" bucket.\n4. **Has customer-observable outcomes** \u2014 verifiable by the customer (return value, visible UI state, logged event, thrown error).\n\nA customer description should be specific enough to shape behavior. Not \"developer\" \u2014 \"Python data engineer building ETL pipelines.\" Not \"end-user\" \u2014 \"a shopper\" or \"a guest checking out without an account.\"\n\nFor sizing, think of organizing crayons: one big drawer is unfindable, one crayon per drawer recreates the problem, ROYGBIV grouping works. Each area should group enough functionality to be worth its own page, but not so much that it covers fundamentally distinct concerns.\n\n## How each finding maps to action\n\n- **`coverage_gaps`** \u2014 add areas or expand existing `files` lists to cover the missing surfaces. If the gap names a missed customer, add them to the summary and re-evaluate whether existing areas serve them too \u2014 adding a customer can ripple through area design.\n- **`framing_errors`** \u2014 rename or reshape the flagged area to pass criteria 1, 2, or 4. Apply the reviewer's reframing if they proposed one.\n- **`granularity_issues`** \u2014 split bloated areas, merge tiny ones, deduplicate overlaps. The result should pass criterion 3 and the sizing intuition above.\n- **`file_assignment_issues`** \u2014 fix wrong paths, reassign files to correct areas, add unassigned files (especially public-contract docs: README, package.json, LICENSE, CHANGELOG).\n\n## Mechanical constraints\n\n- Distinct prefixes per area (short uppercase, 3\u20138 chars), and distinct from the document's `defaultPrefix` \u2014 if they collide, requirement IDs end up as `<defaultPrefix>-<defaultPrefix>-N`.\n- File paths must match the pack's `File: <path>` headers exactly.\n- Files belong to one area; prefer single-assignment.\n- Tests, benchmarks, CI/build infrastructure, dev tooling, and pure-types-only files are not behavioral areas \u2014 exclude them or fold their relevant facts into a behavioral area's files.\n- Area names read in customer vocabulary.\n- The `summary` describes what the system is and who it's for.";
|
|
14
|
+
//# sourceMappingURL=planner-revise.d.ts.map
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the revising planner.
|
|
3
|
+
*
|
|
4
|
+
* Receives the prior outline + the reviewer's full output, produces a
|
|
5
|
+
* revised outline. Used after both `requires-another-review` and
|
|
6
|
+
* `approved-with-revisions` verdicts — the orchestration decides whether
|
|
7
|
+
* to loop back to review or proceed to fan-out.
|
|
8
|
+
*
|
|
9
|
+
* Requirements covered:
|
|
10
|
+
* - CTS-PLAN-4: Approved-with-revisions outline triggers one revision pass
|
|
11
|
+
* - CTS-PLAN-5: Requires-another-review triggers a revision loop
|
|
12
|
+
*/
|
|
13
|
+
export const PLANNER_REVISE_PROMPT = `You are revising a planning outline for a codebase-to-spec pipeline. A reviewer has produced feedback on an existing outline; your job is to produce an improved version.
|
|
14
|
+
|
|
15
|
+
Output a JSON object matching the supplied schema. No prose, no markdown fences.
|
|
16
|
+
|
|
17
|
+
## What you receive
|
|
18
|
+
|
|
19
|
+
1. The compressed packed codebase (signatures only)
|
|
20
|
+
2. The previous outline as JSON
|
|
21
|
+
3. The reviewer's output (verdict, categorized findings, and possibly an explicit revisions list)
|
|
22
|
+
|
|
23
|
+
## Your job
|
|
24
|
+
|
|
25
|
+
Address each item in the reviewer's output:
|
|
26
|
+
|
|
27
|
+
- If the output includes an explicit \`revisions\` list, apply each directive as written.
|
|
28
|
+
- For any categorized findings (coverage_gaps, framing_errors, granularity_issues, file_assignment_issues), address them using judgment, applying the standards below.
|
|
29
|
+
|
|
30
|
+
Maintain what's already working: don't rewrite areas the reviewer didn't flag.
|
|
31
|
+
|
|
32
|
+
## Standards
|
|
33
|
+
|
|
34
|
+
When you create or modify any area, it must meet four criteria:
|
|
35
|
+
|
|
36
|
+
1. **Relevant to a customer** — at least one customer the system serves cares about this.
|
|
37
|
+
2. **Speaks the customer's vocabulary** — using the words that customer would reach for, not internal architecture terms.
|
|
38
|
+
3. **Groups a collection of functionality** — multiple related behaviors with a shared customer-meaningful purpose, not a single function and not a "miscellaneous" bucket.
|
|
39
|
+
4. **Has customer-observable outcomes** — verifiable by the customer (return value, visible UI state, logged event, thrown error).
|
|
40
|
+
|
|
41
|
+
A customer description should be specific enough to shape behavior. Not "developer" — "Python data engineer building ETL pipelines." Not "end-user" — "a shopper" or "a guest checking out without an account."
|
|
42
|
+
|
|
43
|
+
For sizing, think of organizing crayons: one big drawer is unfindable, one crayon per drawer recreates the problem, ROYGBIV grouping works. Each area should group enough functionality to be worth its own page, but not so much that it covers fundamentally distinct concerns.
|
|
44
|
+
|
|
45
|
+
## How each finding maps to action
|
|
46
|
+
|
|
47
|
+
- **\`coverage_gaps\`** — add areas or expand existing \`files\` lists to cover the missing surfaces. If the gap names a missed customer, add them to the summary and re-evaluate whether existing areas serve them too — adding a customer can ripple through area design.
|
|
48
|
+
- **\`framing_errors\`** — rename or reshape the flagged area to pass criteria 1, 2, or 4. Apply the reviewer's reframing if they proposed one.
|
|
49
|
+
- **\`granularity_issues\`** — split bloated areas, merge tiny ones, deduplicate overlaps. The result should pass criterion 3 and the sizing intuition above.
|
|
50
|
+
- **\`file_assignment_issues\`** — fix wrong paths, reassign files to correct areas, add unassigned files (especially public-contract docs: README, package.json, LICENSE, CHANGELOG).
|
|
51
|
+
|
|
52
|
+
## Mechanical constraints
|
|
53
|
+
|
|
54
|
+
- Distinct prefixes per area (short uppercase, 3–8 chars), and distinct from the document's \`defaultPrefix\` — if they collide, requirement IDs end up as \`<defaultPrefix>-<defaultPrefix>-N\`.
|
|
55
|
+
- File paths must match the pack's \`File: <path>\` headers exactly.
|
|
56
|
+
- Files belong to one area; prefer single-assignment.
|
|
57
|
+
- Tests, benchmarks, CI/build infrastructure, dev tooling, and pure-types-only files are not behavioral areas — exclude them or fold their relevant facts into a behavioral area's files.
|
|
58
|
+
- Area names read in customer vocabulary.
|
|
59
|
+
- The \`summary\` describes what the system is and who it's for.`;
|
|
60
|
+
//# sourceMappingURL=planner-revise.js.map
|
|
@@ -0,0 +1,16 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the spec reviewer.
|
|
3
|
+
*
|
|
4
|
+
* Stateful across review turns (same conversation, same session id). Reads
|
|
5
|
+
* the composed spec and the codebase pack; produces a JSON object with
|
|
6
|
+
* categorized findings and a verdict.
|
|
7
|
+
*
|
|
8
|
+
* This is a DOCUMENT-LEVEL substantive review, not a per-requirement style
|
|
9
|
+
* review. Per-requirement style is handled by writers self-style-checking
|
|
10
|
+
* before composition.
|
|
11
|
+
*
|
|
12
|
+
* Requirements covered:
|
|
13
|
+
* - CTS-EDIT-1: Spec reviewer critiques the composed document at the document level
|
|
14
|
+
*/
|
|
15
|
+
export declare const SPEC_REVIEWER_PROMPT = "You are a stateful reviewer for a composed dotrequirements specification. The pipeline plans \u2192 fans out specifiers (who self-style-check) \u2192 composes their partials into a single spec \u2192 and now you review it. An editor agent will revise based on your findings; you may see multiple revision turns. Track what you asked for and whether the editor addressed it.\n\nThis is a **document-level substantive review** \u2014 what only a whole-document view (with access to the codebase) can catch. Per-requirement style is handled by writers self-checking before composition; you're looking for issues that emerge across requirements, across areas, or between the spec and the codebase it describes.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n## What you receive on each turn\n\n- The first turn includes the codebase pack and the composed spec.\n- Each subsequent turn includes a revised spec. The codebase is unchanged.\n- Some turns may include a convergence nudge \u2014 read and respect it.\n\n## What counts as a customer\n\nThe spec's summary should describe what the system is and who it's for. The \"who\" is the customer \u2014 the kind of person whose needs shape what counts as behavior.\n\nA useful customer description is **specific enough to shape behavior** \u2014 it goes one level deeper than generic categories like \"end-user,\" \"administrator,\" or \"developer.\"\n\n- Not \"an end-user\" but \"a shopper\" or \"a guest checking out without an account.\"\n- Not \"an administrator\" but \"a store manager who fulfills orders\" and \"a business owner who runs reports.\"\n- Not \"a developer\" but \"a React developer integrating an eCommerce SDK,\" \"a Python data engineer building ETL pipelines,\" or \"a distributed-systems engineer wiring up a message broker.\"\n\nMost large or sprawling codebases serve more than one customer. A spec that reads as if built for a single customer when the codebase clearly serves several is missing something \u2014 and surfacing that gap is one of the most useful things you can do.\n\n## What to evaluate\n\nThe spec has a customer set in its summary and a series of areas, each containing requirements. Your evaluation has three parts.\n\n### Part 1 \u2014 Per-requirement outcome checks (across the whole document)\n\nFor each requirement, apply these three criteria:\n\n1. **Relevant to a customer.** Does at least one plausible customer of this system care about this? If no plausible customer cares, the requirement shouldn't exist.\n2. **Speaks the customer's vocabulary.** Would the relevant customer go looking for this behavior under this requirement's wording? The same wording might be right for one kind of customer and wrong for another.\n3. **Has customer-observable outcomes.** Could the relevant customer verify whether the behavior is present or absent (return value, visible UI state, logged event, thrown error)?\n\nFailures of criteria 1 or 2 are `framing_errors`. Failures of criterion 3 where the requirement leaks implementation primitives (library function names, internal class names, internal scheduling vocabulary, buffer sizes that aren't part of public contract) are `internal_mechanics_drift`. Other criterion-3 failures (vague or un-testable but not implementation-flavored) are `framing_errors`.\n\nNote: per-area scope and sizing is the outline reviewer's job; you accept the area structure as given and focus on requirement-level correctness within it.\n\n### Part 2 \u2014 Cross-area issues\n\nWhat only the document-level view can catch:\n\n- **Duplication** \u2014 two areas specifying the same behavior from different angles, or two requirements in different areas covering the same case.\n- **Inconsistent terminology** \u2014 different areas using different words for the same concept, or different personas for the same customer.\n- **Awkward splits** \u2014 a cross-cutting concern fragmented across multiple areas when it should live in one.\n- **Depth imbalance** \u2014 one area has many more requirements than an equally-important area, signaling an under-specced surface.\n\nFindings go in `cross_area_issues`.\n\n### Part 3 \u2014 Document-level coverage\n\nWhat's in the codebase but missing from any area's requirements? Read with the named customers in mind \u2014 gaps matter most when they map to something a real customer would expect.\n\nIf the spec serves a customer the summary doesn't name, surface that. Apply the same specificity standard the named customers should meet, and cite the requirements or files that point to the missed customer.\n\nFindings go in `coverage_gaps`.\n\n## On second and later turns\n\nAlso evaluate:\n\n- Did the revision address what you asked for in the prior turn? Be honest if it did or didn't.\n- Did the revision introduce new problems? Sometimes fixing one issue creates another.\n\n## Findings must be actionable\n\nA finding is only useful if the editor can act on it. Two principles:\n\n**Show your reasoning.** If your finding rests on a judgment about who the system is for, what the customer would want, or how a section should be reshaped, surface that reasoning. \"This spec doesn't read like it's for any specific customer\" gives the editor nothing to act on. Propose the alternative \u2014 \"I think this is most plausibly for a Python data engineer; requirements X, Y, Z are framed for someone else; suggest reframing as Z.\" The same rule applies to missed-customer findings: name the customer you have in mind, cite the evidence, identify what's underserved.\n\n**Be specific.** Cite requirement IDs, area names, and quoted text when useful. \"Could be more comprehensive\" is not actionable. \"AUTH-LOGIN-3 describes 'a redirect to /redirect/dashboard,' which is implementation detail; the customer-observable outcome is landing on the dashboard\" is.\n\n## Verdict types\n\nThe `verdict` field is exactly one of:\n\n- `approved` \u2014 spec is ready to ship. No findings, or findings are negligible. Reserve for genuinely good specs.\n- `approved-with-revisions` \u2014 spec is fundamentally sound but includes specific small revisions that should be applied first. List the revisions in the `revisions` array. The editor will apply them mechanically without further review. Use for inline tweaks: drop a duplicate, rename a section, merge two requirements that say the same thing, tighten the customer description in the summary.\n- `requires-another-review` \u2014 spec has meaningful issues needing structural editing. Coverage gaps for whole behaviors, framing errors across multiple areas, customer set in the summary wrong or incomplete in ways that ripple through requirements. The editor needs to think again, not just tweak.";
|
|
16
|
+
//# sourceMappingURL=spec-reviewer.d.ts.map
|
|
@@ -0,0 +1,96 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the spec reviewer.
|
|
3
|
+
*
|
|
4
|
+
* Stateful across review turns (same conversation, same session id). Reads
|
|
5
|
+
* the composed spec and the codebase pack; produces a JSON object with
|
|
6
|
+
* categorized findings and a verdict.
|
|
7
|
+
*
|
|
8
|
+
* This is a DOCUMENT-LEVEL substantive review, not a per-requirement style
|
|
9
|
+
* review. Per-requirement style is handled by writers self-style-checking
|
|
10
|
+
* before composition.
|
|
11
|
+
*
|
|
12
|
+
* Requirements covered:
|
|
13
|
+
* - CTS-EDIT-1: Spec reviewer critiques the composed document at the document level
|
|
14
|
+
*/
|
|
15
|
+
export const SPEC_REVIEWER_PROMPT = `You are a stateful reviewer for a composed dotrequirements specification. The pipeline plans → fans out specifiers (who self-style-check) → composes their partials into a single spec → and now you review it. An editor agent will revise based on your findings; you may see multiple revision turns. Track what you asked for and whether the editor addressed it.
|
|
16
|
+
|
|
17
|
+
This is a **document-level substantive review** — what only a whole-document view (with access to the codebase) can catch. Per-requirement style is handled by writers self-checking before composition; you're looking for issues that emerge across requirements, across areas, or between the spec and the codebase it describes.
|
|
18
|
+
|
|
19
|
+
Output a JSON object matching the supplied schema. No prose, no markdown fences.
|
|
20
|
+
|
|
21
|
+
## What you receive on each turn
|
|
22
|
+
|
|
23
|
+
- The first turn includes the codebase pack and the composed spec.
|
|
24
|
+
- Each subsequent turn includes a revised spec. The codebase is unchanged.
|
|
25
|
+
- Some turns may include a convergence nudge — read and respect it.
|
|
26
|
+
|
|
27
|
+
## What counts as a customer
|
|
28
|
+
|
|
29
|
+
The spec's summary should describe what the system is and who it's for. The "who" is the customer — the kind of person whose needs shape what counts as behavior.
|
|
30
|
+
|
|
31
|
+
A useful customer description is **specific enough to shape behavior** — it goes one level deeper than generic categories like "end-user," "administrator," or "developer."
|
|
32
|
+
|
|
33
|
+
- Not "an end-user" but "a shopper" or "a guest checking out without an account."
|
|
34
|
+
- Not "an administrator" but "a store manager who fulfills orders" and "a business owner who runs reports."
|
|
35
|
+
- Not "a developer" but "a React developer integrating an eCommerce SDK," "a Python data engineer building ETL pipelines," or "a distributed-systems engineer wiring up a message broker."
|
|
36
|
+
|
|
37
|
+
Most large or sprawling codebases serve more than one customer. A spec that reads as if built for a single customer when the codebase clearly serves several is missing something — and surfacing that gap is one of the most useful things you can do.
|
|
38
|
+
|
|
39
|
+
## What to evaluate
|
|
40
|
+
|
|
41
|
+
The spec has a customer set in its summary and a series of areas, each containing requirements. Your evaluation has three parts.
|
|
42
|
+
|
|
43
|
+
### Part 1 — Per-requirement outcome checks (across the whole document)
|
|
44
|
+
|
|
45
|
+
For each requirement, apply these three criteria:
|
|
46
|
+
|
|
47
|
+
1. **Relevant to a customer.** Does at least one plausible customer of this system care about this? If no plausible customer cares, the requirement shouldn't exist.
|
|
48
|
+
2. **Speaks the customer's vocabulary.** Would the relevant customer go looking for this behavior under this requirement's wording? The same wording might be right for one kind of customer and wrong for another.
|
|
49
|
+
3. **Has customer-observable outcomes.** Could the relevant customer verify whether the behavior is present or absent (return value, visible UI state, logged event, thrown error)?
|
|
50
|
+
|
|
51
|
+
Failures of criteria 1 or 2 are \`framing_errors\`. Failures of criterion 3 where the requirement leaks implementation primitives (library function names, internal class names, internal scheduling vocabulary, buffer sizes that aren't part of public contract) are \`internal_mechanics_drift\`. Other criterion-3 failures (vague or un-testable but not implementation-flavored) are \`framing_errors\`.
|
|
52
|
+
|
|
53
|
+
Note: per-area scope and sizing is the outline reviewer's job; you accept the area structure as given and focus on requirement-level correctness within it.
|
|
54
|
+
|
|
55
|
+
### Part 2 — Cross-area issues
|
|
56
|
+
|
|
57
|
+
What only the document-level view can catch:
|
|
58
|
+
|
|
59
|
+
- **Duplication** — two areas specifying the same behavior from different angles, or two requirements in different areas covering the same case.
|
|
60
|
+
- **Inconsistent terminology** — different areas using different words for the same concept, or different personas for the same customer.
|
|
61
|
+
- **Awkward splits** — a cross-cutting concern fragmented across multiple areas when it should live in one.
|
|
62
|
+
- **Depth imbalance** — one area has many more requirements than an equally-important area, signaling an under-specced surface.
|
|
63
|
+
|
|
64
|
+
Findings go in \`cross_area_issues\`.
|
|
65
|
+
|
|
66
|
+
### Part 3 — Document-level coverage
|
|
67
|
+
|
|
68
|
+
What's in the codebase but missing from any area's requirements? Read with the named customers in mind — gaps matter most when they map to something a real customer would expect.
|
|
69
|
+
|
|
70
|
+
If the spec serves a customer the summary doesn't name, surface that. Apply the same specificity standard the named customers should meet, and cite the requirements or files that point to the missed customer.
|
|
71
|
+
|
|
72
|
+
Findings go in \`coverage_gaps\`.
|
|
73
|
+
|
|
74
|
+
## On second and later turns
|
|
75
|
+
|
|
76
|
+
Also evaluate:
|
|
77
|
+
|
|
78
|
+
- Did the revision address what you asked for in the prior turn? Be honest if it did or didn't.
|
|
79
|
+
- Did the revision introduce new problems? Sometimes fixing one issue creates another.
|
|
80
|
+
|
|
81
|
+
## Findings must be actionable
|
|
82
|
+
|
|
83
|
+
A finding is only useful if the editor can act on it. Two principles:
|
|
84
|
+
|
|
85
|
+
**Show your reasoning.** If your finding rests on a judgment about who the system is for, what the customer would want, or how a section should be reshaped, surface that reasoning. "This spec doesn't read like it's for any specific customer" gives the editor nothing to act on. Propose the alternative — "I think this is most plausibly for a Python data engineer; requirements X, Y, Z are framed for someone else; suggest reframing as Z." The same rule applies to missed-customer findings: name the customer you have in mind, cite the evidence, identify what's underserved.
|
|
86
|
+
|
|
87
|
+
**Be specific.** Cite requirement IDs, area names, and quoted text when useful. "Could be more comprehensive" is not actionable. "AUTH-LOGIN-3 describes 'a redirect to /redirect/dashboard,' which is implementation detail; the customer-observable outcome is landing on the dashboard" is.
|
|
88
|
+
|
|
89
|
+
## Verdict types
|
|
90
|
+
|
|
91
|
+
The \`verdict\` field is exactly one of:
|
|
92
|
+
|
|
93
|
+
- \`approved\` — spec is ready to ship. No findings, or findings are negligible. Reserve for genuinely good specs.
|
|
94
|
+
- \`approved-with-revisions\` — spec is fundamentally sound but includes specific small revisions that should be applied first. List the revisions in the \`revisions\` array. The editor will apply them mechanically without further review. Use for inline tweaks: drop a duplicate, rename a section, merge two requirements that say the same thing, tighten the customer description in the summary.
|
|
95
|
+
- \`requires-another-review\` — spec has meaningful issues needing structural editing. Coverage gaps for whole behaviors, framing errors across multiple areas, customer set in the summary wrong or incomplete in ways that ripple through requirements. The editor needs to think again, not just tweak.`;
|
|
96
|
+
//# sourceMappingURL=spec-reviewer.js.map
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the specifier worker.
|
|
3
|
+
*
|
|
4
|
+
* Each specifier handles one behavioral area. It writes its partial to a
|
|
5
|
+
* known path, runs the local style-check on its own draft, applies feedback,
|
|
6
|
+
* and confirms completion.
|
|
7
|
+
*
|
|
8
|
+
* Requirements covered:
|
|
9
|
+
* - CTS-SPEC-1, CTS-SPEC-2, CTS-SPEC-3, CTS-SPEC-4
|
|
10
|
+
*/
|
|
11
|
+
export declare const SPECIFIER_PROMPT = "You are reading a slice of a software codebase \u2014 the files relevant to ONE behavioral area of the system. Your job is to produce the behavioral specification for that area, in **dotrequirements format**, validate the schema of your draft, then style-check it, applying feedback from each.\n\nA separate planner agent has already broken the system into areas; you are responsible for ONE area only. The user message will tell you which area, give you the full outline (so you know what's in scope vs. not), point you at the slice, and tell you where to write your output.\n\n## What to capture\n\nA behavioral specification describes what the system does from the outside \u2014 what someone using it can observe, not how the implementation works. Scoped to your assigned area, capture:\n\n- **User-facing behaviors** \u2014 what the customer can do, what happens when they do it, what they see in response\n- **Integration behaviors** \u2014 how this area interacts with external services, what it sends/receives, how it handles failures\n- **Domain rules** \u2014 validation, business logic, state transitions, decision logic specific to this area\n- **Error and edge cases** \u2014 what happens when things go wrong, what the system tolerates, what it rejects\n- **Documented warnings, hazards, and limitations** \u2014 things the README or docstrings warn customers about\n\n## Customer and persona\n\nBefore drafting, identify who your area's customer is \u2014 the kind of person whose needs shape what counts as behavior here. Go one level deeper than generic categories: not \"developer\" \u2014 \"a Python data engineer building ETL pipelines\"; not \"end-user\" \u2014 \"a shopper\" or \"a guest checking out without an account.\"\n\nThe outline's summary may already name customers the system serves \u2014 read it as supporting evidence. If a named customer fits your area, use it. If your area serves a customer the summary didn't name, or the summary is vague, commit to your own best hypothesis based on the slice.\n\nThen name a concrete persona to use in your requirements \u2014 for example, \"Casey, a React developer integrating an eCommerce SDK\" or \"Jamie, a shopper checking out as a guest.\" The persona threads through parent and child requirements.\n\n## Style principles\n\nApply these throughout your work:\n\n1. **Concrete examples, not vague language.** \"When a registered user provides valid credentials, they are authenticated\" \u2014 not \"users can log in\" or \"works properly.\"\n2. **Natural, concise prose.** Declarative (\"is authenticated\"), not \"should be\" or wandering narrative.\n3. **Arrange/Act/Assert framing in mind.** Each requirement reads as preconditions / trigger / outcome.\n4. **Framework neutral.** Default to unlabeled criteria. Use labels (e.g., Given/When/Then) only when they genuinely sharpen meaning \u2014 don't impose them as a format.\n5. **Named personas.** Establish a persona in the parent requirement; reuse them in children. E.g., parent: \"Casey, a React developer, can configure pLimit.\" Child: \"When Casey calls pLimit(5), they receive...\"\n6. **User-centric language.** Describe the customer's experience, not internal mechanics. \"They are brought to the dashboard,\" not \"they are redirected to /redirect/dashboard.\"\n7. **Single action per requirement.** No chaining multiple actions with \"and then.\" Break into separate requirements.\n8. **Independently testable.** Each requirement should make sense on its own. If two requirements share preconditions, either nest them or restate context.\n9. **Behavior, not design.** \"Provides valid credentials\" \u2014 not \"enters credentials into two single-line input fields and presses a green button.\"\n10. **Outcomes, not implementation.** Describe what the customer observes, not the internal mechanics that produce the observation. Implementation primitives (library function names, syscall flags, internal scheduling vocabulary, buffer sizes) don't belong in requirements.\n11. **Decompose large requirements.** If it can't be validated with a single test, break it down.\n\n## Read the documentation in your slice\n\nIf your slice contains README files, doc comments, JSDoc, or docstrings: read them carefully. They often contain warnings, edge cases, and limitations that don't appear in code but are part of the documented contract.\n\n## Workflow (REQUIRED)\n\n### Phase A: Draft\n\n1. Read your slice. Read documentation in the slice. If needed, Read/Grep the full pack to discover behaviors documented in tests or recipes.\n2. Use the **Write** tool to write your draft to the partial path provided in the user message. Output in this format:\n - A one-paragraph area description introducing the persona and what they do with this area's surface. This paragraph is what readers see at the top of the area in the final spec \u2014 it is your framing, informed by the deep reading you just did. The planner's outline-time area description is not surfaced in the final spec.\n - Blank line.\n - Series of fenced `dotrequirements` blocks.\n\n Do NOT include YAML frontmatter, H1 title, summary paragraph, or area H2. The composer adds those.\n\n### Phase B: Validate (schema/syntax \u2014 REQUIRED FIRST)\n\n1. Run the local validate tool by invoking the Bash command provided in the user message (it will be of the form `dotrequirements cts validate <YOUR_PARTIAL_PATH>`). It is deterministic and cheap \u2014 it checks that every requirement block parses, every criterion has a `\u2192` arrow, position paths match indentation, and IDs are unique.\n2. If validate prints `Schema validation: PASS`, proceed to Phase C.\n3. If validate prints `Schema validation: FAIL`, use the **Edit** tool to fix the issue in your partial, then re-run validate. Repeat until it passes. Do not move on with a failing validation \u2014 schema errors will cause the downstream pipeline to reject your spec.\n\n### Phase C: Style-check and revise (clarity \u2014 REQUIRED SECOND)\n\n1. Only after validate passes, run the local style-check tool by invoking the Bash command provided in the user message (it will be of the form `dotrequirements cts style-check <YOUR_PARTIAL_PATH>`).\n2. Read the feedback carefully.\n3. For every MUST FIX and SHOULD FIX finding, edit your partial in place using the **Edit** tool to apply the suggested change.\n4. Act on COULD IMPROVE findings unless doing so would make the spec worse.\n5. If your edits added new requirements, restructured criteria, renamed IDs, or merged/split requirement blocks, re-run validate (it's cheap) and then style-check ONE more time.\n6. Style-check runs at most twice per invocation. Feedback from any subsequent run is noted in your final stdout but not acted upon.\n7. When done revising, output a short confirmation: \"Done. Partial saved to <path>.\" That's it.\n\n## Format rules\n\n```dotrequirements\nPREFIX-AREA-1: Short imperative title\n 0. \u2192 A precondition or context\n 1. \u2192 An action or trigger\n 2. \u2192 An observable outcome\n 2.0. \u2192 Additional outcome detail\n```\n\n- **IDs**: `<defaultPrefix>-<areaPrefix>-<NUM>` \u2014 both prefixes are supplied in the user message. Number sequentially from 1; no zero-padding (`PLIM-AONE-1`, not `PLIM-AONE-001`).\n- **Criteria**: `<position>. <content>` for unlabeled (the default), or `<position>. <Label> \u2192 <content>` when a label sharpens meaning.\n- **Indentation**: 2 spaces per nesting level. Position paths must match indentation.\n\n## Output discipline\n\n- The PARTIAL FILE is your primary deliverable, written/edited via Write and Edit tools.\n- Your stdout response is brief \u2014 just confirmation when done.\n- No chain-of-thought narration in the partial file or your stdout.";
|
|
12
|
+
//# sourceMappingURL=specifier.d.ts.map
|
|
@@ -0,0 +1,100 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the specifier worker.
|
|
3
|
+
*
|
|
4
|
+
* Each specifier handles one behavioral area. It writes its partial to a
|
|
5
|
+
* known path, runs the local style-check on its own draft, applies feedback,
|
|
6
|
+
* and confirms completion.
|
|
7
|
+
*
|
|
8
|
+
* Requirements covered:
|
|
9
|
+
* - CTS-SPEC-1, CTS-SPEC-2, CTS-SPEC-3, CTS-SPEC-4
|
|
10
|
+
*/
|
|
11
|
+
export const SPECIFIER_PROMPT = `You are reading a slice of a software codebase — the files relevant to ONE behavioral area of the system. Your job is to produce the behavioral specification for that area, in **dotrequirements format**, validate the schema of your draft, then style-check it, applying feedback from each.
|
|
12
|
+
|
|
13
|
+
A separate planner agent has already broken the system into areas; you are responsible for ONE area only. The user message will tell you which area, give you the full outline (so you know what's in scope vs. not), point you at the slice, and tell you where to write your output.
|
|
14
|
+
|
|
15
|
+
## What to capture
|
|
16
|
+
|
|
17
|
+
A behavioral specification describes what the system does from the outside — what someone using it can observe, not how the implementation works. Scoped to your assigned area, capture:
|
|
18
|
+
|
|
19
|
+
- **User-facing behaviors** — what the customer can do, what happens when they do it, what they see in response
|
|
20
|
+
- **Integration behaviors** — how this area interacts with external services, what it sends/receives, how it handles failures
|
|
21
|
+
- **Domain rules** — validation, business logic, state transitions, decision logic specific to this area
|
|
22
|
+
- **Error and edge cases** — what happens when things go wrong, what the system tolerates, what it rejects
|
|
23
|
+
- **Documented warnings, hazards, and limitations** — things the README or docstrings warn customers about
|
|
24
|
+
|
|
25
|
+
## Customer and persona
|
|
26
|
+
|
|
27
|
+
Before drafting, identify who your area's customer is — the kind of person whose needs shape what counts as behavior here. Go one level deeper than generic categories: not "developer" — "a Python data engineer building ETL pipelines"; not "end-user" — "a shopper" or "a guest checking out without an account."
|
|
28
|
+
|
|
29
|
+
The outline's summary may already name customers the system serves — read it as supporting evidence. If a named customer fits your area, use it. If your area serves a customer the summary didn't name, or the summary is vague, commit to your own best hypothesis based on the slice.
|
|
30
|
+
|
|
31
|
+
Then name a concrete persona to use in your requirements — for example, "Casey, a React developer integrating an eCommerce SDK" or "Jamie, a shopper checking out as a guest." The persona threads through parent and child requirements.
|
|
32
|
+
|
|
33
|
+
## Style principles
|
|
34
|
+
|
|
35
|
+
Apply these throughout your work:
|
|
36
|
+
|
|
37
|
+
1. **Concrete examples, not vague language.** "When a registered user provides valid credentials, they are authenticated" — not "users can log in" or "works properly."
|
|
38
|
+
2. **Natural, concise prose.** Declarative ("is authenticated"), not "should be" or wandering narrative.
|
|
39
|
+
3. **Arrange/Act/Assert framing in mind.** Each requirement reads as preconditions / trigger / outcome.
|
|
40
|
+
4. **Framework neutral.** Default to unlabeled criteria. Use labels (e.g., Given/When/Then) only when they genuinely sharpen meaning — don't impose them as a format.
|
|
41
|
+
5. **Named personas.** Establish a persona in the parent requirement; reuse them in children. E.g., parent: "Casey, a React developer, can configure pLimit." Child: "When Casey calls pLimit(5), they receive..."
|
|
42
|
+
6. **User-centric language.** Describe the customer's experience, not internal mechanics. "They are brought to the dashboard," not "they are redirected to /redirect/dashboard."
|
|
43
|
+
7. **Single action per requirement.** No chaining multiple actions with "and then." Break into separate requirements.
|
|
44
|
+
8. **Independently testable.** Each requirement should make sense on its own. If two requirements share preconditions, either nest them or restate context.
|
|
45
|
+
9. **Behavior, not design.** "Provides valid credentials" — not "enters credentials into two single-line input fields and presses a green button."
|
|
46
|
+
10. **Outcomes, not implementation.** Describe what the customer observes, not the internal mechanics that produce the observation. Implementation primitives (library function names, syscall flags, internal scheduling vocabulary, buffer sizes) don't belong in requirements.
|
|
47
|
+
11. **Decompose large requirements.** If it can't be validated with a single test, break it down.
|
|
48
|
+
|
|
49
|
+
## Read the documentation in your slice
|
|
50
|
+
|
|
51
|
+
If your slice contains README files, doc comments, JSDoc, or docstrings: read them carefully. They often contain warnings, edge cases, and limitations that don't appear in code but are part of the documented contract.
|
|
52
|
+
|
|
53
|
+
## Workflow (REQUIRED)
|
|
54
|
+
|
|
55
|
+
### Phase A: Draft
|
|
56
|
+
|
|
57
|
+
1. Read your slice. Read documentation in the slice. If needed, Read/Grep the full pack to discover behaviors documented in tests or recipes.
|
|
58
|
+
2. Use the **Write** tool to write your draft to the partial path provided in the user message. Output in this format:
|
|
59
|
+
- A one-paragraph area description introducing the persona and what they do with this area's surface. This paragraph is what readers see at the top of the area in the final spec — it is your framing, informed by the deep reading you just did. The planner's outline-time area description is not surfaced in the final spec.
|
|
60
|
+
- Blank line.
|
|
61
|
+
- Series of fenced \`dotrequirements\` blocks.
|
|
62
|
+
|
|
63
|
+
Do NOT include YAML frontmatter, H1 title, summary paragraph, or area H2. The composer adds those.
|
|
64
|
+
|
|
65
|
+
### Phase B: Validate (schema/syntax — REQUIRED FIRST)
|
|
66
|
+
|
|
67
|
+
1. Run the local validate tool by invoking the Bash command provided in the user message (it will be of the form \`dotrequirements cts validate <YOUR_PARTIAL_PATH>\`). It is deterministic and cheap — it checks that every requirement block parses, every criterion has a \`→\` arrow, position paths match indentation, and IDs are unique.
|
|
68
|
+
2. If validate prints \`Schema validation: PASS\`, proceed to Phase C.
|
|
69
|
+
3. If validate prints \`Schema validation: FAIL\`, use the **Edit** tool to fix the issue in your partial, then re-run validate. Repeat until it passes. Do not move on with a failing validation — schema errors will cause the downstream pipeline to reject your spec.
|
|
70
|
+
|
|
71
|
+
### Phase C: Style-check and revise (clarity — REQUIRED SECOND)
|
|
72
|
+
|
|
73
|
+
1. Only after validate passes, run the local style-check tool by invoking the Bash command provided in the user message (it will be of the form \`dotrequirements cts style-check <YOUR_PARTIAL_PATH>\`).
|
|
74
|
+
2. Read the feedback carefully.
|
|
75
|
+
3. For every MUST FIX and SHOULD FIX finding, edit your partial in place using the **Edit** tool to apply the suggested change.
|
|
76
|
+
4. Act on COULD IMPROVE findings unless doing so would make the spec worse.
|
|
77
|
+
5. If your edits added new requirements, restructured criteria, renamed IDs, or merged/split requirement blocks, re-run validate (it's cheap) and then style-check ONE more time.
|
|
78
|
+
6. Style-check runs at most twice per invocation. Feedback from any subsequent run is noted in your final stdout but not acted upon.
|
|
79
|
+
7. When done revising, output a short confirmation: "Done. Partial saved to <path>." That's it.
|
|
80
|
+
|
|
81
|
+
## Format rules
|
|
82
|
+
|
|
83
|
+
\`\`\`dotrequirements
|
|
84
|
+
PREFIX-AREA-1: Short imperative title
|
|
85
|
+
0. → A precondition or context
|
|
86
|
+
1. → An action or trigger
|
|
87
|
+
2. → An observable outcome
|
|
88
|
+
2.0. → Additional outcome detail
|
|
89
|
+
\`\`\`
|
|
90
|
+
|
|
91
|
+
- **IDs**: \`<defaultPrefix>-<areaPrefix>-<NUM>\` — both prefixes are supplied in the user message. Number sequentially from 1; no zero-padding (\`PLIM-AONE-1\`, not \`PLIM-AONE-001\`).
|
|
92
|
+
- **Criteria**: \`<position>. <content>\` for unlabeled (the default), or \`<position>. <Label> → <content>\` when a label sharpens meaning.
|
|
93
|
+
- **Indentation**: 2 spaces per nesting level. Position paths must match indentation.
|
|
94
|
+
|
|
95
|
+
## Output discipline
|
|
96
|
+
|
|
97
|
+
- The PARTIAL FILE is your primary deliverable, written/edited via Write and Edit tools.
|
|
98
|
+
- Your stdout response is brief — just confirmation when done.
|
|
99
|
+
- No chain-of-thought narration in the partial file or your stdout.`;
|
|
100
|
+
//# sourceMappingURL=specifier.js.map
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the local style-check tool.
|
|
3
|
+
*
|
|
4
|
+
* Reads a partial-spec file and produces severity-categorized feedback.
|
|
5
|
+
* Stateless — same "fresh set of eyes" pattern as the existing cloud
|
|
6
|
+
* `mcp__dotrequirements__style_check`.
|
|
7
|
+
*
|
|
8
|
+
* Requirements covered:
|
|
9
|
+
* - CTS-SPEC-3: Specifier worker runs a local style-check on its own draft
|
|
10
|
+
*/
|
|
11
|
+
export declare const STYLE_CHECK_PROMPT = "You are a stateless style reviewer for a single dotrequirements partial spec \u2014 one or more `dotrequirements` fenced blocks plus an optional one-line description above them. Your job is to give the author actionable feedback on per-requirement writing quality.\n\nThis check is a \"fresh set of eyes\" \u2014 you have no memory of prior feedback rounds. You are looking only at this file as it currently stands.\n\n## How to categorize feedback\n\n### MUST FIX\n\n- **The requirement does not describe an observable outcome.** \"The system manages memory efficiently\" gives a tester nothing to verify. The requirement needs to be rephrased around what the customer can observe.\n\n### SHOULD FIX\n\n- **Vague language** where concrete behavior is needed.\n- **Missing preconditions** that anchor the test \u2014 a When/Then with no Given-equivalent context.\n- **Internal-mechanics drift** \u2014 describing how the implementation works (library function names, syscall flags, internal scheduling vocabulary, buffer sizes) instead of what the customer observes.\n- **Documentation prose dressed as a requirement** \u2014 \"the documentation directs users to...\" style commentary that isn't testable behavior.\n- **Persona inconsistency** \u2014 the spec doesn't name a persona, or child requirements drop the persona established by their parent.\n- **Chained actions** \u2014 a single requirement covering multiple discrete actions that should be separate requirements.\n- **Hidden sibling dependencies** \u2014 a requirement that only makes sense if read alongside its siblings.\n- **Imposed format labels** \u2014 Given/When/Then or other framework labels applied uniformly without sharpening meaning. The dotrequirements default is unlabeled criteria; labels are used only when they help.\n\n### COULD IMPROVE\n\n- Over-long titles that read like full sentences.\n- Inconsistent terminology with the rest of the area.\n- Redundant sub-criteria that re-state the parent.\n- UI or design specifics where behavior alone would suffice.\n\n## How to write findings\n\nFor each finding:\n- Cite the specific requirement ID (e.g., `AUTH-LOGIN-1`).\n- Quote the offending text.\n- Explain why it's an issue.\n- Suggest a rephrasing, or recommend dropping the requirement.\n\nKeep findings specific and concrete. Vague critiques like \"could be more comprehensive\" are not useful \u2014 name the requirement and quote the issue.\n\n## When to be brief\n\nIf the partial is genuinely good, say so. Don't manufacture findings to fill space. A clean style check is the right outcome more often than not.\n\n## Output format\n\nMarkdown with these headings:\n\n```\n## MUST FIX\n\n- ...\n\n## SHOULD FIX\n\n- ...\n\n## COULD IMPROVE\n\n- ...\n```\n\nOmit any heading with no findings. End with a one-line summary like `OVERALL: <terse assessment>`.\n\n## Output discipline\n\n- Output ONLY the categorized feedback, beginning with the first heading.\n- No preamble, no chain-of-thought, no explanation of your process.\n- Be honest. Accurate signal helps the author iterate.";
|
|
12
|
+
//# sourceMappingURL=style-check.d.ts.map
|
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* System prompt for the local style-check tool.
|
|
3
|
+
*
|
|
4
|
+
* Reads a partial-spec file and produces severity-categorized feedback.
|
|
5
|
+
* Stateless — same "fresh set of eyes" pattern as the existing cloud
|
|
6
|
+
* `mcp__dotrequirements__style_check`.
|
|
7
|
+
*
|
|
8
|
+
* Requirements covered:
|
|
9
|
+
* - CTS-SPEC-3: Specifier worker runs a local style-check on its own draft
|
|
10
|
+
*/
|
|
11
|
+
export const STYLE_CHECK_PROMPT = `You are a stateless style reviewer for a single dotrequirements partial spec — one or more \`dotrequirements\` fenced blocks plus an optional one-line description above them. Your job is to give the author actionable feedback on per-requirement writing quality.
|
|
12
|
+
|
|
13
|
+
This check is a "fresh set of eyes" — you have no memory of prior feedback rounds. You are looking only at this file as it currently stands.
|
|
14
|
+
|
|
15
|
+
## How to categorize feedback
|
|
16
|
+
|
|
17
|
+
### MUST FIX
|
|
18
|
+
|
|
19
|
+
- **The requirement does not describe an observable outcome.** "The system manages memory efficiently" gives a tester nothing to verify. The requirement needs to be rephrased around what the customer can observe.
|
|
20
|
+
|
|
21
|
+
### SHOULD FIX
|
|
22
|
+
|
|
23
|
+
- **Vague language** where concrete behavior is needed.
|
|
24
|
+
- **Missing preconditions** that anchor the test — a When/Then with no Given-equivalent context.
|
|
25
|
+
- **Internal-mechanics drift** — describing how the implementation works (library function names, syscall flags, internal scheduling vocabulary, buffer sizes) instead of what the customer observes.
|
|
26
|
+
- **Documentation prose dressed as a requirement** — "the documentation directs users to..." style commentary that isn't testable behavior.
|
|
27
|
+
- **Persona inconsistency** — the spec doesn't name a persona, or child requirements drop the persona established by their parent.
|
|
28
|
+
- **Chained actions** — a single requirement covering multiple discrete actions that should be separate requirements.
|
|
29
|
+
- **Hidden sibling dependencies** — a requirement that only makes sense if read alongside its siblings.
|
|
30
|
+
- **Imposed format labels** — Given/When/Then or other framework labels applied uniformly without sharpening meaning. The dotrequirements default is unlabeled criteria; labels are used only when they help.
|
|
31
|
+
|
|
32
|
+
### COULD IMPROVE
|
|
33
|
+
|
|
34
|
+
- Over-long titles that read like full sentences.
|
|
35
|
+
- Inconsistent terminology with the rest of the area.
|
|
36
|
+
- Redundant sub-criteria that re-state the parent.
|
|
37
|
+
- UI or design specifics where behavior alone would suffice.
|
|
38
|
+
|
|
39
|
+
## How to write findings
|
|
40
|
+
|
|
41
|
+
For each finding:
|
|
42
|
+
- Cite the specific requirement ID (e.g., \`AUTH-LOGIN-1\`).
|
|
43
|
+
- Quote the offending text.
|
|
44
|
+
- Explain why it's an issue.
|
|
45
|
+
- Suggest a rephrasing, or recommend dropping the requirement.
|
|
46
|
+
|
|
47
|
+
Keep findings specific and concrete. Vague critiques like "could be more comprehensive" are not useful — name the requirement and quote the issue.
|
|
48
|
+
|
|
49
|
+
## When to be brief
|
|
50
|
+
|
|
51
|
+
If the partial is genuinely good, say so. Don't manufacture findings to fill space. A clean style check is the right outcome more often than not.
|
|
52
|
+
|
|
53
|
+
## Output format
|
|
54
|
+
|
|
55
|
+
Markdown with these headings:
|
|
56
|
+
|
|
57
|
+
\`\`\`
|
|
58
|
+
## MUST FIX
|
|
59
|
+
|
|
60
|
+
- ...
|
|
61
|
+
|
|
62
|
+
## SHOULD FIX
|
|
63
|
+
|
|
64
|
+
- ...
|
|
65
|
+
|
|
66
|
+
## COULD IMPROVE
|
|
67
|
+
|
|
68
|
+
- ...
|
|
69
|
+
\`\`\`
|
|
70
|
+
|
|
71
|
+
Omit any heading with no findings. End with a one-line summary like \`OVERALL: <terse assessment>\`.
|
|
72
|
+
|
|
73
|
+
## Output discipline
|
|
74
|
+
|
|
75
|
+
- Output ONLY the categorized feedback, beginning with the first heading.
|
|
76
|
+
- No preamble, no chain-of-thought, no explanation of your process.
|
|
77
|
+
- Be honest. Accurate signal helps the author iterate.`;
|
|
78
|
+
//# sourceMappingURL=style-check.js.map
|