@popoverai/dotrequirements 0.24.0 → 0.24.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +4 -3
- package/dist/cli.js +0 -0
- package/dist/codebase-to-spec/claude.d.ts +1 -0
- package/dist/codebase-to-spec/claude.js +9 -0
- package/dist/codebase-to-spec/pack.d.ts +16 -0
- package/dist/codebase-to-spec/pack.js +17 -3
- package/dist/codebase-to-spec/present.d.ts +8 -1
- package/dist/codebase-to-spec/present.js +7 -4
- package/dist/codebase-to-spec/progress.d.ts +6 -0
- package/dist/codebase-to-spec/progress.js +34 -0
- package/dist/codebase-to-spec/prompts/outline-reviewer.d.ts +1 -1
- package/dist/codebase-to-spec/prompts/outline-reviewer.js +3 -1
- package/dist/codebase-to-spec/prompts/planner-initial.d.ts +1 -1
- package/dist/codebase-to-spec/prompts/planner-initial.js +4 -0
- package/dist/codebase-to-spec/prompts/planner-revise.d.ts +1 -1
- package/dist/codebase-to-spec/prompts/planner-revise.js +2 -2
- package/dist/codebase-to-spec/prompts/spec-reviewer.d.ts +1 -1
- package/dist/codebase-to-spec/prompts/spec-reviewer.js +6 -1
- package/dist/codebase-to-spec/prompts/specifier.d.ts +1 -1
- package/dist/codebase-to-spec/prompts/specifier.js +6 -4
- package/dist/codebase-to-spec/prompts/style-check.d.ts +10 -2
- package/dist/codebase-to-spec/prompts/style-check.js +76 -46
- package/dist/codebase-to-spec/schemas.d.ts +85 -1
- package/dist/codebase-to-spec/schemas.js +25 -1
- package/dist/codebase-to-spec/specifier.js +6 -0
- package/dist/commands/codebase-to-spec/index.js +13 -0
- package/dist/commands/codebase-to-spec/pack.d.ts +6 -0
- package/dist/commands/codebase-to-spec/pack.js +1 -0
- package/dist/commands/codebase-to-spec/present.d.ts +5 -0
- package/dist/commands/codebase-to-spec/present.js +6 -1
- package/dist/commands/codebase-to-spec/run.js +1 -0
- package/package.json +5 -6
- package/dist/codebase-to-spec/prompts/planner-apply.d.ts +0 -12
- package/dist/codebase-to-spec/prompts/planner-apply.js +0 -32
- package/dist/commands/browsertest.d.ts +0 -6
- package/dist/commands/browsertest.js +0 -212
- package/dist/commands/login.d.ts +0 -12
- package/dist/commands/login.js +0 -117
- package/dist/commands/logout.d.ts +0 -5
- package/dist/commands/logout.js +0 -17
- package/dist/commands/mcp-setup.d.ts +0 -5
- package/dist/commands/mcp-setup.js +0 -441
- package/dist/commands/test.d.ts +0 -6
- package/dist/commands/test.js +0 -72
- package/dist/mcp/grep.d.ts +0 -24
- package/dist/mcp/grep.js +0 -306
- package/dist/mcp/handlers/coverage.d.ts +0 -44
- package/dist/mcp/handlers/coverage.js +0 -103
- package/dist/mcp/requirements.d.ts +0 -57
- package/dist/mcp/requirements.js +0 -155
- package/dist/mcp/testCodeExtractor.d.ts +0 -22
- package/dist/mcp/testCodeExtractor.js +0 -150
- package/dist/mcp/types.d.ts +0 -27
- package/dist/mcp/types.js +0 -2
- package/dist/utils/local-project.d.ts +0 -31
- package/dist/utils/local-project.js +0 -33
- package/dist/utils/token-refresh.d.ts +0 -24
- package/dist/utils/token-refresh.js +0 -69
- package/dist/utils/token-storage.d.ts +0 -31
- package/dist/utils/token-storage.js +0 -57
package/README.md
CHANGED
|
@@ -9,7 +9,7 @@ Tests prove *something* works—but nobody is certain it's the right something.
|
|
|
9
9
|
|
|
10
10
|
**dot•requirements** closes this gap. Write requirements as structured Markdown, reference them directly in tests, and see coverage update automatically. When a requirement changes, the tests that validate it are one click away.
|
|
11
11
|
|
|
12
|
-
> **Alpha Software** — Under active development. Please report issues to support@
|
|
12
|
+
> **Alpha Software** — Under active development. Please report issues to support@dotrequirements.io.
|
|
13
13
|
|
|
14
14
|
## Who Is This For?
|
|
15
15
|
|
|
@@ -248,12 +248,13 @@ Generate behavioral requirements from an existing codebase. Packs the codebase,
|
|
|
248
248
|
```bash
|
|
249
249
|
dotreq cts run --scope src # full pipeline end-to-end
|
|
250
250
|
dotreq cts run --scope src --fresh # discard cache and start over
|
|
251
|
+
dotreq cts run --scope src --ignore-requirements # run against a codebase that already has its own .requirements/
|
|
251
252
|
dotreq cts skill-install # install the conversational skill wrapper
|
|
252
253
|
```
|
|
253
254
|
|
|
254
255
|
`cts` shells out to `claude -p` and uses whatever auth mode you've configured for Claude Code. Each pipeline stage is also runnable on its own (`dotreq cts pack`, `plan-loop`, `fan-out`, `compose`, `edit-loop`, `present`) for partial re-runs and debugging.
|
|
255
256
|
|
|
256
|
-
> **Alpha:** output quality is prompt-sensitive and varies by codebase. See the [Codebase to Spec docs](https://dotrequirements.io/tools/cli/codebase-to-spec) for prerequisites, options, exit codes, and known rough edges. Feedback to support@
|
|
257
|
+
> **Alpha:** output quality is prompt-sensitive and varies by codebase. See the [Codebase to Spec docs](https://docs.dotrequirements.io/tools/cli/codebase-to-spec) for prerequisites, options, exit codes, and known rough edges. Feedback to support@dotrequirements.io welcome.
|
|
257
258
|
|
|
258
259
|
> **Heads up:** as of June 15, 2026, `claude -p` bills against your Claude subscription's API credit instead of the subscription seat (per Anthropic's May 13, 2026 announcement). `cts` runs will draw from that credit.
|
|
259
260
|
|
|
@@ -725,7 +726,7 @@ This file is automatically added to `.gitignore` during initialization.
|
|
|
725
726
|
|
|
726
727
|
- [Documentation](https://docs.dotrequirements.io)
|
|
727
728
|
- [Getting Started Guide](https://docs.dotrequirements.io/getting-started)
|
|
728
|
-
- [Support](mailto:support@
|
|
729
|
+
- [Support](mailto:support@dotrequirements.io)
|
|
729
730
|
|
|
730
731
|
---
|
|
731
732
|
|
package/dist/cli.js
CHANGED
|
File without changes
|
|
@@ -15,6 +15,7 @@
|
|
|
15
15
|
* Requirements covered:
|
|
16
16
|
* - CTS-PLAN-2, CTS-EDIT-1: stateful reviewer sessions via --session-id / --resume
|
|
17
17
|
* - CTS-PLAN-2, CTS-EDIT-1: JSON-schema-validated output via --json-schema
|
|
18
|
+
* - CTS-OBSERVE-1.4: subprocess PID emitted to stderr after spawn
|
|
18
19
|
*/
|
|
19
20
|
export interface ClaudeRunOptions {
|
|
20
21
|
/** System prompt content (passed via --system-prompt). */
|
|
@@ -15,9 +15,11 @@
|
|
|
15
15
|
* Requirements covered:
|
|
16
16
|
* - CTS-PLAN-2, CTS-EDIT-1: stateful reviewer sessions via --session-id / --resume
|
|
17
17
|
* - CTS-PLAN-2, CTS-EDIT-1: JSON-schema-validated output via --json-schema
|
|
18
|
+
* - CTS-OBSERVE-1.4: subprocess PID emitted to stderr after spawn
|
|
18
19
|
*/
|
|
19
20
|
import { spawn } from "node:child_process";
|
|
20
21
|
import { tmpdir } from "node:os";
|
|
22
|
+
import { stderr as processStderr } from "node:process";
|
|
21
23
|
/**
|
|
22
24
|
* Default runner — spawns the `claude` binary as a subprocess.
|
|
23
25
|
*/
|
|
@@ -67,6 +69,13 @@ export async function runClaude(options) {
|
|
|
67
69
|
// mode the user has configured (API key if set, OAuth otherwise).
|
|
68
70
|
stdio: ["pipe", "pipe", "pipe"],
|
|
69
71
|
});
|
|
72
|
+
// Diagnostic: announce the spawned PID to stderr so engineers watching a
|
|
73
|
+
// long-running pipeline can verify the subprocess is alive while we await
|
|
74
|
+
// its response (CTS-OBSERVE-1.4). Stderr keeps the structured progress
|
|
75
|
+
// stream on stdout clean for the skill and CI parsers.
|
|
76
|
+
if (child.pid !== undefined) {
|
|
77
|
+
processStderr.write(`[cts.diag] claude -p spawned (pid=${child.pid})\n`);
|
|
78
|
+
}
|
|
70
79
|
// Close stdin immediately — we pass user message via argv, not stdin.
|
|
71
80
|
child.stdin.end();
|
|
72
81
|
let stdout = "";
|
|
@@ -8,6 +8,7 @@
|
|
|
8
8
|
* Requirements covered:
|
|
9
9
|
* - CTS-PACK-1: Pack stage produces compressed and uncompressed views
|
|
10
10
|
* - CTS-PACK-2: Engineer can extend the ignore list per project
|
|
11
|
+
* - CTS-PRESENT-5.0: --ignore-requirements adds .requirements/** to pack ignores
|
|
11
12
|
*/
|
|
12
13
|
import type { CachePaths } from "./cache.js";
|
|
13
14
|
/**
|
|
@@ -31,6 +32,12 @@ export interface PackOptions {
|
|
|
31
32
|
scope?: string;
|
|
32
33
|
/** Additional ignore patterns (beyond defaults and project-level). */
|
|
33
34
|
extraIgnores?: string[];
|
|
35
|
+
/**
|
|
36
|
+
* When true, add `.requirements/**` to the ignore list so an existing
|
|
37
|
+
* requirements directory does not influence the generated output.
|
|
38
|
+
* See CTS-PRESENT-5.
|
|
39
|
+
*/
|
|
40
|
+
ignoreRequirements?: boolean;
|
|
34
41
|
}
|
|
35
42
|
export interface PackResult {
|
|
36
43
|
overviewPath: string;
|
|
@@ -47,5 +54,14 @@ export interface PackResult {
|
|
|
47
54
|
* NOTE: Repomix's TypeScript API is invoked via `runCli`. We invoke it as the
|
|
48
55
|
* library exposes it; the surface is small enough that this wraps cleanly.
|
|
49
56
|
*/
|
|
57
|
+
/**
|
|
58
|
+
* Build the full ignore list for a pack invocation, combining defaults,
|
|
59
|
+
* project-level patterns, caller-supplied extras, and (when set) the
|
|
60
|
+
* `.requirements/**` exclusion that powers `--ignore-requirements`.
|
|
61
|
+
*/
|
|
62
|
+
export declare function buildIgnoreList(projectRoot: string, options?: {
|
|
63
|
+
extraIgnores?: string[];
|
|
64
|
+
ignoreRequirements?: boolean;
|
|
65
|
+
}): string[];
|
|
50
66
|
export declare function runPack(options: PackOptions): Promise<PackResult>;
|
|
51
67
|
//# sourceMappingURL=pack.d.ts.map
|
|
@@ -8,6 +8,7 @@
|
|
|
8
8
|
* Requirements covered:
|
|
9
9
|
* - CTS-PACK-1: Pack stage produces compressed and uncompressed views
|
|
10
10
|
* - CTS-PACK-2: Engineer can extend the ignore list per project
|
|
11
|
+
* - CTS-PRESENT-5.0: --ignore-requirements adds .requirements/** to pack ignores
|
|
11
12
|
*/
|
|
12
13
|
import { existsSync, readFileSync } from "node:fs";
|
|
13
14
|
import { join } from "node:path";
|
|
@@ -89,13 +90,26 @@ export function readProjectIgnores(projectRoot) {
|
|
|
89
90
|
* NOTE: Repomix's TypeScript API is invoked via `runCli`. We invoke it as the
|
|
90
91
|
* library exposes it; the surface is small enough that this wraps cleanly.
|
|
91
92
|
*/
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
93
|
+
/**
|
|
94
|
+
* Build the full ignore list for a pack invocation, combining defaults,
|
|
95
|
+
* project-level patterns, caller-supplied extras, and (when set) the
|
|
96
|
+
* `.requirements/**` exclusion that powers `--ignore-requirements`.
|
|
97
|
+
*/
|
|
98
|
+
export function buildIgnoreList(projectRoot, options = {}) {
|
|
99
|
+
const { extraIgnores = [], ignoreRequirements = false } = options;
|
|
100
|
+
return [
|
|
95
101
|
...DEFAULT_IGNORES,
|
|
96
102
|
...readProjectIgnores(projectRoot),
|
|
97
103
|
...extraIgnores,
|
|
104
|
+
...(ignoreRequirements ? [".requirements/**"] : []),
|
|
98
105
|
];
|
|
106
|
+
}
|
|
107
|
+
export async function runPack(options) {
|
|
108
|
+
const { projectRoot, paths, scope, extraIgnores = [], ignoreRequirements = false, } = options;
|
|
109
|
+
const ignores = buildIgnoreList(projectRoot, {
|
|
110
|
+
extraIgnores,
|
|
111
|
+
ignoreRequirements,
|
|
112
|
+
});
|
|
99
113
|
const target = scope ? join(projectRoot, scope) : projectRoot;
|
|
100
114
|
// Import dynamically — Repomix is heavy and we don't want to load it at CLI
|
|
101
115
|
// boot time for unrelated commands.
|
|
@@ -11,6 +11,7 @@
|
|
|
11
11
|
* - CTS-PRESENT-1: Final spec written to .requirements/, split when 5+ areas
|
|
12
12
|
* - CTS-PRESENT-2: Interactive overwrite prompt with three choices
|
|
13
13
|
* - CTS-PRESENT-3: Non-interactive overwrites require --overwrite or --skip-existing
|
|
14
|
+
* - CTS-PRESENT-5.1: --ignore-requirements redirects output to .requirements/cts/
|
|
14
15
|
*/
|
|
15
16
|
import type { Outline } from "./schemas.js";
|
|
16
17
|
export declare const AREA_SPLIT_THRESHOLD = 5;
|
|
@@ -21,6 +22,12 @@ export interface PresentOptions {
|
|
|
21
22
|
/** Project root (.requirements/ will be created under this). */
|
|
22
23
|
projectRoot: string;
|
|
23
24
|
overwritePolicy: OverwritePolicy;
|
|
25
|
+
/**
|
|
26
|
+
* Optional subdirectory under `.requirements/` to write output into.
|
|
27
|
+
* Used by `--ignore-requirements` to redirect output to `.requirements/cts/`.
|
|
28
|
+
* See CTS-PRESENT-5.
|
|
29
|
+
*/
|
|
30
|
+
outputSubdir?: string;
|
|
24
31
|
}
|
|
25
32
|
export interface PresentResult {
|
|
26
33
|
/** Paths written under .requirements/ (or those that would have been written, in fail cases). */
|
|
@@ -39,7 +46,7 @@ export interface PresentAction {
|
|
|
39
46
|
/**
|
|
40
47
|
* Resolve the output path(s) the present stage would write to.
|
|
41
48
|
*/
|
|
42
|
-
export declare function planPresentPaths(outline: Outline, projectRoot: string): {
|
|
49
|
+
export declare function planPresentPaths(outline: Outline, projectRoot: string, outputSubdir?: string): {
|
|
43
50
|
mode: "single" | "split";
|
|
44
51
|
paths: {
|
|
45
52
|
area?: string;
|
|
@@ -11,6 +11,7 @@
|
|
|
11
11
|
* - CTS-PRESENT-1: Final spec written to .requirements/, split when 5+ areas
|
|
12
12
|
* - CTS-PRESENT-2: Interactive overwrite prompt with three choices
|
|
13
13
|
* - CTS-PRESENT-3: Non-interactive overwrites require --overwrite or --skip-existing
|
|
14
|
+
* - CTS-PRESENT-5.1: --ignore-requirements redirects output to .requirements/cts/
|
|
14
15
|
*/
|
|
15
16
|
import { existsSync, mkdirSync, readFileSync, writeFileSync } from "node:fs";
|
|
16
17
|
import { dirname, join } from "node:path";
|
|
@@ -21,8 +22,10 @@ export const AREA_SPLIT_THRESHOLD = 5;
|
|
|
21
22
|
/**
|
|
22
23
|
* Resolve the output path(s) the present stage would write to.
|
|
23
24
|
*/
|
|
24
|
-
export function planPresentPaths(outline, projectRoot) {
|
|
25
|
-
const requirementsDir =
|
|
25
|
+
export function planPresentPaths(outline, projectRoot, outputSubdir) {
|
|
26
|
+
const requirementsDir = outputSubdir
|
|
27
|
+
? join(projectRoot, ".requirements", outputSubdir)
|
|
28
|
+
: join(projectRoot, ".requirements");
|
|
26
29
|
if (outline.areas.length >= AREA_SPLIT_THRESHOLD) {
|
|
27
30
|
return {
|
|
28
31
|
mode: "split",
|
|
@@ -204,9 +207,9 @@ function validateContent(content) {
|
|
|
204
207
|
* "no file overwritten without choice" guarantee from CTS-PRESENT-2.1).
|
|
205
208
|
*/
|
|
206
209
|
export async function runPresent(options) {
|
|
207
|
-
const { outline, finalSpecPath, projectRoot, overwritePolicy } = options;
|
|
210
|
+
const { outline, finalSpecPath, projectRoot, overwritePolicy, outputSubdir } = options;
|
|
208
211
|
const composedSpec = readFileSync(finalSpecPath, "utf-8");
|
|
209
|
-
const plan = planPresentPaths(outline, projectRoot);
|
|
212
|
+
const plan = planPresentPaths(outline, projectRoot, outputSubdir);
|
|
210
213
|
// Compute the content for each output up front.
|
|
211
214
|
const planned = [];
|
|
212
215
|
if (plan.mode === "single") {
|
|
@@ -11,6 +11,12 @@
|
|
|
11
11
|
* - CTS-CLI-5: CLI emissions are part of the public contract (stable format)
|
|
12
12
|
* - CTS-OBSERVE-1: CLI emits structured progress as the pipeline runs
|
|
13
13
|
*/
|
|
14
|
+
export declare function installCtsAbortHandlers(): void;
|
|
15
|
+
/**
|
|
16
|
+
* Reset the install-once flag. Only intended for tests — production code
|
|
17
|
+
* should never call this.
|
|
18
|
+
*/
|
|
19
|
+
export declare function __resetCtsAbortHandlersForTests(): void;
|
|
14
20
|
export type Stage = "pack" | "plan" | "outline-review" | "specify" | "compose" | "spec-review" | "present";
|
|
15
21
|
export interface ProgressEvent {
|
|
16
22
|
stage: Stage;
|
|
@@ -12,6 +12,40 @@
|
|
|
12
12
|
* - CTS-OBSERVE-1: CLI emits structured progress as the pipeline runs
|
|
13
13
|
*/
|
|
14
14
|
import { stdout } from "node:process";
|
|
15
|
+
/**
|
|
16
|
+
* Install SIGTERM/SIGINT handlers that emit a `pipeline/aborted` progress line
|
|
17
|
+
* before exiting. Idempotent — repeat calls are no-ops so multiple cts
|
|
18
|
+
* subcommands invoked in sequence don't stack handlers.
|
|
19
|
+
*
|
|
20
|
+
* Requirement: CTS-OBSERVE-1.5
|
|
21
|
+
*/
|
|
22
|
+
let abortHandlersInstalled = false;
|
|
23
|
+
export function installCtsAbortHandlers() {
|
|
24
|
+
if (abortHandlersInstalled)
|
|
25
|
+
return;
|
|
26
|
+
abortHandlersInstalled = true;
|
|
27
|
+
let aborting = false;
|
|
28
|
+
const handle = (signal) => {
|
|
29
|
+
if (aborting)
|
|
30
|
+
return; // de-dupe back-to-back signals
|
|
31
|
+
aborting = true;
|
|
32
|
+
// Write directly — the StdoutProgress emit() requires a known Stage and
|
|
33
|
+
// `pipeline` is a cross-stage lifecycle event, not a pipeline stage.
|
|
34
|
+
stdout.write(`[CTS] pipeline/aborted Received ${signal}, exiting\n`);
|
|
35
|
+
// Standard exit code for terminated-by-signal: 128 + signal number.
|
|
36
|
+
// 128+15=143 (SIGTERM), 128+2=130 (SIGINT).
|
|
37
|
+
process.exit(signal === "SIGINT" ? 130 : 143);
|
|
38
|
+
};
|
|
39
|
+
process.on("SIGTERM", () => handle("SIGTERM"));
|
|
40
|
+
process.on("SIGINT", () => handle("SIGINT"));
|
|
41
|
+
}
|
|
42
|
+
/**
|
|
43
|
+
* Reset the install-once flag. Only intended for tests — production code
|
|
44
|
+
* should never call this.
|
|
45
|
+
*/
|
|
46
|
+
export function __resetCtsAbortHandlersForTests() {
|
|
47
|
+
abortHandlersInstalled = false;
|
|
48
|
+
}
|
|
15
49
|
export class StdoutProgress {
|
|
16
50
|
emit(event) {
|
|
17
51
|
const stagePath = event.step ? `${event.stage}/${event.step}` : event.stage;
|
|
@@ -8,5 +8,5 @@
|
|
|
8
8
|
* Requirements covered:
|
|
9
9
|
* - CTS-PLAN-2: Outline reviewer critiques the outline in a stateful session
|
|
10
10
|
*/
|
|
11
|
-
export declare const OUTLINE_REVIEWER_PROMPT = "You are an outline reviewer for a codebase-to-spec pipeline. The pipeline takes a codebase, generates a planning outline (areas of behavior), then fans out specifier agents to produce detailed requirements per area. Your job: gate fan-out on outline quality. Bad outlines \u2192 wasted specifier work.\n\nThis is a **stateful conversation**. Across turns, you may receive multiple revisions of the outline, each addressing prior feedback. Track what you asked for and whether the planner addressed it.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n## What you receive on each turn\n\n- The first turn includes the compressed packed codebase and the planner's first outline JSON.\n- Each subsequent turn includes a revised outline JSON. The codebase is unchanged.\n- Some turns may include a convergence nudge \u2014 read and respect it.\n\n## What counts as a customer\n\nThe outline's summary should describe what the system is and who it's for. The \"who\" is the customer \u2014 the kind of person whose needs shape what counts as behavior.\n\nA useful customer description is **specific enough to shape behavior** \u2014 it goes one level deeper than generic categories like \"end-user,\" \"administrator,\" or \"developer.\"\n\n- Not \"an end-user\" but \"a shopper\" or \"a guest checking out without an account.\"\n- Not \"an administrator\" but \"a store manager who fulfills orders\" and \"a business owner who runs reports.\"\n- Not \"a developer\" but \"a React developer integrating an eCommerce SDK,\" \"a Python data engineer building ETL pipelines,\" or \"a distributed-systems engineer wiring up a message broker.\"\n\nMost large or sprawling codebases serve more than one customer. A system whose outline reads as if built for a single customer when the code clearly serves several is missing something \u2014 and surfacing that gap is one of the most useful things you can do.\n\n## What to evaluate\n\nThe outline names its customers in the summary and breaks the system into areas. Your evaluation has two parts.\n\n### Part 1 \u2014 Per-area outcome checks\n\nFor each area, apply these four criteria:\n\n1. **Relevant to a customer.**
|
|
11
|
+
export declare const OUTLINE_REVIEWER_PROMPT = "You are an outline reviewer for a codebase-to-spec pipeline. The pipeline takes a codebase, generates a planning outline (areas of behavior), then fans out specifier agents to produce detailed requirements per area. Your job: gate fan-out on outline quality. Bad outlines \u2192 wasted specifier work.\n\nThis is a **stateful conversation**. Across turns, you may receive multiple revisions of the outline, each addressing prior feedback. Track what you asked for and whether the planner addressed it.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n## What you receive on each turn\n\n- The first turn includes the compressed packed codebase and the planner's first outline JSON.\n- Each subsequent turn includes a revised outline JSON. The codebase is unchanged.\n- Some turns may include a convergence nudge \u2014 read and respect it.\n\n## What counts as a customer\n\nThe outline's summary should describe what the system is and who it's for. The \"who\" is the customer \u2014 the kind of person whose needs shape what counts as behavior.\n\nA useful customer description is **specific enough to shape behavior** \u2014 it goes one level deeper than generic categories like \"end-user,\" \"administrator,\" or \"developer.\"\n\n- Not \"an end-user\" but \"a shopper\" or \"a guest checking out without an account.\"\n- Not \"an administrator\" but \"a store manager who fulfills orders\" and \"a business owner who runs reports.\"\n- Not \"a developer\" but \"a React developer integrating an eCommerce SDK,\" \"a Python data engineer building ETL pipelines,\" or \"a distributed-systems engineer wiring up a message broker.\"\n\nMost large or sprawling codebases serve more than one customer. A system whose outline reads as if built for a single customer when the code clearly serves several is missing something \u2014 and surfacing that gap is one of the most useful things you can do.\n\n**Customers consume the software; contributors work on it.** A developer can be a legitimate customer when they consume the software being specified \u2014 a React developer integrating an SDK, a Python data engineer using a library, an operator running a CLI in CI, a contributor to an open-source project who uses it as much as they extend it. But a developer whose only role is to *work on this codebase* \u2014 described as \"a contributor,\" \"an internal maintainer,\" \"a stage author building the next feature,\" or similar \u2014 is not a customer. They're the audience for code comments and architecture docs, not for behavioral requirements. Data contracts between architectural components can still be (and often should be) specified, but think about whose experience they matter for. If an interface mediates business logic between a front-end and a server, that logic should be specified in terms of what it does for the user, not what it does for the front-end developer.\n\n## What to evaluate\n\nThe outline names its customers in the summary and breaks the system into areas. Your evaluation has two parts.\n\n### Part 1 \u2014 Per-area outcome checks\n\nFor each area, apply these four criteria:\n\n1. **Relevant to a customer.** Each area must declare at least one customer in its `customers` field. Apply the consume-vs-work-on test from \"What counts as a customer\" above: a customer is someone who *uses* the software being specified, not someone who works on its codebase. An area whose `customers` field is empty, missing, or contains only developers-of-this-codebase (contributors, maintainers, stage authors) is a `framing_error` \u2014 propose either reframing the area for a real customer or dropping it. When the area's description is written in mechanism-only voice (\"the subprocess wrapper that spawns...\", \"how the pipeline generates...\") with no customer named, that's the same finding even if a customer name was bolted on after the fact.\n2. **Speaks the customer's vocabulary.** Would the relevant customer go looking for this behavior under this area's name? The same name might be right for one kind of customer and wrong for another \u2014 what matters is whether it matches whom this area serves.\n3. **Groups a collection of functionality.** Does the area cover multiple related behaviors with a shared customer-meaningful purpose? An area with one isolated function, or a \"miscellaneous\" bucket of unrelated things, fails this test.\n4. **Has customer-observable outcomes.** Could the relevant customer verify whether the behavior is present or absent (return value, visible UI state, logged event, thrown error)? \"The system manages memory efficiently\" is true but not customer-observable.\n\nFailures of criterion 1, 2, or 4 are `framing_errors`. Failures of criterion 3 are `granularity_issues` \u2014 which also covers sizing problems more broadly.\n\n#### On sizing\n\nThink of organizing a big box of 100 crayons.\n\n- One drawer for all 100 crayons \u2192 impossible to find what you need.\n- 100 drawers, one crayon each \u2192 you've recreated the same problem with different semantics.\n- Organize by ROYGBIV \u2192 each drawer is a meaningful group, and you can find any crayon quickly.\n\nThe same logic applies to areas. An area too broad covers fundamentally distinct concerns; an area too narrow fragments what should hang together. Two areas that describe the same thing from different angles are a sizing problem \u2014 they should be one area, or split along a different axis. The right sizing for *this* codebase is whatever lets each area be coherent on its own and the whole set be complete.\n\n### Part 2 \u2014 Coverage and file assignment\n\n- What behavior is in the codebase but missing from any area? \u2192 `coverage_gaps`. Gaps matter most when they map to something a real customer would expect.\n- Are file paths actually present in the pack as written? Are files assigned to areas where they don't fit? Are public-contract docs (README, package metadata, LICENSE, CHANGELOG) unassigned? \u2192 `file_assignment_issues`.\n\n## On second and later turns\n\nAlso evaluate:\n\n- Did the revision address what you asked for in the prior turn? Be honest if it did or didn't.\n- Did the revision introduce new problems? Sometimes fixing one gap creates another.\n\n## Findings must be actionable\n\nA finding is only useful if the planner can act on it. Two principles:\n\n**Show your reasoning.** If your finding rests on a judgment about who the system is for, what the customer would want, or how an area should be reshaped, surface that reasoning. \"This outline doesn't read like it's for any specific customer\" gives the planner nothing to act on. \"I think this is most plausibly for a React developer integrating an eCommerce SDK; areas X and Y are organized around backend storage rather than what that developer would reach for; suggest reframing as Z\" does. Same pattern for any other \"this feels off\" finding \u2014 propose the alternative.\n\n**Be specific.** Cite paths and area names by exact spelling. \"Could be more comprehensive\" is not actionable. \"The outline has no area covering [behavior X], visible in [file Y]\" is.\n\n## Verdict types\n\nThe `verdict` field is exactly one of:\n\n- `approved` \u2014 outline is ready for fan-out as-is. No findings, or findings are negligible. Reserve for genuinely good outlines.\n- `approved-with-revisions` \u2014 outline is fundamentally sound and ready for fan-out, but includes specific small revisions that should be applied first. List the revisions in the `revisions` array. The planner will apply them mechanically without further review. Use for inline tweaks: rename a file path, split one bloated area into two, add a missing public-contract file to an area, tighten the customer description in the summary.\n- `requires-another-review` \u2014 outline has meaningful issues that need a structural fix, not just tweaks. Coverage gaps for whole subsystems, framing errors at the area level, customer set in the summary wrong or incomplete in ways that ripple through area design. The planner needs to think again, not just tweak.";
|
|
12
12
|
//# sourceMappingURL=outline-reviewer.d.ts.map
|
|
@@ -32,6 +32,8 @@ A useful customer description is **specific enough to shape behavior** — it go
|
|
|
32
32
|
|
|
33
33
|
Most large or sprawling codebases serve more than one customer. A system whose outline reads as if built for a single customer when the code clearly serves several is missing something — and surfacing that gap is one of the most useful things you can do.
|
|
34
34
|
|
|
35
|
+
**Customers consume the software; contributors work on it.** A developer can be a legitimate customer when they consume the software being specified — a React developer integrating an SDK, a Python data engineer using a library, an operator running a CLI in CI, a contributor to an open-source project who uses it as much as they extend it. But a developer whose only role is to *work on this codebase* — described as "a contributor," "an internal maintainer," "a stage author building the next feature," or similar — is not a customer. They're the audience for code comments and architecture docs, not for behavioral requirements. Data contracts between architectural components can still be (and often should be) specified, but think about whose experience they matter for. If an interface mediates business logic between a front-end and a server, that logic should be specified in terms of what it does for the user, not what it does for the front-end developer.
|
|
36
|
+
|
|
35
37
|
## What to evaluate
|
|
36
38
|
|
|
37
39
|
The outline names its customers in the summary and breaks the system into areas. Your evaluation has two parts.
|
|
@@ -40,7 +42,7 @@ The outline names its customers in the summary and breaks the system into areas.
|
|
|
40
42
|
|
|
41
43
|
For each area, apply these four criteria:
|
|
42
44
|
|
|
43
|
-
1. **Relevant to a customer.**
|
|
45
|
+
1. **Relevant to a customer.** Each area must declare at least one customer in its \`customers\` field. Apply the consume-vs-work-on test from "What counts as a customer" above: a customer is someone who *uses* the software being specified, not someone who works on its codebase. An area whose \`customers\` field is empty, missing, or contains only developers-of-this-codebase (contributors, maintainers, stage authors) is a \`framing_error\` — propose either reframing the area for a real customer or dropping it. When the area's description is written in mechanism-only voice ("the subprocess wrapper that spawns...", "how the pipeline generates...") with no customer named, that's the same finding even if a customer name was bolted on after the fact.
|
|
44
46
|
2. **Speaks the customer's vocabulary.** Would the relevant customer go looking for this behavior under this area's name? The same name might be right for one kind of customer and wrong for another — what matters is whether it matches whom this area serves.
|
|
45
47
|
3. **Groups a collection of functionality.** Does the area cover multiple related behaviors with a shared customer-meaningful purpose? An area with one isolated function, or a "miscellaneous" bucket of unrelated things, fails this test.
|
|
46
48
|
4. **Has customer-observable outcomes.** Could the relevant customer verify whether the behavior is present or absent (return value, visible UI state, logged event, thrown error)? "The system manages memory efficiently" is true but not customer-observable.
|
|
@@ -7,5 +7,5 @@
|
|
|
7
7
|
* Requirements covered:
|
|
8
8
|
* - CTS-PLAN-1: Planner produces a behavioral outline from the compressed pack
|
|
9
9
|
*/
|
|
10
|
-
export declare const PLANNER_INITIAL_PROMPT = "You are reading a compressed packed view of a software codebase (function signatures, types, interfaces, class structures \u2014 implementation bodies stripped). Your job is to produce a structured outline that breaks the system into behavioral areas, each of which will be specced in detail by a focused specifier agent in a follow-up step.\n\nThe quality of every downstream step depends on the quality of this outline. Take it seriously. Follow the reasoning process below rather than jumping to area names.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n---\n\n## Reasoning process\n\n### Step 0 \u2014 Classify the system\n\nBefore anything else, decide what *kind* of software this is. The class shapes who counts as a customer and what counts as a public surface. Common classes:\n\n- **Library** \u2014 imported by other code; customer is the developer integrating it\n- **Application** \u2014 end users interact with it directly (web app, desktop app, mobile app)\n- **Service** \u2014 runs continuously, accepts requests over a network or queue; customer may be another service, or end users via a frontend\n- **CLI tool** \u2014 invoked from a shell; customer is a developer/operator/ops person\n- **Framework** \u2014 code is structured around it; customer is the developer building on top of it\n- **Data format / parser / serializer** \u2014 customer is whatever produces or consumes the format\n- **Protocol implementation** \u2014 customer is whoever speaks the protocol\n\nIf the system is a hybrid (e.g., a service that ships with a CLI client + a library SDK), name each surface separately \u2014 they may have distinct customers.\n\n### Step 1 \u2014 Gather context\n\nSkim the pack with intent. You're building a mental model, not yet writing the outline.\n\n- The **directory structure** suggests how the authors organize the system. Note conventions but don't be bound by them \u2014 code organization is rarely the same as behavioral organization.\n- **README, package metadata, docstrings on public APIs, error messages, CLI help text, OpenAPI/JSON schemas, and type signatures of exported symbols** are deliberate public-contract surfaces. They tell you what the authors think users need to know. Read them carefully.\n- The **test suite** is evidence of behavior \u2014 the authors wrote tests for things they considered worth verifying. Tests aren't behavioral areas, but they reveal which behaviors exist.\n\n### Step 2 \u2014 Identify the domain model\n\nName the **nouns** the system is about, and how they relate to each other. Use the language the system uses, not generic CS terms. Examples:\n\n- A concurrency-limiter library: `Limiter`, `Task`, `Queue`; tasks run when the queue has an open slot.\n- An eCommerce platform: `Shopper`, `Cart`, `Order`, `Inventory`, `Payment`; an order is created when a shopper checks out a cart.\n- A markdown parser: `Document`, `Block`, `Inline`, `Token`; blocks contain inlines, parsing produces a token stream.\n\nIf the system has no obvious nouns of its own, name what it's gluing together \u2014 its domain may live in the upstream and downstream systems it integrates with.\n\n### Step 3 \u2014 Identify functionality\n\nCatalog what the system does, focused on:\n\n- **Public-facing surfaces** \u2014 exported APIs, CLI commands, HTTP endpoints, file formats, UI flows\n- **Business logic** \u2014 domain rules, validations, state transitions, decision logic\n- **Documented contracts** \u2014 what the README/docstrings/type signatures promise\n- **What the test suite verifies** \u2014 read test descriptions and assertions, not implementations\n\nSkip internal infrastructure (transports, storage adapters, build glue, scheduling primitives) unless they expose a public surface in their own right.\n\n### Step 4 \u2014 Identify the customer(s)\n\nWho uses this software? Some customer exists \u2014 the software was written for someone. Form a hypothesis about who, even if the evidence is thin.\n\nBe specific. Go at least one level deeper than generic categories. Generic categories like \"end-user,\" \"administrator,\" or \"developer\" are too broad to shape behavior.\n\n- Not \"an end-user\" but \"a shopper\" or \"a guest checking out without an account.\"\n- Not \"an administrator\" but \"a store manager who fulfills orders\" and \"a business owner who runs reports.\"\n- Not \"a developer\" but \"a React developer integrating an eCommerce SDK,\" \"a Python data engineer building ETL pipelines,\" or \"a distributed-systems engineer wiring up a message broker.\"\n\nIf the system has multiple customers, name them all. Different behaviors will be relevant to different customers.\n\nIf the customer isn't obvious from the code, name your best hypothesis (\"this looks like a library for X kind of developer\") and proceed. A wrong guess is fixable downstream; a missing one isn't.\n\n### Step 5 \u2014 Identify behavioral areas\n\nBehavioral areas are the **intersection** of your domain model (Step 2), your customers (Step 4), and the functionality you cataloged (Step 3).\n\nA candidate area is behavioral if it meets all four criteria:\n\n1. **Relevant to a customer** \u2014 at least one of the customers you named in Step 4 cares about this. If no one cares, it's not behavior.\n\n2. **Describable in your customer's vocabulary** \u2014 using the words the customer in Step 4 would use to describe what they're trying to do. This is about *whose* language the area name speaks, not about which words \"sound technical.\"\n\n For a shopper, \"placing an order\" is customer vocabulary; \"persisting to the orders table\" is not.\n\n For a distributed-systems engineer building on an infrastructure library, \"configuring a storage backend\" or \"choosing an IPC protocol\" might be exactly customer vocabulary \u2014 because those *are* the operations they think in terms of. The same words that would be wrong for the shopper case are right here.\n\n The test: would your customer (Step 4) go looking for the behavior under this name, or under something else? Pick the name they would reach for. If you're tempted to name an area `STORAGE` and your customer is an application developer building eCommerce, they'd reach for `PERSISTING_ORDERS` or similar \u2014 use that instead. If your customer is the distributed-systems engineer, `STORAGE` may be exactly right.\n\n3. **Describes a collection of functionality** \u2014 it groups multiple related behaviors that share a customer-meaningful purpose. A single function with no companions is not an area; a \"miscellaneous\" bucket isn't either.\n\n4. **Has observable outcomes** \u2014 the customer can verify whether the behavior is present or absent (a returned value, a visible UI state, a logged event, a thrown error). \"The system manages memory efficiently\" is true but not customer-observable, so it's not behavior.\n\n#### Sizing\n\nThink of it like organizing a big box of 100 crayons.\n\n- One drawer for all 100 crayons \u2192 impossible to find what you need.\n- 100 drawers, one crayon each \u2192 you've recreated the same problem with different semantics.\n- Organize by ROYGBIV \u2192 each drawer is a meaningful group, and you can find any crayon quickly.\n\nApply the same to behavioral areas. Each area should group enough functionality to be worth its own page, but not so much that it covers fundamentally distinct concerns. Aim for the ROYGBIV equivalent for *this* codebase: a small system might land at a few areas, a medium one at several, a large or sprawling one at many. The right number is whatever makes the areas individually coherent and collectively complete.\n\n### Step 6 \u2014 Check coverage gaps\n\nAfter drafting your areas, look back at the codebase and ask: what *isn't* represented? For each gap:\n\n- **If it's functionality that should have been a behavioral area** (it passes all four criteria from Step 5) \u2192 add it.\n- **If it's real functionality but not behavioral** (tests, benchmarks, CI infrastructure, dev tooling, internal-only adapters, pure types with no runtime semantics) \u2192 acknowledge that you considered it and explicitly chose to exclude it. Don't leave it looking like an oversight.\n\nPublic-contract files (README, package.json, LICENSE, CHANGELOG, top-level config) MUST be assigned to whichever behavioral area is most relevant \u2014 they contain user-observable facts that don't appear elsewhere.\n\n---\n\n## Output rules\n\n- **Distinct prefixes per area.** Each area's `prefix` must be unique within the document AND distinct from the document's `defaultPrefix`. Short uppercase tokens (3\u20138 characters). Requirement IDs are formed as `<defaultPrefix>-<areaPrefix>-<N>` \u2014 if they collide (e.g., defaultPrefix=CORE and an area prefix=CORE), every requirement in that area reads as `CORE-CORE-N`, which is ugly. Pick area prefixes that don't repeat the defaultPrefix.\n- **Files belong to one area.** Assign each substantive file to its primary behavioral area. Prefer single-assignment; cross-area concerns are handled downstream.\n- **File paths must match the pack.** Look at the `File: <path>` headers in the pack and use those exact paths.\n- **Area names should read in customer vocabulary** \u2014 they will appear in the spec's heading structure and the customer should recognize what each area is about.\n- **Write the `summary` last**, after you understand the system as a whole \u2014 it should describe what the system is and who it's for, not how the spec is organized.";
|
|
10
|
+
export declare const PLANNER_INITIAL_PROMPT = "You are reading a compressed packed view of a software codebase (function signatures, types, interfaces, class structures \u2014 implementation bodies stripped). Your job is to produce a structured outline that breaks the system into behavioral areas, each of which will be specced in detail by a focused specifier agent in a follow-up step.\n\nThe quality of every downstream step depends on the quality of this outline. Take it seriously. Follow the reasoning process below rather than jumping to area names.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n---\n\n## Reasoning process\n\n### Step 0 \u2014 Classify the system\n\nBefore anything else, decide what *kind* of software this is. The class shapes who counts as a customer and what counts as a public surface. Common classes:\n\n- **Library** \u2014 imported by other code; customer is the developer integrating it\n- **Application** \u2014 end users interact with it directly (web app, desktop app, mobile app)\n- **Service** \u2014 runs continuously, accepts requests over a network or queue; customer may be another service, or end users via a frontend\n- **CLI tool** \u2014 invoked from a shell; customer is a developer/operator/ops person\n- **Framework** \u2014 code is structured around it; customer is the developer building on top of it\n- **Data format / parser / serializer** \u2014 customer is whatever produces or consumes the format\n- **Protocol implementation** \u2014 customer is whoever speaks the protocol\n\nIf the system is a hybrid (e.g., a service that ships with a CLI client + a library SDK), name each surface separately \u2014 they may have distinct customers.\n\n### Step 1 \u2014 Gather context\n\nSkim the pack with intent. You're building a mental model, not yet writing the outline.\n\n- The **directory structure** suggests how the authors organize the system. Note conventions but don't be bound by them \u2014 code organization is rarely the same as behavioral organization.\n- **README, package metadata, docstrings on public APIs, error messages, CLI help text, OpenAPI/JSON schemas, and type signatures of exported symbols** are deliberate public-contract surfaces. They tell you what the authors think users need to know. Read them carefully.\n- The **test suite** is evidence of behavior \u2014 the authors wrote tests for things they considered worth verifying. Tests aren't behavioral areas, but they reveal which behaviors exist.\n\n### Step 2 \u2014 Identify the domain model\n\nName the **nouns** the system is about, and how they relate to each other. Use the language the system uses, not generic CS terms. Examples:\n\n- A concurrency-limiter library: `Limiter`, `Task`, `Queue`; tasks run when the queue has an open slot.\n- An eCommerce platform: `Shopper`, `Cart`, `Order`, `Inventory`, `Payment`; an order is created when a shopper checks out a cart.\n- A markdown parser: `Document`, `Block`, `Inline`, `Token`; blocks contain inlines, parsing produces a token stream.\n\nIf the system has no obvious nouns of its own, name what it's gluing together \u2014 its domain may live in the upstream and downstream systems it integrates with.\n\n### Step 3 \u2014 Identify functionality\n\nCatalog what the system does, focused on:\n\n- **Public-facing surfaces** \u2014 exported APIs, CLI commands, HTTP endpoints, file formats, UI flows\n- **Business logic** \u2014 domain rules, validations, state transitions, decision logic\n- **Documented contracts** \u2014 what the README/docstrings/type signatures promise\n- **What the test suite verifies** \u2014 read test descriptions and assertions, not implementations\n\nSkip internal infrastructure (transports, storage adapters, build glue, scheduling primitives) unless they expose a public surface in their own right.\n\n### Step 4 \u2014 Identify the customer(s)\n\nWho uses this software? Some customer exists \u2014 the software was written for someone. Form a hypothesis about who, even if the evidence is thin.\n\nBe specific. Go at least one level deeper than generic categories. Generic categories like \"end-user,\" \"administrator,\" or \"developer\" are too broad to shape behavior.\n\n- Not \"an end-user\" but \"a shopper\" or \"a guest checking out without an account.\"\n- Not \"an administrator\" but \"a store manager who fulfills orders\" and \"a business owner who runs reports.\"\n- Not \"a developer\" but \"a React developer integrating an eCommerce SDK,\" \"a Python data engineer building ETL pipelines,\" or \"a distributed-systems engineer wiring up a message broker.\"\n\nIf the system has multiple customers, name them all. Different behaviors will be relevant to different customers.\n\nIf the customer isn't obvious from the code, name your best hypothesis (\"this looks like a library for X kind of developer\") and proceed. A wrong guess is fixable downstream; a missing one isn't.\n\n### Step 5 \u2014 Identify behavioral areas\n\nBehavioral areas are the **intersection** of your domain model (Step 2), your customers (Step 4), and the functionality you cataloged (Step 3).\n\n**Each area must declare its customer(s) explicitly.** In the area's `customers` array, name one or more of the customers from Step 4 that this area is *for* \u2014 the people whose experience this area's behavior shapes. If an area has no real customer from your Step 4 list, it isn't a behavioral area and shouldn't exist.\n\nThe area's `description` should be written for the customer(s) it serves \u2014 describing what those customers can do, what they observe, what they care about \u2014 not in mechanism-only voice (\"the subprocess wrapper that spawns...\", \"how the pipeline generates...\"). If you can't write the description in customer terms, that's a signal the area might be internal infrastructure rather than behavior.\n\nA candidate area is behavioral if it meets all four criteria:\n\n1. **Relevant to a customer** \u2014 at least one of the customers you named in Step 4 cares about this. If no one cares, it's not behavior.\n\n2. **Describable in your customer's vocabulary** \u2014 using the words the customer in Step 4 would use to describe what they're trying to do. This is about *whose* language the area name speaks, not about which words \"sound technical.\"\n\n For a shopper, \"placing an order\" is customer vocabulary; \"persisting to the orders table\" is not.\n\n For a distributed-systems engineer building on an infrastructure library, \"configuring a storage backend\" or \"choosing an IPC protocol\" might be exactly customer vocabulary \u2014 because those *are* the operations they think in terms of. The same words that would be wrong for the shopper case are right here.\n\n The test: would your customer (Step 4) go looking for the behavior under this name, or under something else? Pick the name they would reach for. If you're tempted to name an area `STORAGE` and your customer is an application developer building eCommerce, they'd reach for `PERSISTING_ORDERS` or similar \u2014 use that instead. If your customer is the distributed-systems engineer, `STORAGE` may be exactly right.\n\n3. **Describes a collection of functionality** \u2014 it groups multiple related behaviors that share a customer-meaningful purpose. A single function with no companions is not an area; a \"miscellaneous\" bucket isn't either.\n\n4. **Has observable outcomes** \u2014 the customer can verify whether the behavior is present or absent (a returned value, a visible UI state, a logged event, a thrown error). \"The system manages memory efficiently\" is true but not customer-observable, so it's not behavior.\n\n#### Sizing\n\nThink of it like organizing a big box of 100 crayons.\n\n- One drawer for all 100 crayons \u2192 impossible to find what you need.\n- 100 drawers, one crayon each \u2192 you've recreated the same problem with different semantics.\n- Organize by ROYGBIV \u2192 each drawer is a meaningful group, and you can find any crayon quickly.\n\nApply the same to behavioral areas. Each area should group enough functionality to be worth its own page, but not so much that it covers fundamentally distinct concerns. Aim for the ROYGBIV equivalent for *this* codebase: a small system might land at a few areas, a medium one at several, a large or sprawling one at many. The right number is whatever makes the areas individually coherent and collectively complete.\n\n### Step 6 \u2014 Check coverage gaps\n\nAfter drafting your areas, look back at the codebase and ask: what *isn't* represented? For each gap:\n\n- **If it's functionality that should have been a behavioral area** (it passes all four criteria from Step 5) \u2192 add it.\n- **If it's real functionality but not behavioral** (tests, benchmarks, CI infrastructure, dev tooling, internal-only adapters, pure types with no runtime semantics) \u2192 acknowledge that you considered it and explicitly chose to exclude it. Don't leave it looking like an oversight.\n\nPublic-contract files (README, package.json, LICENSE, CHANGELOG, top-level config) MUST be assigned to whichever behavioral area is most relevant \u2014 they contain user-observable facts that don't appear elsewhere.\n\n---\n\n## Output rules\n\n- **Distinct prefixes per area.** Each area's `prefix` must be unique within the document AND distinct from the document's `defaultPrefix`. Short uppercase tokens (3\u20138 characters). Requirement IDs are formed as `<defaultPrefix>-<areaPrefix>-<N>` \u2014 if they collide (e.g., defaultPrefix=CORE and an area prefix=CORE), every requirement in that area reads as `CORE-CORE-N`, which is ugly. Pick area prefixes that don't repeat the defaultPrefix.\n- **Files belong to one area.** Assign each substantive file to its primary behavioral area. Prefer single-assignment; cross-area concerns are handled downstream.\n- **File paths must match the pack.** Look at the `File: <path>` headers in the pack and use those exact paths.\n- **Area names should read in customer vocabulary** \u2014 they will appear in the spec's heading structure and the customer should recognize what each area is about.\n- **Write the `summary` last**, after you understand the system as a whole \u2014 it should describe what the system is and who it's for, not how the spec is organized.";
|
|
11
11
|
//# sourceMappingURL=planner-initial.d.ts.map
|
|
@@ -78,6 +78,10 @@ If the customer isn't obvious from the code, name your best hypothesis ("this lo
|
|
|
78
78
|
|
|
79
79
|
Behavioral areas are the **intersection** of your domain model (Step 2), your customers (Step 4), and the functionality you cataloged (Step 3).
|
|
80
80
|
|
|
81
|
+
**Each area must declare its customer(s) explicitly.** In the area's \`customers\` array, name one or more of the customers from Step 4 that this area is *for* — the people whose experience this area's behavior shapes. If an area has no real customer from your Step 4 list, it isn't a behavioral area and shouldn't exist.
|
|
82
|
+
|
|
83
|
+
The area's \`description\` should be written for the customer(s) it serves — describing what those customers can do, what they observe, what they care about — not in mechanism-only voice ("the subprocess wrapper that spawns...", "how the pipeline generates..."). If you can't write the description in customer terms, that's a signal the area might be internal infrastructure rather than behavior.
|
|
84
|
+
|
|
81
85
|
A candidate area is behavioral if it meets all four criteria:
|
|
82
86
|
|
|
83
87
|
1. **Relevant to a customer** — at least one of the customers you named in Step 4 cares about this. If no one cares, it's not behavior.
|
|
@@ -10,5 +10,5 @@
|
|
|
10
10
|
* - CTS-PLAN-4: Approved-with-revisions outline triggers one revision pass
|
|
11
11
|
* - CTS-PLAN-5: Requires-another-review triggers a revision loop
|
|
12
12
|
*/
|
|
13
|
-
export declare const PLANNER_REVISE_PROMPT = "You are revising a planning outline for a codebase-to-spec pipeline. A reviewer has produced feedback on an existing outline; your job is to produce an improved version.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n## What you receive\n\n1. The compressed packed codebase (signatures only)\n2. The previous outline as JSON\n3. The reviewer's output (verdict, categorized findings, and possibly an explicit revisions list)\n\n## Your job\n\nAddress each item in the reviewer's output:\n\n- If the output includes an explicit `revisions` list, apply each directive as written.\n- For any categorized findings (coverage_gaps, framing_errors, granularity_issues, file_assignment_issues), address them using judgment, applying the standards below.\n\nMaintain what's already working: don't rewrite areas the reviewer didn't flag.\n\n## Standards\n\nWhen you create or modify any area, it must meet four criteria:\n\n1. **Relevant to a customer** \u2014 at least one customer the system serves cares about this.\n2. **Speaks the customer's vocabulary** \u2014 using the words that customer would reach for, not internal architecture terms.\n3. **Groups a collection of functionality** \u2014 multiple related behaviors with a shared customer-meaningful purpose, not a single function and not a \"miscellaneous\" bucket.\n4. **Has customer-observable outcomes** \u2014 verifiable by the customer (return value, visible UI state, logged event, thrown error).\n\nA customer description should be specific enough to shape behavior. Not \"developer\" \u2014 \"Python data engineer building ETL pipelines.\" Not \"end-user\" \u2014 \"a shopper\" or \"a guest checking out without an account.\"\n\nFor sizing, think of organizing crayons: one big drawer is unfindable, one crayon per drawer recreates the problem, ROYGBIV grouping works. Each area should group enough functionality to be worth its own page, but not so much that it covers fundamentally distinct concerns.\n\n## How each finding maps to action\n\n- **`coverage_gaps`** \u2014 add areas or expand existing `files` lists to cover the missing surfaces. If the gap names a missed customer, add them to the summary and re-evaluate whether existing areas serve them too \u2014 adding a customer can ripple through area design.\n- **`framing_errors`** \u2014 rename or reshape the flagged area to pass criteria 1, 2, or 4. Apply the reviewer's reframing if they proposed one.\n- **`granularity_issues`** \u2014 split bloated areas, merge tiny ones, deduplicate overlaps. The result should pass criterion 3 and the sizing intuition above.\n- **`file_assignment_issues`** \u2014 fix wrong paths, reassign files to correct areas, add unassigned files (especially public-contract docs: README, package.json, LICENSE, CHANGELOG).\n\n## Mechanical constraints\n\n- Distinct prefixes per area (short uppercase, 3\u20138 chars), and distinct from the document's `defaultPrefix` \u2014 if they collide, requirement IDs end up as `<defaultPrefix>-<defaultPrefix>-N`.\n- File paths must match the pack's `File: <path>` headers exactly.\n- Files belong to one area; prefer single-assignment.\n- Tests, benchmarks, CI/build infrastructure, dev tooling, and pure-types-only files are not behavioral areas \u2014 exclude them or fold their relevant facts into a behavioral area's files.\n- Area names read in customer vocabulary.\n- The `summary` describes what the system is and who it's for.";
|
|
13
|
+
export declare const PLANNER_REVISE_PROMPT = "You are revising a planning outline for a codebase-to-spec pipeline. A reviewer has produced feedback on an existing outline; your job is to produce an improved version.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n## What you receive\n\n1. The compressed packed codebase (signatures only)\n2. The previous outline as JSON\n3. The reviewer's output (verdict, categorized findings, and possibly an explicit revisions list)\n\n## Your job\n\nAddress each item in the reviewer's output:\n\n- If the output includes an explicit `revisions` list, apply each directive as written.\n- For any categorized findings (coverage_gaps, framing_errors, granularity_issues, file_assignment_issues), address them using judgment, applying the standards below.\n\nMaintain what's already working: don't rewrite areas the reviewer didn't flag.\n\n## Standards\n\nWhen you create or modify any area, it must meet four criteria:\n\n1. **Relevant to a customer** \u2014 at least one customer the system serves cares about this. The area's `customers` array must name one or more such customers explicitly. An area that has no real customer is not behavior; drop it.\n2. **Speaks the customer's vocabulary** \u2014 using the words that customer would reach for, not internal architecture terms. The area's `description` is written for the customer(s) it serves, not in mechanism-only voice.\n3. **Groups a collection of functionality** \u2014 multiple related behaviors with a shared customer-meaningful purpose, not a single function and not a \"miscellaneous\" bucket.\n4. **Has customer-observable outcomes** \u2014 verifiable by the customer (return value, visible UI state, logged event, thrown error).\n\nA customer description should be specific enough to shape behavior. Not \"developer\" \u2014 \"Python data engineer building ETL pipelines.\" Not \"end-user\" \u2014 \"a shopper\" or \"a guest checking out without an account.\"\n\nFor sizing, think of organizing crayons: one big drawer is unfindable, one crayon per drawer recreates the problem, ROYGBIV grouping works. Each area should group enough functionality to be worth its own page, but not so much that it covers fundamentally distinct concerns.\n\n## How each finding maps to action\n\n- **`coverage_gaps`** \u2014 add areas or expand existing `files` lists to cover the missing surfaces. If the gap names a missed customer, add them to the summary and re-evaluate whether existing areas serve them too \u2014 adding a customer can ripple through area design.\n- **`framing_errors`** \u2014 rename or reshape the flagged area to pass criteria 1, 2, or 4. Apply the reviewer's reframing if they proposed one.\n- **`granularity_issues`** \u2014 split bloated areas, merge tiny ones, deduplicate overlaps. The result should pass criterion 3 and the sizing intuition above.\n- **`file_assignment_issues`** \u2014 fix wrong paths, reassign files to correct areas, add unassigned files (especially public-contract docs: README, package.json, LICENSE, CHANGELOG).\n\n## Mechanical constraints\n\n- Distinct prefixes per area (short uppercase, 3\u20138 chars), and distinct from the document's `defaultPrefix` \u2014 if they collide, requirement IDs end up as `<defaultPrefix>-<defaultPrefix>-N`.\n- File paths must match the pack's `File: <path>` headers exactly.\n- Files belong to one area; prefer single-assignment.\n- Tests, benchmarks, CI/build infrastructure, dev tooling, and pure-types-only files are not behavioral areas \u2014 exclude them or fold their relevant facts into a behavioral area's files.\n- Area names read in customer vocabulary.\n- The `summary` describes what the system is and who it's for.";
|
|
14
14
|
//# sourceMappingURL=planner-revise.d.ts.map
|
|
@@ -33,8 +33,8 @@ Maintain what's already working: don't rewrite areas the reviewer didn't flag.
|
|
|
33
33
|
|
|
34
34
|
When you create or modify any area, it must meet four criteria:
|
|
35
35
|
|
|
36
|
-
1. **Relevant to a customer** — at least one customer the system serves cares about this.
|
|
37
|
-
2. **Speaks the customer's vocabulary** — using the words that customer would reach for, not internal architecture terms.
|
|
36
|
+
1. **Relevant to a customer** — at least one customer the system serves cares about this. The area's \`customers\` array must name one or more such customers explicitly. An area that has no real customer is not behavior; drop it.
|
|
37
|
+
2. **Speaks the customer's vocabulary** — using the words that customer would reach for, not internal architecture terms. The area's \`description\` is written for the customer(s) it serves, not in mechanism-only voice.
|
|
38
38
|
3. **Groups a collection of functionality** — multiple related behaviors with a shared customer-meaningful purpose, not a single function and not a "miscellaneous" bucket.
|
|
39
39
|
4. **Has customer-observable outcomes** — verifiable by the customer (return value, visible UI state, logged event, thrown error).
|
|
40
40
|
|
|
@@ -12,5 +12,5 @@
|
|
|
12
12
|
* Requirements covered:
|
|
13
13
|
* - CTS-EDIT-1: Spec reviewer critiques the composed document at the document level
|
|
14
14
|
*/
|
|
15
|
-
export declare const SPEC_REVIEWER_PROMPT = "You are a stateful reviewer for a composed dotrequirements specification. The pipeline plans \u2192 fans out specifiers (who self-style-check) \u2192 composes their partials into a single spec \u2192 and now you review it. An editor agent will revise based on your findings; you may see multiple revision turns. Track what you asked for and whether the editor addressed it.\n\nThis is a **document-level substantive review** \u2014 what only a whole-document view (with access to the codebase) can catch. Per-requirement style is handled by writers self-checking before composition; you're looking for issues that emerge across requirements, across areas, or between the spec and the codebase it describes.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n## What you receive on each turn\n\n- The first turn includes the codebase pack and the composed spec.\n- Each subsequent turn includes a revised spec. The codebase is unchanged.\n- Some turns may include a convergence nudge \u2014 read and respect it.\n\n## What counts as a customer\n\nThe spec's summary should describe what the system is and who it's for. The \"who\" is the customer \u2014 the kind of person whose needs shape what counts as behavior.\n\nA useful customer description is **specific enough to shape behavior** \u2014 it goes one level deeper than generic categories like \"end-user,\" \"administrator,\" or \"developer.\"\n\n- Not \"an end-user\" but \"a shopper\" or \"a guest checking out without an account.\"\n- Not \"an administrator\" but \"a store manager who fulfills orders\" and \"a business owner who runs reports.\"\n- Not \"a developer\" but \"a React developer integrating an eCommerce SDK,\" \"a Python data engineer building ETL pipelines,\" or \"a distributed-systems engineer wiring up a message broker.\"\n\nMost large or sprawling codebases serve more than one customer. A spec that reads as if built for a single customer when the codebase clearly serves several is missing something \u2014 and surfacing that gap is one of the most useful things you can do.\n\n## What to evaluate\n\nThe spec has a customer set in its summary and a series of areas, each containing requirements. Your evaluation has three parts.\n\n### Part 1 \u2014 Per-requirement outcome checks (across the whole document)\n\nFor each requirement, apply these three criteria:\n\n1. **Relevant to a customer.** Does at least one plausible customer of this system care about this? If no plausible customer cares, the requirement shouldn't exist.\n2. **Speaks the customer's vocabulary.** Would the relevant customer go looking for this behavior under this requirement's wording? The same wording might be right for one kind of customer and wrong for another.\n3. **Has customer-observable outcomes.** Could the relevant customer verify whether the behavior is present or absent (return value, visible UI state, logged event, thrown error)?\n\nFailures of criteria 1 or 2 are `framing_errors`. Failures of criterion 3 where the requirement
|
|
15
|
+
export declare const SPEC_REVIEWER_PROMPT = "You are a stateful reviewer for a composed dotrequirements specification. The pipeline plans \u2192 fans out specifiers (who self-style-check) \u2192 composes their partials into a single spec \u2192 and now you review it. An editor agent will revise based on your findings; you may see multiple revision turns. Track what you asked for and whether the editor addressed it.\n\nThis is a **document-level substantive review** \u2014 what only a whole-document view (with access to the codebase) can catch. Per-requirement style is handled by writers self-checking before composition; you're looking for issues that emerge across requirements, across areas, or between the spec and the codebase it describes.\n\nOutput a JSON object matching the supplied schema. No prose, no markdown fences.\n\n## What you receive on each turn\n\n- The first turn includes the codebase pack and the composed spec.\n- Each subsequent turn includes a revised spec. The codebase is unchanged.\n- Some turns may include a convergence nudge \u2014 read and respect it.\n\n## What counts as a customer\n\nThe spec's summary should describe what the system is and who it's for. The \"who\" is the customer \u2014 the kind of person whose needs shape what counts as behavior.\n\nA useful customer description is **specific enough to shape behavior** \u2014 it goes one level deeper than generic categories like \"end-user,\" \"administrator,\" or \"developer.\"\n\n- Not \"an end-user\" but \"a shopper\" or \"a guest checking out without an account.\"\n- Not \"an administrator\" but \"a store manager who fulfills orders\" and \"a business owner who runs reports.\"\n- Not \"a developer\" but \"a React developer integrating an eCommerce SDK,\" \"a Python data engineer building ETL pipelines,\" or \"a distributed-systems engineer wiring up a message broker.\"\n\nMost large or sprawling codebases serve more than one customer. A spec that reads as if built for a single customer when the codebase clearly serves several is missing something \u2014 and surfacing that gap is one of the most useful things you can do.\n\n## What to evaluate\n\nThe spec has a customer set in its summary and a series of areas, each containing requirements. Your evaluation has three parts.\n\n### Part 1 \u2014 Per-requirement outcome checks (across the whole document)\n\nFor each requirement, apply these three criteria:\n\n1. **Relevant to a customer.** Does at least one plausible customer of this system care about this? If no plausible customer cares, the requirement shouldn't exist.\n2. **Speaks the customer's vocabulary.** Would the relevant customer go looking for this behavior under this requirement's wording? The same wording might be right for one kind of customer and wrong for another.\n3. **Has customer-observable outcomes.** Could the relevant customer verify whether the behavior is present or absent (return value, visible UI state, logged event, thrown error)?\n\nFailures of criteria 1 or 2 are `framing_errors`. Failures of criterion 3 where the requirement describes implementation rather than observed behavior are `internal_mechanics_drift`. Two flavors:\n\n- *Implementation primitives leaking out* \u2014 library function names, internal class names, scheduling vocabulary, buffer sizes that aren't part of public contract.\n- *Pinning literals instead of properties* \u2014 when a requirement names a specific value (an exit code, an error string, a format prefix, a field name, a file path), the customer almost always depends on some characteristic (stable, distinguishable, parseable, idiomatic, named-rather-than-anonymous) rather than the value itself.\n\nOther criterion-3 failures (vague or un-testable but not implementation-flavored) are `framing_errors`.\n\nNote: per-area scope and sizing is the outline reviewer's job; you accept the area structure as given and focus on requirement-level correctness within it.\n\n### Part 2 \u2014 Cross-area issues\n\nWhat only the document-level view can catch:\n\n- **Duplication** \u2014 two areas specifying the same behavior from different angles, or two requirements in different areas covering the same case.\n- **Inconsistent terminology** \u2014 different areas using different words for the same concept, or different personas for the same customer.\n- **Awkward splits** \u2014 a cross-cutting concern fragmented across multiple areas when it should live in one.\n- **Depth imbalance** \u2014 one area has many more requirements than an equally-important area, signaling an under-specced surface.\n\nFindings go in `cross_area_issues`.\n\n### Part 3 \u2014 Document-level coverage\n\nWhat's in the codebase but missing from any area's requirements? Read with the named customers in mind \u2014 gaps matter most when they map to something a real customer would expect.\n\nIf the spec serves a customer the summary doesn't name, surface that. Apply the same specificity standard the named customers should meet, and cite the requirements or files that point to the missed customer.\n\nFindings go in `coverage_gaps`.\n\n## On second and later turns\n\nAlso evaluate:\n\n- Did the revision address what you asked for in the prior turn? Be honest if it did or didn't.\n- Did the revision introduce new problems? Sometimes fixing one issue creates another.\n\n## Findings must be actionable\n\nA finding is only useful if the editor can act on it. Two principles:\n\n**Show your reasoning.** If your finding rests on a judgment about who the system is for, what the customer would want, or how a section should be reshaped, surface that reasoning. \"This spec doesn't read like it's for any specific customer\" gives the editor nothing to act on. Propose the alternative \u2014 \"I think this is most plausibly for a Python data engineer; requirements X, Y, Z are framed for someone else; suggest reframing as Z.\" The same rule applies to missed-customer findings: name the customer you have in mind, cite the evidence, identify what's underserved.\n\n**Be specific.** Cite requirement IDs, area names, and quoted text when useful. \"Could be more comprehensive\" is not actionable. \"AUTH-LOGIN-3 describes 'a redirect to /redirect/dashboard,' which is implementation detail; the customer-observable outcome is landing on the dashboard\" is.\n\n## Verdict types\n\nThe `verdict` field is exactly one of:\n\n- `approved` \u2014 spec is ready to ship. No findings, or findings are negligible. Reserve for genuinely good specs.\n- `approved-with-revisions` \u2014 spec is fundamentally sound but includes specific small revisions that should be applied first. List the revisions in the `revisions` array. The editor will apply them mechanically without further review. Use for inline tweaks: drop a duplicate, rename a section, merge two requirements that say the same thing, tighten the customer description in the summary.\n- `requires-another-review` \u2014 spec has meaningful issues needing structural editing. Coverage gaps for whole behaviors, framing errors across multiple areas, customer set in the summary wrong or incomplete in ways that ripple through requirements. The editor needs to think again, not just tweak.";
|
|
16
16
|
//# sourceMappingURL=spec-reviewer.d.ts.map
|
|
@@ -48,7 +48,12 @@ For each requirement, apply these three criteria:
|
|
|
48
48
|
2. **Speaks the customer's vocabulary.** Would the relevant customer go looking for this behavior under this requirement's wording? The same wording might be right for one kind of customer and wrong for another.
|
|
49
49
|
3. **Has customer-observable outcomes.** Could the relevant customer verify whether the behavior is present or absent (return value, visible UI state, logged event, thrown error)?
|
|
50
50
|
|
|
51
|
-
Failures of criteria 1 or 2 are \`framing_errors\`. Failures of criterion 3 where the requirement
|
|
51
|
+
Failures of criteria 1 or 2 are \`framing_errors\`. Failures of criterion 3 where the requirement describes implementation rather than observed behavior are \`internal_mechanics_drift\`. Two flavors:
|
|
52
|
+
|
|
53
|
+
- *Implementation primitives leaking out* — library function names, internal class names, scheduling vocabulary, buffer sizes that aren't part of public contract.
|
|
54
|
+
- *Pinning literals instead of properties* — when a requirement names a specific value (an exit code, an error string, a format prefix, a field name, a file path), the customer almost always depends on some characteristic (stable, distinguishable, parseable, idiomatic, named-rather-than-anonymous) rather than the value itself.
|
|
55
|
+
|
|
56
|
+
Other criterion-3 failures (vague or un-testable but not implementation-flavored) are \`framing_errors\`.
|
|
52
57
|
|
|
53
58
|
Note: per-area scope and sizing is the outline reviewer's job; you accept the area structure as given and focus on requirement-level correctness within it.
|
|
54
59
|
|
|
@@ -8,5 +8,5 @@
|
|
|
8
8
|
* Requirements covered:
|
|
9
9
|
* - CTS-SPEC-1, CTS-SPEC-2, CTS-SPEC-3, CTS-SPEC-4
|
|
10
10
|
*/
|
|
11
|
-
export declare const SPECIFIER_PROMPT = "You are reading a slice of a software codebase \u2014 the files relevant to ONE behavioral area of the system. Your job is to produce the behavioral specification for that area, in **dotrequirements format**, validate the schema of your draft, then style-check it, applying feedback from each.\n\nA separate planner agent has already broken the system into areas; you are responsible for ONE area only. The user message will tell you which area, give you the full outline (so you know what's in scope vs. not), point you at the slice, and tell you where to write your output.\n\n## What to capture\n\nA behavioral specification describes what the system does from the outside \u2014 what someone using it can observe, not how the implementation works. Scoped to your assigned area, capture:\n\n- **User-facing behaviors** \u2014 what the customer can do, what happens when they do it, what they see in response\n- **Integration behaviors** \u2014 how this area interacts with external services, what it sends/receives, how it handles failures\n- **Domain rules** \u2014 validation, business logic, state transitions, decision logic specific to this area\n- **Error and edge cases** \u2014 what happens when things go wrong, what the system tolerates, what it rejects\n- **Documented warnings, hazards, and limitations** \u2014 things the README or docstrings warn customers about\n\n## Customer and persona\n\
|
|
11
|
+
export declare const SPECIFIER_PROMPT = "You are reading a slice of a software codebase \u2014 the files relevant to ONE behavioral area of the system. Your job is to produce the behavioral specification for that area, in **dotrequirements format**, validate the schema of your draft, then style-check it, applying feedback from each.\n\nA separate planner agent has already broken the system into areas; you are responsible for ONE area only. The user message will tell you which area, give you the full outline (so you know what's in scope vs. not), point you at the slice, and tell you where to write your output.\n\n## What to capture\n\nA behavioral specification describes what the system does from the outside \u2014 what someone using it can observe, not how the implementation works. Scoped to your assigned area, capture:\n\n- **User-facing behaviors** \u2014 what the customer can do, what happens when they do it, what they see in response\n- **Integration behaviors** \u2014 how this area interacts with external services, what it sends/receives, how it handles failures\n- **Domain rules** \u2014 validation, business logic, state transitions, decision logic specific to this area\n- **Error and edge cases** \u2014 what happens when things go wrong, what the system tolerates, what it rejects\n- **Documented warnings, hazards, and limitations** \u2014 things the README or docstrings warn customers about\n\n## Customer and persona\n\nThe planner has already identified your area's customer(s) \u2014 they're listed in the area's `customers` field, which is passed to you in the user message. Each customer entry has a `name` and a `description` of who they are and what they care about.\n\nUse those customers as the named personas for your requirements. Pick the customer most relevant to each requirement (or requirement tree). Different requirements in the same area can use different personas if the area genuinely serves multiple customers; just keep each requirement tree (parent + children) grounded in a single persona so the tree reads coherently.\n\nIf the customers handed to you are not real users of the software \u2014 for example, they describe a contributor or stage author of *this codebase* rather than someone who consumes the software \u2014 STOP. Do not invent an alternative customer to make the area work. Instead, write a single short partial that says only \"AREA-LACKS-CUSTOMER: <one-sentence explanation of why no real customer was identified>\" and confirm completion. The pipeline will surface this as a finding for the human to reshape the outline.\n\n## Style principles\n\nApply these throughout your work:\n\n1. **Concrete examples, not vague language.** \"When a registered user provides valid credentials, they are authenticated\" \u2014 not \"users can log in\" or \"works properly.\"\n2. **Natural, concise prose.** Declarative (\"is authenticated\"), not \"should be\" or wandering narrative.\n3. **Arrange/Act/Assert framing in mind.** Each requirement reads as preconditions / trigger / outcome.\n4. **Framework neutral.** Default to unlabeled criteria. Use labels (e.g., Given/When/Then) only when they genuinely sharpen meaning \u2014 don't impose them as a format.\n5. **Named personas.** Establish a persona in the parent requirement; reuse them in children. E.g., parent: \"Casey, a React developer, can configure pLimit.\" Child: \"When Casey calls pLimit(5), they receive...\"\n6. **User-centric language.** Describe the customer's experience, not internal mechanics. \"They are brought to the dashboard,\" not \"they are redirected to /redirect/dashboard.\"\n7. **Single action per requirement.** No chaining multiple actions with \"and then.\" Break into separate requirements.\n8. **Independently testable.** Each requirement should make sense on its own. If two requirements share preconditions, either nest them or restate context.\n9. **Behavior, not design.** \"Provides valid credentials\" \u2014 not \"enters credentials into two single-line input fields and presses a green button.\"\n10. **Outcomes, not implementation.** Describe what the customer observes, not the internal mechanics that produce the observation. Two flavors of drift to watch for:\n - *Implementation primitives leaking out* \u2014 library function names, internal class names, scheduling vocabulary, buffer sizes \u2014 don't belong in requirements.\n - *Pinning literals instead of properties* \u2014 when a requirement names a specific value (an exit code, an error string, a format prefix, a field name, a file path), the customer almost always depends on some characteristic (stable, distinguishable, parseable, idiomatic, named-rather-than-anonymous) rather than the value itself.\n11. **Decompose large requirements.** If it can't be validated with a single test, break it down.\n\n## Read the documentation in your slice\n\nIf your slice contains README files, doc comments, JSDoc, or docstrings: read them carefully. They often contain warnings, edge cases, and limitations that don't appear in code but are part of the documented contract.\n\n## Workflow (REQUIRED)\n\n### Phase A: Draft\n\n1. Read your slice. Read documentation in the slice. If needed, Read/Grep the full pack to discover behaviors documented in tests or recipes.\n2. Use the **Write** tool to write your draft to the partial path provided in the user message. Output in this format:\n - A one-paragraph area description introducing the persona and what they do with this area's surface. This paragraph is what readers see at the top of the area in the final spec \u2014 it is your framing, informed by the deep reading you just did. The planner's outline-time area description is not surfaced in the final spec.\n - Blank line.\n - Series of fenced `dotrequirements` blocks.\n\n Do NOT include YAML frontmatter, H1 title, summary paragraph, or area H2. The composer adds those.\n\n### Phase B: Validate (schema/syntax \u2014 REQUIRED FIRST)\n\n1. Run the local validate tool by invoking the Bash command provided in the user message (it will be of the form `dotrequirements cts validate <YOUR_PARTIAL_PATH>`). It is deterministic and cheap \u2014 it checks that every requirement block parses, every criterion has a `\u2192` arrow, position paths match indentation, and IDs are unique.\n2. If validate prints `Schema validation: PASS`, proceed to Phase C.\n3. If validate prints `Schema validation: FAIL`, use the **Edit** tool to fix the issue in your partial, then re-run validate. Repeat until it passes. Do not move on with a failing validation \u2014 schema errors will cause the downstream pipeline to reject your spec.\n\n### Phase C: Style-check and revise (clarity \u2014 REQUIRED SECOND)\n\n1. Only after validate passes, run the local style-check tool by invoking the Bash command provided in the user message (it will be of the form `dotrequirements cts style-check <YOUR_PARTIAL_PATH>`).\n2. Read the feedback carefully.\n3. For every MUST FIX and SHOULD FIX finding, edit your partial in place using the **Edit** tool to apply the suggested change.\n4. Act on COULD IMPROVE findings unless doing so would make the spec worse.\n5. If your edits added new requirements, restructured criteria, renamed IDs, or merged/split requirement blocks, re-run validate (it's cheap) and then style-check ONE more time.\n6. Style-check runs at most twice per invocation. Feedback from any subsequent run is noted in your final stdout but not acted upon.\n7. When done revising, output a short confirmation: \"Done. Partial saved to <path>.\" That's it.\n\n## Format rules\n\n```dotrequirements\nPREFIX-AREA-1: Short imperative title\n 0. \u2192 A precondition or context\n 1. \u2192 An action or trigger\n 2. \u2192 An observable outcome\n 2.0. \u2192 Additional outcome detail\n```\n\n- **IDs**: `<defaultPrefix>-<areaPrefix>-<NUM>` \u2014 both prefixes are supplied in the user message. Number sequentially from 1; no zero-padding (`PLIM-AONE-1`, not `PLIM-AONE-001`).\n- **Criteria**: `<position>. <content>` for unlabeled (the default), or `<position>. <Label> \u2192 <content>` when a label sharpens meaning.\n- **Indentation**: 2 spaces per nesting level. Position paths must match indentation.\n\n## Output discipline\n\n- The PARTIAL FILE is your primary deliverable, written/edited via Write and Edit tools.\n- Your stdout response is brief \u2014 just confirmation when done.\n- No chain-of-thought narration in the partial file or your stdout.";
|
|
12
12
|
//# sourceMappingURL=specifier.d.ts.map
|
|
@@ -24,11 +24,11 @@ A behavioral specification describes what the system does from the outside — w
|
|
|
24
24
|
|
|
25
25
|
## Customer and persona
|
|
26
26
|
|
|
27
|
-
|
|
27
|
+
The planner has already identified your area's customer(s) — they're listed in the area's \`customers\` field, which is passed to you in the user message. Each customer entry has a \`name\` and a \`description\` of who they are and what they care about.
|
|
28
28
|
|
|
29
|
-
|
|
29
|
+
Use those customers as the named personas for your requirements. Pick the customer most relevant to each requirement (or requirement tree). Different requirements in the same area can use different personas if the area genuinely serves multiple customers; just keep each requirement tree (parent + children) grounded in a single persona so the tree reads coherently.
|
|
30
30
|
|
|
31
|
-
|
|
31
|
+
If the customers handed to you are not real users of the software — for example, they describe a contributor or stage author of *this codebase* rather than someone who consumes the software — STOP. Do not invent an alternative customer to make the area work. Instead, write a single short partial that says only "AREA-LACKS-CUSTOMER: <one-sentence explanation of why no real customer was identified>" and confirm completion. The pipeline will surface this as a finding for the human to reshape the outline.
|
|
32
32
|
|
|
33
33
|
## Style principles
|
|
34
34
|
|
|
@@ -43,7 +43,9 @@ Apply these throughout your work:
|
|
|
43
43
|
7. **Single action per requirement.** No chaining multiple actions with "and then." Break into separate requirements.
|
|
44
44
|
8. **Independently testable.** Each requirement should make sense on its own. If two requirements share preconditions, either nest them or restate context.
|
|
45
45
|
9. **Behavior, not design.** "Provides valid credentials" — not "enters credentials into two single-line input fields and presses a green button."
|
|
46
|
-
10. **Outcomes, not implementation.** Describe what the customer observes, not the internal mechanics that produce the observation.
|
|
46
|
+
10. **Outcomes, not implementation.** Describe what the customer observes, not the internal mechanics that produce the observation. Two flavors of drift to watch for:
|
|
47
|
+
- *Implementation primitives leaking out* — library function names, internal class names, scheduling vocabulary, buffer sizes — don't belong in requirements.
|
|
48
|
+
- *Pinning literals instead of properties* — when a requirement names a specific value (an exit code, an error string, a format prefix, a field name, a file path), the customer almost always depends on some characteristic (stable, distinguishable, parseable, idiomatic, named-rather-than-anonymous) rather than the value itself.
|
|
47
49
|
11. **Decompose large requirements.** If it can't be validated with a single test, break it down.
|
|
48
50
|
|
|
49
51
|
## Read the documentation in your slice
|