@acrasie/dev-flow 0.0.0-stage → 1.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (58) hide show
  1. package/.codex-plugin/plugin.json +20 -0
  2. package/LICENSE +21 -0
  3. package/README.md +181 -2
  4. package/dist/codex-dev-flow.mjs +3 -0
  5. package/dist/dev-flow.mjs +241 -0
  6. package/docs/adr/0001-hybrid-portable-workflow.md +23 -0
  7. package/docs/adr/0002-share-an-invalidable-context-capsule.md +55 -0
  8. package/docs/adr/0004-scale-assurance-lanes-by-applicable-risk.md +36 -0
  9. package/docs/adr/0006-make-intake-adaptive-user-authoritative-and-token-efficient.md +76 -0
  10. package/docs/adr/0007-collect-opt-in-local-benchmark-feedback.md +82 -0
  11. package/docs/adr/0008-automate-maintainer-releases-with-an-interactive-bun-workflow.md +121 -0
  12. package/docs/adr/0009-separate-intake-decisions-from-shape-discovery.md +200 -0
  13. package/docs/adr/0010-choose-quick-or-plan-after-discovery.md +161 -0
  14. package/docs/adr/0011-separate-fast-local-and-authoritative-ci-quality-gates.md +49 -0
  15. package/docs/adr/0012-use-bun-test-and-require-node-24.md +41 -0
  16. package/docs/adr/0013-layer-source-distribution-and-runtime-tests.md +42 -0
  17. package/docs/adr/0014-ratchet-source-coverage-with-bun.md +51 -0
  18. package/docs/adr/0015-split-fast-and-type-aware-linting.md +41 -0
  19. package/docs/adr/0016-use-husky-with-a-tested-bun-staged-file-adapter.md +45 -0
  20. package/docs/adr/0017-format-conservatively-with-oxfmt.md +45 -0
  21. package/docs/adr/0018-use-a-high-signal-oxlint-policy.md +53 -0
  22. package/docs/adr/0019-gate-deterministic-size-and-observe-timing.md +44 -0
  23. package/docs/adr/0020-support-linux-and-macos-with-targeted-ci.md +41 -0
  24. package/docs/adr/0021-randomize-tests-without-retries.md +35 -0
  25. package/docs/adr/0022-use-one-root-bun-workspace.md +41 -0
  26. package/docs/adr/0024-make-gate-a-minimal-plan-approval.md +74 -0
  27. package/docs/adr/0025-end-the-lifecycle-after-assure.md +55 -0
  28. package/docs/adr/0026-keep-intake-product-stable-and-interview-shape-by-dependency.md +151 -0
  29. package/docs/adr/0027-add-agentic-project-init-and-versioned-engineering-profiles.md +147 -0
  30. package/docs/adr/0028-make-public-documentation-user-first-and-current.md +65 -0
  31. package/docs/adr/0029-make-build-a-native-execution-boundary.md +51 -0
  32. package/docs/adr/0030-unify-product-domain-and-technical-design-interviews.md +240 -0
  33. package/docs/adr/0031-make-assure-the-success-boundary.md +205 -0
  34. package/docs/artifacts.md +47 -0
  35. package/docs/baselines/2026-07-18-p0-lifecycle.json +142 -0
  36. package/docs/design.md +101 -0
  37. package/docs/getting-started.md +204 -0
  38. package/docs/glossary/dev-flow.md +527 -0
  39. package/docs/lifecycle-contract.md +189 -0
  40. package/docs/lifecycle-contract.projection.json +931 -0
  41. package/docs/metrics-protocol.md +113 -0
  42. package/docs/project-profile-contract.md +157 -0
  43. package/docs/runbooks/maintainer-release.md +291 -0
  44. package/docs/target-intake-shape-contract.md +416 -0
  45. package/package.json +68 -4
  46. package/schemas/config.schema.json +104 -0
  47. package/schemas/policy.schema.json +17 -0
  48. package/schemas/project-init-state.schema.json +159 -0
  49. package/schemas/project-profile-local.schema.json +53 -0
  50. package/schemas/project-profile.schema.json +285 -0
  51. package/schemas/state.schema.json +826 -0
  52. package/skills/debug-root-cause/SKILL.md +16 -0
  53. package/skills/design-decisions/SKILL.md +24 -0
  54. package/skills/dev-flow/SKILL.md +306 -0
  55. package/skills/dev-flow/agents/openai.yaml +6 -0
  56. package/skills/discover-change/SKILL.md +31 -0
  57. package/skills/plan-change/SKILL.md +29 -0
  58. package/skills/review-change/SKILL.md +21 -0
@@ -0,0 +1,23 @@
1
+ # ADR 0001: Use a Hybrid Portable Workflow
2
+
3
+ ## Status
4
+
5
+ Accepted.
6
+
7
+ ## Context
8
+
9
+ The workflow must replace Superpowers for Codex while remaining portable across repositories and avoiding parallel write conflicts. Prompt-only workflows do not provide reliable state or policy enforcement. A monolithic command-line state machine would be too rigid for varied codebases.
10
+
11
+ ## Decision
12
+
13
+ Publish `codex-dev-flow` as a Codex plugin with a single public `$dev-flow` skill. Internal skills own reasoning. Versioned scripts own deterministic operations: command parsing, configuration resolution, policy enforcement, state validation, fingerprints, and durable artifact handling.
14
+
15
+ The plugin reads repository guidance first, then inspects manifests, CI, and existing tests. It does not require an `AGENTS.md`, but repository guidance wins when present.
16
+
17
+ ## Consequences
18
+
19
+ - One primary workflow avoids duplicated lifecycle rules.
20
+ - Scripts make safety-critical behavior testable and reproducible.
21
+ - Skills remain adaptable to repository-specific architecture.
22
+ - The plugin must keep schemas, scripts, and skill instructions compatible.
23
+ - The plugin cannot infer or grant Git or GitHub permissions.
@@ -0,0 +1,55 @@
1
+ # ADR 0002: Share an Invalidable Context Capsule Across Lifecycle Phases
2
+
3
+ ## Status
4
+
5
+ Accepted.
6
+
7
+ ## Context
8
+
9
+ Discovery, planning, execution, review, and verification can repeatedly inspect and
10
+ restate the same repository facts. This consumes quota and tokens without adding
11
+ assurance. Reusing unconstrained summaries would reduce cost but could also propagate
12
+ stale or unsupported claims across phases.
13
+
14
+ Codex Dev Flow optimizes quota and token consumption subject to two hard constraints:
15
+ reliability and absence of false success.
16
+
17
+ ## Decision
18
+
19
+ Each durable Dev Flow task owns one compact Context Capsule. Lifecycle phases consume
20
+ the capsule and enrich it instead of independently reconstructing all prior context.
21
+
22
+ The capsule carries structured references to:
23
+
24
+ - objective, approved scope, mode, and risk;
25
+ - relevant files and symbols;
26
+ - effective constraints and policy;
27
+ - acceptance criteria and validation commands;
28
+ - decisions and evidence already obtained;
29
+ - Git, configuration, policy, and artifact fingerprints.
30
+
31
+ Every material claim records provenance and freshness inputs. Fingerprint changes
32
+ invalidate affected sections, not necessarily the whole capsule. A phase must reinspect
33
+ invalid, missing, or insufficient context before relying on it. It may never interpret
34
+ stale context as proof.
35
+
36
+ The capsule is a compact workflow projection, not a second owner for detailed evidence,
37
+ architecture decisions, or approved scope. Versioned change records and ADRs retain
38
+ those responsibilities.
39
+
40
+ ## Consequences
41
+
42
+ - Repeated repository reading and prompt restatement decrease.
43
+ - Resume can reconstruct less conversational context.
44
+ - Phase contracts must declare capsule fields consumed and produced.
45
+ - Capsule schema, provenance, fingerprints, invalidation, and compaction require
46
+ deterministic code and tests.
47
+ - Selective invalidation is more complex than a single opaque summary.
48
+ - Missing assurance data blocks or triggers reinspection; it never permits degraded
49
+ success.
50
+
51
+ ## Follow-up Decisions
52
+
53
+ - Define exact schema and storage projection with durable state work in P1.
54
+ - Define invalidation matrix for each fingerprint and phase contract.
55
+ - Define token and cache-hit metrics before dogfooding.
@@ -0,0 +1,36 @@
1
+ # ADR 0004: Scale Assurance Lanes by Applicable Risk
2
+
3
+ ## Status
4
+
5
+ Accepted.
6
+
7
+ ## Context
8
+
9
+ Repeating full repository discovery during review and running every specialist review
10
+ lane for every change wastes quota. Removing independent review from Critical changes,
11
+ however, would weaken assurance and could hide false success.
12
+
13
+ ## Decision
14
+
15
+ ASSURE consumes a compact Diff Capsule derived from approved criteria, applicable risks,
16
+ changed surfaces, dependency impact, and fresh validation evidence.
17
+
18
+ Standard runs one structured review over changed and impacted surfaces. Critical keeps
19
+ independent read-only review, but activates specialist lanes according to applicable
20
+ risk. A lane may be omitted only with a structured, evidence-backed non-applicability
21
+ decision. Ambiguous applicability activates the lane.
22
+
23
+ Findings retain deterministic fingerprints. A resolved finding is not reconsidered
24
+ unless its affected code or supporting assumptions changed. Verification reuses fresh
25
+ evidence and reruns missing, invalidated, or globally mandatory checks.
26
+
27
+ Known deterministic failures block expensive reviewer calls until corrected.
28
+
29
+ ## Consequences
30
+
31
+ - Review tokens track changed risk surface more closely than repository size.
32
+ - Critical independence remains while irrelevant specialist work decreases.
33
+ - Risk-to-lane rules and non-applicability evidence become testable policy.
34
+ - Incorrectly classifying a lane as irrelevant is a new risk; ambiguity therefore fails
35
+ toward review, not omission.
36
+ - Diff and evidence invalidation must share fingerprints with Context Capsule state.
@@ -0,0 +1,76 @@
1
+ # ADR 0006: Make Intake Adaptive, User-Authoritative, and Token-Efficient
2
+
3
+ ## Status
4
+
5
+ Accepted for current runtime. Replacement target is recorded by ADR 0030; ADR 0026 is
6
+ superseded as target design.
7
+
8
+ ## Context
9
+
10
+ INTAKE must establish enough shared understanding to classify a request safely before
11
+ SHAPE explores the repository. A fixed questionnaire wastes tokens on already explicit
12
+ requests. Letting the model silently fill gaps is cheaper in the short term but removes
13
+ user control and can create expensive rework later. Moving full technical discovery into
14
+ INTAKE would duplicate SHAPE and spend quota before mode and risk are confirmed.
15
+
16
+ The workflow therefore needs an interview that can be relentless about unresolved
17
+ decisions without becoming exhaustive about irrelevant context.
18
+
19
+ ## Decision
20
+
21
+ INTAKE uses a two-level interview boundary:
22
+
23
+ - INTAKE resolves product intent, an observable success signal, determining constraints,
24
+ mode/risk, and only the scope boundaries needed to classify the request.
25
+ - SHAPE resolves repository and implementation uncertainty, then turns the success
26
+ signal into testable acceptance criteria and an executable contract.
27
+
28
+ The interview is adaptive. A sufficiently explicit request may require zero questions.
29
+ Every generated question must identify the unresolved decision and the lifecycle outputs
30
+ its answer can change. Questions without determining impact are forbidden.
31
+
32
+ Questions are generated lazily, one at a time. Each question has two or three mutually
33
+ exclusive choices, one clearly justified recommendation, and a free-form answer path.
34
+ The user must answer or cancel. Silence, “decide for me”, or a request to advance cannot
35
+ bypass an unresolved question. Recommendations reduce interaction cost but never become
36
+ decisions without an explicit answer.
37
+
38
+ The model owns semantic sufficiency and question wording. Deterministic runtime checks
39
+ own structure, active-question uniqueness, required answers, contradictions,
40
+ deduplication, redaction, and the exit guard. One model evaluation per user response
41
+ integrates the answer, resolves every explicit decision it contains, evaluates
42
+ completeness, and emits at most the next question. A second model is not called merely
43
+ to verify the first.
44
+
45
+ INTAKE inspects only cheap entry context: repository guidance, Dev Flow configuration
46
+ and policy, primary manifests, workspace/Git identity, and explicitly cited documents.
47
+ Deeper code, dependency, test, and architecture exploration belongs to SHAPE.
48
+
49
+ INTAKE persists a compact Intake Brief and one active Decision Question, never the full
50
+ conversation. SHAPE consumes this structured projection and source references rather
51
+ than the interview transcript. If a later phase discovers missing objective, success,
52
+ determining-constraint, or mode/risk context, the workflow returns to INTAKE and
53
+ selectively invalidates dependent outputs.
54
+
55
+ Standard and Critical move automatically to SHAPE when semantic and deterministic
56
+ completeness checks pass. An explicit “move on” request is refused while context is
57
+ incomplete. Answering INTAKE questions is not implementation Approval; BUILD remains
58
+ guarded by the canonical GATE contract. Quick may bypass SHAPE only when its existing
59
+ bounded-scope guard also proves the request is already executable and verifiable without
60
+ a design decision.
61
+
62
+ After five distinct questions, INTAKE presents a compact checkpoint of resolved and
63
+ remaining decisions. This is a control point, not a hard question limit.
64
+
65
+ ## Consequences
66
+
67
+ - Clear requests pay near-zero interview overhead.
68
+ - Ambiguous requests spend tokens only on decisions capable of changing outcome, risk,
69
+ mode, or downstream work.
70
+ - The user remains the authority for every human decision.
71
+ - Resume can re-present the exact active question without replaying the conversation.
72
+ - SHAPE receives smaller, more stable context and can detect ownership violations.
73
+ - A hybrid question engine and new persisted fields require schema, state-machine,
74
+ migration, validation, and crash/resume tests.
75
+ - Model sufficiency can still be wrong; deterministic guards, downstream reopening, and
76
+ benchmark feedback expose that risk without a routine second model call.
@@ -0,0 +1,82 @@
1
+ # ADR 0007: Collect Opt-In Local Benchmark Feedback
2
+
3
+ ## Status
4
+
5
+ Accepted.
6
+
7
+ ## Context
8
+
9
+ Codex Dev Flow is valuable only when its reliability and avoided rework justify its
10
+ orchestration cost. Running equivalent tasks with multiple models or comparing Dev Flow
11
+ against an additional vanilla Codex execution would multiply token usage and defeat that
12
+ goal. Automatic production telemetry would also create privacy, consent, and operational
13
+ costs disproportionate to early local dogfooding.
14
+
15
+ The maintainer still needs structured evidence about user satisfaction, contract
16
+ conformance, perceived plugin value, and common deviations.
17
+
18
+ ## Decision
19
+
20
+ New runs may opt into local benchmark feedback with the public invocation flag:
21
+
22
+ ```text
23
+ $dev-flow --benchmark <objective>
24
+ ```
25
+
26
+ The flag is disabled by default, accepted only when creating a task, persisted with that
27
+ task, retained automatically across resume, and exposed by status. It never triggers a
28
+ second task, another model, an A/B comparison, or an upload.
29
+
30
+ After `finished` or `failed`, Dev Flow offers three deterministic
31
+ questions without a model call:
32
+
33
+ 1. satisfaction from 1 to 5;
34
+ 2. conformance to the expressed need: yes, partially, or no;
35
+ 3. Dev Flow value: useful, neutral, or counterproductive.
36
+
37
+ Each question permits an optional free-form clarification, followed by an optional
38
+ general comment. The lifecycle result is already final: feedback cannot change task
39
+ status or evidence. No questionnaire is offered for `blocked` or `cancelled`.
40
+
41
+ A benchmark record is written atomically only after all three required answers exist.
42
+ Closing the session before completion creates no benchmark record and contributes no
43
+ benchmark data.
44
+
45
+ Deterministic deviation checks compare:
46
+
47
+ - Quick and Plan against current canonical Change Contract and Plan Approval Receipt.
48
+
49
+ They record structured facts such as unexpected file additions/deletions, unplanned
50
+ surfaces, uncovered criteria, skipped or failed checks, and unapproved scope changes. A
51
+ new file is a deviation only when the authoritative contract did not
52
+ permit it.
53
+
54
+ Completed records live under:
55
+
56
+ ```text
57
+ .codex/benchmarks/<task-id>.json
58
+ ```
59
+
60
+ The directory is gitignored and excluded from product-diff analysis. Records contain
61
+ allowlisted metadata, questionnaire answers, deterministic deviation signals, task and
62
+ artifact references, plugin/policy versions, duration, and token/call metrics when
63
+ available. They contain no prompt, transcript, code, diff, secret, or automatic upload.
64
+
65
+ The deterministic command below summarizes completed records without a model:
66
+
67
+ ```bash
68
+ npx codex-dev-flow benchmark summary
69
+ ```
70
+
71
+ ## Consequences
72
+
73
+ - Maintainer can dogfood one real task and retain comparable feedback without paying for
74
+ duplicate executions.
75
+ - Non-response is intentionally absent rather than treated as negative or positive
76
+ evidence.
77
+ - Public flag and record format become compatibility surfaces requiring parser, schema,
78
+ state, privacy, and migration tests.
79
+ - User comments remain local but still require length bounds and deterministic secret
80
+ redaction before persistence.
81
+ - Benchmark collection stays outside the five lifecycle phases and cannot weaken truthful
82
+ terminality.
@@ -0,0 +1,121 @@
1
+ # ADR 0008: Automate Maintainer Releases with an Interactive Bun Workflow
2
+
3
+ ## Status
4
+
5
+ Accepted.
6
+
7
+ ## Context
8
+
9
+ Releasing Codex Dev Flow currently requires the maintainer to remember unrelated Git,
10
+ plugin-cache, npm-registry, and Codex marketplace commands. Some commands depend on
11
+ machine-specific absolute paths, and the Codex plugin and npm package have different
12
+ publication mechanisms. This makes an otherwise routine release error-prone and leaves
13
+ partial-release recovery implicit.
14
+
15
+ The maintainer needs a self-service release flow with one memorable command, explicit
16
+ authority before irreversible actions, deterministic recommendations, and no model call.
17
+
18
+ ## Decision
19
+
20
+ The repository will expose this command from its root:
21
+
22
+ ```bash
23
+ bun run release
24
+ ```
25
+
26
+ A private root `package.json` delegates to a TypeScript release program in the plugin
27
+ package. The program runs with Bun and uses `@clack/prompts`. V1 guarantees macOS
28
+ support, avoids shell-specific implementation, is expected to work on Linux, and does
29
+ not claim tested Windows support. The repository migrates from `package-lock.json` to a
30
+ single committed `bun.lock`.
31
+
32
+ The first prompt selects one release target:
33
+
34
+ - Codex plugin;
35
+ - npm package;
36
+ - Codex plugin and npm package, selected by default.
37
+
38
+ The next prompt selects publication or read-only preview. Preview performs discovery,
39
+ version recommendation, preflight checks, and summary generation without modifying the
40
+ workspace or any remote system.
41
+
42
+ V1 publishes stable versions only. A deterministic Conventional Commits analysis
43
+ recommends the version increment:
44
+
45
+ - `fix:` recommends patch;
46
+ - `feat:` recommends minor;
47
+ - `!` or `BREAKING CHANGE:` recommends major;
48
+ - unclassified commits recommend patch.
49
+
50
+ The maintainer may choose patch, minor, major, or a greater unpublished custom stable
51
+ version. For a plugin-only release, the semantic base remains unchanged and only the
52
+ Codex cachebuster changes. For an npm-only release, only the package semantic version
53
+ changes. A combined release gives both artifacts the same semantic base and adds the
54
+ cachebuster only to the plugin manifest.
55
+
56
+ Release identities are target-specific:
57
+
58
+ ```text
59
+ combined: v0.2.0
60
+ npm only: npm-v0.2.0
61
+ plugin only: plugin-v0.2.0-codex.20260728...
62
+ ```
63
+
64
+ The corresponding commit uses `chore(release): <release-identity>`. The Git tag is
65
+ annotated and follows the local Git signing configuration.
66
+
67
+ Before changing files, the program requires a clean `main`, fetches `origin/main`, and
68
+ fast-forwards when safe. It refuses local commits, divergence, conflicts, or another
69
+ branch. It never resets, rebases, or merges automatically.
70
+
71
+ All applicable validation is mandatory and has no force-publish bypass:
72
+
73
+ - frozen Bun dependency installation;
74
+ - typecheck, deterministic build, and tests;
75
+ - plugin manifest validation;
76
+ - npm package-content dry run;
77
+ - target-version availability;
78
+ - GitHub and npm authentication;
79
+ - expected clean or allowlisted generated-file state.
80
+
81
+ The release commit may contain only the plugin manifest, affected package metadata,
82
+ `bun.lock` when required, and the deterministic distribution bundle. Any other generated
83
+ change blocks the release.
84
+
85
+ Authentication is checked before mutation. The program never reads, prompts for, logs,
86
+ or stores credentials. Missing authentication produces an exact corrective command and
87
+ requires the maintainer to rerun the release. OTP handling remains inside `bun publish`.
88
+
89
+ The program generates GitHub release notes before publication, displays a non-editable
90
+ preview, and lets the maintainer accept or abort. The final Clack confirmation defaults
91
+ to `No` and summarizes target, versions, tag, files, checks, publication order, and
92
+ planned local installation.
93
+
94
+ Before the first successful push, cancellation, `Ctrl+C`, or failure restores the exact
95
+ original bytes and removes only local commit or tag objects created by the program. Once
96
+ an external publication succeeds, the program never attempts automatic rollback.
97
+
98
+ Publication order is:
99
+
100
+ 1. push the release commit and tag;
101
+ 2. publish the npm package with `bun publish` when applicable;
102
+ 3. create the GitHub Release from the previewed notes;
103
+ 4. optionally refresh marketplace `acrazie`, reinstall `codex-dev-flow`, and remind the
104
+ maintainer to start a new Codex thread.
105
+
106
+ Git and registry state form the recovery record. A later `bun run release` detects a
107
+ partial release and offers the single missing idempotent action instead of creating a
108
+ new version. No local state file or secret is required.
109
+
110
+ ## Consequences
111
+
112
+ - Routine release operation becomes one short command with maintainer-controlled
113
+ decisions and no model dependency.
114
+ - Plugin and npm releases may remain independent without ambiguous Git tags.
115
+ - Strict preflight and generated-file allowlists reduce accidental publication.
116
+ - Git, npm, and GitHub cannot form one atomic transaction, so explicit partial-release
117
+ detection and resume behavior are required.
118
+ - Root release orchestration, Bun lockfile migration, subprocess isolation, cancellation,
119
+ and recovery paths require dedicated deterministic tests.
120
+ - V1 intentionally excludes prereleases, editable release notes, Windows guarantees, and
121
+ validation bypasses.
@@ -0,0 +1,200 @@
1
+ # ADR 0009: Separate Intake Decisions from Shape Discovery
2
+
3
+ ## Status
4
+
5
+ Accepted for current runtime. Replacement target is recorded by ADR 0030; ADR 0026 is
6
+ superseded as target design.
7
+
8
+ ## Context
9
+
10
+ Adaptive INTAKE now owns product intent, observable success, determining constraints,
11
+ and risk before SHAPE performs deeper repository discovery. Without a stronger
12
+ downstream boundary, SHAPE can duplicate the interview, silently choose among outcomes
13
+ that need human authority, or send an unstructured ambiguity back to INTAKE and force
14
+ the same technical investigation to run twice.
15
+
16
+ The boundary cannot depend only on whether a choice first appears technical. Repository
17
+ discovery can reveal alternatives whose effects change user experience, cost, delivery
18
+ time, scope, or the observable definition of success.
19
+
20
+ ## Decision
21
+
22
+ Phase ownership follows decision impact:
23
+
24
+ - INTAKE owns every unresolved choice that can change product intent, user experience,
25
+ cost, delivery time, scope, determining constraints, risk, or observable success.
26
+ - SHAPE owns repository discovery and purely technical uncertainty that does not change
27
+ those product outcomes.
28
+ - SHAPE never silently resolves a product-impacting technical choice. It suspends
29
+ shaping and requests an explicit INTAKE decision.
30
+ - SHAPE does not repeat an INTAKE decision unless fresh evidence invalidates it.
31
+
32
+ The SHAPE-to-INTAKE return uses a rich, structured Decision Escalation rather than a
33
+ plain missing-context signal. It contains:
34
+
35
+ - a stable decision key;
36
+ - technical evidence and provenance;
37
+ - the product impact;
38
+ - two or three feasible, mutually exclusive options;
39
+ - an evidence-backed recommendation;
40
+ - the SHAPE sections that may depend on the answer.
41
+
42
+ INTAKE validates this escalation and presents the explicit Decision Question. It does
43
+ not repeat the underlying repository discovery merely to reconstruct options. The user
44
+ remains the authority for the answer; a recommendation is not a decision.
45
+
46
+ SHAPE invalidation is always selective. Opening an escalation suspends
47
+ `contract_complete` and marks only the declared dependent sections as unreliable.
48
+ Fresh, independent discovery evidence and contract sections remain reusable. After the
49
+ INTAKE answer, the workflow invalidates the transitive dependency closure of the
50
+ decision and resumes targeted discovery or planning.
51
+
52
+ SHAPE is never reset wholesale. If dependency impact is ambiguous, the workflow widens
53
+ targeted discovery to resolve that ambiguity or blocks with the unresolved dependency;
54
+ it does not discard unrelated evidence. This rule prevents already-paid repository
55
+ reasoning from being repeated solely because one product decision changed.
56
+
57
+ Selective invalidation uses an explicit persisted dependency graph. Decisions, source
58
+ fingerprints, discovery evidence, and contract sections have stable identities. Every
59
+ derived discovery item and contract section declares `dependsOn` references to the
60
+ items required for its validity. The runtime validates that references exist, rejects
61
+ cycles, and computes the transitive dependency closure when a root changes.
62
+
63
+ The model proposes dependency links while shaping because it owns their semantic
64
+ meaning. The deterministic runtime owns structural validation and invalidation. It
65
+ never reconstructs the graph by asking a model to reread all SHAPE output during
66
+ resume.
67
+
68
+ SHAPE discovery is adaptive and targeted, not a fixed repository checklist. It starts
69
+ from the Intake Brief, structured decisions, transferred technical unknowns, and fresh
70
+ reusable evidence. Each discovery target has a stable key, a precise technical question,
71
+ its possible material impact, candidate sources, a result with provenance, and explicit
72
+ dependencies.
73
+
74
+ Discovery widens only when fresh evidence exposes an uncertainty capable of changing
75
+ scope, risk, acceptance criteria, implementation plan, validation, or the preparation
76
+ profile recommendation. SHAPE stops
77
+ exploring when no material technical uncertainty remains and every required contract
78
+ section has sufficient fresh supporting evidence.
79
+
80
+ SHAPE retains two strictly separated roles over one shared structured state:
81
+
82
+ - Discovery may read the repository and resolve Discovery Targets into fresh,
83
+ provenance-bound evidence. It does not author the canonical implementation plan.
84
+ - Planning may project the canonical contract from fresh evidence. It does not read the
85
+ repository or silently supplement missing evidence.
86
+
87
+ After Discovery proves `discovery_sufficient`, the workflow obtains the explicit Quick
88
+ or Plan preparation choice before Planning starts. This is a SHAPE projection choice,
89
+ not a product decision. Planning uses the selected profile without repeating Discovery.
90
+
91
+ When Planning lacks evidence, it creates a new Discovery Target and yields to Discovery.
92
+ After that target is resolved, Planning resumes from shared state. The transfer is
93
+ structured and dependency-linked, never a prose summary that causes either role to
94
+ repeat the other's work.
95
+
96
+ The nominal SHAPE path uses one Discovery reasoning session, one profile choice, and one
97
+ Planning reasoning session. One Discovery session may resolve multiple related targets.
98
+ There is no routine semantic re-review call: deterministic runtime verifies structure,
99
+ freshness, coverage, and dependencies. Additional reasoning is allowed only for a new
100
+ material Discovery Target or an invalidated evidence closure, and receives only the
101
+ relevant state subset. Planning never rereads repository sources.
102
+
103
+ SHAPE uses configurable soft budgets for reasoning calls, estimated tokens, and context.
104
+ Defaults must come from benchmark evidence rather than arbitrary hard-coded counts. When
105
+ a budget approaches exhaustion, the workflow persists all current evidence and displays
106
+ a compact checkpoint offering explicit continuation, scope reduction through INTAKE, or
107
+ cancellation. It never discards SHAPE work or silently spends beyond the budget.
108
+
109
+ `ShapeState` is the aggregate root for `discoveryTargets`, `evidence`,
110
+ `sourceFingerprints`, `decisionEscalations`, `preparationProfile`, `contract`,
111
+ `dependencyGraph`, and semantic `evaluatorReceipts`. Its primary progression is:
112
+
113
+ `discovering -> awaiting_profile_choice -> planning -> awaiting_approval`
114
+
115
+ Planning returns to `discovering` only through a material Discovery Target. Discovery or
116
+ Planning enters `awaiting_intake_decision` only through a Decision Escalation. Required
117
+ events include `discovery.target_recorded`, `discovery.evidence_recorded`,
118
+ `discovery.completed`, `profile.choice_requested`, `profile.choice_recorded`,
119
+ `contract.projected`, and `evidence.invalidated`.
120
+
121
+ For the Plan profile, Planning first completes the canonical contract and the parent
122
+ workflow renders a Markdown projection marked `Proposed`. GATE presents that projection
123
+ with the canonical digest. Approval binds logical contract content; changing projection
124
+ status to `Approved` does not alter that content.
125
+
126
+ The projection has its own fingerprint. A manual file edit never silently overwrites
127
+ canonical state and is never silently discarded. It returns to Planning with an
128
+ explicit reconciliation choice: integrate the edit into the canonical contract or
129
+ regenerate the projection. Any integrated material change requires a new Gate digest
130
+ and Approval.
131
+
132
+ SHAPE completion uses two hybrid proofs:
133
+
134
+ - `discovery_sufficient`: the Discovery model declares that no material technical
135
+ uncertainty remains; deterministic checks require every target to be resolved or
136
+ escalated, every supporting item to be fresh and provenance-bound, and no relevant
137
+ invalidation to remain open.
138
+ - `contract_complete`: the Planning model declares the contract implementable;
139
+ deterministic checks require the profile- and policy-applicable contract sections,
140
+ valid acyclic
141
+ dependencies, criterion-to-task coverage, task-to-criterion coverage, and a
142
+ verification method for every criterion. Risk or policy additionally requires threat
143
+ and rollback records when applicable.
144
+
145
+ Exiting SHAPE requires both proofs. A second model does not re-evaluate another model's
146
+ semantic judgment merely for confirmation; deterministic runtime owns structural
147
+ verification.
148
+
149
+ A Decision Escalation moves the task to a distinct `awaiting_intake_decision` state
150
+ rather than overloading `awaiting_mode_confirmation`. The state persists the escalation
151
+ identity, active Decision Question, prior SHAPE substate, affected dependency roots, and
152
+ freshness references for retained evidence.
153
+
154
+ After an explicit answer, INTAKE integrates the decision and recomputes classification
155
+ or routing only when their declared dependencies are affected. Runtime invalidates the
156
+ decision's dependency closure. The task resumes `discovering` when required evidence is
157
+ missing; otherwise it resumes `planning`. It does not rerun initial INTAKE.
158
+
159
+ Freshness is tracked per evidence source rather than through one repository-wide Git
160
+ invalidation. Tracked files use blob identities; untracked files and configuration use
161
+ content digests; command evidence binds the command, relevant inputs, and observed
162
+ result; external evidence uses a stable identity plus policy-defined observation time
163
+ or TTL.
164
+
165
+ Resume checks only sources referenced by live evidence. A changed source invalidates
166
+ its transitive dependency closure and creates a targeted refresh Discovery Target.
167
+ Unchanged source evidence is reused without another model read. A broad Git change
168
+ alone is not sufficient reason to invalidate unrelated SHAPE work.
169
+
170
+ ## Consequences
171
+
172
+ - INTAKE and SHAPE have an effect-based ownership boundary that also covers choices
173
+ discovered late.
174
+ - Technical evidence crosses the boundary without transferring ownership of the human
175
+ decision to SHAPE.
176
+ - Selective invalidation requires explicit dependency links between decisions,
177
+ discovery evidence, and contract sections.
178
+ - Resume reuses fingerprint-valid independent evidence and spends tokens only on the
179
+ invalidated dependency closure.
180
+ - Persisted dependency metadata grows with the contract, but makes invalidation
181
+ deterministic and auditable.
182
+ - Small changes inspect few surfaces; complex changes widen only through recorded
183
+ material reasons.
184
+ - Discovery and Planning can be tested and routed independently without paying for
185
+ duplicate repository reads.
186
+ - Planning may require multiple resumable passes, but each additional pass must be
187
+ justified by a new or invalidated Discovery Target.
188
+ - The common path pays for two SHAPE reasoning sessions rather than repeated discovery,
189
+ planning, and model-verification passes.
190
+ - Budget exhaustion is resumable and user-authoritative rather than a lossy hard stop.
191
+ - Semantic sufficiency remains model-owned while structural completion and coverage are
192
+ deterministic and testable.
193
+ - The lifecycle gains an internal INTAKE state whose name reflects late product
194
+ decisions rather than mode confirmation.
195
+ - Crash/resume can re-present the exact escalated question and return to the correct
196
+ SHAPE role.
197
+ - Source-granular fingerprints cost more metadata but prevent unrelated Git changes
198
+ from causing expensive rediscovery.
199
+ - Persisted state and schemas need a resumable Decision Escalation linked to the active
200
+ Decision Question.