@acrasie/dev-flow 0.0.0-stage → 1.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.codex-plugin/plugin.json +20 -0
- package/LICENSE +21 -0
- package/README.md +181 -2
- package/dist/codex-dev-flow.mjs +3 -0
- package/dist/dev-flow.mjs +241 -0
- package/docs/adr/0001-hybrid-portable-workflow.md +23 -0
- package/docs/adr/0002-share-an-invalidable-context-capsule.md +55 -0
- package/docs/adr/0004-scale-assurance-lanes-by-applicable-risk.md +36 -0
- package/docs/adr/0006-make-intake-adaptive-user-authoritative-and-token-efficient.md +76 -0
- package/docs/adr/0007-collect-opt-in-local-benchmark-feedback.md +82 -0
- package/docs/adr/0008-automate-maintainer-releases-with-an-interactive-bun-workflow.md +121 -0
- package/docs/adr/0009-separate-intake-decisions-from-shape-discovery.md +200 -0
- package/docs/adr/0010-choose-quick-or-plan-after-discovery.md +161 -0
- package/docs/adr/0011-separate-fast-local-and-authoritative-ci-quality-gates.md +49 -0
- package/docs/adr/0012-use-bun-test-and-require-node-24.md +41 -0
- package/docs/adr/0013-layer-source-distribution-and-runtime-tests.md +42 -0
- package/docs/adr/0014-ratchet-source-coverage-with-bun.md +51 -0
- package/docs/adr/0015-split-fast-and-type-aware-linting.md +41 -0
- package/docs/adr/0016-use-husky-with-a-tested-bun-staged-file-adapter.md +45 -0
- package/docs/adr/0017-format-conservatively-with-oxfmt.md +45 -0
- package/docs/adr/0018-use-a-high-signal-oxlint-policy.md +53 -0
- package/docs/adr/0019-gate-deterministic-size-and-observe-timing.md +44 -0
- package/docs/adr/0020-support-linux-and-macos-with-targeted-ci.md +41 -0
- package/docs/adr/0021-randomize-tests-without-retries.md +35 -0
- package/docs/adr/0022-use-one-root-bun-workspace.md +41 -0
- package/docs/adr/0024-make-gate-a-minimal-plan-approval.md +74 -0
- package/docs/adr/0025-end-the-lifecycle-after-assure.md +55 -0
- package/docs/adr/0026-keep-intake-product-stable-and-interview-shape-by-dependency.md +151 -0
- package/docs/adr/0027-add-agentic-project-init-and-versioned-engineering-profiles.md +147 -0
- package/docs/adr/0028-make-public-documentation-user-first-and-current.md +65 -0
- package/docs/adr/0029-make-build-a-native-execution-boundary.md +51 -0
- package/docs/adr/0030-unify-product-domain-and-technical-design-interviews.md +240 -0
- package/docs/adr/0031-make-assure-the-success-boundary.md +205 -0
- package/docs/artifacts.md +47 -0
- package/docs/baselines/2026-07-18-p0-lifecycle.json +142 -0
- package/docs/design.md +101 -0
- package/docs/getting-started.md +204 -0
- package/docs/glossary/dev-flow.md +527 -0
- package/docs/lifecycle-contract.md +189 -0
- package/docs/lifecycle-contract.projection.json +931 -0
- package/docs/metrics-protocol.md +113 -0
- package/docs/project-profile-contract.md +157 -0
- package/docs/runbooks/maintainer-release.md +291 -0
- package/docs/target-intake-shape-contract.md +416 -0
- package/package.json +68 -4
- package/schemas/config.schema.json +104 -0
- package/schemas/policy.schema.json +17 -0
- package/schemas/project-init-state.schema.json +159 -0
- package/schemas/project-profile-local.schema.json +53 -0
- package/schemas/project-profile.schema.json +285 -0
- package/schemas/state.schema.json +826 -0
- package/skills/debug-root-cause/SKILL.md +16 -0
- package/skills/design-decisions/SKILL.md +24 -0
- package/skills/dev-flow/SKILL.md +306 -0
- package/skills/dev-flow/agents/openai.yaml +6 -0
- package/skills/discover-change/SKILL.md +31 -0
- package/skills/plan-change/SKILL.md +29 -0
- package/skills/review-change/SKILL.md +21 -0
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
# ADR 0001: Use a Hybrid Portable Workflow
|
|
2
|
+
|
|
3
|
+
## Status
|
|
4
|
+
|
|
5
|
+
Accepted.
|
|
6
|
+
|
|
7
|
+
## Context
|
|
8
|
+
|
|
9
|
+
The workflow must replace Superpowers for Codex while remaining portable across repositories and avoiding parallel write conflicts. Prompt-only workflows do not provide reliable state or policy enforcement. A monolithic command-line state machine would be too rigid for varied codebases.
|
|
10
|
+
|
|
11
|
+
## Decision
|
|
12
|
+
|
|
13
|
+
Publish `codex-dev-flow` as a Codex plugin with a single public `$dev-flow` skill. Internal skills own reasoning. Versioned scripts own deterministic operations: command parsing, configuration resolution, policy enforcement, state validation, fingerprints, and durable artifact handling.
|
|
14
|
+
|
|
15
|
+
The plugin reads repository guidance first, then inspects manifests, CI, and existing tests. It does not require an `AGENTS.md`, but repository guidance wins when present.
|
|
16
|
+
|
|
17
|
+
## Consequences
|
|
18
|
+
|
|
19
|
+
- One primary workflow avoids duplicated lifecycle rules.
|
|
20
|
+
- Scripts make safety-critical behavior testable and reproducible.
|
|
21
|
+
- Skills remain adaptable to repository-specific architecture.
|
|
22
|
+
- The plugin must keep schemas, scripts, and skill instructions compatible.
|
|
23
|
+
- The plugin cannot infer or grant Git or GitHub permissions.
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# ADR 0002: Share an Invalidable Context Capsule Across Lifecycle Phases
|
|
2
|
+
|
|
3
|
+
## Status
|
|
4
|
+
|
|
5
|
+
Accepted.
|
|
6
|
+
|
|
7
|
+
## Context
|
|
8
|
+
|
|
9
|
+
Discovery, planning, execution, review, and verification can repeatedly inspect and
|
|
10
|
+
restate the same repository facts. This consumes quota and tokens without adding
|
|
11
|
+
assurance. Reusing unconstrained summaries would reduce cost but could also propagate
|
|
12
|
+
stale or unsupported claims across phases.
|
|
13
|
+
|
|
14
|
+
Codex Dev Flow optimizes quota and token consumption subject to two hard constraints:
|
|
15
|
+
reliability and absence of false success.
|
|
16
|
+
|
|
17
|
+
## Decision
|
|
18
|
+
|
|
19
|
+
Each durable Dev Flow task owns one compact Context Capsule. Lifecycle phases consume
|
|
20
|
+
the capsule and enrich it instead of independently reconstructing all prior context.
|
|
21
|
+
|
|
22
|
+
The capsule carries structured references to:
|
|
23
|
+
|
|
24
|
+
- objective, approved scope, mode, and risk;
|
|
25
|
+
- relevant files and symbols;
|
|
26
|
+
- effective constraints and policy;
|
|
27
|
+
- acceptance criteria and validation commands;
|
|
28
|
+
- decisions and evidence already obtained;
|
|
29
|
+
- Git, configuration, policy, and artifact fingerprints.
|
|
30
|
+
|
|
31
|
+
Every material claim records provenance and freshness inputs. Fingerprint changes
|
|
32
|
+
invalidate affected sections, not necessarily the whole capsule. A phase must reinspect
|
|
33
|
+
invalid, missing, or insufficient context before relying on it. It may never interpret
|
|
34
|
+
stale context as proof.
|
|
35
|
+
|
|
36
|
+
The capsule is a compact workflow projection, not a second owner for detailed evidence,
|
|
37
|
+
architecture decisions, or approved scope. Versioned change records and ADRs retain
|
|
38
|
+
those responsibilities.
|
|
39
|
+
|
|
40
|
+
## Consequences
|
|
41
|
+
|
|
42
|
+
- Repeated repository reading and prompt restatement decrease.
|
|
43
|
+
- Resume can reconstruct less conversational context.
|
|
44
|
+
- Phase contracts must declare capsule fields consumed and produced.
|
|
45
|
+
- Capsule schema, provenance, fingerprints, invalidation, and compaction require
|
|
46
|
+
deterministic code and tests.
|
|
47
|
+
- Selective invalidation is more complex than a single opaque summary.
|
|
48
|
+
- Missing assurance data blocks or triggers reinspection; it never permits degraded
|
|
49
|
+
success.
|
|
50
|
+
|
|
51
|
+
## Follow-up Decisions
|
|
52
|
+
|
|
53
|
+
- Define exact schema and storage projection with durable state work in P1.
|
|
54
|
+
- Define invalidation matrix for each fingerprint and phase contract.
|
|
55
|
+
- Define token and cache-hit metrics before dogfooding.
|
|
@@ -0,0 +1,36 @@
|
|
|
1
|
+
# ADR 0004: Scale Assurance Lanes by Applicable Risk
|
|
2
|
+
|
|
3
|
+
## Status
|
|
4
|
+
|
|
5
|
+
Accepted.
|
|
6
|
+
|
|
7
|
+
## Context
|
|
8
|
+
|
|
9
|
+
Repeating full repository discovery during review and running every specialist review
|
|
10
|
+
lane for every change wastes quota. Removing independent review from Critical changes,
|
|
11
|
+
however, would weaken assurance and could hide false success.
|
|
12
|
+
|
|
13
|
+
## Decision
|
|
14
|
+
|
|
15
|
+
ASSURE consumes a compact Diff Capsule derived from approved criteria, applicable risks,
|
|
16
|
+
changed surfaces, dependency impact, and fresh validation evidence.
|
|
17
|
+
|
|
18
|
+
Standard runs one structured review over changed and impacted surfaces. Critical keeps
|
|
19
|
+
independent read-only review, but activates specialist lanes according to applicable
|
|
20
|
+
risk. A lane may be omitted only with a structured, evidence-backed non-applicability
|
|
21
|
+
decision. Ambiguous applicability activates the lane.
|
|
22
|
+
|
|
23
|
+
Findings retain deterministic fingerprints. A resolved finding is not reconsidered
|
|
24
|
+
unless its affected code or supporting assumptions changed. Verification reuses fresh
|
|
25
|
+
evidence and reruns missing, invalidated, or globally mandatory checks.
|
|
26
|
+
|
|
27
|
+
Known deterministic failures block expensive reviewer calls until corrected.
|
|
28
|
+
|
|
29
|
+
## Consequences
|
|
30
|
+
|
|
31
|
+
- Review tokens track changed risk surface more closely than repository size.
|
|
32
|
+
- Critical independence remains while irrelevant specialist work decreases.
|
|
33
|
+
- Risk-to-lane rules and non-applicability evidence become testable policy.
|
|
34
|
+
- Incorrectly classifying a lane as irrelevant is a new risk; ambiguity therefore fails
|
|
35
|
+
toward review, not omission.
|
|
36
|
+
- Diff and evidence invalidation must share fingerprints with Context Capsule state.
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# ADR 0006: Make Intake Adaptive, User-Authoritative, and Token-Efficient
|
|
2
|
+
|
|
3
|
+
## Status
|
|
4
|
+
|
|
5
|
+
Accepted for current runtime. Replacement target is recorded by ADR 0030; ADR 0026 is
|
|
6
|
+
superseded as target design.
|
|
7
|
+
|
|
8
|
+
## Context
|
|
9
|
+
|
|
10
|
+
INTAKE must establish enough shared understanding to classify a request safely before
|
|
11
|
+
SHAPE explores the repository. A fixed questionnaire wastes tokens on already explicit
|
|
12
|
+
requests. Letting the model silently fill gaps is cheaper in the short term but removes
|
|
13
|
+
user control and can create expensive rework later. Moving full technical discovery into
|
|
14
|
+
INTAKE would duplicate SHAPE and spend quota before mode and risk are confirmed.
|
|
15
|
+
|
|
16
|
+
The workflow therefore needs an interview that can be relentless about unresolved
|
|
17
|
+
decisions without becoming exhaustive about irrelevant context.
|
|
18
|
+
|
|
19
|
+
## Decision
|
|
20
|
+
|
|
21
|
+
INTAKE uses a two-level interview boundary:
|
|
22
|
+
|
|
23
|
+
- INTAKE resolves product intent, an observable success signal, determining constraints,
|
|
24
|
+
mode/risk, and only the scope boundaries needed to classify the request.
|
|
25
|
+
- SHAPE resolves repository and implementation uncertainty, then turns the success
|
|
26
|
+
signal into testable acceptance criteria and an executable contract.
|
|
27
|
+
|
|
28
|
+
The interview is adaptive. A sufficiently explicit request may require zero questions.
|
|
29
|
+
Every generated question must identify the unresolved decision and the lifecycle outputs
|
|
30
|
+
its answer can change. Questions without determining impact are forbidden.
|
|
31
|
+
|
|
32
|
+
Questions are generated lazily, one at a time. Each question has two or three mutually
|
|
33
|
+
exclusive choices, one clearly justified recommendation, and a free-form answer path.
|
|
34
|
+
The user must answer or cancel. Silence, “decide for me”, or a request to advance cannot
|
|
35
|
+
bypass an unresolved question. Recommendations reduce interaction cost but never become
|
|
36
|
+
decisions without an explicit answer.
|
|
37
|
+
|
|
38
|
+
The model owns semantic sufficiency and question wording. Deterministic runtime checks
|
|
39
|
+
own structure, active-question uniqueness, required answers, contradictions,
|
|
40
|
+
deduplication, redaction, and the exit guard. One model evaluation per user response
|
|
41
|
+
integrates the answer, resolves every explicit decision it contains, evaluates
|
|
42
|
+
completeness, and emits at most the next question. A second model is not called merely
|
|
43
|
+
to verify the first.
|
|
44
|
+
|
|
45
|
+
INTAKE inspects only cheap entry context: repository guidance, Dev Flow configuration
|
|
46
|
+
and policy, primary manifests, workspace/Git identity, and explicitly cited documents.
|
|
47
|
+
Deeper code, dependency, test, and architecture exploration belongs to SHAPE.
|
|
48
|
+
|
|
49
|
+
INTAKE persists a compact Intake Brief and one active Decision Question, never the full
|
|
50
|
+
conversation. SHAPE consumes this structured projection and source references rather
|
|
51
|
+
than the interview transcript. If a later phase discovers missing objective, success,
|
|
52
|
+
determining-constraint, or mode/risk context, the workflow returns to INTAKE and
|
|
53
|
+
selectively invalidates dependent outputs.
|
|
54
|
+
|
|
55
|
+
Standard and Critical move automatically to SHAPE when semantic and deterministic
|
|
56
|
+
completeness checks pass. An explicit “move on” request is refused while context is
|
|
57
|
+
incomplete. Answering INTAKE questions is not implementation Approval; BUILD remains
|
|
58
|
+
guarded by the canonical GATE contract. Quick may bypass SHAPE only when its existing
|
|
59
|
+
bounded-scope guard also proves the request is already executable and verifiable without
|
|
60
|
+
a design decision.
|
|
61
|
+
|
|
62
|
+
After five distinct questions, INTAKE presents a compact checkpoint of resolved and
|
|
63
|
+
remaining decisions. This is a control point, not a hard question limit.
|
|
64
|
+
|
|
65
|
+
## Consequences
|
|
66
|
+
|
|
67
|
+
- Clear requests pay near-zero interview overhead.
|
|
68
|
+
- Ambiguous requests spend tokens only on decisions capable of changing outcome, risk,
|
|
69
|
+
mode, or downstream work.
|
|
70
|
+
- The user remains the authority for every human decision.
|
|
71
|
+
- Resume can re-present the exact active question without replaying the conversation.
|
|
72
|
+
- SHAPE receives smaller, more stable context and can detect ownership violations.
|
|
73
|
+
- A hybrid question engine and new persisted fields require schema, state-machine,
|
|
74
|
+
migration, validation, and crash/resume tests.
|
|
75
|
+
- Model sufficiency can still be wrong; deterministic guards, downstream reopening, and
|
|
76
|
+
benchmark feedback expose that risk without a routine second model call.
|
|
@@ -0,0 +1,82 @@
|
|
|
1
|
+
# ADR 0007: Collect Opt-In Local Benchmark Feedback
|
|
2
|
+
|
|
3
|
+
## Status
|
|
4
|
+
|
|
5
|
+
Accepted.
|
|
6
|
+
|
|
7
|
+
## Context
|
|
8
|
+
|
|
9
|
+
Codex Dev Flow is valuable only when its reliability and avoided rework justify its
|
|
10
|
+
orchestration cost. Running equivalent tasks with multiple models or comparing Dev Flow
|
|
11
|
+
against an additional vanilla Codex execution would multiply token usage and defeat that
|
|
12
|
+
goal. Automatic production telemetry would also create privacy, consent, and operational
|
|
13
|
+
costs disproportionate to early local dogfooding.
|
|
14
|
+
|
|
15
|
+
The maintainer still needs structured evidence about user satisfaction, contract
|
|
16
|
+
conformance, perceived plugin value, and common deviations.
|
|
17
|
+
|
|
18
|
+
## Decision
|
|
19
|
+
|
|
20
|
+
New runs may opt into local benchmark feedback with the public invocation flag:
|
|
21
|
+
|
|
22
|
+
```text
|
|
23
|
+
$dev-flow --benchmark <objective>
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
The flag is disabled by default, accepted only when creating a task, persisted with that
|
|
27
|
+
task, retained automatically across resume, and exposed by status. It never triggers a
|
|
28
|
+
second task, another model, an A/B comparison, or an upload.
|
|
29
|
+
|
|
30
|
+
After `finished` or `failed`, Dev Flow offers three deterministic
|
|
31
|
+
questions without a model call:
|
|
32
|
+
|
|
33
|
+
1. satisfaction from 1 to 5;
|
|
34
|
+
2. conformance to the expressed need: yes, partially, or no;
|
|
35
|
+
3. Dev Flow value: useful, neutral, or counterproductive.
|
|
36
|
+
|
|
37
|
+
Each question permits an optional free-form clarification, followed by an optional
|
|
38
|
+
general comment. The lifecycle result is already final: feedback cannot change task
|
|
39
|
+
status or evidence. No questionnaire is offered for `blocked` or `cancelled`.
|
|
40
|
+
|
|
41
|
+
A benchmark record is written atomically only after all three required answers exist.
|
|
42
|
+
Closing the session before completion creates no benchmark record and contributes no
|
|
43
|
+
benchmark data.
|
|
44
|
+
|
|
45
|
+
Deterministic deviation checks compare:
|
|
46
|
+
|
|
47
|
+
- Quick and Plan against current canonical Change Contract and Plan Approval Receipt.
|
|
48
|
+
|
|
49
|
+
They record structured facts such as unexpected file additions/deletions, unplanned
|
|
50
|
+
surfaces, uncovered criteria, skipped or failed checks, and unapproved scope changes. A
|
|
51
|
+
new file is a deviation only when the authoritative contract did not
|
|
52
|
+
permit it.
|
|
53
|
+
|
|
54
|
+
Completed records live under:
|
|
55
|
+
|
|
56
|
+
```text
|
|
57
|
+
.codex/benchmarks/<task-id>.json
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
The directory is gitignored and excluded from product-diff analysis. Records contain
|
|
61
|
+
allowlisted metadata, questionnaire answers, deterministic deviation signals, task and
|
|
62
|
+
artifact references, plugin/policy versions, duration, and token/call metrics when
|
|
63
|
+
available. They contain no prompt, transcript, code, diff, secret, or automatic upload.
|
|
64
|
+
|
|
65
|
+
The deterministic command below summarizes completed records without a model:
|
|
66
|
+
|
|
67
|
+
```bash
|
|
68
|
+
npx codex-dev-flow benchmark summary
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
## Consequences
|
|
72
|
+
|
|
73
|
+
- Maintainer can dogfood one real task and retain comparable feedback without paying for
|
|
74
|
+
duplicate executions.
|
|
75
|
+
- Non-response is intentionally absent rather than treated as negative or positive
|
|
76
|
+
evidence.
|
|
77
|
+
- Public flag and record format become compatibility surfaces requiring parser, schema,
|
|
78
|
+
state, privacy, and migration tests.
|
|
79
|
+
- User comments remain local but still require length bounds and deterministic secret
|
|
80
|
+
redaction before persistence.
|
|
81
|
+
- Benchmark collection stays outside the five lifecycle phases and cannot weaken truthful
|
|
82
|
+
terminality.
|
|
@@ -0,0 +1,121 @@
|
|
|
1
|
+
# ADR 0008: Automate Maintainer Releases with an Interactive Bun Workflow
|
|
2
|
+
|
|
3
|
+
## Status
|
|
4
|
+
|
|
5
|
+
Accepted.
|
|
6
|
+
|
|
7
|
+
## Context
|
|
8
|
+
|
|
9
|
+
Releasing Codex Dev Flow currently requires the maintainer to remember unrelated Git,
|
|
10
|
+
plugin-cache, npm-registry, and Codex marketplace commands. Some commands depend on
|
|
11
|
+
machine-specific absolute paths, and the Codex plugin and npm package have different
|
|
12
|
+
publication mechanisms. This makes an otherwise routine release error-prone and leaves
|
|
13
|
+
partial-release recovery implicit.
|
|
14
|
+
|
|
15
|
+
The maintainer needs a self-service release flow with one memorable command, explicit
|
|
16
|
+
authority before irreversible actions, deterministic recommendations, and no model call.
|
|
17
|
+
|
|
18
|
+
## Decision
|
|
19
|
+
|
|
20
|
+
The repository will expose this command from its root:
|
|
21
|
+
|
|
22
|
+
```bash
|
|
23
|
+
bun run release
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
A private root `package.json` delegates to a TypeScript release program in the plugin
|
|
27
|
+
package. The program runs with Bun and uses `@clack/prompts`. V1 guarantees macOS
|
|
28
|
+
support, avoids shell-specific implementation, is expected to work on Linux, and does
|
|
29
|
+
not claim tested Windows support. The repository migrates from `package-lock.json` to a
|
|
30
|
+
single committed `bun.lock`.
|
|
31
|
+
|
|
32
|
+
The first prompt selects one release target:
|
|
33
|
+
|
|
34
|
+
- Codex plugin;
|
|
35
|
+
- npm package;
|
|
36
|
+
- Codex plugin and npm package, selected by default.
|
|
37
|
+
|
|
38
|
+
The next prompt selects publication or read-only preview. Preview performs discovery,
|
|
39
|
+
version recommendation, preflight checks, and summary generation without modifying the
|
|
40
|
+
workspace or any remote system.
|
|
41
|
+
|
|
42
|
+
V1 publishes stable versions only. A deterministic Conventional Commits analysis
|
|
43
|
+
recommends the version increment:
|
|
44
|
+
|
|
45
|
+
- `fix:` recommends patch;
|
|
46
|
+
- `feat:` recommends minor;
|
|
47
|
+
- `!` or `BREAKING CHANGE:` recommends major;
|
|
48
|
+
- unclassified commits recommend patch.
|
|
49
|
+
|
|
50
|
+
The maintainer may choose patch, minor, major, or a greater unpublished custom stable
|
|
51
|
+
version. For a plugin-only release, the semantic base remains unchanged and only the
|
|
52
|
+
Codex cachebuster changes. For an npm-only release, only the package semantic version
|
|
53
|
+
changes. A combined release gives both artifacts the same semantic base and adds the
|
|
54
|
+
cachebuster only to the plugin manifest.
|
|
55
|
+
|
|
56
|
+
Release identities are target-specific:
|
|
57
|
+
|
|
58
|
+
```text
|
|
59
|
+
combined: v0.2.0
|
|
60
|
+
npm only: npm-v0.2.0
|
|
61
|
+
plugin only: plugin-v0.2.0-codex.20260728...
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
The corresponding commit uses `chore(release): <release-identity>`. The Git tag is
|
|
65
|
+
annotated and follows the local Git signing configuration.
|
|
66
|
+
|
|
67
|
+
Before changing files, the program requires a clean `main`, fetches `origin/main`, and
|
|
68
|
+
fast-forwards when safe. It refuses local commits, divergence, conflicts, or another
|
|
69
|
+
branch. It never resets, rebases, or merges automatically.
|
|
70
|
+
|
|
71
|
+
All applicable validation is mandatory and has no force-publish bypass:
|
|
72
|
+
|
|
73
|
+
- frozen Bun dependency installation;
|
|
74
|
+
- typecheck, deterministic build, and tests;
|
|
75
|
+
- plugin manifest validation;
|
|
76
|
+
- npm package-content dry run;
|
|
77
|
+
- target-version availability;
|
|
78
|
+
- GitHub and npm authentication;
|
|
79
|
+
- expected clean or allowlisted generated-file state.
|
|
80
|
+
|
|
81
|
+
The release commit may contain only the plugin manifest, affected package metadata,
|
|
82
|
+
`bun.lock` when required, and the deterministic distribution bundle. Any other generated
|
|
83
|
+
change blocks the release.
|
|
84
|
+
|
|
85
|
+
Authentication is checked before mutation. The program never reads, prompts for, logs,
|
|
86
|
+
or stores credentials. Missing authentication produces an exact corrective command and
|
|
87
|
+
requires the maintainer to rerun the release. OTP handling remains inside `bun publish`.
|
|
88
|
+
|
|
89
|
+
The program generates GitHub release notes before publication, displays a non-editable
|
|
90
|
+
preview, and lets the maintainer accept or abort. The final Clack confirmation defaults
|
|
91
|
+
to `No` and summarizes target, versions, tag, files, checks, publication order, and
|
|
92
|
+
planned local installation.
|
|
93
|
+
|
|
94
|
+
Before the first successful push, cancellation, `Ctrl+C`, or failure restores the exact
|
|
95
|
+
original bytes and removes only local commit or tag objects created by the program. Once
|
|
96
|
+
an external publication succeeds, the program never attempts automatic rollback.
|
|
97
|
+
|
|
98
|
+
Publication order is:
|
|
99
|
+
|
|
100
|
+
1. push the release commit and tag;
|
|
101
|
+
2. publish the npm package with `bun publish` when applicable;
|
|
102
|
+
3. create the GitHub Release from the previewed notes;
|
|
103
|
+
4. optionally refresh marketplace `acrazie`, reinstall `codex-dev-flow`, and remind the
|
|
104
|
+
maintainer to start a new Codex thread.
|
|
105
|
+
|
|
106
|
+
Git and registry state form the recovery record. A later `bun run release` detects a
|
|
107
|
+
partial release and offers the single missing idempotent action instead of creating a
|
|
108
|
+
new version. No local state file or secret is required.
|
|
109
|
+
|
|
110
|
+
## Consequences
|
|
111
|
+
|
|
112
|
+
- Routine release operation becomes one short command with maintainer-controlled
|
|
113
|
+
decisions and no model dependency.
|
|
114
|
+
- Plugin and npm releases may remain independent without ambiguous Git tags.
|
|
115
|
+
- Strict preflight and generated-file allowlists reduce accidental publication.
|
|
116
|
+
- Git, npm, and GitHub cannot form one atomic transaction, so explicit partial-release
|
|
117
|
+
detection and resume behavior are required.
|
|
118
|
+
- Root release orchestration, Bun lockfile migration, subprocess isolation, cancellation,
|
|
119
|
+
and recovery paths require dedicated deterministic tests.
|
|
120
|
+
- V1 intentionally excludes prereleases, editable release notes, Windows guarantees, and
|
|
121
|
+
validation bypasses.
|
|
@@ -0,0 +1,200 @@
|
|
|
1
|
+
# ADR 0009: Separate Intake Decisions from Shape Discovery
|
|
2
|
+
|
|
3
|
+
## Status
|
|
4
|
+
|
|
5
|
+
Accepted for current runtime. Replacement target is recorded by ADR 0030; ADR 0026 is
|
|
6
|
+
superseded as target design.
|
|
7
|
+
|
|
8
|
+
## Context
|
|
9
|
+
|
|
10
|
+
Adaptive INTAKE now owns product intent, observable success, determining constraints,
|
|
11
|
+
and risk before SHAPE performs deeper repository discovery. Without a stronger
|
|
12
|
+
downstream boundary, SHAPE can duplicate the interview, silently choose among outcomes
|
|
13
|
+
that need human authority, or send an unstructured ambiguity back to INTAKE and force
|
|
14
|
+
the same technical investigation to run twice.
|
|
15
|
+
|
|
16
|
+
The boundary cannot depend only on whether a choice first appears technical. Repository
|
|
17
|
+
discovery can reveal alternatives whose effects change user experience, cost, delivery
|
|
18
|
+
time, scope, or the observable definition of success.
|
|
19
|
+
|
|
20
|
+
## Decision
|
|
21
|
+
|
|
22
|
+
Phase ownership follows decision impact:
|
|
23
|
+
|
|
24
|
+
- INTAKE owns every unresolved choice that can change product intent, user experience,
|
|
25
|
+
cost, delivery time, scope, determining constraints, risk, or observable success.
|
|
26
|
+
- SHAPE owns repository discovery and purely technical uncertainty that does not change
|
|
27
|
+
those product outcomes.
|
|
28
|
+
- SHAPE never silently resolves a product-impacting technical choice. It suspends
|
|
29
|
+
shaping and requests an explicit INTAKE decision.
|
|
30
|
+
- SHAPE does not repeat an INTAKE decision unless fresh evidence invalidates it.
|
|
31
|
+
|
|
32
|
+
The SHAPE-to-INTAKE return uses a rich, structured Decision Escalation rather than a
|
|
33
|
+
plain missing-context signal. It contains:
|
|
34
|
+
|
|
35
|
+
- a stable decision key;
|
|
36
|
+
- technical evidence and provenance;
|
|
37
|
+
- the product impact;
|
|
38
|
+
- two or three feasible, mutually exclusive options;
|
|
39
|
+
- an evidence-backed recommendation;
|
|
40
|
+
- the SHAPE sections that may depend on the answer.
|
|
41
|
+
|
|
42
|
+
INTAKE validates this escalation and presents the explicit Decision Question. It does
|
|
43
|
+
not repeat the underlying repository discovery merely to reconstruct options. The user
|
|
44
|
+
remains the authority for the answer; a recommendation is not a decision.
|
|
45
|
+
|
|
46
|
+
SHAPE invalidation is always selective. Opening an escalation suspends
|
|
47
|
+
`contract_complete` and marks only the declared dependent sections as unreliable.
|
|
48
|
+
Fresh, independent discovery evidence and contract sections remain reusable. After the
|
|
49
|
+
INTAKE answer, the workflow invalidates the transitive dependency closure of the
|
|
50
|
+
decision and resumes targeted discovery or planning.
|
|
51
|
+
|
|
52
|
+
SHAPE is never reset wholesale. If dependency impact is ambiguous, the workflow widens
|
|
53
|
+
targeted discovery to resolve that ambiguity or blocks with the unresolved dependency;
|
|
54
|
+
it does not discard unrelated evidence. This rule prevents already-paid repository
|
|
55
|
+
reasoning from being repeated solely because one product decision changed.
|
|
56
|
+
|
|
57
|
+
Selective invalidation uses an explicit persisted dependency graph. Decisions, source
|
|
58
|
+
fingerprints, discovery evidence, and contract sections have stable identities. Every
|
|
59
|
+
derived discovery item and contract section declares `dependsOn` references to the
|
|
60
|
+
items required for its validity. The runtime validates that references exist, rejects
|
|
61
|
+
cycles, and computes the transitive dependency closure when a root changes.
|
|
62
|
+
|
|
63
|
+
The model proposes dependency links while shaping because it owns their semantic
|
|
64
|
+
meaning. The deterministic runtime owns structural validation and invalidation. It
|
|
65
|
+
never reconstructs the graph by asking a model to reread all SHAPE output during
|
|
66
|
+
resume.
|
|
67
|
+
|
|
68
|
+
SHAPE discovery is adaptive and targeted, not a fixed repository checklist. It starts
|
|
69
|
+
from the Intake Brief, structured decisions, transferred technical unknowns, and fresh
|
|
70
|
+
reusable evidence. Each discovery target has a stable key, a precise technical question,
|
|
71
|
+
its possible material impact, candidate sources, a result with provenance, and explicit
|
|
72
|
+
dependencies.
|
|
73
|
+
|
|
74
|
+
Discovery widens only when fresh evidence exposes an uncertainty capable of changing
|
|
75
|
+
scope, risk, acceptance criteria, implementation plan, validation, or the preparation
|
|
76
|
+
profile recommendation. SHAPE stops
|
|
77
|
+
exploring when no material technical uncertainty remains and every required contract
|
|
78
|
+
section has sufficient fresh supporting evidence.
|
|
79
|
+
|
|
80
|
+
SHAPE retains two strictly separated roles over one shared structured state:
|
|
81
|
+
|
|
82
|
+
- Discovery may read the repository and resolve Discovery Targets into fresh,
|
|
83
|
+
provenance-bound evidence. It does not author the canonical implementation plan.
|
|
84
|
+
- Planning may project the canonical contract from fresh evidence. It does not read the
|
|
85
|
+
repository or silently supplement missing evidence.
|
|
86
|
+
|
|
87
|
+
After Discovery proves `discovery_sufficient`, the workflow obtains the explicit Quick
|
|
88
|
+
or Plan preparation choice before Planning starts. This is a SHAPE projection choice,
|
|
89
|
+
not a product decision. Planning uses the selected profile without repeating Discovery.
|
|
90
|
+
|
|
91
|
+
When Planning lacks evidence, it creates a new Discovery Target and yields to Discovery.
|
|
92
|
+
After that target is resolved, Planning resumes from shared state. The transfer is
|
|
93
|
+
structured and dependency-linked, never a prose summary that causes either role to
|
|
94
|
+
repeat the other's work.
|
|
95
|
+
|
|
96
|
+
The nominal SHAPE path uses one Discovery reasoning session, one profile choice, and one
|
|
97
|
+
Planning reasoning session. One Discovery session may resolve multiple related targets.
|
|
98
|
+
There is no routine semantic re-review call: deterministic runtime verifies structure,
|
|
99
|
+
freshness, coverage, and dependencies. Additional reasoning is allowed only for a new
|
|
100
|
+
material Discovery Target or an invalidated evidence closure, and receives only the
|
|
101
|
+
relevant state subset. Planning never rereads repository sources.
|
|
102
|
+
|
|
103
|
+
SHAPE uses configurable soft budgets for reasoning calls, estimated tokens, and context.
|
|
104
|
+
Defaults must come from benchmark evidence rather than arbitrary hard-coded counts. When
|
|
105
|
+
a budget approaches exhaustion, the workflow persists all current evidence and displays
|
|
106
|
+
a compact checkpoint offering explicit continuation, scope reduction through INTAKE, or
|
|
107
|
+
cancellation. It never discards SHAPE work or silently spends beyond the budget.
|
|
108
|
+
|
|
109
|
+
`ShapeState` is the aggregate root for `discoveryTargets`, `evidence`,
|
|
110
|
+
`sourceFingerprints`, `decisionEscalations`, `preparationProfile`, `contract`,
|
|
111
|
+
`dependencyGraph`, and semantic `evaluatorReceipts`. Its primary progression is:
|
|
112
|
+
|
|
113
|
+
`discovering -> awaiting_profile_choice -> planning -> awaiting_approval`
|
|
114
|
+
|
|
115
|
+
Planning returns to `discovering` only through a material Discovery Target. Discovery or
|
|
116
|
+
Planning enters `awaiting_intake_decision` only through a Decision Escalation. Required
|
|
117
|
+
events include `discovery.target_recorded`, `discovery.evidence_recorded`,
|
|
118
|
+
`discovery.completed`, `profile.choice_requested`, `profile.choice_recorded`,
|
|
119
|
+
`contract.projected`, and `evidence.invalidated`.
|
|
120
|
+
|
|
121
|
+
For the Plan profile, Planning first completes the canonical contract and the parent
|
|
122
|
+
workflow renders a Markdown projection marked `Proposed`. GATE presents that projection
|
|
123
|
+
with the canonical digest. Approval binds logical contract content; changing projection
|
|
124
|
+
status to `Approved` does not alter that content.
|
|
125
|
+
|
|
126
|
+
The projection has its own fingerprint. A manual file edit never silently overwrites
|
|
127
|
+
canonical state and is never silently discarded. It returns to Planning with an
|
|
128
|
+
explicit reconciliation choice: integrate the edit into the canonical contract or
|
|
129
|
+
regenerate the projection. Any integrated material change requires a new Gate digest
|
|
130
|
+
and Approval.
|
|
131
|
+
|
|
132
|
+
SHAPE completion uses two hybrid proofs:
|
|
133
|
+
|
|
134
|
+
- `discovery_sufficient`: the Discovery model declares that no material technical
|
|
135
|
+
uncertainty remains; deterministic checks require every target to be resolved or
|
|
136
|
+
escalated, every supporting item to be fresh and provenance-bound, and no relevant
|
|
137
|
+
invalidation to remain open.
|
|
138
|
+
- `contract_complete`: the Planning model declares the contract implementable;
|
|
139
|
+
deterministic checks require the profile- and policy-applicable contract sections,
|
|
140
|
+
valid acyclic
|
|
141
|
+
dependencies, criterion-to-task coverage, task-to-criterion coverage, and a
|
|
142
|
+
verification method for every criterion. Risk or policy additionally requires threat
|
|
143
|
+
and rollback records when applicable.
|
|
144
|
+
|
|
145
|
+
Exiting SHAPE requires both proofs. A second model does not re-evaluate another model's
|
|
146
|
+
semantic judgment merely for confirmation; deterministic runtime owns structural
|
|
147
|
+
verification.
|
|
148
|
+
|
|
149
|
+
A Decision Escalation moves the task to a distinct `awaiting_intake_decision` state
|
|
150
|
+
rather than overloading `awaiting_mode_confirmation`. The state persists the escalation
|
|
151
|
+
identity, active Decision Question, prior SHAPE substate, affected dependency roots, and
|
|
152
|
+
freshness references for retained evidence.
|
|
153
|
+
|
|
154
|
+
After an explicit answer, INTAKE integrates the decision and recomputes classification
|
|
155
|
+
or routing only when their declared dependencies are affected. Runtime invalidates the
|
|
156
|
+
decision's dependency closure. The task resumes `discovering` when required evidence is
|
|
157
|
+
missing; otherwise it resumes `planning`. It does not rerun initial INTAKE.
|
|
158
|
+
|
|
159
|
+
Freshness is tracked per evidence source rather than through one repository-wide Git
|
|
160
|
+
invalidation. Tracked files use blob identities; untracked files and configuration use
|
|
161
|
+
content digests; command evidence binds the command, relevant inputs, and observed
|
|
162
|
+
result; external evidence uses a stable identity plus policy-defined observation time
|
|
163
|
+
or TTL.
|
|
164
|
+
|
|
165
|
+
Resume checks only sources referenced by live evidence. A changed source invalidates
|
|
166
|
+
its transitive dependency closure and creates a targeted refresh Discovery Target.
|
|
167
|
+
Unchanged source evidence is reused without another model read. A broad Git change
|
|
168
|
+
alone is not sufficient reason to invalidate unrelated SHAPE work.
|
|
169
|
+
|
|
170
|
+
## Consequences
|
|
171
|
+
|
|
172
|
+
- INTAKE and SHAPE have an effect-based ownership boundary that also covers choices
|
|
173
|
+
discovered late.
|
|
174
|
+
- Technical evidence crosses the boundary without transferring ownership of the human
|
|
175
|
+
decision to SHAPE.
|
|
176
|
+
- Selective invalidation requires explicit dependency links between decisions,
|
|
177
|
+
discovery evidence, and contract sections.
|
|
178
|
+
- Resume reuses fingerprint-valid independent evidence and spends tokens only on the
|
|
179
|
+
invalidated dependency closure.
|
|
180
|
+
- Persisted dependency metadata grows with the contract, but makes invalidation
|
|
181
|
+
deterministic and auditable.
|
|
182
|
+
- Small changes inspect few surfaces; complex changes widen only through recorded
|
|
183
|
+
material reasons.
|
|
184
|
+
- Discovery and Planning can be tested and routed independently without paying for
|
|
185
|
+
duplicate repository reads.
|
|
186
|
+
- Planning may require multiple resumable passes, but each additional pass must be
|
|
187
|
+
justified by a new or invalidated Discovery Target.
|
|
188
|
+
- The common path pays for two SHAPE reasoning sessions rather than repeated discovery,
|
|
189
|
+
planning, and model-verification passes.
|
|
190
|
+
- Budget exhaustion is resumable and user-authoritative rather than a lossy hard stop.
|
|
191
|
+
- Semantic sufficiency remains model-owned while structural completion and coverage are
|
|
192
|
+
deterministic and testable.
|
|
193
|
+
- The lifecycle gains an internal INTAKE state whose name reflects late product
|
|
194
|
+
decisions rather than mode confirmation.
|
|
195
|
+
- Crash/resume can re-present the exact escalated question and return to the correct
|
|
196
|
+
SHAPE role.
|
|
197
|
+
- Source-granular fingerprints cost more metadata but prevent unrelated Git changes
|
|
198
|
+
from causing expensive rediscovery.
|
|
199
|
+
- Persisted state and schemas need a resumable Decision Escalation linked to the active
|
|
200
|
+
Decision Question.
|