codex-orchestrator 2.0.1 → 2.0.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +22 -0
- package/README.md +12 -9
- package/dist/src/index.d.ts +10 -0
- package/dist/src/index.d.ts.map +1 -1
- package/dist/src/index.js +5 -0
- package/dist/src/index.js.map +1 -1
- package/dist/src/v2/acceptance-proof.d.ts +3 -0
- package/dist/src/v2/acceptance-proof.d.ts.map +1 -1
- package/dist/src/v2/acceptance-proof.js +2 -8
- package/dist/src/v2/acceptance-proof.js.map +1 -1
- package/dist/src/v2/adapters/gh-issue-adapter.d.ts +5 -3
- package/dist/src/v2/adapters/gh-issue-adapter.d.ts.map +1 -1
- package/dist/src/v2/adapters/gh-issue-adapter.js +63 -7
- package/dist/src/v2/adapters/gh-issue-adapter.js.map +1 -1
- package/dist/src/v2/adapters/issues.d.ts +16 -2
- package/dist/src/v2/adapters/issues.d.ts.map +1 -1
- package/dist/src/v2/adapters/issues.js +15 -5
- package/dist/src/v2/adapters/issues.js.map +1 -1
- package/dist/src/v2/adapters/mission-coordinator-lock.d.ts +1 -0
- package/dist/src/v2/adapters/mission-coordinator-lock.d.ts.map +1 -1
- package/dist/src/v2/adapters/mission-coordinator-lock.js +5 -1
- package/dist/src/v2/adapters/mission-coordinator-lock.js.map +1 -1
- package/dist/src/v2/candidate-cli.d.ts +4 -0
- package/dist/src/v2/candidate-cli.d.ts.map +1 -1
- package/dist/src/v2/candidate-cli.js +26 -11
- package/dist/src/v2/candidate-cli.js.map +1 -1
- package/dist/src/v2/cli-contract.d.ts +1 -1
- package/dist/src/v2/cli-contract.d.ts.map +1 -1
- package/dist/src/v2/cli-contract.js +10 -0
- package/dist/src/v2/cli-contract.js.map +1 -1
- package/dist/src/v2/code-review-report.d.ts +66 -0
- package/dist/src/v2/code-review-report.d.ts.map +1 -0
- package/dist/src/v2/code-review-report.js +259 -0
- package/dist/src/v2/code-review-report.js.map +1 -0
- package/dist/src/v2/codex-process.d.ts +8 -1
- package/dist/src/v2/codex-process.d.ts.map +1 -1
- package/dist/src/v2/codex-process.js +11 -0
- package/dist/src/v2/codex-process.js.map +1 -1
- package/dist/src/v2/config.d.ts +2 -1
- package/dist/src/v2/config.d.ts.map +1 -1
- package/dist/src/v2/config.js +8 -3
- package/dist/src/v2/config.js.map +1 -1
- package/dist/src/v2/contained-report-operation.d.ts +100 -0
- package/dist/src/v2/contained-report-operation.d.ts.map +1 -0
- package/dist/src/v2/contained-report-operation.js +200 -0
- package/dist/src/v2/contained-report-operation.js.map +1 -0
- package/dist/src/v2/containment.d.ts +6 -0
- package/dist/src/v2/containment.d.ts.map +1 -1
- package/dist/src/v2/containment.js +40 -1
- package/dist/src/v2/containment.js.map +1 -1
- package/dist/src/v2/direct-delivery.d.ts +101 -0
- package/dist/src/v2/direct-delivery.d.ts.map +1 -0
- package/dist/src/v2/direct-delivery.js +547 -0
- package/dist/src/v2/direct-delivery.js.map +1 -0
- package/dist/src/v2/immutable-workflow-publisher.d.ts +40 -0
- package/dist/src/v2/immutable-workflow-publisher.d.ts.map +1 -0
- package/dist/src/v2/immutable-workflow-publisher.js +218 -0
- package/dist/src/v2/immutable-workflow-publisher.js.map +1 -0
- package/dist/src/v2/implementation-reviewer.d.ts +81 -0
- package/dist/src/v2/implementation-reviewer.d.ts.map +1 -0
- package/dist/src/v2/implementation-reviewer.js +157 -0
- package/dist/src/v2/implementation-reviewer.js.map +1 -0
- package/dist/src/v2/owner-control-lock.d.ts +41 -0
- package/dist/src/v2/owner-control-lock.d.ts.map +1 -0
- package/dist/src/v2/owner-control-lock.js +174 -0
- package/dist/src/v2/owner-control-lock.js.map +1 -0
- package/dist/src/v2/route-continuations.d.ts +32 -0
- package/dist/src/v2/route-continuations.d.ts.map +1 -0
- package/dist/src/v2/route-continuations.js +2 -0
- package/dist/src/v2/route-continuations.js.map +1 -0
- package/dist/src/v2/route-coordinator.d.ts +77 -0
- package/dist/src/v2/route-coordinator.d.ts.map +1 -0
- package/dist/src/v2/route-coordinator.js +370 -0
- package/dist/src/v2/route-coordinator.js.map +1 -0
- package/dist/src/v2/route-decision.d.ts +129 -0
- package/dist/src/v2/route-decision.d.ts.map +1 -0
- package/dist/src/v2/route-decision.js +400 -0
- package/dist/src/v2/route-decision.js.map +1 -0
- package/dist/src/v2/run-issue.d.ts +63 -2
- package/dist/src/v2/run-issue.d.ts.map +1 -1
- package/dist/src/v2/run-issue.js +906 -91
- package/dist/src/v2/run-issue.js.map +1 -1
- package/dist/src/v2/run-store.d.ts +25 -1
- package/dist/src/v2/run-store.d.ts.map +1 -1
- package/dist/src/v2/run-store.js +143 -3
- package/dist/src/v2/run-store.js.map +1 -1
- package/dist/src/v2/runtime-assets.d.ts +15 -13
- package/dist/src/v2/runtime-assets.d.ts.map +1 -1
- package/dist/src/v2/runtime-assets.js +263 -416
- package/dist/src/v2/runtime-assets.js.map +1 -1
- package/dist/src/v2/runtime.d.ts +14 -6
- package/dist/src/v2/runtime.d.ts.map +1 -1
- package/dist/src/v2/runtime.js +478 -56
- package/dist/src/v2/runtime.js.map +1 -1
- package/dist/src/v2/setup-cli.d.ts.map +1 -1
- package/dist/src/v2/setup-cli.js +1 -0
- package/dist/src/v2/setup-cli.js.map +1 -1
- package/dist/src/v2/setup-runtime.d.ts.map +1 -1
- package/dist/src/v2/setup-runtime.js +20 -72
- package/dist/src/v2/setup-runtime.js.map +1 -1
- package/dist/src/v2/setup.d.ts +4 -1
- package/dist/src/v2/setup.d.ts.map +1 -1
- package/dist/src/v2/setup.js +104 -1
- package/dist/src/v2/setup.js.map +1 -1
- package/dist/src/v2/spec-coordinator.d.ts +85 -0
- package/dist/src/v2/spec-coordinator.d.ts.map +1 -0
- package/dist/src/v2/spec-coordinator.js +88 -0
- package/dist/src/v2/spec-coordinator.js.map +1 -0
- package/dist/src/v2/spec-delivery.d.ts +143 -0
- package/dist/src/v2/spec-delivery.d.ts.map +1 -0
- package/dist/src/v2/spec-delivery.js +401 -0
- package/dist/src/v2/spec-delivery.js.map +1 -0
- package/dist/src/v2/triage-route.d.ts +68 -0
- package/dist/src/v2/triage-route.d.ts.map +1 -0
- package/dist/src/v2/triage-route.js +223 -0
- package/dist/src/v2/triage-route.js.map +1 -0
- package/dist/src/v2/waiting-human-coordinator.d.ts +49 -0
- package/dist/src/v2/waiting-human-coordinator.d.ts.map +1 -0
- package/dist/src/v2/waiting-human-coordinator.js +509 -0
- package/dist/src/v2/waiting-human-coordinator.js.map +1 -0
- package/dist/src/v2/waiting-human.d.ts +143 -0
- package/dist/src/v2/waiting-human.d.ts.map +1 -0
- package/dist/src/v2/waiting-human.js +408 -0
- package/dist/src/v2/waiting-human.js.map +1 -0
- package/dist/src/v2/workflow-assets.d.ts +90 -0
- package/dist/src/v2/workflow-assets.d.ts.map +1 -0
- package/dist/src/v2/workflow-assets.js +554 -0
- package/dist/src/v2/workflow-assets.js.map +1 -0
- package/docs/deep-dive.md +15 -8
- package/internal-workflow/docs/agents/artifact-review-loop.md +267 -0
- package/internal-workflow/docs/agents/bug-workflow-routing.md +24 -0
- package/internal-workflow/docs/agents/coding-skill-routing.md +203 -0
- package/internal-workflow/docs/agents/confidence-rubric.md +65 -0
- package/internal-workflow/docs/agents/contract-test-ledger.md +60 -0
- package/internal-workflow/docs/agents/implementation-review-loop.md +302 -0
- package/internal-workflow/docs/agents/review-gates.md +49 -0
- package/internal-workflow/docs/agents/review-protocol.md +170 -0
- package/internal-workflow/docs/agents/tool-usage.md +88 -0
- package/internal-workflow/manifest.json +1 -0
- package/internal-workflow/operations/acceptance-proof/SKILL.md +3 -0
- package/internal-workflow/operations/ambiguity-review/SKILL.md +3 -0
- package/internal-workflow/operations/cleanup-review/SKILL.md +3 -0
- package/internal-workflow/operations/code-review/SKILL.md +3 -0
- package/internal-workflow/operations/implementation/SKILL.md +3 -0
- package/internal-workflow/operations/spec-author/SKILL.md +3 -0
- package/internal-workflow/operations/spec-implementation/SKILL.md +3 -0
- package/internal-workflow/operations/spec-review/SKILL.md +3 -0
- package/internal-workflow/operations/triage/SKILL.md +3 -0
- package/internal-workflow/profiles/analyst_deep.toml +9 -0
- package/internal-workflow/profiles/implementer_deep.toml +9 -0
- package/internal-workflow/profiles/implementer_standard.toml +9 -0
- package/internal-workflow/profiles/proof_agent.toml +8 -0
- package/internal-workflow/profiles/researcher_standard.toml +9 -0
- package/internal-workflow/profiles/reviewer_deep.toml +9 -0
- package/internal-workflow/profiles/reviewer_fast.toml +9 -0
- package/internal-workflow/profiles/reviewer_standard.toml +9 -0
- package/internal-workflow/schemas/ambiguity-review-v1.json +1 -0
- package/internal-workflow/schemas/code-review-v1.json +1 -0
- package/internal-workflow/schemas/implementation-report-v1.json +1 -0
- package/internal-workflow/schemas/proof-report-v1.json +1 -0
- package/internal-workflow/schemas/spec-author-v1.json +1 -0
- package/internal-workflow/schemas/spec-review-v1.json +30 -0
- package/internal-workflow/schemas/triage-route-v1.json +1 -0
- package/internal-workflow/skills/acceptance-proof/agents/openai.yaml +6 -0
- package/internal-workflow/skills/agent-auto/agents/openai.yaml +6 -0
- package/internal-workflow/skills/cleanup-review/SKILL.md +84 -0
- package/internal-workflow/skills/cleanup-review/agents/openai.yaml +6 -0
- package/internal-workflow/skills/code-review/SKILL.md +257 -0
- package/internal-workflow/skills/code-review/agents/openai.yaml +4 -0
- package/internal-workflow/skills/code-review/references/bug-classes.md +56 -0
- package/internal-workflow/skills/code-review/references/framework-lenses.md +34 -0
- package/internal-workflow/skills/code-review/references/targeted-recipes.md +49 -0
- package/internal-workflow/skills/codebase-design/DEEPENING.md +35 -0
- package/internal-workflow/skills/codebase-design/DESIGN-IT-TWICE.md +50 -0
- package/internal-workflow/skills/codebase-design/SKILL.md +82 -0
- package/internal-workflow/skills/codebase-design/agents/openai.yaml +6 -0
- package/internal-workflow/skills/diagnosing-bugs/SKILL.md +138 -0
- package/internal-workflow/skills/diagnosing-bugs/agents/openai.yaml +6 -0
- package/internal-workflow/skills/diagnosing-bugs/scripts/hitl-loop.template.sh +41 -0
- package/internal-workflow/skills/implementation-spec-maker/SKILL.md +93 -0
- package/internal-workflow/skills/implementation-spec-maker/agents/openai.yaml +6 -0
- package/internal-workflow/skills/implementation-spec-maker/references/source-modes.md +31 -0
- package/internal-workflow/skills/implementation-spec-maker/references/spec-template.md +146 -0
- package/internal-workflow/skills/implementation-spec-review/SKILL.md +211 -0
- package/internal-workflow/skills/implementation-spec-review/agents/openai.yaml +6 -0
- package/internal-workflow/skills/research/SKILL.md +107 -0
- package/internal-workflow/skills/research/agents/openai.yaml +6 -0
- package/internal-workflow/skills/small-task-implementer/SKILL.md +97 -0
- package/internal-workflow/skills/small-task-implementer/agents/openai.yaml +6 -0
- package/internal-workflow/skills/spec-implementer/SKILL.md +197 -0
- package/internal-workflow/skills/spec-implementer/agents/openai.yaml +6 -0
- package/internal-workflow/skills/tdd/SKILL.md +59 -0
- package/internal-workflow/skills/tdd/agents/openai.yaml +6 -0
- package/internal-workflow/skills/tdd/interface-design.md +31 -0
- package/internal-workflow/skills/tdd/mocking.md +59 -0
- package/internal-workflow/skills/tdd/refactoring.md +10 -0
- package/internal-workflow/skills/tdd/tests.md +77 -0
- package/internal-workflow/skills/triage/AGENT-BRIEF.md +192 -0
- package/internal-workflow/skills/triage/OUT-OF-SCOPE.md +101 -0
- package/internal-workflow/skills/triage/SKILL.md +134 -0
- package/internal-workflow/skills/triage/agents/openai.yaml +6 -0
- package/internal-workflow/skills/ui-evidence-proof/SKILL.md +123 -0
- package/internal-workflow/skills/ui-evidence-proof/agents/openai.yaml +6 -0
- package/package.json +6 -3
- /package/{internal-skills → internal-workflow/skills}/acceptance-proof/SKILL.md +0 -0
- /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/android.md +0 -0
- /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/browser.md +0 -0
- /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/ios.md +0 -0
- /package/{internal-skills → internal-workflow/skills}/acceptance-proof/tools/android-lease.mjs +0 -0
- /package/{internal-skills → internal-workflow/skills}/acceptance-proof/tools/ios-lease.mjs +0 -0
- /package/{internal-skills → internal-workflow/skills}/agent-auto/SKILL.md +0 -0
|
@@ -0,0 +1,138 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: diagnosing-bugs
|
|
3
|
+
description: Debug hard, flaky, unclear, or performance bugs through reproduce, minimize, hypothesize, instrument, fix, and regression-test. Trigger for nondeterministic failures, unclear breakage, or performance regressions.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Diagnosing Bugs
|
|
7
|
+
|
|
8
|
+
A discipline for hard bugs. Skip phases only when explicitly justified.
|
|
9
|
+
|
|
10
|
+
Routing precedence: use `$CODEX_ORCHESTRATOR_WORKFLOW_ROOT/docs/agents/bug-workflow-routing.md`. This skill owns the feedback loop; after the loop proves the bug, return to the original intent: diagnosis-only output or implementation through `code-debugger`.
|
|
11
|
+
|
|
12
|
+
Use `$CODEX_ORCHESTRATOR_WORKFLOW_ROOT/docs/agents/confidence-rubric.md` when deciding whether a hypothesis, root cause, or fix is high-confidence enough to act on. Low-confidence concerns are questions or verification gaps, not proven causes.
|
|
13
|
+
|
|
14
|
+
When exploring the codebase, read `CONTEXT.md` (if it exists) to get a clear mental model of the relevant modules, and check ADRs in the area you're touching.
|
|
15
|
+
|
|
16
|
+
## Phase 1 — Build a feedback loop
|
|
17
|
+
|
|
18
|
+
**This is the skill.** Everything else is mechanical. If you have a **tight** pass/fail signal for the bug — one that goes red on _this_ bug — you will find the cause; bisection, hypothesis-testing, and instrumentation all just consume it. If you don't have one, no amount of staring at code will save you.
|
|
19
|
+
|
|
20
|
+
Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**
|
|
21
|
+
|
|
22
|
+
### Ways to construct one — try them in roughly this order
|
|
23
|
+
|
|
24
|
+
1. **Failing test** at whatever seam reaches the bug — unit, integration, e2e.
|
|
25
|
+
2. **Curl / HTTP script** against a running dev server.
|
|
26
|
+
3. **CLI invocation** with a fixture input, diffing stdout against a known-good snapshot.
|
|
27
|
+
4. **Headless browser script** (Playwright / Puppeteer) — drives the UI, asserts on DOM/console/network.
|
|
28
|
+
5. **Replay a captured trace.** Save a real network request / payload / event log to disk; replay it through the code path in isolation.
|
|
29
|
+
6. **Throwaway harness.** Spin up a minimal subset of the system (one service, mocked deps) that exercises the bug code path with a single function call.
|
|
30
|
+
7. **Property / fuzz loop.** If the bug is "sometimes wrong output", run 1000 random inputs and look for the failure mode.
|
|
31
|
+
8. **Bisection harness.** If the bug appeared between two known states (commit, dataset, version), automate "boot at state X, check, repeat" so you can `git bisect run` it.
|
|
32
|
+
9. **Differential loop.** Run the same input through old-version vs new-version (or two configs) and diff outputs.
|
|
33
|
+
10. **HITL bash script.** Last resort. If a human must click, drive _them_ with `scripts/hitl-loop.template.sh` so the loop is still structured. Captured output feeds back to you.
|
|
34
|
+
|
|
35
|
+
Build the right feedback loop, and the bug is 90% fixed.
|
|
36
|
+
|
|
37
|
+
### Tighten the loop
|
|
38
|
+
|
|
39
|
+
Treat the loop as a product. Once you have _a_ loop, **tighten** it:
|
|
40
|
+
|
|
41
|
+
- Can I make it faster? (Cache setup, skip unrelated init, narrow the test scope.)
|
|
42
|
+
- Can I make the signal sharper? (Assert on the specific symptom, not "didn't crash".)
|
|
43
|
+
- Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, freeze network.)
|
|
44
|
+
|
|
45
|
+
A 30-second flaky loop is barely better than no loop; a 2-second deterministic one is tight — a debugging superpower.
|
|
46
|
+
|
|
47
|
+
### Non-deterministic bugs
|
|
48
|
+
|
|
49
|
+
The goal is not a clean repro but a **higher reproduction rate**. Loop the trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not — keep raising the rate until it's debuggable.
|
|
50
|
+
|
|
51
|
+
### When you genuinely cannot build a loop
|
|
52
|
+
|
|
53
|
+
Stop and say so explicitly. List what you tried. Ask the user for: (a) access to whatever environment reproduces it, (b) a captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. Do **not** proceed to hypothesise without a loop.
|
|
54
|
+
|
|
55
|
+
### Completion criterion — a tight loop that goes red
|
|
56
|
+
|
|
57
|
+
Phase 1 is done when the loop is **tight** and **red-capable**: you can name **one command** — a script path, a test invocation, a curl — that you have **already run at least once** (paste the invocation and its output), and that is:
|
|
58
|
+
|
|
59
|
+
- [ ] **Red-capable** — it drives the actual bug code path and asserts the **user's exact symptom**, so it can go red on this bug and green once fixed. Not "runs without erroring" — it must be able to _catch this specific bug_.
|
|
60
|
+
- [ ] **Deterministic** — same verdict every run (flaky bugs: a pinned, high reproduction rate, per above).
|
|
61
|
+
- [ ] **Fast** — seconds, not minutes.
|
|
62
|
+
- [ ] **Agent-runnable** — you can run it unattended; a human in the loop only via `scripts/hitl-loop.template.sh`.
|
|
63
|
+
|
|
64
|
+
If you catch yourself reading code to build a theory before this command exists, **stop — jumping straight to a hypothesis is the exact failure this skill prevents.** No red-capable command, no Phase 2.
|
|
65
|
+
|
|
66
|
+
## Phase 2 — Reproduce + minimise
|
|
67
|
+
|
|
68
|
+
Run the loop. Watch it go red — the bug appears.
|
|
69
|
+
|
|
70
|
+
Confirm:
|
|
71
|
+
|
|
72
|
+
- [ ] The loop produces the failure mode the **user** described — not a different failure that happens to be nearby. Wrong bug = wrong fix.
|
|
73
|
+
- [ ] The failure is reproducible across multiple runs (or, for non-deterministic bugs, reproducible at a high enough rate to debug against).
|
|
74
|
+
- [ ] You have captured the exact symptom (error message, wrong output, slow timing) so later phases can verify the fix actually addresses it.
|
|
75
|
+
|
|
76
|
+
### Minimise
|
|
77
|
+
|
|
78
|
+
Once it's red, shrink the repro to the **smallest scenario that still goes red**. Cut inputs, callers, config, data, and steps **one at a time**, re-running the loop after each cut — keep only what's load-bearing for the failure.
|
|
79
|
+
|
|
80
|
+
Why bother: a minimal repro shrinks the hypothesis space in Phase 3 (fewer moving parts left to suspect) and becomes the clean regression test in Phase 5.
|
|
81
|
+
|
|
82
|
+
Done when **every remaining element is load-bearing** — removing any one of them makes the loop go green.
|
|
83
|
+
|
|
84
|
+
Do not proceed until you have reproduced **and** minimised.
|
|
85
|
+
|
|
86
|
+
## Phase 3 — Hypothesise
|
|
87
|
+
|
|
88
|
+
Generate **3–5 ranked hypotheses** before testing any of them. Single-hypothesis generation anchors on the first plausible idea.
|
|
89
|
+
|
|
90
|
+
Each hypothesis must be **falsifiable**: state the prediction it makes.
|
|
91
|
+
|
|
92
|
+
> Format: "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."
|
|
93
|
+
|
|
94
|
+
If you cannot state the prediction, the hypothesis is a vibe — discard or sharpen it.
|
|
95
|
+
|
|
96
|
+
**Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly ("we just deployed a change to #3"), or know hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't block on it — proceed with your ranking if the user is AFK.
|
|
97
|
+
|
|
98
|
+
## Phase 4 — Instrument
|
|
99
|
+
|
|
100
|
+
Each probe must map to a specific prediction from Phase 3. **Change one variable at a time.**
|
|
101
|
+
|
|
102
|
+
Tool preference:
|
|
103
|
+
|
|
104
|
+
1. **Debugger / REPL inspection** if the env supports it. One breakpoint beats ten logs.
|
|
105
|
+
2. **Targeted logs** at the boundaries that distinguish hypotheses.
|
|
106
|
+
3. Never "log everything and grep".
|
|
107
|
+
|
|
108
|
+
**Tag every debug log** with a unique prefix, e.g. `[DEBUG-a4f2]`. Cleanup at the end becomes a single grep. Untagged logs survive; tagged logs die.
|
|
109
|
+
|
|
110
|
+
**Perf branch.** For performance regressions, logs are usually wrong. Instead: establish a baseline measurement (timing harness, `performance.now()`, profiler, query plan), then bisect. Measure first, fix second.
|
|
111
|
+
|
|
112
|
+
## Phase 5 — Fix + regression test
|
|
113
|
+
|
|
114
|
+
Write the regression test **before the fix** — but only if there is a **correct seam** for it.
|
|
115
|
+
|
|
116
|
+
A correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site. If the only available seam is too shallow (single-caller test when the bug needs multiple callers, unit test that can't replicate the chain that triggered the bug), a regression test there gives false confidence.
|
|
117
|
+
|
|
118
|
+
**If no correct seam exists, that itself is the finding.** Note it. The codebase architecture is preventing the bug from being locked down. Flag this for the next phase.
|
|
119
|
+
|
|
120
|
+
If a correct seam exists:
|
|
121
|
+
|
|
122
|
+
1. Turn the minimised repro into a failing test at that seam.
|
|
123
|
+
2. Watch it fail.
|
|
124
|
+
3. Apply the fix.
|
|
125
|
+
4. Watch it pass.
|
|
126
|
+
5. Re-run the Phase 1 feedback loop against the original (un-minimised) scenario.
|
|
127
|
+
|
|
128
|
+
## Phase 6 — Cleanup + post-mortem
|
|
129
|
+
|
|
130
|
+
Required before declaring done:
|
|
131
|
+
|
|
132
|
+
- [ ] Original repro no longer reproduces (re-run the Phase 1 loop)
|
|
133
|
+
- [ ] Regression test passes (or absence of seam is documented)
|
|
134
|
+
- [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix)
|
|
135
|
+
- [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug location)
|
|
136
|
+
- [ ] The hypothesis that turned out correct is stated in the commit / PR message — so the next debugger learns
|
|
137
|
+
|
|
138
|
+
**Then ask: what would have prevented this bug?** If the answer involves architectural change (no good test seam, tangled callers, hidden coupling) hand off to the `$improve-codebase-architecture` skill with the specifics. Make the recommendation **after** the fix is in, not before — you have more information now than when you started.
|
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# Human-in-the-loop reproduction loop.
|
|
3
|
+
# Copy this file, edit the steps below, and run it.
|
|
4
|
+
# The agent runs the script; the user follows prompts in their terminal.
|
|
5
|
+
#
|
|
6
|
+
# Usage:
|
|
7
|
+
# bash hitl-loop.template.sh
|
|
8
|
+
#
|
|
9
|
+
# Two helpers:
|
|
10
|
+
# step "<instruction>" → show instruction, wait for Enter
|
|
11
|
+
# capture VAR "<question>" → show question, read response into VAR
|
|
12
|
+
#
|
|
13
|
+
# At the end, captured values are printed as KEY=VALUE for the agent to parse.
|
|
14
|
+
|
|
15
|
+
set -euo pipefail
|
|
16
|
+
|
|
17
|
+
step() {
|
|
18
|
+
printf '\n>>> %s\n' "$1"
|
|
19
|
+
read -r -p " [Enter when done] " _
|
|
20
|
+
}
|
|
21
|
+
|
|
22
|
+
capture() {
|
|
23
|
+
local var="$1" question="$2" answer
|
|
24
|
+
printf '\n>>> %s\n' "$question"
|
|
25
|
+
read -r -p " > " answer
|
|
26
|
+
printf -v "$var" '%s' "$answer"
|
|
27
|
+
}
|
|
28
|
+
|
|
29
|
+
# --- edit below ---------------------------------------------------------
|
|
30
|
+
|
|
31
|
+
step "Open the app at http://localhost:3000 and sign in."
|
|
32
|
+
|
|
33
|
+
capture ERRORED "Click the 'Export' button. Did it throw an error? (y/n)"
|
|
34
|
+
|
|
35
|
+
capture ERROR_MSG "Paste the error message (or 'none'):"
|
|
36
|
+
|
|
37
|
+
# --- edit above ---------------------------------------------------------
|
|
38
|
+
|
|
39
|
+
printf '\n--- Captured ---\n'
|
|
40
|
+
printf 'ERRORED=%s\n' "$ERRORED"
|
|
41
|
+
printf 'ERROR_MSG=%s\n' "$ERROR_MSG"
|
|
@@ -0,0 +1,93 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: "implementation-spec-maker"
|
|
3
|
+
description: "Turn an approved plan, implementation issue, contract-discovery task, or existing spec into the smallest deterministic implementation spec. Use when downstream coding needs an executable checklist with confirmed scope, targets, commands, contracts, validation, and review evidence; do not use for product discovery or implementation."
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Implementation Spec Maker
|
|
7
|
+
|
|
8
|
+
Create or revise an execution-ready specification for a downstream coding agent. Do not implement code or reopen approved product scope.
|
|
9
|
+
|
|
10
|
+
## Core Contract
|
|
11
|
+
|
|
12
|
+
- Treat the supplied plan, issue, discovery task, or existing spec as source authority.
|
|
13
|
+
- Preserve approved scope, exclusions, guardrails, rejected approaches, blockers, validation, and required docs.
|
|
14
|
+
- Confirm execution-critical facts from repository evidence or trusted external contracts. Never invent paths, symbols, commands, fixtures, env vars, schemas, ownership, or API behavior.
|
|
15
|
+
- Produce the smallest spec another agent can execute without guessing. Save a useful `blocked` spec when a material unknown cannot be resolved.
|
|
16
|
+
- Reference approved source content instead of repeating it; write only the missing execution delta.
|
|
17
|
+
|
|
18
|
+
## Preflight
|
|
19
|
+
|
|
20
|
+
1. Read the source authority, applicable repository instructions, and only the evidence needed to confirm targets, commands, contracts, consumers, fixtures, and validation.
|
|
21
|
+
2. Reuse valid Evidence Maps and `$research` artifacts. Refresh only claims invalidated by changed files, versions, dates, contracts, or conflicts.
|
|
22
|
+
3. Read the relevant section of [source modes](references/source-modes.md). Stop or mark the spec blocked when its source-specific requirements are not satisfied.
|
|
23
|
+
4. Classify and record these independent facts:
|
|
24
|
+
- `spec_mode`: `compact | full` — document and coordination density.
|
|
25
|
+
- `implementation_size`: `small | medium | large` — expected delivery shape.
|
|
26
|
+
- `review_profile`: `simple | medium | high` — consequence and uncertainty, resolved through the Artifact Review Module.
|
|
27
|
+
- `expected_repositories`: exact positive integer from approved scope.
|
|
28
|
+
|
|
29
|
+
Do not infer one classification from another. For ticket work, `direct` returns
|
|
30
|
+
to `$tdd`, `compact spec` requests compact mode, and `standard spec` asks the
|
|
31
|
+
maker to choose the smallest deterministic shape. Start a standard ticket in
|
|
32
|
+
compact mode and expand to full only when repository evidence proves a concrete
|
|
33
|
+
ambiguity that compact form cannot remove safely.
|
|
34
|
+
|
|
35
|
+
## Choose The Smallest Shape
|
|
36
|
+
|
|
37
|
+
Default to `compact`, including coherent high-risk or cross-repository work, when ownership, sequencing, stop conditions, and proof fit clearly.
|
|
38
|
+
|
|
39
|
+
Use `full` only when compact form would leave a concrete ambiguity in safety, contract, ownership, sequencing, validation, revision history, or multi-agent integration. Add only the conditional controls that resolve that ambiguity. Risk alone does not require a long document.
|
|
40
|
+
|
|
41
|
+
Keep `execution_model: "single-agent"` unless write scopes are perfectly disjoint and one integrator contract is necessary. Never assign overlapping ownership of files, schemas, generated artifacts, migrations, source-of-truth rules, or shared contracts.
|
|
42
|
+
|
|
43
|
+
## Minimum Solution Gate
|
|
44
|
+
|
|
45
|
+
Before drafting slices:
|
|
46
|
+
|
|
47
|
+
1. Reduce the approved outcome to required behavior, material invariants, and proof.
|
|
48
|
+
2. State the direct `Minimum Solution` through existing owners, public seams, and repository patterns.
|
|
49
|
+
3. Set `Added Complexity: None` unless the minimum solution cannot satisfy a named requirement or evidenced failure path.
|
|
50
|
+
4. For every added mechanism, including a new service, helper, adapter, layer, schema object, transaction, retry policy, job, cache, flag, compatibility path, or coordination boundary, record the exact invariant or failure that requires it and what breaks without it.
|
|
51
|
+
5. Run the deletion challenge: if removing a proposed mechanism still satisfies all approved behavior, invariants, and proof, remove it from the spec.
|
|
52
|
+
|
|
53
|
+
Judge simplicity by the fewest necessary concepts, owners, states, and integration points, not by line or file count. Do not require complexity scores or alternative-solution essays.
|
|
54
|
+
|
|
55
|
+
## Draft The Execution Contract
|
|
56
|
+
|
|
57
|
+
Read [the spec template](references/spec-template.md) before drafting, then remove every unused placeholder and optional block.
|
|
58
|
+
|
|
59
|
+
- Name exact source material, approved scope, exclusions, preconditions, confirmed targets, commands, and observable done criteria.
|
|
60
|
+
- Organize behavior-changing work as narrow vertical slices. Start each slice with the first failing behavior test or exact observable proof, then implementation targets and a slice exit gate.
|
|
61
|
+
- For contract-heavy behavior, use `../../docs/agents/contract-test-ledger.md` and include only material invariants with their first RED test or proof.
|
|
62
|
+
- For UI/app-facing behavior, invoke `$ui-evidence-proof` and embed its task-specific workflow, expected screen state, viewport coverage, fresh artifacts, and criterion-to-artifact mapping in the relevant slice.
|
|
63
|
+
- State exact manual/live proof when automation is not applicable.
|
|
64
|
+
- Name one source of truth when behavior or data can drift. Reuse existing owners and public seams; invoke `$codebase-design` only when ownership or a public seam changes.
|
|
65
|
+
- Add task-specific review checkpoints only when a risky slice becomes stable before later work. Otherwise assign its mandatory lenses, applicable targeted recipes, and concrete bug classes to final review coverage.
|
|
66
|
+
- For medium/high-risk specs, point `Final Handoff Requirements` to the standard `$spec-implementer` Final Risk Handoff and add only task-specific deviations; do not copy its field list.
|
|
67
|
+
- Keep optional cleanup, compatibility logic, feature flags, telemetry, rollout machinery, generic fallbacks, and speculative abstractions out of the spec unless source authority or a proven failure path requires them.
|
|
68
|
+
|
|
69
|
+
## Review And Save
|
|
70
|
+
|
|
71
|
+
1. Save the draft at `docs/implementation-specs/YYYY-MM-DD/HHMM-<slug>.md` with temporary `status: "draft"` and `review_outcome: "Pending"` so review applies to a stable artifact path without presenting it as approved.
|
|
72
|
+
2. Invoke `../../docs/agents/artifact-review-loop.md` with `$implementation-spec-review` as its Adapter. Supply the saved spec, source authority, approved decisions, and evidence; do not restate or replace the Module's topology, defect lifecycle, or convergence rules.
|
|
73
|
+
3. Apply consolidated, scope-preserving repairs in place until the Module returns `Approved`, `Blocked`, or an eligible user-authorized `Waived` outcome.
|
|
74
|
+
4. A preflight-blocked spec may be saved with zero reviews and `review_verdict: "Not run"`. Never fabricate approval or use `Not required`.
|
|
75
|
+
5. Replace temporary lifecycle metadata with the Module outcome, last real Adapter verdict, review coverage/counts, accepted risks, and open stable IDs. Any substantive post-approval edit invalidates approval until reviewed again.
|
|
76
|
+
|
|
77
|
+
## Final Response
|
|
78
|
+
|
|
79
|
+
Return only:
|
|
80
|
+
|
|
81
|
+
```text
|
|
82
|
+
Spec Status: Ready | Blocked
|
|
83
|
+
Saved Path: <path>
|
|
84
|
+
Execution: <single-agent | multi-agent>; <compact | full>; <small | medium | large>; <n> repository/repositories
|
|
85
|
+
Review: <Approved | Blocked | Waived>; <simple | medium | high>; <pass counts and coverage>
|
|
86
|
+
Adapter Verdict: <Approved | Needs Work | Rejected | Not run>
|
|
87
|
+
Verified Defects: <stable IDs or None>
|
|
88
|
+
Accepted Risks: <stable IDs, authority, and reason or None>
|
|
89
|
+
Open Defects: <stable IDs or None>
|
|
90
|
+
Blockers: <unresolved blockers or None>
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
Do not repeat the specification or downstream implementation/signoff procedure in chat.
|
|
@@ -0,0 +1,6 @@
|
|
|
1
|
+
interface:
|
|
2
|
+
display_name: "Implementation Spec Maker"
|
|
3
|
+
short_description: "Create lean deterministic implementation specs"
|
|
4
|
+
default_prompt: "Use $implementation-spec-maker to create the smallest deterministic spec while classifying document mode, implementation size, repository count, and review risk independently."
|
|
5
|
+
policy:
|
|
6
|
+
allow_implicit_invocation: true
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
# Source Modes
|
|
2
|
+
|
|
3
|
+
Read the section matching the active source. Apply the common source-authority and evidence rules from `SKILL.md` in every mode.
|
|
4
|
+
|
|
5
|
+
## Plan-Based
|
|
6
|
+
|
|
7
|
+
- Treat the approved plan as architectural authority.
|
|
8
|
+
- Preserve its scope, vertical-slice boundaries, guardrails, rejected paths, required docs, validation, and blocking assumptions.
|
|
9
|
+
- Block when missing guardrails would let implementation drift or require redesign.
|
|
10
|
+
|
|
11
|
+
## Issue-Based
|
|
12
|
+
|
|
13
|
+
- Treat one issue's acceptance criteria as the execution contract; parent material supplies product context, not sibling-ticket authority.
|
|
14
|
+
- Read comments for changed decisions, blockers, credentials, external contracts, live prerequisites, and rejected approaches.
|
|
15
|
+
- Preserve relevant `Implementation preparation`, `External contracts`, `Verification`, and `Blocked by` content.
|
|
16
|
+
- Use `source_type: "issue"` when no plan exists. Do not block merely because `source_plan` is absent.
|
|
17
|
+
- Block when acceptance criteria are ambiguous, non-verifiable, or contradicted by repository evidence.
|
|
18
|
+
|
|
19
|
+
## Contract Discovery
|
|
20
|
+
|
|
21
|
+
- Specify discovery only: exact sources/tools to inspect, evidence to collect, decision record to update, and issue fields/comments to update.
|
|
22
|
+
- Confirm the API surface, auth/secret source, license or terms constraints, acquisition path, deterministic fixture strategy, live-validation prerequisite, and rejected acquisition paths that matter to later implementation.
|
|
23
|
+
- Do not include downstream implementation slices while material external behavior remains unconfirmed.
|
|
24
|
+
|
|
25
|
+
## Revision
|
|
26
|
+
|
|
27
|
+
- Reconcile the entire existing spec against new authority and current repository evidence.
|
|
28
|
+
- Preserve still-valid completed `[x]` items exactly.
|
|
29
|
+
- Reopen invalid completed items to `[ ]` and add a short `Revision Note:` with evidence.
|
|
30
|
+
- Never silently delete progress or defect history.
|
|
31
|
+
- Mark the spec blocked when completed history or its contract ledger cannot be trusted.
|
|
@@ -0,0 +1,146 @@
|
|
|
1
|
+
# Implementation Spec Template
|
|
2
|
+
|
|
3
|
+
Use the base template for every spec. Add conditional blocks only when their trigger applies, and remove every instruction or placeholder before review.
|
|
4
|
+
|
|
5
|
+
## Base Template
|
|
6
|
+
|
|
7
|
+
```markdown
|
|
8
|
+
---
|
|
9
|
+
title: "<title>"
|
|
10
|
+
created_at: "<ISO timestamp>"
|
|
11
|
+
source_type: "plan | issue | contract-discovery | revised-spec"
|
|
12
|
+
source_plan: "<absolute path or None>"
|
|
13
|
+
source_issues:
|
|
14
|
+
- "<URL/reference or None>"
|
|
15
|
+
status: "draft | ready | blocked"
|
|
16
|
+
execution_model: "single-agent | multi-agent"
|
|
17
|
+
spec_mode: "compact | full"
|
|
18
|
+
implementation_size: "small | medium | large"
|
|
19
|
+
expected_repositories: <positive integer>
|
|
20
|
+
review_profile: "simple | medium | high"
|
|
21
|
+
review_reasons:
|
|
22
|
+
- "<signal: evidence>"
|
|
23
|
+
review_outcome: "Pending"
|
|
24
|
+
review_verdict: "Not run"
|
|
25
|
+
review_coverage: "Not reviewed"
|
|
26
|
+
review_passes: "0"
|
|
27
|
+
---
|
|
28
|
+
|
|
29
|
+
## 1. Execution Context
|
|
30
|
+
- **Goal:** <one observable outcome>
|
|
31
|
+
- **Source Material:** <exact references>
|
|
32
|
+
- **Approved Scope:** <strict allowed work>
|
|
33
|
+
- **Out of Scope:** <explicit exclusions or None>
|
|
34
|
+
- **Minimum Solution:** <direct path through existing owners and public seams>
|
|
35
|
+
- **Added Complexity:** None | <repeat one entry per mechanism: `<mechanism>` — required for `<invariant or evidenced failure>`; without it `<concrete breakage>`>
|
|
36
|
+
- **Primary Risk:** <main correctness or coordination risk>
|
|
37
|
+
|
|
38
|
+
## 2. Preconditions And Evidence
|
|
39
|
+
- **Required Services / Env / Fixtures:** <exact requirements or None>
|
|
40
|
+
- **Blocking Unknowns:** <exact unknowns when blocked, otherwise None>
|
|
41
|
+
- **Confirmed Targets:** <minimal evidence-backed paths and symbols>
|
|
42
|
+
- **Confirmed Commands:** <exact commands>
|
|
43
|
+
- **Protected Paths / Rejected Approaches:** <items or None>
|
|
44
|
+
- **Source of Truth:** <existing owner whenever behavior/data can drift; otherwise omit>
|
|
45
|
+
- **New Boundaries:** <only when ownership or a public seam changes; otherwise omit>
|
|
46
|
+
|
|
47
|
+
## 3. Execution Slices
|
|
48
|
+
|
|
49
|
+
### Slice 1 — <narrow end-to-end behavior>
|
|
50
|
+
- [ ] **Test/Proof First:** <failing behavior test or exact observable proof>
|
|
51
|
+
- [ ] **Target:** `<exact/path:symbol>` — <specific action>
|
|
52
|
+
- [ ] **Validation:** <target-level check>
|
|
53
|
+
- [ ] **Exit Gate:** <command or proof that the slice works end-to-end>
|
|
54
|
+
|
|
55
|
+
<repeat only for independently verifiable behavior slices>
|
|
56
|
+
|
|
57
|
+
## 4. Validation And Done Criteria
|
|
58
|
+
- [ ] **Lint/Format:** <exact command or Not applicable with reason>
|
|
59
|
+
- [ ] **Typecheck/Build:** <exact command or Not applicable with reason>
|
|
60
|
+
- [ ] **Tests:** <exact command or Not applicable with reason>
|
|
61
|
+
- [ ] **Architecture Check:** <exact command or Not applicable with reason>
|
|
62
|
+
- [ ] **Live/Manual Proof:** <exact flow or Not applicable with reason>
|
|
63
|
+
- [ ] **Behavior Proof:** <observable acceptance proof>
|
|
64
|
+
- [ ] **Reconciliation:** every unchecked item is unfinished, blocked with evidence, or intentionally not applicable.
|
|
65
|
+
- [ ] **Final Handoff Requirements:** <medium/high only: standard `$spec-implementer` Final Risk Handoff plus task-specific deviations or None>
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
## Conditional Blocks
|
|
69
|
+
|
|
70
|
+
### Contract Test Ledger
|
|
71
|
+
|
|
72
|
+
Add for contract-heavy behavior using the shared Contract Test Ledger referenced by `SKILL.md`. Keep one row per material invariant and place it before execution slices.
|
|
73
|
+
|
|
74
|
+
### Review Checkpoint And Focus
|
|
75
|
+
|
|
76
|
+
Add a checkpoint only when the risky target becomes stable before later slices. Otherwise put this compact block in final review coverage:
|
|
77
|
+
|
|
78
|
+
```markdown
|
|
79
|
+
## Review Focus
|
|
80
|
+
- **Mandatory Lenses:** <applicable lenses>
|
|
81
|
+
- **Targeted Recipes:** <applicable recipes or None>
|
|
82
|
+
- **Bug Classes:** <concrete failures to hunt>
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
### Risk Controls
|
|
86
|
+
|
|
87
|
+
Add in `full` mode only for applicable ambiguity:
|
|
88
|
+
|
|
89
|
+
```markdown
|
|
90
|
+
## Risk Controls
|
|
91
|
+
- **Source of Truth:** <owner>
|
|
92
|
+
- **Safety / Contract / State Constraints:** <only applicable constraints>
|
|
93
|
+
- **Forbidden Scope:** <tempting but rejected paths>
|
|
94
|
+
- **Review Timing:** <stable early checkpoint or concrete final-review focus>
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
### Write Scope Summary
|
|
98
|
+
|
|
99
|
+
Add for multi-agent work, generated artifacts, broad runtime changes, or when phase targets do not make the write set auditable.
|
|
100
|
+
|
|
101
|
+
```markdown
|
|
102
|
+
## Write Scope Summary
|
|
103
|
+
- `<path>` — <Create | Update | Delete>; <responsibility>
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
### Integrator Coordination Contract
|
|
107
|
+
|
|
108
|
+
Require only when `execution_model: "multi-agent"`:
|
|
109
|
+
|
|
110
|
+
```markdown
|
|
111
|
+
## Integrator Coordination Contract
|
|
112
|
+
| Agent | Exclusive Write Scope | Handoff | Merge Phase |
|
|
113
|
+
| --- | --- | --- | --- |
|
|
114
|
+
| <agent> | <disjoint paths> | <artifact> | <order> |
|
|
115
|
+
|
|
116
|
+
- **Integrator Owner:** <owner>
|
|
117
|
+
- **Forbidden Overlap:** <paths/contracts>
|
|
118
|
+
- **Final Duties:** <integration, validation, reconciliation>
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
### Halt Conditions
|
|
122
|
+
|
|
123
|
+
Add only when the common contradiction/guessing stop rule is insufficient. Use 3–6 task-specific conditions.
|
|
124
|
+
|
|
125
|
+
### Defect Closure Notes
|
|
126
|
+
|
|
127
|
+
Add only when review returns defects:
|
|
128
|
+
|
|
129
|
+
```markdown
|
|
130
|
+
## Defect Closure Notes
|
|
131
|
+
- **Review Summary:** <pass counts and coverage>
|
|
132
|
+
- **Verified Defects:** <stable IDs or None>
|
|
133
|
+
- **Accepted Risks:** <stable IDs, authority, and reason or None>
|
|
134
|
+
- **Open Defects:** <stable IDs or None>
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
## Terminal Review Metadata
|
|
138
|
+
|
|
139
|
+
Replace the temporary frontmatter values after the Artifact Review Module returns a real terminal outcome:
|
|
140
|
+
|
|
141
|
+
```yaml
|
|
142
|
+
review_outcome: "Approved | Blocked | Waived"
|
|
143
|
+
review_verdict: "Approved | Needs Work | Rejected | Not run"
|
|
144
|
+
review_coverage: "<covered lenses or Not reviewed>"
|
|
145
|
+
review_passes: "<total; full/closure/fresh counts>"
|
|
146
|
+
```
|