codex-orchestrator 2.0.3 → 2.0.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (131) hide show
  1. package/CHANGELOG.md +28 -427
  2. package/README.md +135 -37
  3. package/dist/src/index.d.ts +1 -1
  4. package/dist/src/index.d.ts.map +1 -1
  5. package/dist/src/v2/adapters/gh-issue-adapter.d.ts.map +1 -1
  6. package/dist/src/v2/adapters/gh-issue-adapter.js +6 -7
  7. package/dist/src/v2/adapters/gh-issue-adapter.js.map +1 -1
  8. package/dist/src/v2/cli-contract.d.ts +3 -3
  9. package/dist/src/v2/cli-contract.d.ts.map +1 -1
  10. package/dist/src/v2/cli-contract.js +1 -3
  11. package/dist/src/v2/cli-contract.js.map +1 -1
  12. package/dist/src/v2/cli.d.ts +24 -0
  13. package/dist/src/v2/cli.d.ts.map +1 -0
  14. package/dist/src/v2/{candidate-cli.js → cli.js} +18 -29
  15. package/dist/src/v2/cli.js.map +1 -0
  16. package/dist/src/v2/code-review-report.d.ts +1 -1
  17. package/dist/src/v2/code-review-report.d.ts.map +1 -1
  18. package/dist/src/v2/code-review-report.js +2 -2
  19. package/dist/src/v2/code-review-report.js.map +1 -1
  20. package/dist/src/v2/codex-process.d.ts.map +1 -1
  21. package/dist/src/v2/codex-process.js +12 -1
  22. package/dist/src/v2/codex-process.js.map +1 -1
  23. package/dist/src/v2/config.d.ts +2 -2
  24. package/dist/src/v2/config.d.ts.map +1 -1
  25. package/dist/src/v2/config.js.map +1 -1
  26. package/dist/src/v2/contained-report-operation.d.ts +2 -2
  27. package/dist/src/v2/contained-report-operation.d.ts.map +1 -1
  28. package/dist/src/v2/contained-report-operation.js +1 -1
  29. package/dist/src/v2/contained-report-operation.js.map +1 -1
  30. package/dist/src/v2/containment.d.ts +4 -0
  31. package/dist/src/v2/containment.d.ts.map +1 -1
  32. package/dist/src/v2/containment.js +9 -0
  33. package/dist/src/v2/containment.js.map +1 -1
  34. package/dist/src/v2/direct-delivery.d.ts +5 -10
  35. package/dist/src/v2/direct-delivery.d.ts.map +1 -1
  36. package/dist/src/v2/direct-delivery.js +25 -90
  37. package/dist/src/v2/direct-delivery.js.map +1 -1
  38. package/dist/src/v2/proof-report.d.ts.map +1 -1
  39. package/dist/src/v2/proof-report.js +55 -29
  40. package/dist/src/v2/proof-report.js.map +1 -1
  41. package/dist/src/v2/run-issue.d.ts +6 -9
  42. package/dist/src/v2/run-issue.d.ts.map +1 -1
  43. package/dist/src/v2/run-issue.js +98 -43
  44. package/dist/src/v2/run-issue.js.map +1 -1
  45. package/dist/src/v2/run-store.d.ts +4 -4
  46. package/dist/src/v2/run-store.d.ts.map +1 -1
  47. package/dist/src/v2/run-store.js +25 -40
  48. package/dist/src/v2/run-store.js.map +1 -1
  49. package/dist/src/v2/runtime.d.ts +3 -3
  50. package/dist/src/v2/runtime.d.ts.map +1 -1
  51. package/dist/src/v2/runtime.js +125 -47
  52. package/dist/src/v2/runtime.js.map +1 -1
  53. package/dist/src/v2/setup-cli.d.ts.map +1 -1
  54. package/dist/src/v2/setup-cli.js +4 -11
  55. package/dist/src/v2/setup-cli.js.map +1 -1
  56. package/dist/src/v2/setup-runtime.d.ts.map +1 -1
  57. package/dist/src/v2/setup-runtime.js +1 -61
  58. package/dist/src/v2/setup-runtime.js.map +1 -1
  59. package/dist/src/v2/setup-store.d.ts +0 -5
  60. package/dist/src/v2/setup-store.d.ts.map +1 -1
  61. package/dist/src/v2/setup-store.js +3 -106
  62. package/dist/src/v2/setup-store.js.map +1 -1
  63. package/dist/src/v2/setup.d.ts +6 -46
  64. package/dist/src/v2/setup.d.ts.map +1 -1
  65. package/dist/src/v2/setup.js +11 -293
  66. package/dist/src/v2/setup.js.map +1 -1
  67. package/dist/src/v2/workflow-assets.d.ts +19 -11
  68. package/dist/src/v2/workflow-assets.d.ts.map +1 -1
  69. package/dist/src/v2/workflow-assets.js +132 -40
  70. package/dist/src/v2/workflow-assets.js.map +1 -1
  71. package/docs/deep-dive.md +272 -56
  72. package/internal-workflow/docs/agents/bugfix-quality-gate.md +11 -0
  73. package/internal-workflow/docs/agents/coding-skill-routing.md +116 -196
  74. package/internal-workflow/docs/agents/review-gates.md +32 -39
  75. package/internal-workflow/docs/agents/review-protocol.md +75 -147
  76. package/internal-workflow/evals/coding-skill-evals.json +66 -0
  77. package/internal-workflow/manifest.json +1 -1
  78. package/internal-workflow/operations/acceptance-proof/SKILL.md +7 -1
  79. package/internal-workflow/operations/ambiguity-review/SKILL.md +2 -0
  80. package/internal-workflow/operations/code-review/SKILL.md +21 -1
  81. package/internal-workflow/operations/implementation/SKILL.md +22 -1
  82. package/internal-workflow/operations/spec-author/SKILL.md +10 -1
  83. package/internal-workflow/operations/spec-review/SKILL.md +10 -1
  84. package/internal-workflow/operations/triage/SKILL.md +10 -1
  85. package/internal-workflow/schemas/code-review-v1.json +1 -1
  86. package/internal-workflow/schemas/proof-report-v1.json +1 -1
  87. package/internal-workflow/skills/agent-auto/SKILL.md +6 -1
  88. package/internal-workflow/skills/code-debugger/SKILL.md +122 -0
  89. package/internal-workflow/skills/code-debugger/agents/openai.yaml +7 -0
  90. package/internal-workflow/skills/code-review/SKILL.md +33 -11
  91. package/internal-workflow/skills/code-review/references/cleanup-lens.md +52 -0
  92. package/internal-workflow/skills/implementation-spec-maker/SKILL.md +15 -6
  93. package/internal-workflow/skills/implementation-spec-maker/references/spec-template.md +2 -2
  94. package/internal-workflow/skills/implementation-spec-review/SKILL.md +108 -204
  95. package/internal-workflow/skills/implementation-spec-review/evals/evals.json +24 -0
  96. package/internal-workflow/skills/implementation-spec-review/references/review-loop.md +93 -0
  97. package/internal-workflow/skills/small-task-implementer/SKILL.md +15 -8
  98. package/internal-workflow/skills/spec-implementer/SKILL.md +101 -172
  99. package/internal-workflow/skills/spec-implementer/evals/evals.json +30 -0
  100. package/internal-workflow/skills/spec-implementer/references/review-loop.md +94 -0
  101. package/internal-workflow/skills/tdd/SKILL.md +15 -2
  102. package/internal-workflow/skills/tdd/agents/openai.yaml +2 -2
  103. package/package.json +9 -6
  104. package/dist/src/v2/adapters/target-activity-fence.d.ts +0 -23
  105. package/dist/src/v2/adapters/target-activity-fence.d.ts.map +0 -1
  106. package/dist/src/v2/adapters/target-activity-fence.js +0 -249
  107. package/dist/src/v2/adapters/target-activity-fence.js.map +0 -1
  108. package/dist/src/v2/candidate-cli.d.ts +0 -26
  109. package/dist/src/v2/candidate-cli.d.ts.map +0 -1
  110. package/dist/src/v2/candidate-cli.js.map +0 -1
  111. package/dist/src/v2/legacy-cutover.d.ts +0 -52
  112. package/dist/src/v2/legacy-cutover.d.ts.map +0 -1
  113. package/dist/src/v2/legacy-cutover.js +0 -87
  114. package/dist/src/v2/legacy-cutover.js.map +0 -1
  115. package/internal-workflow/docs/agents/artifact-review-loop.md +0 -267
  116. package/internal-workflow/docs/agents/implementation-review-loop.md +0 -302
  117. package/internal-workflow/operations/cleanup-review/SKILL.md +0 -3
  118. package/internal-workflow/operations/spec-implementation/SKILL.md +0 -3
  119. package/internal-workflow/profiles/implementer_deep.toml +0 -9
  120. package/internal-workflow/profiles/researcher_standard.toml +0 -9
  121. package/internal-workflow/profiles/reviewer_fast.toml +0 -9
  122. package/internal-workflow/skills/cleanup-review/SKILL.md +0 -84
  123. package/internal-workflow/skills/cleanup-review/agents/openai.yaml +0 -6
  124. package/internal-workflow/skills/codebase-design/DEEPENING.md +0 -35
  125. package/internal-workflow/skills/codebase-design/DESIGN-IT-TWICE.md +0 -50
  126. package/internal-workflow/skills/codebase-design/SKILL.md +0 -82
  127. package/internal-workflow/skills/codebase-design/agents/openai.yaml +0 -6
  128. package/internal-workflow/skills/research/SKILL.md +0 -107
  129. package/internal-workflow/skills/research/agents/openai.yaml +0 -6
  130. package/internal-workflow/skills/ui-evidence-proof/SKILL.md +0 -123
  131. package/internal-workflow/skills/ui-evidence-proof/agents/openai.yaml +0 -6
@@ -1,203 +1,123 @@
1
1
  # Coding Skill Routing
2
2
 
3
- Global coding skills own reusable engineering process. Local repo skills and docs own repository-specific facts.
4
-
5
- The single personal-skill root is `../../skills` for root
6
- agents and subagents in every repository. Invocation policy lives in each
7
- skill's `agents/openai.yaml`; do not encode a second policy in `SKILL.md`
8
- frontmatter or copy global workflow skills into repositories.
9
-
10
- Use this routing policy before adding more instructions to a global skill.
3
+ This file is the normative global route and ownership policy. Keep repository
4
+ commands, domain facts, credentials, deployment assumptions, and product
5
+ behavior in repository `AGENTS.md`, `CONTEXT.md`, ADRs, or local skills.
11
6
 
12
7
  ## Ownership
13
8
 
14
- Global skills own reusable process:
15
-
16
- - planning, external research, diagnosis, implementation, review, cleanup,
17
- smoke-test, and commit workflows
18
- - reusable gates such as TDD, contract test ledger, confidence handling, evidence standards, and review order
19
- - deterministic fact collection that does not encode product knowledge
20
- - progressive references for framework-agnostic or broadly reusable guidance
21
-
22
- Local repo skills and docs own repository facts:
23
-
24
- - package manager and project commands
25
- - domain language, product behavior, fixtures, and smoke scenarios
26
- - environment variables, credential sources, service topology, and deployment assumptions
27
- - local architecture decisions, ADRs, and repo-specific validation rules
28
-
29
- Do not move those local facts into a global skill.
30
-
31
- ## Execution Modes
32
-
33
- - **Inline:** root performs non-review work. This is the default for small and ordinary authoring or implementation work.
34
- - **Delegate:** use one named custom agent when independent review or deeper isolated analysis is required.
35
- - **Parallel:** use multiple agents only for independent tracks with disjoint responsibilities.
36
-
37
- Automatic routing never authorizes automatic spawning. Spawn only when the user, an invoked skill, or applicable repo instructions authorize delegation. Root owns decisions, integration, user communication, and the critical path.
38
-
39
- Root must never perform review inline. Every review gate launches the reviewer
40
- role selected by the review profile. Because `agents.max_depth = 1`, a review
41
- Adapter executes inline only after it is already inside that assigned reviewer
42
- child; this is child execution, not root self-review.
43
-
44
- ## Platform UI QA Routing
45
-
46
- Load the platform QA skill first; use `$flutter-attach-session` only as the
47
- runtime ownership and reload/restart layer. Detailed safety rules live in
48
- [`tool-usage.md`](tool-usage.md).
49
-
50
- - Android: `$flutter-android-debug` plus
51
- `test-android-apps:android-emulator-qa`.
52
- - iOS: `$flutter-ios-debug`.
53
- - Live, IDE-owned, machine-owned, or ambiguous runtimes are user-owned. UI
54
- inspection alone never authorizes replacement, termination, install, or launch.
55
-
56
- ## Named Agent Profiles
57
-
58
- | Role | Contract |
9
+ - Personal skills live only in `../../skills`.
10
+ - `agents/openai.yaml` owns invocation metadata; `agents/*.toml` owns named-role
11
+ model and effort.
12
+ - A shared rule has one owner. Callers link to it instead of copying its prose.
13
+ - Skill-specific detail belongs in that skill's `references/` and loads only
14
+ when its branch is active.
15
+ - Root owns the user dialogue, decisions, critical path, integration, and final
16
+ handoff. A reviewer child owns independent review.
17
+
18
+ ## Default Implementation Route
19
+
20
+ `medium` is the default for behavior-changing implementation. Use:
21
+
22
+ - `simple` only for a tiny local change with one obvious proof;
23
+ - `medium` for a clear coherent outcome with settled authority, one ownership
24
+ path, and credible affected validation—even across several files, modules,
25
+ API, persistence, or shared state;
26
+ - `high` only when a sensitive mechanism has both a material failure
27
+ consequence and an uncertainty amplifier such as unclear ownership,
28
+ cross-trust effects, non-local recovery, an unproven external contract, or
29
+ proof that cannot isolate the dangerous state.
30
+
31
+ Prefer direct root implementation for `simple` and ordinary `medium` work. Use
32
+ `$implementation-spec-maker` only for a real execution decision or coordination
33
+ gap. Use `$tickets-orchestrator` only for an approved ticket graph, real delivery
34
+ dependencies or disjoint parallel slices, or an explicit orchestration request.
35
+ Do not manufacture PRDs, tickets, specs, agents, or review checkpoints from
36
+ file count or generic risk labels.
37
+
38
+ ## Core Routes
39
+
40
+ | Situation | Route |
59
41
  | --- | --- |
60
- | `explorer_quick` | Mechanical inventory of files, symbols, usages, tests, and registrations |
61
- | `explorer_fast` | Bounded cross-module execution tracing with file:symbol evidence |
62
- | `analyst_deep` | Read-only architecture, causal, contract, and root-cause synthesis |
63
- | `researcher_standard` | Read-only external primary-source research with claim-level citations |
64
- | `reviewer_fast` | Fast independent review for `simple` profiles |
65
- | `reviewer_standard` | Independent review for `medium` profiles |
66
- | `reviewer_deep` | Deep independent review for `high` profiles and security |
67
- | `implementer_standard` | Write-capable worker for one approved bounded ticket slice |
68
- | `implementer_deep` | Write-capable worker for a rare isolated slice with material technical uncertainty |
69
-
70
- The role files in `agents/*.toml` own model, effort, nickname, and instructions.
71
- Skills request exact role names and never override model or effort. Artifact and
72
- implementation review Modules own Full/Closure topology; Approval Packet review
73
- owns its axis split. Analysis routing remains independent of review profile.
74
-
75
- ## Routing Table
76
-
77
- | Skill or workflow | Default mode | Named-agent routing |
78
- | --- | --- | --- |
79
- | `$grilling` | Root owns the interactive session | Optional `explorer_fast` for bounded evidence and `analyst_deep` for proven ambiguity; children never conduct the user dialogue |
80
- | `$codebase-design` | Bounded reference lens for a named Module Interface or Seam; never a standalone scan, mutation, or implementation workflow | None for bounded reference use; Design It Twice runs 3+ isolated `analyst_deep` children in parallel for a selected consequential candidate, with any inline fallback labelled non-independent |
81
- | `$research` | Root prepares the Research Capsule, verifies decision-driving claims, and saves one artifact | One `researcher_standard`; root may use the documented inline fallback when the role is unavailable |
82
- | `$plans-maker` | Explicit-only root-authored Architecture RFC | Profile-selected reviewer topology from `artifact-review-loop.md`; optional `analyst_deep` only for unresolved architecture |
83
- | `$plan-review` | Inline only inside the assigned reviewer child | Root launches the profile-selected reviewer |
84
- | `$implementation-spec-maker` | Root authors inline | Profile-selected reviewer topology from `artifact-review-loop.md`; optional `analyst_deep` only for proven ambiguity |
85
- | `$implementation-spec-review` | Inline only inside the assigned reviewer child | Root launches the profile-selected reviewer |
86
- | `$to-spec` | Root authors a human-readable PRD inline; combined flow keeps it in context until final publication | Standalone reviewed PRD uses `$tickets-breakdown-review` PRD-only mode |
87
- | `$to-tickets` | Root drafts the ticket graph, obtains one approval, and publishes generated children directly in final AFK/HITL states | One direct low-risk ticket skips independent review; other packets use `$tickets-breakdown-review` |
88
- | `$tickets-breakdown-review` | Root prepares and aggregates only when the caller's review gate applies | One profile-selected child covers both axes by default; two `reviewer_deep` children only when both axes are independently high-risk; PRD-only uses one child |
89
- | `$triage` | Root verifies raw incoming issues or configured external PRs and prepares durable briefs | No implementation worker; never post-processes generated `$to-tickets` children |
90
- | `$small-task-implementer` | Always inline after Fit Gate | None by default |
91
- | `$spec-implementer` | Inline for compact and normal specs | Profile-selected reviewers at required checkpoints; parallel workers only for an explicit multi-agent spec |
92
- | `$tdd` | Inline in the active implementation flow | Never spawns its own agent |
93
- | `$code-debugger` | Inline for reproduced, bounded bugs | `analyst_deep` only while causal or contract ambiguity remains unresolved |
94
- | `$bug-root-cause-explainer` | Root coordinates read-only diagnosis | Optional `explorer_fast`; `analyst_deep` for ambiguous causal synthesis |
95
- | `$diagnosing-bugs` | Root owns the feedback loop | Optional `explorer_fast`; no implementation worker |
96
- | `$code-review` | Root prepares and aggregates | One reviewer for `simple`/`medium`; two disjoint `reviewer_deep` tracks for `high` |
97
- | `$cleanup-review` | Root prepares and integrates | One profile-selected reviewer child |
98
- | `$security-best-practices` | Root coordinates explicit security review | One `reviewer_deep` |
99
- | `$improve-codebase-architecture` | Inline for bounded analysis | `analyst_deep` only for broad or ambiguous architecture |
100
- | `$commit` | Inline | No agent unless another policy already requires review |
101
- | `$tickets-orchestrator` | Root owns ticket graph, user decisions, integration, and delivery | One ready ticket stays root-owned; launch at most two independent disjoint implementers; prepare later tickets only after blockers settle; reuse issue authority or route through maker/spec Modules |
102
- | `$smoke-test-orchestrator` | Inline unless the scenario explicitly requires workers | Follow the scenario's disjoint ownership contract |
103
-
104
- ## Artifact Review Loop
105
-
106
- Plan and implementation-spec authoring use
107
- [`artifact-review-loop.md`](artifact-review-loop.md) as the single review Module.
108
- It owns artifact risk, scope conservation, reviewer topology, and outcome
109
- mapping while applying [`review-protocol.md`](review-protocol.md) for common
110
- review mechanics. Maker skills supply authority and repairs; review Adapters
111
- supply artifact-specific lenses and output.
112
-
113
- ## Implementation Review Loop
114
-
115
- Approved implementation-spec execution uses
116
- [`implementation-review-loop.md`](implementation-review-loop.md) as the single
117
- review Module. It owns approved-spec authority, durable state, whole-spec
118
- topology, validation reuse, gate ordering, final coverage, and audit epochs while
119
- applying [`review-protocol.md`](review-protocol.md). `spec-implementer`, review
120
- Adapters, repo policy, and specs may define lenses and applicability but never
121
- another review protocol or retry loop.
122
-
123
- `$tickets-orchestrator` is an outer delivery caller. Each ticket selects
124
- `direct`, `compact spec`, or `standard spec`. Direct tickets keep the `$tdd` and
125
- repo-review flow without manufacturing Implementation Review State. For each
126
- compact/standard ticket, root invokes `$implementation-spec-maker` and then
127
- `$spec-implementer`; the orchestrator must not reproduce either review Module
128
- or combine approved tickets into a wave-level implementation spec.
129
-
130
- ## External Research Preflight
131
-
132
- Read local evidence before searching externally. Keep one narrow documentation
133
- lookup inline with the owning specialized docs skill or tool unless the user
134
- explicitly requests delegation or a durable artifact. Invoke `$research` for
135
- either explicit request, or when a material coding decision requires
136
- multi-source comparison, freshness checking, or external contract synthesis.
137
-
138
- `$research` authorizes one `researcher_standard` child. Root supplies a bounded
139
- Research Capsule, verifies every claim that drives architecture, scope,
140
- implementation, security, cost, or compatibility, and saves one artifact under
141
- the repository convention or `docs/research/YYYY-MM-DD/HHMM-<slug>.md`.
142
- Research is evidence, not implementation authority. Downstream plans, PRDs,
143
- tickets, and specs cite the artifact; behavior-changing work still follows the
144
- normal TDD, implementation, and review routes.
145
-
146
- ## Precedence
147
-
148
- 1. If the user asks not to edit code, use diagnosis or review skills and stop before implementation.
149
- 2. If the user asks to fix, implement, or build, apply `$tdd` before planning or editing.
150
- 3. If a bug is hard, flaky, or performance-related, use `$diagnosing-bugs` to build a feedback loop before fixing.
151
- 4. If a fix path has already been approved, use `$code-debugger` to implement and verify it.
152
- 5. Run one final `$code-review` wave with bounded cleanup in its spec/standards lens. High-risk work uses two disjoint reviewers in parallel. Run separate `$cleanup-review` only for an explicit concrete evidenced reason that cannot fit that lens; size or risk alone is insufficient.
153
-
154
- Reviewer repairs inside an active authorized implementation/TDD flow follow
155
- [`bug-workflow-routing.md`](bug-workflow-routing.md) and do not automatically
156
- start `code-debugger`; standalone or ambiguous fixes retain the normal route.
157
-
158
- ## Local Fact Rule
159
-
160
- When a global skill needs a repo fact, it should read local evidence first: `AGENTS.md`, `CONTEXT.md`, ADRs, package manifests, lockfiles, existing tests, scripts, and local skill docs. If the fact is not present, say it is not confirmed instead of inventing it.
161
-
162
- ## Availability, Depth, And Fallback
163
-
164
- - Request the exact role name and always start it without inherited conversation; put the necessary verified context in a self-contained brief because full-history forks can inherit the parent profile.
165
- - `agents.max_depth = 1`: children never spawn grandchildren. Root directly owns mandatory reviewer and worker launches.
166
- - After collecting a child's result, root must close it in a finally-equivalent path. Parallel launch must preserve every fulfilled handle (use `allSettled` or equivalent), then close partial launches after timeout, cancellation, or error so completed agents do not consume `max_threads` slots.
167
- - If a required reviewer role is unavailable, do not silently substitute a generic child, inherit root settings, or self-review inline. Report the gate as unavailable or blocked. Non-review skills may keep their explicit inline fallback rules.
168
- - Never override a named role's model or effort at spawn time. Change and revalidate the role file instead.
169
-
170
- ## Broad Exploration Delegation
171
-
172
- Route discovery by required output:
173
-
174
- | Need | Route |
42
+ | Tiny, clear, low-risk edit | `$small-task-implementer` after its Fit Gate |
43
+ | Clear feature or fix | Apply the TDD Fit Gate; when it fits, use Root + one `$tdd` activation, otherwise affected validation |
44
+ | Missing execution detail | `$implementation-spec-maker` -> artifact review -> `$spec-implementer` |
45
+ | Approved implementation spec | `$spec-implementer` |
46
+ | Approved dependency graph or explicit orchestration | `$tickets-orchestrator` |
47
+ | Product discovery or ticket decomposition | `$to-spec`, `$spec-to-tickets`, or `$wayfinder` as applicable; stop before delivery |
48
+ | Explain-only bug | `$bug-root-cause-explainer`; no edits |
49
+ | Confirmed bounded bug fix | Apply the TDD Fit Gate, then `$tdd` + `$code-debugger` when it fits; otherwise `$code-debugger` + affected validation |
50
+ | Hard, flaky, unclear, or performance bug | `$diagnosing-bugs` before the explain/fix route |
51
+ | Review request | `$code-review` in the profile-selected reviewer child |
52
+ | External multi-source uncertainty | `$research`; narrow documentation lookup stays inline |
53
+ | Commit request | `$commit`; push/PR still require separate authority |
54
+
55
+ Generated planning artifacts and labels never authorize implementation. One
56
+ deterministic approved ticket may run directly; a graph follows the authorized
57
+ delivery workflow.
58
+
59
+ ## TDD And Review
60
+
61
+ Apply `$tdd` only when all three Fit Gate conditions hold:
62
+
63
+ 1. The change alters observable behavior.
64
+ 2. A natural public seam can prove that behavior.
65
+ 3. A new test will fail before the change for the intended behavioral reason.
66
+
67
+ Otherwise use existing regression tests plus affected validation. Do not invoke
68
+ TDD merely because implementation files change, and do not manufacture RED
69
+ tests for behavior-preserving cleanup, dead-code deletion, documentation, copy,
70
+ formatting, generated assets, package maintenance, simple config, builds, or
71
+ read-only work. For mixed tasks, activate `$tdd` only for the behavioral slice.
72
+ An absence or architecture guard added after cleanup is validation, not a TDD
73
+ cycle.
74
+
75
+ Review applicability lives in [`review-gates.md`](review-gates.md). Shared
76
+ Full/Closure mechanics live in [`review-protocol.md`](review-protocol.md).
77
+ Artifact review is owned by
78
+ [`implementation-spec-review/references/review-loop.md`](../../skills/implementation-spec-review/references/review-loop.md);
79
+ approved-spec implementation review is owned by
80
+ [`spec-implementer/references/review-loop.md`](../../skills/spec-implementer/references/review-loop.md).
81
+
82
+ Root never substitutes self-review for a required independent reviewer. Use one
83
+ `reviewer_fast` for `simple`, one `reviewer_standard` for `medium`, and two
84
+ disjoint `reviewer_deep` tracks for `high`. A reviewer Adapter executes inline
85
+ only after it is already inside that assigned child.
86
+
87
+ ## Delegation
88
+
89
+ Run work inline by default. Delegate only when the user, an invoked skill, or
90
+ repository policy authorizes it and the task benefits from independent review,
91
+ isolated deep analysis, or disjoint implementation ownership.
92
+
93
+ | Need | Named role |
175
94
  | --- | --- |
176
- | Known path or one narrow execution path | Root reads inline |
177
- | Mechanical inventory or large diff/log scan | `explorer_quick` |
178
- | Bounded cross-module execution trace | `explorer_fast` |
179
- | Material external docs/API/spec question | `$research` with `researcher_standard` |
180
- | Ambiguous architecture, root cause, or contract synthesis | `analyst_deep` |
181
-
182
- Give each child one **Discovery Capsule**: question, known entrypoint, scope,
183
- excluded areas, and expected `answer -> execution path -> file:symbol evidence ->
184
- uncertainty`. Stop when the question is answered or missing evidence is proven.
185
- Use at most two explorer children, only for disjoint questions, and reuse the
186
- same child for follow-up instead of restarting discovery.
187
-
188
- Root verifies evidence that drives edits and keeps a compact **Evidence Map**
189
- for reuse by plan, spec, and implementation. Re-read only entries invalidated by
190
- changed files, contracts, or external sources. `analyst_deep` remains a final
191
- synthesis escalation, not a substitute for evidence collection.
192
-
193
- ## Contract Test Ledger Rule
194
-
195
- Use [`contract-test-ledger.md`](contract-test-ledger.md) for behavior-changing
196
- tasks with contract risk. It maps each invariant to the first failing test or
197
- observable proof before implementation.
198
-
199
- ## Progressive Disclosure Rule
200
-
201
- Keep main skill files focused on routing, workflow, safety, and output. Load
202
- long checklists, framework lenses, recipes, examples, and rubrics from references
203
- only when their trigger applies.
95
+ | Mechanical inventory | `explorer_quick` |
96
+ | Bounded cross-module trace | `explorer_fast` |
97
+ | Ambiguous architecture, contract, or cause | `analyst_deep` |
98
+ | Primary-source external research | `researcher_standard` |
99
+ | Independent review | `reviewer_fast`, `reviewer_standard`, or `reviewer_deep` by profile |
100
+ | Approved isolated implementation slice | `implementer_standard`; `implementer_deep` only for material uncertainty |
101
+
102
+ Keep the root critical path local. Use at most two explorers for disjoint
103
+ questions and at most two parallel implementers with disjoint write scopes.
104
+ Children do not conduct user dialogue or spawn grandchildren.
105
+
106
+ ## Validation And Runtime Safety
107
+
108
+ Use targeted behavior proof plus the smallest affected integration check for
109
+ simple and medium work. Run a full repository suite only when repository policy
110
+ requires it, a broad shared contract cannot be isolated, or the task is `high`.
111
+
112
+ For Flutter UI, follow [`tool-usage.md`](tool-usage.md): platform QA owns UI
113
+ work and `$flutter-attach-session` is only the attach-safe runtime layer. Treat
114
+ live app, IDE, VM Service, and `flutter run` sessions as user-owned.
115
+
116
+ Read local evidence before external search: applicable `AGENTS.md`, `CONTEXT.md`,
117
+ ADRs, manifests, lockfiles, tests, scripts, and code owners. Mark missing facts
118
+ unconfirmed rather than inventing them.
119
+
120
+ Contract-risk implementation uses
121
+ [`contract-test-ledger.md`](contract-test-ledger.md) only for material
122
+ invariants. Long framework lenses, examples, and recipes remain skill-local and
123
+ load on demand.
@@ -1,49 +1,42 @@
1
1
  # Review Gates
2
2
 
3
- This file owns applicability only. Approved-spec execution follows
4
- [`implementation-review-loop.md`](implementation-review-loop.md), which owns
5
- checkpoint order, Full/Closure topology, durable state, and final coverage.
6
- Direct work uses the gates below without manufacturing Module state.
3
+ This file owns review applicability. Review execution mechanics live in
4
+ [`review-protocol.md`](review-protocol.md); approved-spec review shape lives in
5
+ [`spec-implementer/references/review-loop.md`](../../skills/spec-implementer/references/review-loop.md).
7
6
 
8
- ## Cleanup Review
7
+ ## Final Code Review
9
8
 
10
- Do not launch a separate `$cleanup-review` from size or risk classification
11
- alone. The final `$code-review` spec/standards lens owns bounded cleanup for
12
- simple, medium, large, and high-risk changes; high runs that lens in its own
13
- parallel reviewer track.
9
+ Run final `$code-review` for behavior-changing work that affects:
14
10
 
15
- Use separate `$cleanup-review` only when the user, approved source, or repo
16
- policy names a concrete evidenced simplification risk that cannot fit the
17
- bounded spec/standards lens. It runs before final `$code-review`; approved-spec
18
- execution lets the shared Module schedule it.
11
+ - medium/large shared business behavior;
12
+ - API/DTO/schema, migration, persistence, auth, permission, payment, cache,
13
+ concurrency, background jobs, or shared-state contracts;
14
+ - shared UI/navigation/middleware/core flows;
15
+ - runtime logic across three or more files when it crosses an owner or
16
+ validation seam.
19
17
 
20
- Treat a change as large only when it contains several independently verifiable
21
- runtime workflows or material cross-owner, cross-repo, release-sequencing, or
22
- rollback coordination. File count, module count, or a broad mechanical diff is
23
- not enough. Default a coherent feature with one behavior and one validation
24
- path to medium even when it touches several files.
18
+ Do not invoke it automatically for docs, copy, comments, tests-only or
19
+ styling-only changes, formatting, renames, mechanical refactors, or isolated
20
+ low-risk one-file fixes.
25
21
 
26
- Cleanup is not a correctness review. Its skill owns the detailed lens,
27
- confidence handling, repair integration, and output contract.
22
+ `medium` is the normal review profile. API, persistence, statefulness, file
23
+ count, or orchestration strengthens the review focus only when it creates an
24
+ affected contract; none independently selects `high`.
28
25
 
29
- ## Final Code Review
26
+ ## Cleanup
30
27
 
31
- Run final `$code-review` when implementation changes include any of:
28
+ Cleanup is a lens inside the same final `$code-review`, never a separate gate.
29
+ Use bounded cleanup by default. Amplify it only when the user, approved source,
30
+ or repository policy names a concrete evidenced simplification risk; follow
31
+ `../../skills/code-review/references/cleanup-lens.md` for that branch.
32
32
 
33
- - a medium/large feature or shared multi-module business behavior;
34
- - API/DTO/schema, migration, persistence, auth, permission, payment, cache,
35
- concurrency, background-job, or shared-state contracts;
36
- - shared UI/navigation/middleware/core flows;
37
- - runtime logic across three or more files when the change is behavioral rather
38
- than mechanical and crosses one owner or validation seam.
39
-
40
- Do not invoke it automatically for documentation/copy/comments, tests-only or
41
- styling-only changes, formatting/renames, mechanical refactors, or isolated
42
- one-file fixes with low regression risk.
43
-
44
- `$code-review` owns reviewer topology, confidence, auto-fix, and output rules.
45
- Its spec/standards reviewer owns the bounded cleanup lens for every profile; do
46
- not launch `$cleanup-review` for the same settled diff without an explicit
47
- concrete reason beyond size or risk labels.
48
- For approved specs, required final coverage remains part of the shared
49
- Implementation Review State.
33
+ After one consolidated repair, coordinator verification plus affected
34
+ validation closes ordinary medium/low behavior-preserving findings. Use
35
+ Closure only for the triggers in `review-protocol.md`.
36
+
37
+ ## Validation Depth
38
+
39
+ For simple and medium work, run targeted behavior proof and the smallest
40
+ affected integration check. Run a full repository suite only when explicitly
41
+ required by repository policy, when broad contract fan-out cannot be isolated,
42
+ or for a genuinely `high` task.