qfai 1.9.2 → 1.10.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +48 -2
- package/assets/init/.qfai/assistant/agents/acceptance-test-engineer.md +1 -0
- package/assets/init/.qfai/assistant/agents/backend-engineer.md +1 -0
- package/assets/init/.qfai/assistant/agents/completion-reviewer.md +13 -2
- package/assets/init/.qfai/assistant/agents/delivery-planner.md +9 -0
- package/assets/init/.qfai/assistant/agents/frontend-engineer.md +1 -0
- package/assets/init/.qfai/assistant/agents/implementation-reviewer.md +6 -0
- package/assets/init/.qfai/assistant/agents/orchestrator.md +2 -2
- package/assets/init/.qfai/assistant/agents/qa-gatekeeper.md +96 -3
- package/assets/init/.qfai/assistant/agents/test-design-analyst.md +20 -3
- package/assets/init/.qfai/assistant/catalog/cli-ux-guidelines.md +2 -2
- package/assets/init/.qfai/assistant/catalog/spec_required_files.json +2 -1
- package/assets/init/.qfai/assistant/catalog/test-layers.md +355 -14
- package/assets/init/.qfai/assistant/catalog/worklog-entry.schema.md +165 -0
- package/assets/init/.qfai/assistant/constitution/communication.md +1 -1
- package/assets/init/.qfai/assistant/constitution/drift-protocol.md +304 -10
- package/assets/init/.qfai/assistant/constitution/quality.md +35 -5
- package/assets/init/.qfai/assistant/constitution/requirements-decomposition.md +37 -0
- package/assets/init/.qfai/assistant/constitution/shared-skill-delegation-baseline.md +244 -8
- package/assets/init/.qfai/assistant/constitution/shared-skill-operating-baseline.md +122 -5
- package/assets/init/.qfai/assistant/constitution/workflow.md +53 -7
- package/assets/init/.qfai/assistant/manifest/agent-catalog.yml +316 -945
- package/assets/init/.qfai/assistant/manifest/agent-routing.yml +50 -4
- package/assets/init/.qfai/assistant/manifest/review-profiles.yml +9 -0
- package/assets/init/.qfai/assistant/process/migrations/v1.4.27-atdd-alignment.md +1 -1
- package/assets/init/.qfai/assistant/skills/qfai-atdd/SKILL.md +61 -21
- package/assets/init/.qfai/assistant/skills/qfai-atdd/references/test-case-depth-checklist.md +23 -4
- package/assets/init/.qfai/assistant/skills/qfai-configure/SKILL.md +15 -7
- package/assets/init/.qfai/assistant/skills/qfai-discussion/SKILL.md +8 -4
- package/assets/init/.qfai/assistant/skills/qfai-discussion/references/design-md-brand-catalog.md +2 -2
- package/assets/init/.qfai/assistant/skills/qfai-discussion/references/discussion-completion-matrix.md +17 -7
- package/assets/init/.qfai/assistant/skills/qfai-discussion/references/rcp_footer.md +10 -4
- package/assets/init/.qfai/assistant/skills/qfai-discussion/references/review-cycle-playbook.md +1 -1
- package/assets/init/.qfai/assistant/skills/qfai-discussion/references/ui-bearing-playbook.md +4 -4
- package/assets/init/.qfai/assistant/skills/qfai-discussion/references/ui_ux/review_audit_playbook.md +1 -1
- package/assets/init/.qfai/assistant/skills/qfai-discussion/references/ui_ux/trend_scan_playbook.md +1 -1
- package/assets/init/.qfai/assistant/skills/qfai-discussion/references/ui_ux_best_practices.md +17 -7
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/01_Context.md +1 -1
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/02_Inception-Deck.md +1 -1
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/03_Story-Workshop.md +11 -3
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/05_Scope.md +5 -2
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/07_NFR.md +1 -1
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/09_Constraints.md +7 -4
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/10_Policy.md +1 -1
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/11_OQ-Register.md +1 -1
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/12_OQ-Resolution-Log.md +1 -1
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/14_Review-Request.md +14 -7
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/99_delta.md +1 -1
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/review/Rxx_reviewer.md +16 -7
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/review/review_request.md +9 -6
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/uiux/40_screen_contracts.md +3 -2
- package/assets/init/.qfai/assistant/skills/qfai-discussion/templates/uiux/50_review_input_bundle.md +4 -2
- package/assets/init/.qfai/assistant/skills/qfai-implement/SKILL.md +250 -128
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/change-request-reset.md +93 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/checkpoint-verification.md +106 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/cross-spec-ownership.md +73 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/evidence-revision.md +77 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/execution-ledger.md +303 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/final-checklist.md +19 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/finding-classification.md +49 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/ledger-preconditions.md +55 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/oracle-strength.md +81 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/parallelization-policy.md +235 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/red-admissibility.md +91 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/red-not-observable.md +70 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/relevant-test-suite.md +87 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/review-artifact-layout.md +31 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/round-evidence.md +102 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/selector-granularity.md +24 -0
- package/assets/init/.qfai/assistant/skills/qfai-implement/references/volume-policy.md +149 -0
- package/assets/init/.qfai/assistant/skills/qfai-prototyping/SKILL.md +35 -18
- package/assets/init/.qfai/assistant/skills/qfai-prototyping/references/evidence-requirements.md +1 -1
- package/assets/init/.qfai/assistant/skills/qfai-prototyping/references/generator-prompt.md +106 -7
- package/assets/init/.qfai/assistant/skills/qfai-prototyping/references/handoff.md +40 -11
- package/assets/init/.qfai/assistant/skills/qfai-prototyping/references/iteration-loop.md +48 -1
- package/assets/init/.qfai/assistant/skills/qfai-prototyping/templates/DESIGN.md.sample +6 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/SKILL.md +82 -22
- package/assets/init/.qfai/assistant/skills/qfai-sdd/references/contract-artifact-rules.md +148 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/references/rcp_footer.md +10 -4
- package/assets/init/.qfai/assistant/skills/qfai-sdd/references/review-cycle-playbook.md +1 -1
- package/assets/init/.qfai/assistant/skills/qfai-sdd/references/sdd-execution-playbook.md +4 -2
- package/assets/init/.qfai/assistant/skills/qfai-sdd/references/sdd-phase-checklists.md +18 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/references/sdd-quality-gate.md +44 -3
- package/assets/init/.qfai/assistant/skills/qfai-sdd/references/sdd-triage.md +60 -7
- package/assets/init/.qfai/assistant/skills/qfai-sdd/references/spec-traceability-rules.md +157 -5
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/change-request.md +125 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/contracts/db-contract.sample.sql +6 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/evidence/sdd-spec.md +92 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/_policies/01_Objective.md +27 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/_policies/02_Initiative.md +30 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/_policies/05_Contracts.md +16 -6
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/_policies/06_Glossary.md +19 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/_policies/07_Constraints.md +25 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/_policies/08_Decisions.md +22 -2
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/_policies/11_Slice-Policy.md +37 -10
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/spec/02_User-stories.md +20 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/spec/03_Acceptance-Criteria.md +19 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/spec/04_Business-Rules.md +32 -3
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/spec/06_Test-Cases.md +56 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/spec/07_Decisions.md +29 -2
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/spec/10_Plan.md +41 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/spec/16_Traceability-ledger.md +58 -0
- package/assets/init/.qfai/assistant/skills/qfai-sdd/templates/specs/spec/tdd/test-list.md +56 -0
- package/assets/init/.qfai/assistant/skills/qfai-verify/SKILL.md +57 -102
- package/assets/init/.qfai/assistant/skills/qfai-verify/references/articles.md +24 -0
- package/assets/init/.qfai/assistant/skills/qfai-verify/references/context-load.md +22 -0
- package/assets/init/.qfai/assistant/skills/qfai-verify/references/verify-output-contract.md +48 -0
- package/assets/init/.qfai/assistant/skills/qfai-verify/templates/verify-evidence.md +48 -0
- package/assets/init/.qfai/assistant/skills/web-research/SKILL.md +25 -7
- package/assets/init/.qfai/waivers.yml +11 -5
- package/assets/init/root/DESIGN.md +6 -0
- package/assets/init/root/qfai.config.yaml +15 -12
- package/dist/cli/index.cjs +11023 -7139
- package/dist/cli/index.cjs.map +1 -1
- package/dist/cli/index.mjs +10963 -7080
- package/dist/cli/index.mjs.map +1 -1
- package/dist/index.cjs +8782 -5257
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +280 -8
- package/dist/index.d.ts +280 -8
- package/dist/index.mjs +11776 -8264
- package/dist/index.mjs.map +1 -1
- package/package.json +18 -19
|
@@ -17,7 +17,7 @@ roles:
|
|
|
17
17
|
completion-reviewer,
|
|
18
18
|
product-surface-reviewer,
|
|
19
19
|
]
|
|
20
|
-
routing-profile:
|
|
20
|
+
routing-profile: implementation-heavy
|
|
21
21
|
mode: approval-gated
|
|
22
22
|
---
|
|
23
23
|
|
|
@@ -31,14 +31,21 @@ QFAI Skill Body (SSOT)
|
|
|
31
31
|
|
|
32
32
|
[DRIFT-PROTOCOL:MANDATORY]
|
|
33
33
|
|
|
34
|
+
## Preconditions
|
|
35
|
+
|
|
36
|
+
- **`.qfai/specs/<spec-id>/tdd/test-list.md` must exist and contain the eight required columns.** It is the ledger every step of this skill reads.
|
|
37
|
+
- **Producer**: `/qfai-sdd` Phase 2b seeds it. Do **not** proceed with an absent ledger and do **not** invent rows that no TC backs.
|
|
38
|
+
- **An empty ledger is a fault only when `06_Test-Cases.md` disagrees.** Never read a header-only table as "nothing to do" on its own. The recovery procedure and the coverage-target test that separates a truthfully empty ledger from an incomplete one are in
|
|
39
|
+
`references/ledger-preconditions.md`; read it before exiting on an empty ledger.
|
|
40
|
+
|
|
34
41
|
## Spec Auto-Discovery Protocol
|
|
35
42
|
|
|
36
|
-
When no explicit argument is given, detect the candidate
|
|
43
|
+
When no explicit argument is given, detect the candidate specs. Execution is constrained to one spec at a time. Auto-discovery MAY present several specs as a queue to be processed sequentially (see Volume Policy > Multi-spec queue); this protocol does NOT enable multi-spec parallel execution.
|
|
37
44
|
|
|
38
45
|
### User Selection Flow
|
|
39
46
|
|
|
40
47
|
- Single spec: announce the detected spec; ask for confirmation when scope is ambiguous.
|
|
41
|
-
- Multiple specs: display the candidates and require the user to choose one spec.
|
|
48
|
+
- Multiple specs: display the candidates and require the user to choose one spec, or to confirm an ordered queue processed one spec at a time.
|
|
42
49
|
- Zero specs: stop and ask the user to provide the target spec explicitly.
|
|
43
50
|
|
|
44
51
|
## User Questions (AskUserQuestion Protocol)
|
|
@@ -52,15 +59,22 @@ Skill-specific examples:
|
|
|
52
59
|
|
|
53
60
|
## CRITICAL CONSTRAINTS (Read First)
|
|
54
61
|
|
|
55
|
-
- This skill processes **one test at a time** from `test-list.md
|
|
62
|
+
- This skill processes **one test at a time** from `test-list.md`: at most one row is in `red` or `green` at any moment, except under an item-level parallel dispatch authorized by `## Parallelization Policy` below. A T1 row parked in `refactor` waiting for its review group (see Volume Policy) does not violate this.
|
|
56
63
|
- Each item goes through the full TDD micro-cycle: write a **failing test** first, then make it pass, then refactor.
|
|
57
64
|
- The execution ledger is located at `.qfai/specs/<spec-id>/tdd/test-list.md`.
|
|
58
|
-
-
|
|
59
|
-
-
|
|
65
|
+
- Write a `.qfai/steering/<id>.md` work-log entry when this stage hits a condition in the `kind` trigger table of `.qfai/assistant/catalog/worklog-entry.schema.md` — `blocker`, `handoff`, `consultation-needed` and `decision` are the ones it reaches most. `npx qfai validate` polices that surface but nothing else asks for an entry, so an unwritten one is simply lost.
|
|
66
|
+
- Items are processed **serially** by default. Item-level parallel processing inside one spec is allowed only under `## Parallelization Policy` below — both its technical gate and its consent gate must hold, and user approval cannot override a technical DENY. Cross-spec parallelism is never allowed.
|
|
67
|
+
- Status transitions follow a strict forward-only lifecycle: `todo` -> `red` -> `green` -> `refactor` -> `done`. The single re-entry is `refactor` -> `red` after a `qa-gatekeeper` `REVISE` on the row's RED/GREEN evidence.
|
|
60
68
|
- The `exception` status can be reached from any active status when an anomaly is detected.
|
|
61
|
-
- Backward transitions are prohibited (e.g., `green` -> `red` is not allowed).
|
|
62
|
-
- Completed items (`done`) are skipped on re-execution.
|
|
63
|
-
- When
|
|
69
|
+
- Backward transitions are prohibited (e.g., `green` -> `red` is not allowed). The only exception is an approved Change Request reset (see Status Lifecycle).
|
|
70
|
+
- Completed items (`done`) are skipped on re-execution, unless an approved Change Request reset them.
|
|
71
|
+
- When every item is terminal (`done` or a valid `exception`) **and the mandatory Change Request
|
|
72
|
+
preflight (see Required Process) reset nothing**, the per-item work is finished — but the
|
|
73
|
+
**spec-level checkpoint boundary** may still be owed — an interrupted run, or a re-run of an
|
|
74
|
+
already-terminal ledger, leaves it unrecorded. Before reporting "nothing to do" and exiting,
|
|
75
|
+
confirm fresh spec-level checkpoint verification evidence exists for this ledger state; run the
|
|
76
|
+
per-spec verification first when it is missing or stale. See `references/checkpoint-verification.md#spec-level-boundary-on-an-already-complete-ledger`. Only
|
|
77
|
+
then report "nothing to do" for that spec, then advance to the next spec of a confirmed queue; exit when the queue is empty (Volume Policy > Advancing the queue).
|
|
64
78
|
|
|
65
79
|
## Goal
|
|
66
80
|
|
|
@@ -76,78 +90,88 @@ Execute the TDD micro-cycle for each pending item in `test-list.md`, transitioni
|
|
|
76
90
|
|
|
77
91
|
## Non-goals
|
|
78
92
|
|
|
79
|
-
- Writing spec artifacts (use `/qfai-sdd`).
|
|
80
|
-
- Writing acceptance tests (use `/qfai-atdd`).
|
|
93
|
+
- Writing spec artifacts other than this skill's own `tdd/test-list.md` ledger (use `/qfai-sdd`). The ledger's `Status` / `DR-ID` / `Evidence` cells are the one carve-out the Drift Protocol grants (`constitution/drift-protocol.md#allowed-exceptions-minimal-whitelist`); its rows are still upstream.
|
|
94
|
+
- Writing acceptance tests (use `/qfai-atdd`). `Layer = E2E` / `Layer = API` ledger rows are tracked here but their tests are authored there.
|
|
81
95
|
- Running validation gates (use `/qfai-verify`).
|
|
82
|
-
- Parallel execution across multiple specs simultaneously.
|
|
96
|
+
- Parallel execution across multiple **specs** simultaneously. (Item-level parallelism _within_ one spec is a separate question, governed by
|
|
97
|
+
`## Parallelization Policy` below.)
|
|
83
98
|
|
|
84
99
|
## Execution Ledger: test-list.md
|
|
85
100
|
|
|
86
|
-
The execution ledger at `.qfai/specs/<spec-id>/tdd/test-list.md`
|
|
87
|
-
|
|
88
|
-
| Column | Description |
|
|
89
|
-
| --------- | -------------------------------------------------------- |
|
|
90
|
-
| TDD-ID | Unique identifier for the TDD item (e.g., TDD-0001) |
|
|
91
|
-
| TC-Refs | References to test cases from `06_Test-Cases.md` |
|
|
92
|
-
| Layer | Test layer (Unit, Integration, etc.) |
|
|
93
|
-
| Test file | Path to the test file |
|
|
94
|
-
| Selector | Test selector/description for targeted execution |
|
|
95
|
-
| Status | Current lifecycle status |
|
|
96
|
-
| DR-ID | Decision Record ID for exception items (blank otherwise) |
|
|
97
|
-
| Evidence | RED/GREEN command+result pairs proving the TDD cycle |
|
|
98
|
-
|
|
99
|
-
### Status Lifecycle
|
|
100
|
-
|
|
101
|
-
Valid status values: `todo`, `red`, `green`, `refactor`, `done`, `exception`.
|
|
102
|
-
|
|
103
|
-
Allowed transitions:
|
|
104
|
-
|
|
105
|
-
- `todo` -> `red` (write a failing test)
|
|
106
|
-
- `red` -> `green` (make the test pass with minimal code)
|
|
107
|
-
- `green` -> `refactor` (improve code quality while keeping tests green)
|
|
108
|
-
- `refactor` -> `done` (item complete)
|
|
109
|
-
- Any active status -> `exception` (anomaly detected; record DR-ID in DR-ID column)
|
|
110
|
-
|
|
111
|
-
Backward transitions are prohibited. Attempting `green` -> `red` must produce:
|
|
112
|
-
`"Backward transition prohibited: green -> red"`.
|
|
101
|
+
The execution ledger at `.qfai/specs/<spec-id>/tdd/test-list.md` is the single record of what this skill has done and may still do. Status values are `todo`, `blocked`, `red`, `green`, `refactor`, `done`, `exception`;
|
|
102
|
+
the lifecycle is forward-only, an `exception` requires a DR-ID, and a `blocked` row requires a `Blocked-By` and is never selected.
|
|
113
103
|
|
|
114
|
-
|
|
104
|
+
The eight required columns, the allowed transitions and the exception rules are in `references/execution-ledger.md`. Read it before writing to the ledger.
|
|
115
105
|
|
|
116
|
-
|
|
106
|
+
## Required Process
|
|
117
107
|
|
|
118
|
-
|
|
119
|
-
- If the DR-ID column is empty, emit error: `"exception status requires DR-ID in DR-ID column"`.
|
|
108
|
+
### Phase: Stage 0 + Preflight — MANDATORY, runs first
|
|
120
109
|
|
|
121
|
-
|
|
110
|
+
1. Follow `.qfai/assistant/constitution/shared-skill-operating-baseline.md#stage-0---steering-completion-refresh-mandatory`, `#format-ssot-mandatory` and `#delta-rejected-guard-mandatory`, then read `catalog/tech.md` + `catalog/structure.md` and take every Test / Lint / Typecheck / Build command below from `tech.md#standard-commands-copy-paste` rather than inventing one. Refresh both files when the repository contradicts them; do not continue on stale steering. This stage was bound by neither inheritance route while being the one that creates production source trees.
|
|
111
|
+
2. Enumerate the in-scope `.qfai/decisions/CR-*.md` and apply every approved reset per `references/change-request-reset.md` **before** any other ledger judgement — including the all-`done` "nothing to do" exit, which an approved reset invalidates.
|
|
122
112
|
|
|
123
113
|
### Phase: Red (Write Failing Test)
|
|
124
114
|
|
|
125
|
-
1. Read `test-list.md` and select the first
|
|
126
|
-
2. Transition status to `red
|
|
127
|
-
3. Write a **failing test**
|
|
128
|
-
|
|
129
|
-
|
|
115
|
+
1. Read `test-list.md`. **Rework first**: if any row is at `review-fix`, select the first such row and resume its rework (`references/round-evidence.md`) before any `todo` row — one left by an interrupted session is otherwise never picked up. Otherwise select the first row with `Status = todo`, skipping any `blocked` row — it cannot be started, and re-issuing it is what made the determination get re-derived every pass (`references/execution-ledger.md#blocked-rows`).
|
|
116
|
+
2. Transition status to `red` — **only for a `todo` row**. A `review-fix` row **stays at `review-fix`** for the whole rework: it runs steps 3-5 and the Green phase in place, and `review-fix -> red` is not an allowed transition.
|
|
117
|
+
3. Write a **failing test** from the row's obligation column, selected by `Layer`: `TC-Refs` for `Unit` / `Component` / `Integration`, `US-Refs` for `E2E`, `CON-API-Refs` for `API`. An `E2E` or `API` row's test is authored by `/qfai-atdd` (Non-goals); this skill only drives that row's status and evidence once the acceptance test exists, and stops with a handoff note if it does not.
|
|
118
|
+
3a. Create the **minimal seam** the test imports — module, export or signature with no behaviour — so the test module loads. This is not Phase Green's production code: it implements no predicate. Without it, the first failure of any new-symbol row is a resolution error by construction.
|
|
119
|
+
4. Run the test and **watch it fail**. Admissible only when an assertion — or an expected-exception check — inside this row's `Selector` raised the failure and its message names the predicate the row owns; a collection / import / syntax / fixture error, or an unasserted throw, is a **missing seam**, not a RED (`references/red-admissibility.md`). Observe each `Selector` entry's failure separately; one aggregate run is not a valid RED observation. Submit that run to `qa-gatekeeper` and obtain confirmation **before** any production code exists — routing phase `red`.
|
|
120
|
+
5. If the test unexpectedly passes, classify **why** before doing anything else. An obligation
|
|
121
|
+
already satisfied by a sibling row is **not an anomaly** and does **not** go to `exception`;
|
|
122
|
+
anything else transitions to `exception` and records the anomaly as `.qfai/decisions/DR-<id>-<slug>.md` — never in
|
|
123
|
+
`07_Decisions.md` / `09_delta.md`, which are upstream SSOT this skill may not patch. Never weaken a correct
|
|
124
|
+
test until it fails in order to manufacture a RED. See `references/red-not-observable.md`.
|
|
125
|
+
> **RED observation is only as good as the selector's granularity.** A single test function can fail
|
|
126
|
+
> only once, so if one selector entry carries an entire obligation matrix, "the expected reason" is
|
|
127
|
+
> whichever assert happens to execute first — every assertion after it is unobserved on every RED
|
|
128
|
+
> run, and a non-deterministic assertion placed early silently disables everything below it. A TDD
|
|
129
|
+
> row whose selector accumulates unrelated boundaries therefore **invalidates its own RED
|
|
130
|
+
> observation**. Split the row per `#selector-granularity-must` before continuing; do not proceed to
|
|
131
|
+
> Green.
|
|
130
132
|
|
|
131
133
|
### Phase: Green (Make It Pass)
|
|
132
134
|
|
|
133
|
-
1. Write the **minimum production code** to make the failing test pass.
|
|
134
|
-
2. Run the test and **watch it pass
|
|
135
|
-
3. Transition status to `green
|
|
135
|
+
1. Write the **minimum production code** to make the failing test pass. On the _RED not observable_ path there is none to write — the `Satisfied-by` row already implements the predicate — so go straight to step 2.
|
|
136
|
+
2. Run the test and **watch it pass**, and submit that run to `qa-gatekeeper` for the GREEN confirmation — routing phase `build`, which is blocking for the same reason `red` is.
|
|
137
|
+
3. Transition status to `green` — **only for a row that entered from `todo`**. A `review-fix` row stays at `review-fix` here too; `review-fix -> green` is not an allowed transition and must not be written to the ledger.
|
|
136
138
|
4. If the test still fails after implementation, investigate and fix. Do not skip to refactor.
|
|
137
139
|
|
|
138
140
|
### Phase: Refactor
|
|
139
141
|
|
|
140
|
-
1. Improve code quality (naming, structure, duplication removal) while keeping all tests green.
|
|
141
|
-
2. Run the
|
|
142
|
+
1. Improve code quality (naming, structure, duplication removal) while keeping all tests green. **Before editing a production file**, check whether another spec's `tdd/test-list.md` names it in `Test file`; if so the edit is cross-spec — record it in this item's evidence and re-run `completion-reviewer` against that spec's obligations too (`references/cross-spec-ownership.md`).
|
|
143
|
+
2. Run the **relevant test suite** to confirm nothing broke. "Relevant" means the
|
|
144
|
+
smallest selector that covers the module you touched **plus its reverse
|
|
145
|
+
dependency closure** — walk the production import graph backwards, not just the
|
|
146
|
+
test files that import the module directly. Fall back to the package containing
|
|
147
|
+
the touched module whenever that walk cannot be completed; never "every test in
|
|
148
|
+
the repository" at this step. Cadence: **narrow suite per item, full suite at
|
|
149
|
+
each checkpoint boundary**, because a full run per item is quadratic in ledger
|
|
150
|
+
size. Full rules, the fallback triggers and the boundary list:
|
|
151
|
+
`references/relevant-test-suite.md`.
|
|
142
152
|
3. Transition status to `refactor`.
|
|
143
|
-
4. Submit for completion review (`completion-reviewer`) and code quality review
|
|
144
|
-
|
|
153
|
+
4. Submit for completion review (`completion-reviewer`) and code quality review
|
|
154
|
+
(`implementation-reviewer`). A T1 row submits with its coherent group and stays in
|
|
155
|
+
`refactor` until the group closes (Volume Policy > Group formation).
|
|
156
|
+
5. After all routed blocking reviewers return PASS, run checkpoint verification
|
|
157
|
+
**while the item is still `refactor`** (see `#checkpoint-verification`). On a
|
|
158
|
+
checkpoint boundary that means the full suite. Off a boundary it is already
|
|
159
|
+
satisfied by step 2's narrow suite — nothing is re-run. Transition to `done`
|
|
160
|
+
only on PASS; on failure transition to `exception` with a DR-ID (legal from
|
|
161
|
+
`refactor`, whereas re-opening a `done` row is not). For a T1 group every member
|
|
162
|
+
transitions in the same ledger write.
|
|
145
163
|
|
|
146
164
|
### Completion
|
|
147
165
|
|
|
148
|
-
1. After processing all items, update `test-list.md` with final
|
|
166
|
+
1. After processing all items, update `test-list.md` with final Status, DR-ID and Evidence values — the three cells the Drift Protocol carve-out covers, and the ones gate item 10 reads.
|
|
149
167
|
2. If all items are `done`, report "All items complete".
|
|
150
|
-
3. If some items are `exception`, report them
|
|
168
|
+
3. If some items are `exception`, report them as **blocking output**, not as an
|
|
169
|
+
informational list: for each, the `TDD-ID`, the `DR-ID`, and whether that DR
|
|
170
|
+
is a user-approved accepted-risk waiver. Completion cannot be declared while
|
|
171
|
+
any `exception` row lacks such a waiver.
|
|
172
|
+
4. If a multi-spec queue was confirmed, announce the next queued spec and restart at
|
|
173
|
+
Phase: Red with its ledger; exit only after the last entry (Volume Policy >
|
|
174
|
+
Advancing the queue).
|
|
151
175
|
|
|
152
176
|
## Sub-agent Delegation (MANDATORY)
|
|
153
177
|
|
|
@@ -158,18 +182,21 @@ Follow `.qfai/assistant/constitution/shared-skill-delegation-baseline.md`.
|
|
|
158
182
|
- Orchestrator MUST NOT write test or production code directly; delegate every TDD phase to the routed implementation agents.
|
|
159
183
|
- Additional implement-specific overrides:
|
|
160
184
|
- read `test-list.md`, determine the next pending item, and delegate each TDD phase;
|
|
161
|
-
- update `test-list.md`
|
|
185
|
+
- update `test-list.md` **Status and Evidence** after each phase completes, recording the delegated agent's one-word RED/GREEN outcome plus the anchor into `.qfai/evidence/implement-<spec-id>.md`, and that agent's command+result verbatim in the evidence file itself — a GFM cell cannot hold either a newline or a bare `|` (`references/execution-ledger.md#evidence-cell-contract`). Gate item 10 requires both columns, the protocol permits both, and the orchestrator is the only role permitted to write this file (`references/parallelization-policy.md#ledger-ownership`).
|
|
162
186
|
|
|
163
187
|
### Formal Sub-agent Roster
|
|
164
188
|
|
|
165
|
-
This skill delegates through the centralized routing policy in `.qfai/assistant/manifest/agent-routing.yml`.
|
|
189
|
+
This skill delegates through the centralized routing policy in `.qfai/assistant/manifest/agent-routing.yml`. Its `red`, `build`, `test` and `review` phases carry `iteration: per-ledger-item` — they run once **per row**, not once per invocation, which is what puts `qa-gatekeeper` in a phase where a RED state still exists to observe.
|
|
166
190
|
|
|
167
191
|
- `delivery-planner`
|
|
168
|
-
- reads `test-list.md`, selects the next pending item,
|
|
192
|
+
- reads `test-list.md`, selects the next pending item, and is the sole authority for **item selection and item scope** — whether this row's selector is a sufficient slice of its `TC-*` obligation
|
|
193
|
+
- enforces Red-Green-Refactor **ordering** (which phase may run next), not the RED/GREEN observation itself
|
|
194
|
+
- is the sole authority for parallel dispatch decisions
|
|
169
195
|
- `frontend-engineer` / `backend-engineer`
|
|
170
196
|
- implement the selected item only, write the failing test first, write minimal passing code, and refactor without unrelated changes
|
|
171
197
|
- `qa-gatekeeper`
|
|
172
|
-
- is the sole authority for validating RED/GREEN observation evidence and completion gate evidence
|
|
198
|
+
- is the sole authority for validating **RED/GREEN observation evidence** — did the test fail (or pass) for the expected reason — and completion gate evidence. Routed per row in `red` (before production code) and `build` (after), not only in `review`
|
|
199
|
+
- does not adjudicate item scope; a scope objection is `delivery-planner`'s call
|
|
173
200
|
- `implementation-reviewer`
|
|
174
201
|
- reviews code quality, maintainability, backend correctness, and hidden coupling
|
|
175
202
|
- `completion-reviewer`
|
|
@@ -177,16 +204,42 @@ This skill delegates through the centralized routing policy in `.qfai/assistant/
|
|
|
177
204
|
- `product-surface-reviewer`
|
|
178
205
|
- reviews UI-affecting implementation when the item changes surface behavior
|
|
179
206
|
|
|
207
|
+
## Volume Policy (MUST)
|
|
208
|
+
|
|
209
|
+
Scale the ceremony to the ledger: derive a **risk tier** per row, **batch** T1 gatekeeping and
|
|
210
|
+
reviews per coherent BR/AC group, process multiple specs as a **sequential queue**, state the
|
|
211
|
+
implied **cost** before starting. The tier scales how **often** a gate runs, never **whether** it
|
|
212
|
+
runs: `agent-routing.yml` keeps `qa-gatekeeper`, `completion-reviewer` and
|
|
213
|
+
`implementation-reviewer` all mandatory (only the first two are in `blocking_agents`, but item 8 of
|
|
214
|
+
the 11-point gate makes an `implementation-reviewer` REVISE block `done` anyway), and criticality
|
|
215
|
+
(authz, crypto, money, data integrity) forces T2 regardless of layer. Why this exists, the tier
|
|
216
|
+
table, the group-formation transitions and the queue-advance steps: `references/volume-policy.md`.
|
|
217
|
+
|
|
180
218
|
### Handoff Contracts
|
|
181
219
|
|
|
182
220
|
All agent-to-agent transitions follow these contracts:
|
|
183
221
|
|
|
184
222
|
1. `delivery-planner` selects the next item and assigns it to the appropriate implementation agent.
|
|
185
|
-
2. Implementation agent submits RED
|
|
186
|
-
3. `qa-gatekeeper` confirms or rejects
|
|
187
|
-
4. After
|
|
223
|
+
2. Implementation agent submits the RED run to `qa-gatekeeper` **while no production code exists** (routing phase `red`), then the GREEN run after it (routing phase `build`). One combined post-hoc submission is not a substitute: RED is unrecoverable once Green begins, so it leaves nothing but the implementer's own account of a destroyed state — the self-attestation this gate exists to prevent. It **returns it to the orchestrator in the per-item evidence contract's form** so it can be written to the ledger. The implementation agent never writes `test-list.md` itself.
|
|
224
|
+
3. `qa-gatekeeper` confirms or rejects each observation. A RED rejection stops the row before Green; it does not wait for the `review` phase.
|
|
225
|
+
4. After the item reaches `refactor`, implementation agent submits it to `completion-reviewer` for spec alignment and to `implementation-reviewer` for code quality review. Review is requested from `refactor`, never from `green`, so a `REVISE` always lands on the one status with an outbound `review-fix` edge.
|
|
188
226
|
5. `product-surface-reviewer` is added when the item affects UI behavior or rendered output.
|
|
189
|
-
6. Only after
|
|
227
|
+
6. Only after every required reviewer passes may the item transition to `done`. "Required" is wider than `blocking_agents`: `implementation-reviewer` is mandatory and its `REVISE` blocks `done` independently of that list, and `product-surface-reviewer` joins for UI-affecting items. The authority for an item transition is `#item-completion-checklist-12-point-gate`, not the routing list, which governs phase progression (`references/volume-policy.md#routing-is-unchanged`).
|
|
228
|
+
7. For T1 rows the submitted unit in steps 2-4 is the coherent group, not the row; every required reviewer still runs, once per group. T2/T3 rows submit alone.
|
|
229
|
+
|
|
230
|
+
#### Precedence between `delivery-planner` and `qa-gatekeeper`
|
|
231
|
+
|
|
232
|
+
The two roles answer different questions and are ordered, not concurrent:
|
|
233
|
+
|
|
234
|
+
- `delivery-planner` answers _is this item's scope sufficient to be the whole of its `TC-*` obligation_.
|
|
235
|
+
- `qa-gatekeeper` answers _did the test fail (or pass) for the expected reason_.
|
|
236
|
+
|
|
237
|
+
Precedence rules:
|
|
238
|
+
|
|
239
|
+
- A `delivery-planner` REVISE on item scope MUST be resolved **before** RED evidence is submitted to `qa-gatekeeper` (step 2). Do not run step 2 while a scope REVISE is open.
|
|
240
|
+
- Once `qa-gatekeeper` PASSes the observation for a RED round, item scope MUST NOT be re-litigated for that round. A newly discovered scope gap opens a **new** `test-list.md` row rather than reopening the passed one.
|
|
241
|
+
- If a scope objection nonetheless arrives after step 3, it is treated as a new-row request; the existing PASS stands and is not discarded.
|
|
242
|
+
- Neither role may overrule the other inside the other's domain: a `qa-gatekeeper` PASS never widens item scope, and a `delivery-planner` verdict never substitutes for RED/GREEN observation evidence.
|
|
190
243
|
|
|
191
244
|
### Capability Probe (MUST)
|
|
192
245
|
|
|
@@ -195,11 +248,11 @@ All agent-to-agent transitions follow these contracts:
|
|
|
195
248
|
### Delegation Failure (Hard Stop)
|
|
196
249
|
|
|
197
250
|
- No additional overrides.
|
|
198
|
-
- Do not simulate roles.
|
|
251
|
+
- Do not simulate roles. Classify the failure per the baseline taxonomy first: `unavailable` stops the stage with a remediation report; `saturated` uses the bounded retry branch and keeps the stage open.
|
|
199
252
|
|
|
200
253
|
## Work Orders Summary
|
|
201
254
|
|
|
202
|
-
Use the shared schema (per-row `Status (PASS/REVISE)` column, reviewer response `Result: PASS | REVISE`).
|
|
255
|
+
Use the shared schema (per-row `Status (PASS/REVISE/PENDING)` column, reviewer response `Result: PASS | REVISE`).
|
|
203
256
|
|
|
204
257
|
### Reviewer Gate (MUST)
|
|
205
258
|
|
|
@@ -209,34 +262,42 @@ Use the shared schema (per-row `Status (PASS/REVISE)` column, reviewer response
|
|
|
209
262
|
- Test volume floors/ratios are not gates; they are signals.
|
|
210
263
|
- Do not declare DONE until Reviewer returns `PASS`; otherwise apply `REVISE`.
|
|
211
264
|
|
|
212
|
-
|
|
213
|
-
|
|
214
|
-
- **Default**: Serial execution. Items are processed one test at a time in `test-list.md` order.
|
|
215
|
-
- **Exception**: When items target completely independent SUT modules with no shared state, parallel processing may be used with explicit user approval.
|
|
216
|
-
- Serial execution ensures that each test is written and verified in isolation before moving to the next.
|
|
217
|
-
- `delivery-planner` is the sole authority for authorizing parallel dispatch.
|
|
265
|
+
#### Blocking vs advisory findings
|
|
218
266
|
|
|
219
|
-
|
|
267
|
+
Follow `shared-skill-delegation-baseline.md#finding-provenance-must`.
|
|
220
268
|
|
|
221
|
-
-
|
|
222
|
-
|
|
223
|
-
|
|
224
|
-
-
|
|
225
|
-
-
|
|
226
|
-
|
|
269
|
+
- Only **blocking** findings force `REVISE` and hold an item out of `done`. An **advisory**
|
|
270
|
+
finding (`Traces to: none`) MUST NOT be implemented as production code or pinned as a test
|
|
271
|
+
assertion; route it per `drift-protocol.md#reviewer-originated-obligations`.
|
|
272
|
+
- Do **not** edit `08_Open-questions.md` here — it is upstream SSOT under the Drift Protocol and
|
|
273
|
+
creating spec artifacts is a non-goal of this skill; the owner phase (`/qfai-sdd`) records and
|
|
274
|
+
adjudicates it.
|
|
275
|
+
- What each class cites, when an advisory takes the Change Request path, and why an advisory-only
|
|
276
|
+
review still returns `PASS`: `references/finding-classification.md`.
|
|
227
277
|
|
|
228
|
-
|
|
278
|
+
## Parallelization Policy
|
|
229
279
|
|
|
230
|
-
-
|
|
231
|
-
|
|
232
|
-
-
|
|
233
|
-
|
|
234
|
-
|
|
280
|
+
- **Cross-spec parallelism is barred.** One spec per invocation, always. This
|
|
281
|
+
is the Non-goal above and it is not approvable.
|
|
282
|
+
- **Item-level parallelism inside one spec** may be authorized. Two gates apply
|
|
283
|
+
and **both must hold**: a technical gate adjudicated by `delivery-planner`
|
|
284
|
+
(the sole authority), and explicit user approval. **User approval cannot
|
|
285
|
+
override a technical DENY.**
|
|
286
|
+
- **Default**: Serial execution, one test at a time in `test-list.md` order.
|
|
287
|
+
- The allow/deny conditions are stated as **concurrent write conflicts** —
|
|
288
|
+
including one item writing a module another item's test or implementation
|
|
289
|
+
reads — not as the existence of shared things, and authorized parallel runs
|
|
290
|
+
use the coordinated mode in which the orchestrator owns every `test-list.md`
|
|
291
|
+
write. Under RED-first the source modules do not exist when `delivery-planner` must judge, so the conditions are evaluated over each row's declared `Owning module` (`references/execution-ledger.md`); a ledger without that column supports parallel dispatch only for seams that already exist. Full rules: `references/parallelization-policy.md`.
|
|
292
|
+
- `parallel_groups: []` in `agent-routing.yml` describes **role fan-out within
|
|
293
|
+
a phase**, not item dispatch.
|
|
235
294
|
|
|
236
295
|
### Post-parallel integration verify
|
|
237
296
|
|
|
297
|
+
- **Reconcile the ledger first.** Under worktree separation each worker holds a private copy of `test-list.md`, so the merged trunk carries none of their transitions. Write Status + Evidence for every merged item from the worker reports **before** integration verify, and fail the verify if any merged item's row is still `todo` — an unreconciled ledger reports finished work as unstarted (`references/parallelization-policy.md#ledger-ownership`).
|
|
238
298
|
- After parallel slices complete and merge, run integration verify on the merged result
|
|
239
|
-
-
|
|
299
|
+
- Then reconcile the seams: diff each slice's touched `src/` paths against its declared `Owning module` and report undeclared or overlapping paths as a deny-condition breach — **independently of whether the merged suite is green** (`references/parallelization-policy.md#seam-reconciliation-after-a-parallel-run`)
|
|
300
|
+
- If integration verify fails, re-run it once with no intervening change before acting. A failure that does **not** reproduce is an `environment/tooling` finding (`shared-skill-operating-baseline.md#nondeterministic-gates`), reported with every run — not a rollback trigger. For a reproducible failure, **classify before acting** per `shared-skill-operating-baseline.md#gate-failure-autorepair-protocol`, attributing it to one slice, to the merge resolution, or to code outside every slice. Remedies by class: `references/parallelization-policy.md#failed-integration-verify`. Unconditional rollback is not one of them — the protocol classifies this as a local, non-destructive defect to fix and re-run, and reserves stopping for destructive changes. Only a **reproducible** failure flags all slices for re-examination and rolls back the merge, and only where the classification calls for it
|
|
240
301
|
- If integration verify passes, state transitions back to `delivery-planner` for sequential flow
|
|
241
302
|
|
|
242
303
|
## Completion Contract (Shared)
|
|
@@ -244,42 +305,74 @@ Use the shared schema (per-row `Status (PASS/REVISE)` column, reviewer response
|
|
|
244
305
|
Follow `.qfai/assistant/constitution/shared-skill-operating-baseline.md#completion-contract-shared`.
|
|
245
306
|
Follow `.qfai/assistant/constitution/shared-skill-operating-baseline.md#gate-failure-autorepair-protocol` for validate, doctor, and quality-gate failures.
|
|
246
307
|
|
|
247
|
-
### Item completion checklist (
|
|
308
|
+
### Item completion checklist (12-point gate)
|
|
248
309
|
|
|
249
|
-
An item in `test-list.md` may transition to `done` only when ALL of the following are satisfied
|
|
310
|
+
An item in `test-list.md` may transition to `done` only when ALL of the following are satisfied. For T1 rows, items 3, 5, 7 and 8 are satisfied by the confirmation covering the row's coherent group; they are never waived.
|
|
250
311
|
|
|
251
312
|
1. Corresponding `TDD-ID` has been selected and is in progress
|
|
252
|
-
2. A failing test was added first (test-first)
|
|
253
|
-
3. RED was observed — `qa-gatekeeper` confirmed
|
|
254
|
-
4. Minimal production code was written to make the test pass
|
|
255
|
-
5. GREEN was observed — `qa-gatekeeper` confirmed the test passes after implementation (watch it pass)
|
|
313
|
+
2. A failing test was added first (test-first) — **or**, on the _RED not observable_ path, the correct test was added first and proven falsifiable by mutation instead of by a natural failure
|
|
314
|
+
3. RED was observed — `qa-gatekeeper` confirmed an **admissible** failure: an assertion or expected-exception check inside the row's `Selector`, not a load or fixture error (`references/red-admissibility.md`), **or** the row carries falsifiability evidence per _RED not observable_
|
|
315
|
+
4. Minimal production code was written to make the test pass — **waived** on the _RED not observable_ path, where the `Satisfied-by` row already implements the predicate; do not manufacture a change to satisfy this item
|
|
316
|
+
5. GREEN was observed — `qa-gatekeeper` confirmed the test passes after implementation (watch it pass) **and** that the pass depends on this item's behaviour: `Oracle proof` records a production mutation that made the test fail again, or `equivalent-mutant` naming the weaker contract clause (`references/oracle-strength.md`). Exit code 0 alone does not distinguish a discriminating test from one that cannot fail
|
|
256
317
|
6. Refactor was performed and GREEN was re-confirmed after refactor
|
|
257
318
|
7. `completion-reviewer` returned PASS (spec / completion review gate)
|
|
258
319
|
8. `implementation-reviewer` returned PASS (code quality review gate)
|
|
259
320
|
9. UI-affecting items have prototype parity PASS from `product-surface-reviewer`
|
|
260
|
-
10. `test-list.md` Status and Evidence
|
|
261
|
-
11.
|
|
321
|
+
10. `test-list.md` Status is current and its Evidence cell's anchor resolves to a fresh per-item entry in `.qfai/evidence/implement-<spec-id>.md` (the cell is a pointer, not the payload — `references/execution-ledger.md#evidence-cell-contract`), and the item's four sub-agent observations (items 3, 5, 7, 8) all name the **same** revision (`references/evidence-revision.md`)
|
|
322
|
+
11. `.qfai/evidence/implement-<spec-id>.md` is appended with both reviewer verdicts after items 7-8 returned PASS
|
|
323
|
+
12. Checkpoint verification passed (see `#checkpoint-verification`). The **full** suite is required here only when the item sits on a checkpoint boundary; a row between boundaries satisfies this with the narrow relevant suite from Phase: Refactor step 2, which is also what items 6, 7 and 8 are evaluated against.
|
|
324
|
+
|
|
325
|
+
Sequencing note: the phase-authored part of `.qfai/evidence/implement-<spec-id>.md` (RED / GREEN /
|
|
326
|
+
Refactor commands and results) is written **before** items 7-8, because it is the evidence the
|
|
327
|
+
reviewers audit. The verdict fields are appended **after** items 7-8. A phase-authored evidence file
|
|
328
|
+
whose only gap is the verdict fields is NOT a blocking finding at review time — see
|
|
329
|
+
`Per-item evidence contract`.
|
|
330
|
+
|
|
331
|
+
### Review artifact layout (MUST)
|
|
332
|
+
|
|
333
|
+
Gate items 7-9 are evidence-bearing: reviewer verdicts must be written to a review pack, not left in
|
|
334
|
+
conversation. There is exactly **one** `.qfai/review/**` layout — `review-<17-digit-timestamp>/`
|
|
335
|
+
holding `review_request.md`, `R01_<reviewer-id>.md` (at least one) and `summary.json`. Do not nest
|
|
336
|
+
`<scope>/<layer>/attempt-NN/` directories: packs written there are invisible to `npx qfai validate`.
|
|
337
|
+
Each review round creates a new pack. Full schema and the `REVISE` -> `status: "REVISE"` mapping:
|
|
338
|
+
`references/review-artifact-layout.md`.
|
|
262
339
|
|
|
263
340
|
### Spec completion conditions
|
|
264
341
|
|
|
265
342
|
The skill may declare "this spec's implementation is complete" only when:
|
|
266
343
|
|
|
267
|
-
- All TC-\* from `06_Test-Cases.md` with applicable layer are present in `test-list.md`
|
|
268
|
-
-
|
|
344
|
+
- All TC-\* from `06_Test-Cases.md` with applicable layer are present in `test-list.md`. "Applicable layer" is decided by `.qfai/assistant/catalog/test-layers.md#layer-derivation-procedure-normative`
|
|
345
|
+
- Every `US-*` the spec declares has a `Layer = E2E` row whose `US-Refs` names it,
|
|
346
|
+
and every declared `CON-API-*` has a `Layer = API` row whose `CON-API-Refs`
|
|
347
|
+
names it. Without these rows an all-`done` ledger can sit alongside a
|
|
348
|
+
`QFAI-ATDD-111` / `QFAI-ATDD-113` hard gate at 0%- Each item reached `done` or valid `exception` (with DR-ID)
|
|
269
349
|
- 0 blocking reviewer issues remain
|
|
270
|
-
- Checkpoint verification passed
|
|
271
|
-
- No unresolved Change Request or waiver dependency exists
|
|
350
|
+
- Checkpoint verification passed at the spec-level boundary (see `#checkpoint-verification`)
|
|
351
|
+
- No unresolved Change Request or waiver dependency exists. The gate covers only the
|
|
352
|
+
`.qfai/decisions/CR-*.md` **in scope for this spec**; a CR confined to another spec never blocks
|
|
353
|
+
this one. An in-scope CR is **resolved** only when every condition in
|
|
354
|
+
`references/change-request-reset.md#when-an-in-scope-cr-counts-as-resolved` holds — `Status` is
|
|
355
|
+
`approved`, `rejected` or `superseded` (never `open`), the approval fields are populated,
|
|
356
|
+
`Resolution` records what was done, and when `Status` is `approved`, `Applied at` is populated —
|
|
357
|
+
approval alone does not release the gate. A CR failing any one of them, a half-filled record
|
|
358
|
+
included, is **unresolved** and blocks completion.
|
|
272
359
|
|
|
273
360
|
### Completion prohibition conditions
|
|
274
361
|
|
|
275
362
|
Completion MUST NOT be declared when any of the following are true:
|
|
276
363
|
|
|
277
|
-
- No RED fresh evidence exists for the item
|
|
364
|
+
- No RED fresh evidence exists for the item, and no falsifiability evidence replaces it
|
|
278
365
|
- No GREEN fresh evidence exists for the item
|
|
279
|
-
- Either reviewer (`completion-reviewer` or `implementation-reviewer`) has not been run or returned
|
|
280
|
-
-
|
|
366
|
+
- Either reviewer (`completion-reviewer` or `implementation-reviewer`) has not been run or returned REVISE
|
|
367
|
+
- `.qfai/evidence/implement-<spec-id>.md` does not exist, or does not record both reviewer verdicts for the item (this is the single blocking statement about the evidence file; its absence of _verdicts_ is never blocking before items 7-8)
|
|
368
|
+
- A `## Cross-spec obligations` entry in this spec's evidence file is still open — the change it names has not landed, or the blocked spec's obligation is still untested. A clean completion here would certify an obligation this run knowingly left unmet (`references/cross-spec-ownership.md`)
|
|
369
|
+
- Items with `todo`, `red`, `green`, `refactor`, or `review-fix` status still exist (for spec-level completion)
|
|
370
|
+
- Items with `exception` status still exist, **unless** the row's `DR-ID` names
|
|
371
|
+
a Decision Record explicitly recorded as a **user-approved accepted-risk
|
|
372
|
+
waiver** (a `TDDLIST-001` entry in `.qfai/waivers.yml`). An `exception` whose
|
|
373
|
+
DR only describes the anomaly is a parked defect, not a completed item.
|
|
281
374
|
- Parallel slices were used but integration verify has not been run post-merge
|
|
282
|
-
-
|
|
375
|
+
- A checkpoint boundary was reached (see `#checkpoint-verification`) but the verification command set was not executed, or any command in it exited non-zero — the last row a run completes is always a boundary, not the physical last row of the file, which is often already `done` and skipped, so every spec runs the full suite at least once
|
|
283
376
|
- `it.todo(...)` / `test.todo(...)` / `describe.todo(...)` stubs remain in any file covered by `validation.traceability.testFileGlobs` (`QFAI-TEST-001`). Implement the body or delete the stub — an opt-out via `validation.testStrategy.forbidTestTodoStubs: false` is permitted only with an accompanying waiver DR-ID.
|
|
284
377
|
|
|
285
378
|
## Evidence (MANDATORY)
|
|
@@ -290,45 +383,74 @@ Required sections:
|
|
|
290
383
|
|
|
291
384
|
- Objective
|
|
292
385
|
- Items processed (TDD-ID, TC-Refs, final status)
|
|
386
|
+
- **Per item, one `### TDD-NNNN` section** carrying the contract below — the single home for the RED/GREEN commands and output. The ledger's `Evidence` cell anchors here and holds only the one-word outcomes, because a GFM cell cannot hold a newline or a bare `|` (`references/execution-ledger.md#evidence-cell-contract`)
|
|
293
387
|
- Test results summary
|
|
294
388
|
- Exception items (if any) with DR-IDs
|
|
389
|
+
- `## Cross-spec obligations` (if any): per affected spec, the TDD-ID that forced the change, the blocked spec and its TDD-IDs, the file, the change required, the obligation left unverified, and the resolution (`re-reviewed` or a `CR-*`). Fields and the rule: `references/cross-spec-ownership.md`
|
|
295
390
|
- Commands executed
|
|
296
391
|
|
|
297
392
|
### Per-item evidence contract (fresh evidence required)
|
|
298
393
|
|
|
299
|
-
Each TDD item MUST have fresh evidence containing at minimum
|
|
394
|
+
Each TDD item MUST have fresh evidence containing at minimum the fields below. The contract has two
|
|
395
|
+
parts with different write points; the fields are the same, the sequencing is not.
|
|
396
|
+
|
|
397
|
+
**Phase-authored (written before the reviewer gate, items 7-8):**
|
|
300
398
|
|
|
301
399
|
- `TDD-ID` — the item identifier
|
|
302
|
-
- `TC-ref` — reference to the test case(s)
|
|
400
|
+
- `TC-ref` — reference to the test case(s). On a `Layer = E2E` row read `US-ref` (the row's `US-Refs`) instead, and on a `Layer = API` row read `CON-API-ref` (the row's `CON-API-Refs`): exactly one obligation reference is required, the one the row's `Layer` selects
|
|
401
|
+
- `Revision` — the state the observation was made against: `git rev-parse HEAD`, or `working-tree+<porcelain digest>` for an uncommitted tree. One per round block, and one for the refactor-verify pair (`references/evidence-revision.md`)
|
|
303
402
|
- `RED command` — the exact command executed to observe failure
|
|
304
|
-
- `RED result` — the failure output
|
|
403
|
+
- `RED result` — the failure output. Truncation is acceptable for the stack tail, never for the assertion message and its location: that is what demonstrates admissibility
|
|
404
|
+
- `RED failure mode` — `assertion` | `expected-error` | `falsifiability`. There is no admissible value for a load error (`references/red-admissibility.md`)
|
|
405
|
+
- **Exclusive alternative to the RED pair**: a row on the _RED not observable_ path carries
|
|
406
|
+
`Satisfied-by`, `Falsifiability command` and `Falsifiability result` in place of the two
|
|
407
|
+
RED fields above. Exactly one of the two forms must be present — never both, never
|
|
408
|
+
neither (`references/red-not-observable.md`).
|
|
305
409
|
- `GREEN command` — the exact command executed to observe success
|
|
306
410
|
- `GREEN result` — the success output
|
|
307
|
-
-
|
|
308
|
-
- `Refactor verify
|
|
309
|
-
- `
|
|
310
|
-
- `
|
|
411
|
+
- Each RED/GREEN cycle is one **round block** and every field above carries a `Round N:` prefix; numbering, the two rework paths and the full field list are in `references/round-evidence.md`
|
|
412
|
+
- `Refactor verify command` — the exact command re-executed after refactor. Written once for the item as a whole, so it takes no `Round N:` prefix
|
|
413
|
+
- `Refactor verify result` — the output confirming GREEN is maintained (likewise once per item)
|
|
414
|
+
- `Oracle proof` — the smallest production change that makes this item's test fail again, its command and its failing output, reverted immediately; or `equivalent-mutant` naming the contract clause weaker than the obligation. A row on the _RED not observable_ path satisfies this with its falsifiability fields (`references/oracle-strength.md`)
|
|
415
|
+
|
|
416
|
+
These exist _for_ the reviewers: they are the evidence items 7-8 audit. They MUST be present when a
|
|
417
|
+
review is requested.
|
|
418
|
+
|
|
419
|
+
**Gate-completed (appended after items 7-8 return PASS):**
|
|
420
|
+
|
|
421
|
+
- `Spec review` — completion-reviewer result (PASS or REVISE) with its `Reviewed revision`
|
|
422
|
+
- `Code quality review` — implementation-reviewer result (PASS or REVISE) with its `Reviewed revision`
|
|
311
423
|
- `Prototype parity` — product-surface-reviewer result for UI-affecting items (PASS or REVISE)
|
|
424
|
+
- `Checkpoint verification command` — the exact command set executed at the checkpoint boundary
|
|
425
|
+
- `Checkpoint verification result` — the outcome of that command set (PASS only when every command exits 0)
|
|
426
|
+
|
|
427
|
+
These record verdicts that do not exist until the reviews have run. A reviewer MUST NOT treat their
|
|
428
|
+
absence as a blocking gap during review — an evidence file complete in its phase-authored part and
|
|
429
|
+
missing only the verdict fields is the expected state at review time. It becomes blocking only at
|
|
430
|
+
the completion gate (see `Completion prohibition conditions`).
|
|
312
431
|
|
|
313
432
|
### Evidence hard rules
|
|
314
433
|
|
|
315
|
-
- Status-only evidence (e.g., "Status: PASS" with no command) is invalid and MUST be rejected
|
|
316
|
-
-
|
|
317
|
-
- Stale evidence from a previous run MUST NOT be reused to claim completion for a new cycle
|
|
318
|
-
-
|
|
434
|
+
- Status-only evidence (e.g., "Status: PASS" with no command) is invalid and MUST be rejected; both command and result are required, and "should pass" or "looks good" alone is not acceptable — `TDDLIST_EVIDENCE_STATUS_ONLY` (warning, waivable as `TDDLIST-004`: ledgers predating the check carry prose verdicts)
|
|
435
|
+
- Empty evidence entries are rejected: minimum evidence per TDD item must be met — `TDDLIST_EVIDENCE_EMPTY` (error)
|
|
436
|
+
- Stale evidence from a previous run MUST NOT be reused to claim completion for a new cycle. **Stale is mechanical**: evidence whose named `Revision` differs from the revision the item finally landed at (`references/evidence-revision.md`). **Reviewer obligation, not a machine gate** — why, and the full rules: `references/execution-ledger.md`.
|
|
437
|
+
- **Selective reporting of repeated runs of the same gate is invalid.** When a gate was run more than once, report every run in order — a clean rerun after a red one is an `environment/tooling` finding, not a pass (`.qfai/assistant/constitution/shared-skill-operating-baseline.md#nondeterministic-gates`)
|
|
438
|
+
|
|
439
|
+
## Checkpoint Verification
|
|
440
|
+
|
|
441
|
+
"Checkpoint verification" is the whole-repository regression check run at a checkpoint boundary. It
|
|
442
|
+
is what item 12 of the 12-point gate refers to and the only thing it refers to. A boundary is
|
|
443
|
+
reached **per item** (after all routed blocking reviewers return PASS, before `refactor` -> `done`)
|
|
444
|
+
and **per spec** (after the last ledger row is terminal). There is no "every N items" rule.
|
|
445
|
+
|
|
446
|
+
It PASSES only when **every** command in the verification command set exits 0; a partial run is not
|
|
447
|
+
a pass. The boundary definition, command set, pass criteria and evidence fields are in
|
|
448
|
+
`references/checkpoint-verification.md`.
|
|
319
449
|
|
|
320
450
|
## FINAL CHECKLIST (Check Last)
|
|
321
451
|
|
|
322
|
-
-
|
|
323
|
-
|
|
324
|
-
- [ ] Red phase: test was written and confirmed to fail.
|
|
325
|
-
- [ ] Green phase: minimal code was written and test confirmed to pass.
|
|
326
|
-
- [ ] Refactor phase: code improved with tests still passing.
|
|
327
|
-
- [ ] `test-list.md` statuses are accurate.
|
|
328
|
-
- [ ] No backward transitions occurred.
|
|
329
|
-
- [ ] Exception items have DR-IDs recorded.
|
|
330
|
-
- [ ] All tests pass.
|
|
331
|
-
- [ ] `qfai validate --profile tdd --fail-on error` passes with zero `QFAI-TEST-001` findings (no `it.todo` / `test.todo` / `describe.todo` stubs remain).
|
|
452
|
+
Work through `references/final-checklist.md` immediately before the completion message. Every box
|
|
453
|
+
must be ticked; a box that cannot be ticked is a reason not to declare completion.
|
|
332
454
|
|
|
333
455
|
## Completion Checklist (MUST)
|
|
334
456
|
|
|
@@ -343,7 +465,7 @@ Each TDD item MUST have fresh evidence containing at minimum:
|
|
|
343
465
|
When this skill is complete, provide a final user-facing completion message and enumerate all actionable next steps.
|
|
344
466
|
|
|
345
467
|
- Verify gates: `/qfai-verify`.
|
|
346
|
-
Action: run `qfai validate --profile tdd --fail-on error` for this skill, then `/qfai-verify` for full-scan approval.
|
|
468
|
+
Action: run `npx qfai validate --profile tdd --fail-on error` for this skill, then `/qfai-verify` for full-scan approval.
|
|
347
469
|
- Spec updates needed: `/qfai-sdd`.
|
|
348
470
|
Action: update spec artifacts if implementation revealed scope changes.
|
|
349
471
|
- Acceptance tests: `/qfai-atdd`.
|
|
@@ -372,6 +494,6 @@ A skill MAY narrow the auto-decide bucket (drop entries) but MUST NOT widen it.
|
|
|
372
494
|
|
|
373
495
|
project_memory:
|
|
374
496
|
|
|
375
|
-
- One TDD item at a time from test-list.md; status lifecycle is forward-only (todo → red → green → refactor → done); exception requires DR-ID.
|
|
376
|
-
- Fresh RED + GREEN command/result evidence is mandatory per item; status-only evidence (e.g. "Status: PASS") is rejected.
|
|
497
|
+
- One TDD item at a time from test-list.md by default; item-level parallelism inside one spec only when the Parallelization Policy technical gate and user consent both pass; status lifecycle is forward-only (todo → red → green → refactor → done) with one recorded re-entry, refactor → red on a qa-gatekeeper REVISE of the row's RED/GREEN evidence; exception requires DR-ID.
|
|
498
|
+
- Fresh RED + GREEN command/result evidence is mandatory per item, except on the _RED not observable_ path where `Satisfied-by` + falsifiability command/result replace the RED pair (exclusive alternative, never both); status-only evidence (e.g. "Status: PASS") is rejected.
|
|
377
499
|
- UI-affecting items require product-surface-reviewer prototype-parity PASS before the item can transition to done.
|