@thebackstoryis/engineering-with-ai 0.3.4-beta.6 → 0.3.5-beta.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/Docs/README.md CHANGED
@@ -8,7 +8,7 @@ For feature delivery, **`ewai-deliver` coordinates the full fourteen-stage workf
8
8
 
9
9
  If you want EWAI to act on an approved, exact pool of registered intents, read [governed autonomous intent delivery](autonomous-intent-delivery.md). It starts off, needs a named and bounded grant, and leaves human Build approval and Manual QA separate.
10
10
 
11
- Read how [EWAI now supports Jev and the OpenAI Decisions API](releases/0.3.4-beta.6.md) in the latest beta.
11
+ Read [what's new in beta 0.3.5-beta.1](releases/0.3.5-beta.1.md), the [stable 0.3.4 release notes](releases/0.3.4.md), or the [earlier Jev and OpenAI Decisions beta notes](releases/0.3.4-beta.6.md).
12
12
 
13
13
  ## New to EWAI?
14
14
 
@@ -35,3 +35,23 @@ Answering an autonomy question records a private owner response for the run. It
35
35
  Use `ewai autonomy status --project . --run RUN_ID --json` to inspect a run. `ewai autonomy pause|resume|cancel|recover --project . --run RUN_ID --expected-revision N --yes --json` applies a revision-bound control; a stale revision must be inspected again. For a human question, use the dashboard's guarded answer form or `ewai autonomy answer --project . --input PROJECT_RELATIVE_JSON --yes --json`, with a private project-local input file. Do not put private answers or credentials in chat or command arguments.
36
36
 
37
37
  The package tests exercise a tarball-installed consumer with preseeded offline dependencies; they do not prove a fresh registry dependency bootstrap or migration from every historical version. A controlled test simulates a lost response and verifies no repeated canonical effect. A separate installed-worker test substitutes only the provider boundary with a local child process; it does not establish live provider conformance. Human Manual QA, external acceptance, deployment and npm release remain separate decisions and evidence.
38
+
39
+ ## Continuing incomplete Build workers
40
+
41
+ An approved task can opt into continuation through `ralph_loop.allowed: true` and `max_iterations` between 1 and 100; a usual starting limit is 3. Put this bound in the task graph before Build approval. Each iteration is a provider invocation and consumes the supervisor's existing provider-attempt budget when autonomy owns the run. No new grant or larger budget is inferred.
42
+
43
+ If a worker exits successfully before its declared green check passes, the AFK conductor reads the actual result and starts another implementation turn in the same isolated worktree, with a reference to the captured failure output. It retains partial work and records each iteration. One task timeout covers the turns, verification, review and post-merge checks; it does not reset on each continuation. Disabled loops retain one implementation turn.
44
+
45
+ A worker summary or completion promise cannot finish the task. Tests, file scope, conductor-owned commits, fresh-context review and integration checks still apply. The loop stops on an `EWAI_BLOCKED:` worker decision, provider failure or timeout, unknown termination, scope breach, changed authority or identity, failed review, or exhausted limits. Pause prevents another implementation turn while preserving the stopped task's worktree; explicit resume follows the existing recovery/revalidation route and can create a new task attempt. Cancellation never assumes a child has stopped.
46
+
47
+ The ordinary interactive companion does not supervise individual chat-turn endings. Use the existing AFK executor for eligible approved Build work that should continue across worker turns. This capability does not autonomously complete every lifecycle phase: pre-Build drafts need review, and human approvals remain separate.
48
+
49
+ ## Lightweight delivery-claim validation
50
+
51
+ The existing fresh-context reviewer assesses every named completion claim separately and supplies a specific source/test/check citation and pass/fail result. The conductor rejects omitted, duplicate, unknown, unsupported or failed results before integration; a generic verdict or green exit is insufficient. No additional model call is needed. New task evidence uses v3 and attaches each claim to hashed review evidence rather than automatically citing the same green output. New task contracts declare claim_validation.required=true so removing required claim metadata cannot satisfy them with legacy evidence. Historical v1/v2 remains readable. This is model-assisted semantic review, not proof of production deployment or human acceptance; those require their separate evidence and authority.
52
+
53
+ ## Challenge, observe and recover
54
+
55
+ The existing reviewer tries to falsify every completion claim and records a concrete attempt/result, distinguishing source tracing from execution. v3 claim results include `challenge: {attempt, result}` and `observation: {evidence_type, expected, actual, evidence}`. Planned `completion_requirements` assign an evidence type to each exact completion name; behaviour is the default. A configured switch proves configuration, an executed command proves its own check, a reviewed artefact proves review, and delivered behaviour needs evidence of the claimed outcome. Missing, skipped or unavailable checks stay unverified. This remains model-assisted judgement with host-enforced structure and hashes. New contracts require claim_validation.version=3; historical v1/v2 evidence remains readable but cannot satisfy a v3 requirement.
56
+
57
+ AFK status returns an evidence-derived checkpoint: repository integration branches, declared/actual task branches and worktrees, verified completed claim labels, concrete next action and open questions. The projection is persisted atomically with run.json and recomputed on status reads. It never grants authority or certifies human acceptance. Tampered evidence removes completion from the refreshed view; uncertain worker termination directs recovery before resume. Use the existing resume preflight and canonical phase gates.
@@ -46,3 +46,9 @@ EWAI saves the intent and phase evidence in SPECS. Ask it to show what's complet
46
46
  If the outcome changes materially, revisit the intent and plan. A new Build approval may be needed. If only an evidence reference needs a correction, follow the [amendment guide](../completed-phase-evidence-amendments.md), which explains the narrow permitted operations.
47
47
 
48
48
  Try the [first-delivery tutorial](../tutorials/first-delivery.md). For exact commands and evidence contracts, use the [developer delivery guide](../developer-delivery-guide.md).
49
+
50
+ ## A worker turn is not task completion
51
+
52
+ The delivery workflow governs what may happen next. For eligible approved Build tasks, the AFK conductor also keeps execution moving between worker turns. A task can approve a finite continuation loop (`ralph_loop.allowed: true`, `max_iterations: 3`, for example). A normal worker exit followed by failing green verification then leads to another turn in the same worktree, within the original deadline and scope. Host verification decides when the task is ready for review. Failed reviews and genuine blockers require attention.
53
+
54
+ This is distinct from merely instructing the interactive companion to keep going: EWAI cannot restart that host's conversation after a chat turn ends. The conductor supplies the persistent execution path for approved Build tasks, while canonical phase gates continue to govern the wider delivery.
@@ -0,0 +1,7 @@
1
+ # What's new in this version
2
+
3
+ Version `0.3.4` adds optional **Jev and OpenAI Decisions** support to rank choices, compare AI recommendations, prioritise context and suggest personas or additional checks. Both start switched off. Configure them privately through the CLI or dashboard, compare advice in shadow mode or apply supported recommendations in active mode, and track usage, estimated cost and call limits.
4
+
5
+ **Concise answers and Grok Build.** EWAI puts the result or recommendation first, with clearer choices and less repetition. Grok Build is available as a coding companion, with dashboard and CLI controls for choosing your primary assistant and independent reviewers.
6
+
7
+ **Pick up the next ready piece of work.** Approve a list of work items and EWAI can select the next ready item using your recorded priorities. Set time and attempt limits, follow progress, and pause or recover a run. Build approval, required reviews, Manual QA and release decisions stay with you.
@@ -0,0 +1,7 @@
1
+ ## What's new in this version
2
+
3
+ Version `0.3.5-beta.1` adds **bounded Build continuation**. When an approved autonomous Build worker finishes before its required checks pass, EWAI can give it the failed-check feedback and continue in the same isolated worktree. Continuation is opt-in, with configured iteration and time limits, and respects pause, cancellation and blocked-work signals.
4
+
5
+ **Check the delivery claim.** The existing independent Build review now checks each declared completion claim against supporting evidence and attempts to challenge it. Reviews distinguish observed behaviour from configuration, command results and review results. Missing, failed or incomplete claim evidence prevents integration, without adding another review call.
6
+
7
+ **Recover with a clearer checkpoint.** Saved runs record verified completed work, branches and worktrees, the next action and unresolved questions. Status checks revalidate the evidence before listing work as complete. These changes support continued execution within approved Build boundaries; Manual QA, acceptance and release decisions remain human checkpoints.
package/README.md CHANGED
@@ -1,33 +1,18 @@
1
1
  # Engineering With AI
2
2
 
3
- ## New in 0.3.4 Beta 6 release
4
-
5
- **EWAI now supports Jev and the OpenAI Decisions API.**
6
-
7
- - **Rank options and recommend a choice.** Evaluate available approaches and return an ordered list with a recommended answer.
8
- - **Check AI recommendations.** Compare supported options returned by a larger language model with a second recommendation.
9
- - **Recommend suitable personas.** Select from personas available to your project, including core, project, personal and installed premium personas.
10
- - **Reduce larger-model calls.** Evaluate supported choices through focused decision calls instead of sending the same options to a larger language model.
11
- - **Prioritise context.** Help order supporting context for coding and review tasks.
12
- - **Support additional checks.** Assess supplied claim evidence, classify sanitised failures and recommend supplementary impact reviews and tests.
13
- - **Choose your provider.** Select Jev, OpenAI Decisions or Automatic. Automatic prefers configured and authorised OpenAI access, then Jev.
14
- - **Configure through the CLI or dashboard.** Save API keys privately, enable either integration and select the uses you want.
15
- - **Choose shadow or active mode.** Shadow shows recommendations for comparison. Active can apply supported recommendations to context ordering and persona selection.
16
- - **Track usage and estimated cost.** View decision-call usage and set call and input-token limits.
3
+ Engineering With AI (EWAI) helps you plan, build and review software with an AI assistant. It gives the assistant a shared record of the project, a delivery workflow and checks against your engineering standards. You keep control of the decisions and approve implementation before it starts.
17
4
 
18
- Both integrations are optional and start disabled. EWAI continues to work without either service.
5
+ [![Intent Studio in the EWAI dashboard: describing a feature, agreeing its acceptance criteria and engaging specialist personas before any code is written](https://www.conversationalcoding.dev/wp-content/uploads/sites/5/2026/09/intent-studio-full-be7758ffd224-1536x704.webp)](https://www.conversationalcoding.dev/engineering-with-ai-harness/?utm_source=readme&utm_medium=referral&utm_campaign=ewai)
19
6
 
20
- Install or update from the npm beta channel:
7
+ **Website, guides and books:** [conversationalcoding.dev](https://www.conversationalcoding.dev/engineering-with-ai-harness/?utm_source=readme&utm_medium=referral&utm_campaign=ewai) · [Online documentation](https://www.conversationalcoding.dev/engineering-with-ai-harness/docs/?utm_source=readme&utm_medium=referral&utm_campaign=ewai) · [Persona library](https://www.conversationalcoding.dev/personas/?utm_source=readme&utm_medium=referral&utm_campaign=ewai)
21
8
 
22
- ```sh
23
- npm install --save-dev @thebackstoryis/engineering-with-ai@beta
24
- ```
9
+ ## What's new in this version
25
10
 
26
- Engineering With AI (EWAI) helps you plan, build and review software with an AI assistant. It gives the assistant a shared record of the project, a delivery workflow and checks against your engineering standards. You keep control of the decisions and approve implementation before it starts.
11
+ Version `0.3.5-beta.1` adds **bounded Build continuation**. When an approved autonomous Build worker finishes before its required checks pass, EWAI can give it the failed-check feedback and continue in the same isolated worktree. Continuation is opt-in, with configured iteration and time limits, and respects pause, cancellation and blocked-work signals.
27
12
 
28
- [![Intent Studio in the EWAI dashboard: describing a feature, agreeing its acceptance criteria and engaging specialist personas before any code is written](https://www.conversationalcoding.dev/wp-content/uploads/sites/5/2026/09/intent-studio-full-be7758ffd224-1536x704.webp)](https://www.conversationalcoding.dev/engineering-with-ai-harness/?utm_source=readme&utm_medium=referral&utm_campaign=ewai)
13
+ **Check the delivery claim.** The existing independent Build review now checks each declared completion claim against supporting evidence and attempts to challenge it. Reviews distinguish observed behaviour from configuration, command results and review results. Missing, failed or incomplete claim evidence prevents integration, without adding another review call.
29
14
 
30
- **Website, guides and books:** [conversationalcoding.dev](https://www.conversationalcoding.dev/engineering-with-ai-harness/?utm_source=readme&utm_medium=referral&utm_campaign=ewai) · [Online documentation](https://www.conversationalcoding.dev/engineering-with-ai-harness/docs/?utm_source=readme&utm_medium=referral&utm_campaign=ewai) · [Persona library](https://www.conversationalcoding.dev/personas/?utm_source=readme&utm_medium=referral&utm_campaign=ewai)
15
+ **Recover with a clearer checkpoint.** Saved runs record verified completed work, branches and worktrees, the next action and unresolved questions. Status checks revalidate the evidence before listing work as complete. These changes support continued execution within approved Build boundaries; Manual QA, acceptance and release decisions remain human checkpoints.
31
16
 
32
17
  ## Quick start
33
18
 
@@ -41,27 +26,7 @@ ewai
41
26
 
42
27
  `ewai` opens your assistant and walks you through setting up the project. See [Install and start](#install-and-start) for details.
43
28
 
44
- Versions `0.3.2` and `0.3.3` update the package's links and documentation. The features below arrived in `0.3.1`.
45
-
46
- ## New in 0.3.1: concise answers and Grok Build
47
-
48
- Version `0.3.1` makes EWAI's guidance more concise and adds Grok Build as a coding provider.
49
-
50
- **Concise answers and guided decisions.** EWAI's managed instructions now prioritise correctness and usefulness, then brevity. Expect the result or recommendation first, with less repetition and routine narration. Decision requests explain the action, options and consequences, then give a recommendation with its reason and a suggested response when useful. When a request is unclear, EWAI leads with its recommended interpretation and states the assumptions that matter. The guidance is designed to cut unnecessary output; tool-result compaction is still planned, and token or cost savings haven't been measured yet. See [concise answers and guided decisions](Docs/context-management-and-token-efficiency.md#concise-answers-and-guided-decisions).
51
-
52
- **Grok Build and coding provider settings.** EWAI can open Grok Build as a native companion, install its skills and set up project MCP. **Configuration → Coding providers** in the dashboard, or `ewai providers` in a terminal, lets you choose a primary coding CLI, independent reviewers and an eligible pool for unattended work. Each CLI keeps its own model choice by default. Grok can run isolated coding, read-only review and restricted proposal workers. Each mode passes an offline conformance check before dispatch, and your xAI key is saved privately outside the project. Build approval, required review and Manual QA stay with you. See [provider settings](Docs/reference/cli-and-configuration.md#coding-provider-settings) and [Grok Build setup](Docs/operations/installation-updating-and-entitlements.md#grok-build).
53
-
54
- After updating, initialisation or the next check-in refreshes EWAI's managed instructions in `AGENTS.md` and `CLAUDE.md`, preserving project-authored guidance outside that block. Start a fresh host conversation after the refresh.
55
-
56
- ## EWAI can now pick up the next ready piece of work
57
-
58
- If you’ve prepared several work items, you can choose which ones EWAI is allowed to take on. EWAI checks what’s ready, uses the priorities you’ve recorded to pick the next item, and starts its delivery workflow. Once you’ve separately approved the Build, it can run the approved build tasks, their tests and a fresh review.
59
-
60
- This `0.3.0` release brings the capability out of beta. It adds dashboard and command-line controls to preview the work, approve the exact list, set time and attempt limits, follow progress, and pause, cancel or recover a run. New work isn’t added to the list automatically. EWAI stops when it needs a decision from you.
61
-
62
- The final whole-delivery test stage, Manual QA and release preparation still happen through the normal EWAI workflow. EWAI doesn’t approve or complete those steps for you.
63
-
64
- To use the release in a project without replacing a global installation, run:
29
+ To use EWAI in a project without replacing a global installation, run:
65
30
 
66
31
  ```bash
67
32
  npm install --save-dev @thebackstoryis/engineering-with-ai
@@ -70,7 +35,7 @@ npx ewai
70
35
 
71
36
  To install or update EWAI globally, run `npm install --global @thebackstoryis/engineering-with-ai`.
72
37
 
73
- ## Install and start
38
+ ### Install and start
74
39
 
75
40
  You'll need Node.js 22.5 or newer, npm, and a supported AI command-line tool such as Codex or Claude Code, installed and signed in.
76
41
 
@@ -88,12 +53,12 @@ It first asks where to keep your project's SPECS records, then helps you describ
88
53
 
89
54
  To return to the project, run `ewai` in the same folder again. It picks up the saved project state rather than starting setup from scratch. Keep your normal Git workflow for your application's code; EWAI itself is installed and updated through npm.
90
55
 
91
- If npm returns `E404`, follow [installation troubleshooting](Docs/operations/installation-updating-and-entitlements.md). That's a package-acquisition failure, not a premium-persona licence error. Don't paste licence keys into npm commands or chat.
56
+ If npm returns `E404`, follow [installation troubleshooting](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/operations/installation-updating-and-entitlements.md). That's a package-acquisition failure, not a premium-persona licence error. Don't paste licence keys into npm commands or chat.
92
57
 
93
- - [First session: set up a small project](Docs/tutorials/first-session.md)
94
- - [First delivery: work through a CSV export](Docs/tutorials/first-delivery.md)
95
- - [Already have a codebase?](Docs/existing-project-onboarding-guide.md) Archaeology can reconstruct missing documentation, but it's your choice whether to run it.
96
- - [Installation, updates and host options](Docs/operations/installation-updating-and-entitlements.md)
58
+ - [First session: set up a small project](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/tutorials/first-session.md)
59
+ - [First delivery: work through a CSV export](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/tutorials/first-delivery.md)
60
+ - [Already have a codebase?](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/existing-project-onboarding-guide.md) Archaeology can reconstruct missing documentation, but it's your choice whether to run it.
61
+ - [Installation, updates and host options](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/operations/installation-updating-and-entitlements.md)
97
62
 
98
63
  ## Start with a conversation
99
64
 
@@ -113,7 +78,7 @@ The supporting CLI commands are available for direct control, troubleshooting an
113
78
 
114
79
  **SPECS** means Scope, Purpose, Evidence, Constraints and Strategy. These readable project records hold the purpose, requirements, decisions and evidence the team has agreed. They remain useful outside an AI session.
115
80
 
116
- The dashboard runs locally and lets you inspect work and make supported choices. Its loopback address isn't a shared team website. The AI host does the guided work; the runtime records progress and checks the conditions for moving on. [How these parts fit together](Docs/explanation/core-concepts.md).
81
+ The dashboard runs locally and lets you inspect work and make supported choices. Its loopback address isn't a shared team website. The AI host does the guided work; the runtime records progress and checks the conditions for moving on. [How these parts fit together](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/explanation/core-concepts.md).
117
82
 
118
83
  Human judgement remains essential. A persona isn't a real stakeholder, green tests aren't human acceptance, and a delivery handoff isn't permission to deploy.
119
84
 
@@ -127,21 +92,21 @@ The source is published for transparency and review, but EWAI is not open
127
92
  source. You may not modify, repackage, redistribute, rebrand or commercially
128
93
  exploit the EWAI core without separate written permission from Backstory Group.
129
94
  The licence does not restrict the project content or output you create by using
130
- EWAI. Read the [Backstory Group Source-Available Licence](LICENSE) for the full
95
+ EWAI. Read the [Backstory Group Source-Available Licence](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/LICENSE) for the full
131
96
  terms.
132
97
 
133
98
  ## Choose what you need
134
99
 
135
100
  | When you want to… | Start here |
136
101
  | --- | --- |
137
- | Decide which work needs attention next | [Companion](Docs/context-aware-delivery-companion-user-guide.md) |
138
- | Describe a feature and agree its boundaries | [Intent Studio](Docs/guided-intent-workspace-guide.md) |
139
- | Understand how delivery moves through its stages | [The fourteen-stage workflow](Docs/explanation/delivery-workflow.md) |
140
- | Investigate dependencies before a change | [Source Map](Docs/repository-source-map-guide.md) and [Blast Radius](Docs/blast-radius-and-impact-routing-guide.md) |
141
- | Turn a meeting into reviewed project evidence | [Meeting evidence](Docs/meeting-evidence-user-guide.md) |
142
- | Check an implemented feature with a person | [Manual QA](Docs/quality/manual-qa-and-acceptance.md) |
143
- | Adapt the dashboard to your work | [Dashboard configuration](Docs/operations/dashboard-configuration.md) |
144
- | Use shared organisational guidance | [Blueprints](Docs/designing-organisation-blueprint-packs.md) and [rollout](Docs/organisation-rollout-guide.md) |
102
+ | Decide which work needs attention next | [Companion](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/context-aware-delivery-companion-user-guide.md) |
103
+ | Describe a feature and agree its boundaries | [Intent Studio](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/guided-intent-workspace-guide.md) |
104
+ | Understand how delivery moves through its stages | [The fourteen-stage workflow](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/explanation/delivery-workflow.md) |
105
+ | Investigate dependencies before a change | [Source Map](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/repository-source-map-guide.md) and [Blast Radius](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/blast-radius-and-impact-routing-guide.md) |
106
+ | Turn a meeting into reviewed project evidence | [Meeting evidence](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/meeting-evidence-user-guide.md) |
107
+ | Check an implemented feature with a person | [Manual QA](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/quality/manual-qa-and-acceptance.md) |
108
+ | Adapt the dashboard to your work | [Dashboard configuration](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/operations/dashboard-configuration.md) |
109
+ | Use shared organisational guidance | [Blueprints](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/designing-organisation-blueprint-packs.md) and [rollout](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/organisation-rollout-guide.md) |
145
110
 
146
111
  Start with the parts your project needs. Portfolio, Team Hub and other advanced dashboard views are optional; hiding a view doesn't disable mandatory project checks.
147
112
 
@@ -149,18 +114,18 @@ Start with the parts your project needs. Portfolio, Team Hub and other advanced
149
114
 
150
115
  The included personas support the normal workflow. You can also create project-specific personas and use your own personal library.
151
116
 
152
- Premium personas are optional specialist perspectives. If you have a subscription, [enter your key privately and install the pack](Docs/operations/premium-personas-setup.md) before the analysis you want it to support. Installing a persona doesn't give it authority to approve a requirement, bypass a check or speak for a real user.
117
+ Premium personas are optional specialist perspectives. If you have a subscription, [enter your key privately and install the pack](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/operations/premium-personas-setup.md) before the analysis you want it to support. Installing a persona doesn't give it authority to approve a requirement, bypass a check or speak for a real user.
153
118
 
154
119
  ## Go deeper
155
120
 
156
- - [User guides and learning routes](Docs/README.md)
157
- - [Complete guide catalogue](Docs/guide-catalogue.md)
158
- - [Capabilities and project layout](Docs/reference/capabilities-and-project-layout.md)
159
- - [Commands and configuration](Docs/reference/cli-and-configuration.md)
160
- - [Troubleshooting and recovery](Docs/operations/troubleshooting-and-recovery.md)
121
+ - [User guides and learning routes](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/README.md)
122
+ - [Complete guide catalogue](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/guide-catalogue.md)
123
+ - [Capabilities and project layout](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/reference/capabilities-and-project-layout.md)
124
+ - [Commands and configuration](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/reference/cli-and-configuration.md)
125
+ - [Troubleshooting and recovery](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/operations/troubleshooting-and-recovery.md)
161
126
  - [EWAI on the web: harness overview, online docs and changelog](https://www.conversationalcoding.dev/engineering-with-ai-harness/?utm_source=readme&utm_medium=referral&utm_campaign=ewai)
162
127
  - [Engineering With AI, the book behind the method](https://www.conversationalcoding.dev/books/?utm_source=readme&utm_medium=referral&utm_campaign=ewai)
163
128
 
164
129
  ## Contributing to EWAI
165
130
 
166
- If you're changing the harness itself, use the [contributor guide](Docs/maintainers/contributing.md) and [verification walkthroughs](Docs/maintainers/verification-walkthroughs.md). Those source-checkout and regression-test instructions aren't part of setting up your own application.
131
+ If you're changing the harness itself, use the [contributor guide](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/maintainers/contributing.md) and [verification walkthroughs](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/maintainers/verification-walkthroughs.md). Those source-checkout and regression-test instructions aren't part of setting up your own application.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@thebackstoryis/engineering-with-ai",
3
- "version": "0.3.4-beta.6",
3
+ "version": "0.3.5-beta.1",
4
4
  "description": "A human-centred, AI-augmented engineering pipeline",
5
5
  "license": "SEE LICENSE IN LICENSE",
6
6
  "author": "The Backstory Is",
@@ -77,6 +77,8 @@ Use guarded operations only:
77
77
 
78
78
  Never mutate a raw phase or status. A phase completes only with its passing EWAI gate ledger and fresh, hashed project evidence. Record meaningful progress events for the dashboard.
79
79
 
80
+ New task contracts must set `claim_validation.required=true` and `claim_validation.version=3`. Declare one `completion_requirements` entry per completion name with `evidence_type`: `behaviour`, `configuration`, `command-result` or `review-result`. Default evidence type for undeclared historical tasks is behaviour. Challenge every claim: try a counterexample and record `challenge.attempt` and `challenge.result`, stating whether executed or source-traced. Record `observation.evidence_type`, `expected`, `actual` and `evidence`; configuration does not establish delivered behaviour and a green command proves only its exercised assertions. Unknown/skipped outcomes cannot pass their obligation. During the existing fresh-context review, assess each `completion_evidence` claim individually against specific source, test assertions or observed check results. Emit one `COMPLETION_CHECK:` JSON line per exact name with `status` and `support`; do not infer all claims from a generic passing test. Unsupported, missing or failed claims block integration. New evidence uses `ewai.task-evidence/v3` and binds completion claims to captured review evidence. Historical v1/v2 evidence remains readable, but cannot satisfy a task requiring claim validation. This adds no separate model invocation; it is model-assisted evidence assessment, not Manual QA or release authority.
81
+
80
82
  Plan must produce `build-plan.md`, `destination.md`, the Claim Ledger, Plan Contract, and complete `task-graph.json`. The deterministic task-graph validator checks source records, claim and slice coverage, dependencies, cycles, write-set isolation, branch uniqueness, task contracts, waves, evidence paths, and review/merge ownership. Do not author its passing output manually.
81
83
 
82
84
  ## Persona-led test scenario integration
@@ -109,13 +111,15 @@ Start the intent-level active session with a stable `ownerId`. A different orche
109
111
 
110
112
  ## Unattended Build conductor
111
113
 
112
- When the user explicitly asks EWAI to continue approved Build work while they are away, use the AFK conductor instead of inventing a host-specific background loop:
114
+ When the user explicitly asks EWAI to continue approved Build work to completion without repeated prompts, or while they are away, use the AFK conductor for eligible tasks instead of inventing a host-specific background loop:
113
115
 
114
116
  1. Call `ewai_afk_preflight` and explain every blocker. Never weaken the task graph to make preflight pass.
115
117
  2. Confirm the user has already approved Build and understands the bounded scope, provider choice, maximum parallel tasks, and timeout.
116
118
  3. Call `ewai_afk_start`. Report the durable run ID and how to pause, resume, cancel, and inspect it; do not expose lease tokens or raw prompts.
117
119
  4. Use `ewai_afk_status` for updates. Translate semantic run/task states into human language rather than streaming model internals.
118
- 5. A blocked run requires attention; diagnose its preserved log/evidence before `ewai_afk_resume`. Never silently retry a stop condition, failed review, scope breach, merge conflict, or post-merge failure.
120
+ 5. For tasks expected to need several implementation turns, propose an explicit `ralph_loop.allowed: true` with a finite `max_iterations` from 1 to 100 (normally 3) before Build approval. Existing disabled contracts remain single-shot; never expand a recorded task or grant to obtain more attempts.
121
+ 6. The conductor can continue only a successfully exited, confirmed stopped implementation worker whose declared green check still fails. It keeps the same task worktree, captures each verification output, and rechecks provider/authority/control boundaries on every invocation. The task timeout is shared across its turns, verification, review and post-merge checks. A completion promise never replaces host verification.
122
+ 7. A blocked run requires attention; diagnose its preserved log/evidence before `ewai_afk_resume`. Never silently retry a worker-reported `EWAI_BLOCKED:` decision, provider error/timeout, unknown termination, stop condition, failed review, scope breach, merge conflict, or post-merge failure.
119
123
 
120
124
  The conductor is a local Build executor, not a replacement for the fourteen-stage harness. It detects simple and multi-repository topology from the configured `pipeline.yaml`; every task's `repo` must match a configured repository, and exactly one configured repository must contain the configured SPECS root to own durable evidence. The workspace root need not be a Git repository. It creates a missing integration branch from each selected repository's clean current branch, then creates a real task branch/worktree there, leases only ready AFK tasks, limits concurrency to the validated graph, requires a fresh-context review, merges into that repository's declared parent branch in graph order, and runs post-merge checks. It blocks detached heads and refuses to silently switch when the declared integration branch already exists elsewhere. Use `task-graph.json.repository_branches` when repositories have different parent branches. Completion means all eligible Build tasks were integrated; it does not complete the Build phase, Standards Sweep, Test Execute, external validation, Delivery, Manual QA, or Retro.
121
125
 
@@ -124,3 +128,15 @@ The conductor is a local Build executor, not a replacement for the fourteen-stag
124
128
  Translate control-plane work into warm, semantic progress such as “Validating the plan against project standards” or “Checking what changed since this work was shelved.” Do not make the user learn the command surface. Ask only decisions the evidence cannot answer.
125
129
 
126
130
  When blocked, retain the exact phase, evidence, questions, and recovery point. When complete, Retro must capture learning and route accepted improvements into project-owned SPECS assets.
131
+
132
+ ## Execution continuity
133
+
134
+ After an authorised action completes, inspect the saved next action and continue all remaining permitted work. A successful tool call, completed task, progress summary or model turn is not delivery completion. Do not end the interaction just to offer to continue work the owner has already authorised.
135
+
136
+ Stop for a real blocker, a required human decision, exhausted configured execution budget, cancellation, or the actual agreed completion boundary. Preserve the exact recovery point and explain the reason. Never interpret an earlier Build instruction as Manual QA, deployment, release or acceptance approval.
137
+
138
+ For eligible approved Build tasks, the conductor provides execution continuity outside individual worker turns. These instructions improve interactive behaviour but cannot restart a native host conversation after its model ends a turn; do not advertise guaranteed unattended completion for the ordinary interactive companion. Phase preparation remains draft-only and stops for its required human review.
139
+
140
+ ## Evidence-derived recovery checkpoint
141
+
142
+ Use the checkpoint returned by AFK status to present integration/task branches, evidence-validated completed claims, the exact next action and open questions. It is stored atomically in the run record and recomputed on status reads, so changed evidence is never accepted from a stale checkpoint. Reports and guarded action/preflight surfaces remain authoritative; checkpoint output never grants resume, Build, release or acceptance authority. Unknown worker termination requires recovery, and missing/tampered completion evidence requires diagnosis. Present the checkpoint at interruptions and handoffs rather than replaying raw logs.
@@ -1,6 +1,6 @@
1
1
  import { randomUUID } from 'node:crypto';
2
2
  import {
3
- copyFileSync, cpSync, existsSync, mkdirSync, readFileSync, readdirSync, realpathSync, rmSync, writeFileSync,
3
+ copyFileSync, cpSync, existsSync, mkdirSync, readFileSync, readdirSync, realpathSync, rmSync, statSync, writeFileSync,
4
4
  } from 'node:fs';
5
5
  import { dirname, relative, resolve } from 'node:path';
6
6
  import { execFileSync, spawn } from 'node:child_process';
@@ -8,7 +8,7 @@ import { fileURLToPath } from 'node:url';
8
8
  import { deliveryPaths, atomicJson, isWithin, sha256, now } from '../delivery-documents.mjs';
9
9
  import { projectPaths } from '../paths.mjs';
10
10
  import { loadProjectConfig, validationStatus } from '../project.mjs';
11
- import { validateTaskGraph, validateTaskReport } from '../task-graph.mjs';
11
+ import { validateTaskGraph, validateTaskReport, parseCompletionClaimReview } from '../task-graph.mjs';
12
12
  import { syncIntentIndex } from './intents.mjs';
13
13
  import {
14
14
  abandonExecutionLeasesForRun,
@@ -74,6 +74,49 @@ function requireRun(projectRoot, runId) {
74
74
  return { paths, run };
75
75
  }
76
76
 
77
+ export function deriveAfkCheckpoint(projectRoot, run) {
78
+ try { return deriveCheckpointFromEvidence(projectRoot,run); }
79
+ catch {
80
+ // A derived convenience view must never obstruct durable control writes.
81
+ return {schema:'ewai.afk-checkpoint/v1',runId:run.id,status:'unavailable',updatedAt:run.updatedAt,
82
+ repositories:[],tasks:[],done:[],openQuestions:['Recovery checkpoint unavailable: task or run evidence is malformed; inspect preserved records.'],
83
+ next:run.executionStopped!==true || run.unconfirmedExecution
84
+ ? {action:'confirm-worker-termination',instruction:'Establish owned worker termination before a human recovery decision.'}
85
+ : {action:'inspect-evidence',instruction:'Repair or reconcile malformed saved evidence through existing guarded operations.'}};
86
+ }
87
+ }
88
+
89
+ function deriveCheckpointFromEvidence(projectRoot, run) {
90
+ const checkpoint = {schema:'ewai.afk-checkpoint/v1',runId:run.id,status:run.status,updatedAt:run.updatedAt,
91
+ repositories:(run.topology?.repositories ?? []).map(repo=>({name:repo.name,integrationBranch:repo.parentBranch})),tasks:[],done:[],openQuestions:[],next:null};
92
+ let graph;
93
+ try { graph = readGraph(projectRoot,run.slug); if (!Array.isArray(graph.tasks)) throw new Error('Malformed task graph'); }
94
+ catch { graph=null; checkpoint.openQuestions.push('The task graph cannot be read; inspect saved delivery evidence.'); }
95
+ for (const task of graph?.tasks ?? []) {
96
+ const state=run.tasks?.[task.id] ?? {};
97
+ let report;
98
+ try { report=validateTaskReport(deliveryPaths(projectRoot,run.slug).deliveryRoot,task); }
99
+ catch { report={status:'incomplete'}; }
100
+ checkpoint.tasks.push({id:task.id,repository:task.repo,declaredBranch:task.branch,actualBranch:state.actualBranch ?? null,
101
+ worktree:state.worktree ?? null,status:state.status ?? 'pending',evidenceStatus:report.status});
102
+ if (report.status === 'complete') checkpoint.done.push({taskId:task.id,claims:task.completion_evidence ?? []});
103
+ else if (state.status === 'completed') checkpoint.openQuestions.push(`${task.id} recorded completion has missing or invalid evidence; inspect it before continuing.`);
104
+ if (state.error) checkpoint.openQuestions.push(`${task.id}: ${state.error}`);
105
+ }
106
+ if (['blocked','cancel-unknown'].includes(run.status)) for(const message of (run.messages ?? []).slice(-1)) checkpoint.openQuestions.push(message.message);
107
+ const stopped=run.executionStopped===true && !run.unconfirmedExecution && !(run.activeWorkers ?? []).length;
108
+ const orphaned = run.status==='running' && !processAlive(run.pid);
109
+ if (!stopped && (!['running','starting'].includes(run.status) || orphaned)) checkpoint.next={action:'confirm-worker-termination',instruction:'Establish that owned workers stopped; inspect preserved work before a human recovery decision.'};
110
+ else if (checkpoint.openQuestions.some(question=>question.includes('evidence') || question.includes('task graph'))) checkpoint.next={action:'inspect-evidence',instruction:'Resolve missing or changed task evidence before claiming completion or resuming.'};
111
+ else if (run.status==='cancelled') checkpoint.next={action:'cancelled',instruction:'Work is preserved; starting further work needs a new permitted action.'};
112
+ else if (run.status==='completed' && graph?.tasks?.length && checkpoint.done.length===graph.tasks.length) checkpoint.next={action:'complete-build-gate',instruction:'All task evidence validates; use the canonical Build gate and subsequent phases. Human acceptance is still separate.'};
113
+ else if (orphaned) checkpoint.next={action:'guarded-resume',instruction:'The saved conductor is no longer running; inspect preserved work and use the existing guarded recovery preflight.'};
114
+ else if (run.status==='paused') checkpoint.next={action:'guarded-resume',instruction:'Inspect the preserved checkpoint, then use the existing resume preflight; this checkpoint grants no authority.'};
115
+ else if (run.status==='blocked') checkpoint.next={action:'inspect-blocker',instruction:'Diagnose the recorded blocker and preserved logs; resume only through existing guarded recovery.'};
116
+ else checkpoint.next={action:'continue-bounded-work',instruction:'The conductor continues only permitted tasks under its saved controls; inspect status before any manual intervention.'};
117
+ return checkpoint;
118
+ }
119
+
77
120
  function saveRun(paths, run, resumeFrom = null) {
78
121
  if (existsSync(paths.runPath)) {
79
122
  const current = JSON.parse(readFileSync(paths.runPath, 'utf8'));
@@ -84,6 +127,7 @@ function saveRun(paths, run, resumeFrom = null) {
84
127
  }
85
128
  }
86
129
  run.updatedAt = now();
130
+ run.checkpoint = deriveAfkCheckpoint(paths.projectRoot,run);
87
131
  atomicJson(paths.runPath, run);
88
132
  return run;
89
133
  }
@@ -420,7 +464,7 @@ export function afkRunStatus(projectRoot, runId = '') {
420
464
  const root = resolve(projectRoot);
421
465
  if (runId) {
422
466
  const { run } = requireRun(root, runId);
423
- return { ...run, alive: processAlive(run.pid) };
467
+ return { ...run, checkpoint:deriveAfkCheckpoint(root,run), alive: processAlive(run.pid) };
424
468
  }
425
469
  const { runsRoot } = pathsFor(root);
426
470
  if (!existsSync(runsRoot)) return [];
@@ -428,7 +472,7 @@ export function afkRunStatus(projectRoot, runId = '') {
428
472
  .filter((entry) => entry.isDirectory() && existsSync(resolve(runsRoot, entry.name, 'run.json')))
429
473
  .map((entry) => JSON.parse(readFileSync(resolve(runsRoot, entry.name, 'run.json'), 'utf8')))
430
474
  .sort((left, right) => right.createdAt.localeCompare(left.createdAt))
431
- .map((run) => ({ ...run, alive: processAlive(run.pid) }));
475
+ .map((run) => ({ ...run, checkpoint:deriveAfkCheckpoint(root,run), alive: processAlive(run.pid) }));
432
476
  }
433
477
 
434
478
  export function pauseAfkRun(projectRoot, runId) {
@@ -509,7 +553,7 @@ export function prepareAfkContext(projectRoot, task, repository, options = {}) {
509
553
  const implementationCommit = String(options.implementationCommit ?? '').trim();
510
554
  const taskContent = JSON.stringify(task);
511
555
  let rules = review
512
- ? `TESTS_FIRST Read tests before implementation. Read cited standards from ${specsRoot}. Review implementation commit ${implementationCommit}. Check the diff, exact task contract, write_set, desired outcome, stop conditions, test evidence, correctness, standards, security and scope. Do not edit files. End with exactly one line: VERDICT: PASS or VERDICT: FAIL.`
556
+ ? `TESTS_FIRST Read tests before implementation. Read cited standards from ${specsRoot}. Review implementation commit ${implementationCommit}. Check the diff, exact task contract, write_set, desired outcome, stop conditions, test evidence, correctness, standards, security and scope. Do not edit files. For EVERY completion_evidence name, assess whether the claimed result is actually delivered. Cite specific test assertions, source symbols or observed check results that support it; a generic green exit is insufficient. Never infer deployment, release or human acceptance from code/tests. Emit one line per name: COMPLETION_CHECK: followed by JSON with name (exact approved label), status (pass or fail), and support (specific evidence and reasoning), challenge (object with attempt and result), and observation (object with evidence_type, expected, actual and evidence). Try to falsify each claim using a concrete counterexample, boundary, missing input, retry or state sequence; report the attempt and what it showed, including if you could only trace source rather than execute it. Observation must describe what was actually observed, not an intention or a configured setting. Evidence type must match completion_requirements for that name; default is behaviour. Configuration proves only configuration; a command-result proves only its check; behaviour requires the claimed outcome; review-result proves only a completed review. Keep skipped or unavailable checks unverified, never invent human acceptance or deployment. Missing or unsupported delivery must be fail. End with exactly one line: VERDICT: PASS or VERDICT: FAIL.`
513
557
  : `You are an EWAI Build worker in an isolated Git worktree. Read standards from ${specsRoot} and tests before implementation. Repository ${repository.name} (${repository.role}). Work only inside write_set. Run only allowed_commands. Stop on every stop_condition. Make the smallest code and test change. Do not create SPECS records, commit, merge, change delivery state or update phase status. AUTHORITY_NONE. Finish with files changed and commands run.`;
514
558
  if(options.evidenceCandidates?.length)rules += ' For this isolated Grok worker, read the complete source-standard and recorded-check sections supplied in this context. Host paths and Git history are inaccessible. Use the supplied exact commit diff for review. Do not execute commands: the conductor runs approved checks and owns commits. Read tests from the declared file snapshot before implementation.';
515
559
  const ruleMarkers = review
@@ -712,7 +756,7 @@ async function executeCommand(command, cwd, outputPath, timeoutMs, projectRoot,
712
756
  const result = await invokeControlled(projectRoot, run, task, { provider: 'command', command: '/bin/sh',
713
757
  args: ['-lc', command], cwd, prompt: '', timeoutMs, logPath: outputPath });
714
758
  if (result.cancelled || result.timedOut || requireRun(projectRoot, run.id).run.desiredState === 'cancelled') throw new Error('AFK command stopped before completion.');
715
- return { command, exitCode: result.exitCode, outputPath };
759
+ return { command, exitCode: result.exitCode, outputPath, outputSha256: sha256(readFileSync(outputPath)) };
716
760
  }
717
761
 
718
762
  function writeTaskEvidence(projectRoot, run, result, postMerge) {
@@ -722,14 +766,15 @@ function writeTaskEvidence(projectRoot, run, result, postMerge) {
722
766
  const delivery = deliveryPaths(projectRoot, run.slug);
723
767
  const evidencePath = resolve(delivery.deliveryRoot, task.evidence_path);
724
768
  const reportPath = resolve(delivery.deliveryRoot, task.report_path);
725
- const reviewRelative = `tasks/${task.id}/evidence/fresh-context-review.md`;
769
+ const reviewRelative = `tasks/${task.id}/evidence/fresh-context-review.json`;
726
770
  const reviewPath = resolve(delivery.deliveryRoot, reviewRelative);
727
771
  mkdirSync(dirname(reviewPath), { recursive: true });
728
- writeFileSync(reviewPath, `${review.output.trim()}\n`, 'utf8');
772
+ writeFileSync(reviewPath, JSON.stringify({schema:'ewai.claim-review/v1',reviews:review.reviewResults.map(item=>({reviewer:item.reviewer,output:item.output}))},null,2)+'\n', 'utf8');
729
773
  const commands = verifiedCommands.map((commandResult) => {
730
774
  const relativePath = `tasks/${task.id}/evidence/${commandResult.stage}.txt`;
731
775
  const destination = resolve(delivery.deliveryRoot, relativePath);
732
776
  mkdirSync(dirname(destination), { recursive: true });
777
+ if (sha256(readFileSync(commandResult.outputPath)) !== commandResult.outputSha256) throw new Error('Captured command output changed before acceptance.');
733
778
  copyFileSync(commandResult.outputPath, destination);
734
779
  return {
735
780
  stage: commandResult.stage,
@@ -740,9 +785,22 @@ function writeTaskEvidence(projectRoot, run, result, postMerge) {
740
785
  output_sha256: sha256(readFileSync(destination)),
741
786
  };
742
787
  });
743
- const green = commands.find((command) => command.stage === 'green');
788
+ const continuation = result.continuation && {
789
+ enabled: result.continuation.enabled,
790
+ max_iterations: result.continuation.maxIterations,
791
+ iterations: result.continuation.iterations.map(entry => {
792
+ const verification = entry.verification;
793
+ if (!verification) return { iteration: entry.iteration, provider_exit_code: entry.providerExitCode, execution_stopped: entry.executionStopped };
794
+ const outputPath = `tasks/${task.id}/evidence/green-iteration-${entry.iteration}.txt`;
795
+ const destination = resolve(delivery.deliveryRoot, outputPath);
796
+ if (sha256(readFileSync(verification.outputPath)) !== verification.outputSha256) throw new Error('Continuation verification output changed before acceptance.');
797
+ copyFileSync(verification.outputPath, destination);
798
+ return { iteration: entry.iteration, provider_exit_code: entry.providerExitCode, execution_stopped: entry.executionStopped,
799
+ verification: { command: verification.command, exit_code: verification.exitCode, output_path: outputPath, output_sha256: sha256(readFileSync(destination)) } };
800
+ }),
801
+ };
744
802
  const evidence = {
745
- schema: 'ewai.task-evidence/v1',
803
+ schema: 'ewai.task-evidence/v3',
746
804
  task_id: task.id,
747
805
  repository: repository.name,
748
806
  task_branch: task.branch,
@@ -751,19 +809,22 @@ function writeTaskEvidence(projectRoot, run, result, postMerge) {
751
809
  commit: implementationCommit,
752
810
  changed_files: changedFiles,
753
811
  commands,
812
+ ...(continuation ? { continuation } : {}),
754
813
  review: {
755
814
  fresh_context: true,
756
815
  tests_reviewed_first: true,
757
816
  status: 'pass',
758
817
  reviewer,
818
+ claim_results: review.claimResults,
759
819
  evidence_path: reviewRelative,
760
820
  evidence_sha256: sha256(readFileSync(reviewPath)),
761
821
  },
762
822
  completion_checks: (task.completion_evidence ?? []).map((name) => ({
763
823
  name,
764
824
  status: 'pass',
765
- evidence_path: green.output_path,
766
- evidence_sha256: green.output_sha256,
825
+ ...review.claimResults.find(result => result.name === name),
826
+ evidence_path: reviewRelative,
827
+ evidence_sha256: sha256(readFileSync(reviewPath)),
767
828
  })),
768
829
  post_merge_checks: postMerge,
769
830
  };
@@ -778,7 +839,7 @@ Implemented ${task.name} in repository ${repository.name} on branch \`${actualBr
778
839
  The declared red command failed as expected; the green and refactor verification commands passed. Exact outputs and hashes are recorded in the evidence sidecar.
779
840
 
780
841
  ## Feedback loops run
781
- The conductor ran the task's bounded verification commands, a fresh-context review, and ${postMerge.length} post-merge check(s).
842
+ The conductor ran ${result.continuation?.iterations.length ?? 1} implementation turn(s), the task's bounded verification commands, a fresh-context review, and ${postMerge.length} post-merge check(s). Per-turn verification outputs are retained in the evidence sidecar.
782
843
 
783
844
  ## Files changed
784
845
  ${changedFiles.map((file) => `- ${file}`).join('\n')}
@@ -798,7 +859,24 @@ No declared stop condition occurred.
798
859
  Automated task and integration evidence passed. Human QA remains governed by the later EWAI delivery gate.
799
860
  `;
800
861
  writeFileSync(reportPath, report, 'utf8');
801
- return [reportPath, evidencePath, reviewPath, ...commands.map(command => resolve(delivery.deliveryRoot, command.output_path))];
862
+ return [reportPath, evidencePath, reviewPath, ...commands.map(command => resolve(delivery.deliveryRoot, command.output_path)),
863
+ ...(continuation?.iterations ?? []).filter(entry => entry.verification).map(entry => resolve(delivery.deliveryRoot, entry.verification.output_path))];
864
+ }
865
+
866
+ // Only assistant/result messages can signal a blocker. Tool output may quote it.
867
+ function workerReportedBlocker(output) {
868
+ const blocked = value => typeof value === 'string' && /^EWAI_BLOCKED:/m.test(value);
869
+ if (blocked(output)) return true;
870
+ for (const line of String(output ?? '').split('\n')) {
871
+ let event;
872
+ try { event = JSON.parse(line); } catch { continue; }
873
+ if (event?.type === 'item.completed' && event.item?.type === 'agent_message' && blocked(event.item.text)) return true;
874
+ if (event?.type === 'result' && blocked(event.result)) return true;
875
+ const message = event?.type === 'assistant' ? event.message : event?.role === 'assistant' ? event : null;
876
+ if (message && (blocked(message.content) || Array.isArray(message.content)
877
+ && message.content.some(part => part?.type === 'text' && blocked(part.text)))) return true;
878
+ }
879
+ return false;
802
880
  }
803
881
 
804
882
  async function implementTask(projectRoot, run, task, provider, dependencies = {}) {
@@ -806,9 +884,20 @@ async function implementTask(projectRoot, run, task, provider, dependencies = {}
806
884
  const paths = pathsFor(projectRoot, run.id);
807
885
  const repository = repositoryForTask(run, task);
808
886
  const attempt = Number(run.tasks?.[task.id]?.attempt ?? 0) + 1;
887
+ const loop = task.ralph_loop?.allowed === true;
888
+ const maxIterations = loop ? task.ralph_loop.max_iterations : 1;
889
+ const monotonicNow = dependencies.monotonicNow ?? (() => performance.now());
890
+ const deadline = monotonicNow() + run.timeoutMs;
891
+ const remainingTaskMs = () => {
892
+ if (!loop) return run.timeoutMs;
893
+ const remaining = Math.floor(deadline - monotonicNow());
894
+ if (remaining < 1) throw new Error(`${task.id} exhausted its task deadline.`);
895
+ return remaining;
896
+ };
897
+ const continuation = { enabled: loop, maxIterations, iterations: [] };
809
898
  const worktree = resolve(paths.worktreesRoot, run.id, repository.name, `${task.id}-attempt-${attempt}`);
810
899
  const actualBranch = taskBranch(repository.root, run, task, attempt);
811
- const logRoot = resolve(paths.runRoot, 'tasks', task.id);
900
+ const logRoot = resolve(paths.runRoot, 'tasks', task.id, `attempt-${attempt}`);
812
901
  let lease = null;
813
902
  try {
814
903
  lease = acquireExecutionLease(projectRoot, run.intentId, {
@@ -832,50 +921,65 @@ async function implementTask(projectRoot, run, task, provider, dependencies = {}
832
921
  task.red_green_refactor.red_command,
833
922
  worktree,
834
923
  resolve(logRoot, 'red.txt'),
835
- run.timeoutMs, projectRoot, run, task,
924
+ remainingTaskMs(), projectRoot, run, task,
836
925
  );
837
926
  red.stage = 'red';
838
927
  if (red.exitCode === 0) throw new Error(`${task.id} red command already passes; the task contract is stale and must be reconciled.`);
839
- taskState(run, task, { status: 'implementing' });
840
- saveRun(paths, run);
841
- event(projectRoot, run, `Agent implementing ${task.id}.`, `${provider} · ${task.name}`);
842
- const implementationEvidence=provider==='grok'?prepareGrokTaskEvidence(projectRoot,worktree,task,{mode:'implementation',commands:[red]}):null;
843
- const invocation = buildProviderInvocation(provider, {
844
- cwd: worktree,
845
- prompt: taskPrompt(projectRoot, task, repository,implementationEvidence),
846
- grokEvidence:implementationEvidence,
847
- task,
848
- timeoutMs: run.timeoutMs,
849
- mode: 'implementation',
850
- codingPolicy:run.codingPolicy,policyDigest:run.codingPolicyDigest,policyRoot:projectRoot,
851
- });
852
- invocation.logPath = resolve(logRoot, 'implementation.log');
853
- const result = await invokeControlled(projectRoot, run, task, invocation, invoke);
854
- if (requireRun(projectRoot, run.id).run.desiredState === 'cancelled') {
855
- throw new Error('AFK run was cancelled while the task agent was active.');
856
- }
857
- if (git(worktree, ['rev-parse', 'HEAD']) !== baseHead) {
858
- throw new Error(`${task.id} created a commit; commits and merges belong to the conductor.`);
928
+ let green;
929
+ for (let iteration = 1; iteration <= maxIterations; iteration++) {
930
+ const control = requireRun(projectRoot, run.id).run;
931
+ if (control.desiredState === 'cancelled' || iteration > 1 && control.desiredState === 'paused') throw new Error(`AFK continuation stopped: ${control.desiredState}.`);
932
+ const timeoutMs = remainingTaskMs();
933
+ heartbeatExecutionLease(projectRoot, lease.id, {
934
+ token: lease.token, durationMs: Math.min(86_400_000, timeoutMs + 10 * 60 * 1000),
935
+ });
936
+ taskState(run, task, { status: 'implementing', continuation });
937
+ saveRun(paths, run);
938
+ event(projectRoot, run, `Agent implementing ${task.id}.`, `${provider} · iteration ${iteration}/${maxIterations}`);
939
+ const prior = continuation.iterations.at(-1)?.verification;
940
+ const feedback = prior
941
+ ? `\nCONTINUATION ${iteration}/${maxIterations}: the previous turn ended, but the task is incomplete. Continue from the files already in this worktree. The declared green command exited ${prior.exitCode}. Read its captured output at ${prior.outputPath}, repair only within write_set and recheck the declared tests. A summary or completion promise is not completion.`
942
+ : '';
943
+ if (prior && sha256(readFileSync(prior.outputPath)) !== prior.outputSha256) throw new Error('Continuation verification output changed; inspect preserved evidence.');
944
+ const continuationPrompt = (loop ? '\nWork to completion within this task. If a stop condition or owner decision prevents completion, end with EWAI_BLOCKED: followed by the reason.' : '') + feedback;
945
+ const implementationEvidence = provider === 'grok' ? prepareGrokTaskEvidence(projectRoot, worktree, task, { mode: 'implementation', commands: [red, ...(green ? [green] : [])] }) : null;
946
+ const invocation = buildProviderInvocation(provider, {
947
+ cwd: worktree, prompt: taskPrompt(projectRoot, task, repository, implementationEvidence) + continuationPrompt,
948
+ grokEvidence: implementationEvidence, task, timeoutMs: remainingTaskMs(), mode: 'implementation',
949
+ codingPolicy: run.codingPolicy, policyDigest: run.codingPolicyDigest, policyRoot: projectRoot,
950
+ });
951
+ invocation.logPath = resolve(logRoot, loop ? `implementation-iteration-${iteration}.log` : 'implementation.log');
952
+ const result = await invokeControlled(projectRoot, run, task, invocation, invoke);
953
+ const entry = { iteration, providerExitCode: result.exitCode, executionStopped: result.executionStopped === true };
954
+ continuation.iterations.push(entry);
955
+ taskState(run, task, { continuation });
956
+ atomicJson(resolve(logRoot, 'continuation.json'), continuation);
957
+ saveRun(paths, run);
958
+ if (requireRun(projectRoot, run.id).run.desiredState === 'cancelled' || result.cancelled) throw new Error('AFK run was cancelled while the task agent was active.');
959
+ if (git(worktree, ['rev-parse', 'HEAD']) !== baseHead) throw new Error(`${task.id} created a commit; commits and merges belong to the conductor.`);
960
+ if (result.timedOut) throw new Error(`${provider} exceeded the task timeout.`);
961
+ if (result.exitCode !== 0) throw new Error(`${provider} exited ${result.exitCode}; see ${relative(projectRoot, result.logPath)}.`);
962
+ // Check scope before running tests or giving another worker turn authority.
963
+ const outside = changedPaths(worktree).filter(path => !allowedPath(path, task));
964
+ if (outside.length) throw new Error(`${task.id} changed files outside its write_set: ${outside.join(', ')}.`);
965
+ if (loop && workerReportedBlocker(result.output)) throw new Error(`${task.id} worker reported a blocker; inspect its preserved log.`);
966
+ green = await executeCommand(task.red_green_refactor.green_command, worktree,
967
+ resolve(logRoot, loop ? `green-iteration-${iteration}.txt` : 'green.txt'), remainingTaskMs(), projectRoot, run, task);
968
+ green.stage = 'green';
969
+ entry.verification = green;
970
+ atomicJson(resolve(logRoot, 'continuation.json'), continuation);
971
+ taskState(run, task, { continuation });
972
+ saveRun(paths, run);
973
+ if (green.exitCode === 0) break;
974
+ if (!loop) throw new Error(`${task.id} green command failed after implementation.`);
975
+ if (iteration === maxIterations) throw new Error(`${task.id} reached its continuation iteration limit (${maxIterations}).`);
976
+ event(projectRoot, run, `Continuing incomplete task ${task.id}.`, `Green verification failed on iteration ${iteration}; remaining work stays in its isolated worktree.`);
859
977
  }
860
- if (result.timedOut) throw new Error(`${provider} exceeded the ${run.timeoutMs}ms task timeout.`);
861
- if (result.exitCode !== 0) throw new Error(`${provider} exited ${result.exitCode}; see ${relative(projectRoot, result.logPath)}.`);
862
- heartbeatExecutionLease(projectRoot, lease.id, {
863
- token: lease.token,
864
- durationMs: Math.min(86_400_000, run.timeoutMs + 10 * 60 * 1000),
865
- });
866
- const green = await executeCommand(
867
- task.red_green_refactor.green_command,
868
- worktree,
869
- resolve(logRoot, 'green.txt'),
870
- run.timeoutMs, projectRoot, run, task,
871
- );
872
- green.stage = 'green';
873
- if (green.exitCode !== 0) throw new Error(`${task.id} green command failed after implementation.`);
874
978
  const refactor = await executeCommand(
875
979
  task.red_green_refactor.green_command,
876
980
  worktree,
877
981
  resolve(logRoot, 'refactor.txt'),
878
- run.timeoutMs, projectRoot, run, task,
982
+ remainingTaskMs(), projectRoot, run, task,
879
983
  );
880
984
  refactor.stage = 'refactor';
881
985
  if (refactor.exitCode !== 0) throw new Error(`${task.id} refactor verification failed.`);
@@ -903,7 +1007,7 @@ async function implementTask(projectRoot, run, task, provider, dependencies = {}
903
1007
  prompt: reviewPrompt(projectRoot, task, implementationCommit,reviewEvidence),
904
1008
  grokEvidence:reviewEvidence,
905
1009
  task,
906
- timeoutMs: run.timeoutMs,
1010
+ timeoutMs: remainingTaskMs(),
907
1011
  mode: 'review',
908
1012
  codingPolicy:run.codingPolicy,policyDigest:run.codingPolicyDigest,policyRoot:projectRoot,
909
1013
  });
@@ -916,6 +1020,7 @@ async function implementTask(projectRoot, run, task, provider, dependencies = {}
916
1020
  if (review.timedOut || review.exitCode !== 0 || !/^VERDICT:\s*PASS\s*$/im.test(review.output)) {
917
1021
  throw new Error(`Fresh-context review did not pass; see ${relative(projectRoot, review.logPath)}.`);
918
1022
  }
1023
+ review.claimResults = parseCompletionClaimReview(review.output, task.completion_evidence ?? [], {observed:true,requirements:task.completion_requirements ?? []});
919
1024
  authorityChecks.get(run)?.({ stage: 'accept', mode: 'review' });
920
1025
  if (git(worktree, ['rev-parse', 'HEAD']) !== implementationCommit || gitDirty(worktree)) {
921
1026
  throw new Error('Fresh-context review changed the inspected implementation.');
@@ -927,6 +1032,7 @@ async function implementTask(projectRoot, run, task, provider, dependencies = {}
927
1032
  token: lease.token,
928
1033
  durationMs: Math.min(86_400_000, run.timeoutMs + 10 * 60 * 1000),
929
1034
  });
1035
+ review.reviewResults = reviewResults;
930
1036
  const branchHead = git(worktree, ['rev-parse', 'HEAD']);
931
1037
  const reviewerIdentity = reviewers.map(reviewer=>`${reviewer}/${reviewer === provider ? 'fresh-session' : 'independent-cli'}`).join(',');
932
1038
  taskState(run, task, { status: 'ready-to-integrate', implementationCommit, branchHead, reviewer: reviewerIdentity });
@@ -934,7 +1040,7 @@ async function implementTask(projectRoot, run, task, provider, dependencies = {}
934
1040
  return {
935
1041
  task, lease, worktree, branchHead, implementationCommit, review,
936
1042
  reviewer: reviewerIdentity, verifiedCommands: [red, green, refactor], changedFiles: changed,
937
- actualBranch, repository,
1043
+ actualBranch, repository, continuation, remainingTaskMs,
938
1044
  };
939
1045
  } catch (error) {
940
1046
  const control = requireRun(projectRoot, run.id).run;
@@ -960,6 +1066,7 @@ async function integrateTask(projectRoot, run, result) {
960
1066
  const specsRepository = run.topology.repositories.find((candidate) => candidate.name === run.topology.specsRepository);
961
1067
  if (!specsRepository) throw new Error('AFK run has no configured SPECS repository.');
962
1068
  const paths = pathsFor(projectRoot, run.id);
1069
+ result.remainingTaskMs?.();
963
1070
  taskState(run, task, { status: 'integrating' });
964
1071
  saveRun(paths, run);
965
1072
  event(projectRoot, run, `Integrating ${task.id}.`, 'The orchestrator is merging sequentially and running post-merge checks.');
@@ -982,7 +1089,7 @@ async function integrateTask(projectRoot, run, result) {
982
1089
  const untracked = git(root, ['ls-files', '--others', '--exclude-standard', '-z']).split('\0').filter(Boolean);
983
1090
  if (git(root, ['rev-parse', 'HEAD']) !== baseline.head || git(root, ['branch', '--show-current']) !== baseline.branch
984
1091
  || stagedDelta.some(path => !staged || !owned(path) || sha256(execFileSync('git', ['show', `:${path}`],
985
- { cwd: root, stdio: ['ignore', 'pipe', 'pipe'] })) !== ownedEvidence.get(resolve(root, path)))
1092
+ { cwd: root, stdio: ['ignore', 'pipe', 'pipe'], maxBuffer: Math.max(1024 * 1024, statSync(resolve(root, path)).size + 1024) })) !== ownedEvidence.get(resolve(root, path)))
986
1093
  || [...workingDelta, ...untracked].some(path => !owned(path))) {
987
1094
  integrationConflict = true;
988
1095
  throw new Error('Repository changed outside the owned integration; preserve the checkout for review.');
@@ -1008,7 +1115,7 @@ async function integrateTask(projectRoot, run, result) {
1008
1115
  const postMerge = [];
1009
1116
  for (const [index, command] of (task.merge?.post_merge_checks ?? []).entries()) {
1010
1117
  const outputPath = resolve(delivery.deliveryRoot, 'tasks', task.id, 'evidence', `post-merge-${index + 1}.txt`);
1011
- const check = await executeCommand(command, repository.root, outputPath, run.timeoutMs, projectRoot, run, task);
1118
+ const check = await executeCommand(command, repository.root, outputPath, result.remainingTaskMs?.() ?? run.timeoutMs, projectRoot, run, task);
1012
1119
  const evidencePath = relative(delivery.deliveryRoot, outputPath).replaceAll('\\', '/');
1013
1120
  postMerge.push({ command, exit_code: check.exitCode, output_path: evidencePath, output_sha256: sha256(readFileSync(outputPath)) });
1014
1121
  ownedEvidence.set(realpathSync(outputPath), postMerge.at(-1).output_sha256);
@@ -1145,14 +1252,15 @@ export async function executeAfkRun(projectRoot, runId, dependencies = {}) {
1145
1252
  } catch (error) {
1146
1253
  const control = requireRun(paths.projectRoot, run.id).run;
1147
1254
  run.desiredState = control.desiredState;
1148
- run.status = control.desiredState === 'cancelled' ? run.unconfirmedExecution ? 'cancel-unknown' : 'cancelled' : 'blocked';
1255
+ run.status = control.desiredState === 'cancelled' ? run.unconfirmedExecution ? 'cancel-unknown' : 'cancelled'
1256
+ : error.message === 'AFK continuation stopped: paused.' && !run.unconfirmedExecution ? 'paused' : 'blocked';
1149
1257
  if (control.desiredState === 'cancelled') run.cancellation = { ...control.cancellation,
1150
1258
  status: run.unconfirmedExecution ? 'unknown' : 'confirmed', settledAt: now() };
1151
1259
  run.pid = null;
1152
1260
  run.messages.push({ at: now(), message: error.message });
1153
1261
  run.executionStopped = !run.unconfirmedExecution;
1154
1262
  if (!run.unconfirmedExecution) abandonExecutionLeasesForRun(paths.projectRoot, run.id, { outcome: 'conductor-blocked' });
1155
- event(paths.projectRoot, run, 'Unattended Build is blocked.', error.message, 'blocked', true);
1263
+ event(paths.projectRoot, run, run.status === 'paused' ? 'Unattended Build safely paused.' : 'Unattended Build is blocked.', error.message, run.status === 'paused' ? 'progress' : 'blocked', run.status !== 'paused');
1156
1264
  saveRun(paths, run);
1157
1265
  return run;
1158
1266
  } finally { authorityChecks.delete(run); }
@@ -152,6 +152,7 @@ export function invokeProvider(invocation, options = {}) {
152
152
  const maxCapture = Number(options.maxCaptureBytes ?? 1024 * 1024);
153
153
 
154
154
  return new Promise((resolvePromise, reject) => {
155
+ output.once('error', reject);
155
156
  const child = spawn(invocation.command, invocation.args, {
156
157
  cwd: invocation.cwd,
157
158
  env: options.env ?? process.env,
@@ -204,8 +205,9 @@ export function invokeProvider(invocation, options = {}) {
204
205
  try { process.kill(-child.pid, 0); kill('SIGKILL'); }
205
206
  catch (error) { processGroupStopped = error.code === 'ESRCH'; }
206
207
  }
207
- output.end();
208
- resolvePromise({
208
+ // Return only after the captured log has finished flushing. Callers hash
209
+ // this file immediately; a process exit alone does not settle its stream.
210
+ output.end(() => resolvePromise({
209
211
  provider: invocation.provider,
210
212
  exitCode: exitCode ?? -1,
211
213
  signal: signal ?? '',
@@ -223,7 +225,7 @@ export function invokeProvider(invocation, options = {}) {
223
225
  invocation.provider,
224
226
  options.providerUsage ?? invocation.providerUsage,
225
227
  ),
226
- });
228
+ }));
227
229
  });
228
230
  options.signal?.addEventListener('abort', cancel, { once: true });
229
231
  if (options.signal?.aborted) cancel();
@@ -110,6 +110,41 @@ function evidenceFile(root, configuredPath, expectedHash, label, errors) {
110
110
  }
111
111
  }
112
112
 
113
+ // Semantic support is assessed by the existing fresh reviewer; the host enforces
114
+ // complete, explicit results rather than inferring every claim from one green exit.
115
+ export function validateCompletionClaims(results, names, options = {}) {
116
+ const errors = [], expected = new Set(names), seen = new Set();
117
+ if (!Array.isArray(results)) return ['completion claim validation results are missing'];
118
+ for (const result of results) {
119
+ if (!result || !expected.has(result.name)) errors.push('completion claim validation contains an unknown claim');
120
+ if (seen.has(result?.name)) errors.push('completion claim validation contains a duplicate claim');
121
+ seen.add(result?.name);
122
+ if (result?.status !== 'pass' || !text(result?.support)) errors.push('completion claim validation requires a passing result with specific supporting evidence');
123
+ }
124
+ if (options.observed) for (const result of results) {
125
+ const requiredType = options.requirements?.find(item => item.name === result?.name)?.evidence_type ?? 'behaviour';
126
+ const observation = result?.observation, challenge = result?.challenge;
127
+ if (!['behaviour','configuration','command-result','review-result'].includes(requiredType) || observation?.evidence_type !== requiredType) errors.push('completion claim evidence type does not match its planned obligation');
128
+ if (!['expected','actual','evidence'].every(key => typeof observation?.[key] === 'string' && observation[key].trim())) errors.push('completion claim requires expected and actual observed result with evidence');
129
+ if (!['attempt','result'].every(key => typeof challenge?.[key] === 'string' && challenge[key].trim())) errors.push('completion claim requires a concrete challenge attempt and result');
130
+ }
131
+ for (const name of expected) if (!seen.has(name)) errors.push(`completion claim validation is missing: ${name}`);
132
+ return errors;
133
+ }
134
+
135
+ export function parseCompletionClaimReview(output, names, options = {}) {
136
+ const results = [];
137
+ for (const line of String(output).split(/\r?\n/)) {
138
+ const match = /^COMPLETION_CHECK:\s*(.+)$/.exec(line.trim());
139
+ if (!match) continue;
140
+ try { results.push(JSON.parse(match[1])); }
141
+ catch { throw new Error('Completion claim validation contains malformed JSON.'); }
142
+ }
143
+ const errors = validateCompletionClaims(results, names, options);
144
+ if (errors.length) throw new Error(errors.join('; '));
145
+ return results;
146
+ }
147
+
113
148
  export function validateTaskEvidence(deliveryRoot, task) {
114
149
  const root = resolve(deliveryRoot);
115
150
  const configured = text(task?.evidence_path);
@@ -129,7 +164,12 @@ export function validateTaskEvidence(deliveryRoot, task) {
129
164
  errors.push(`structured task evidence is not valid JSON: ${error.message}`);
130
165
  }
131
166
  if (evidence) {
132
- if (evidence.schema !== 'ewai.task-evidence/v1') errors.push('schema must be ewai.task-evidence/v1');
167
+ if (!['ewai.task-evidence/v1', 'ewai.task-evidence/v2', 'ewai.task-evidence/v3'].includes(evidence.schema)) errors.push('unsupported task evidence schema');
168
+ const observedValidation = evidence.schema === 'ewai.task-evidence/v3' || task.claim_validation?.version === 3;
169
+ const claimOptions = {observed:observedValidation,requirements:task.completion_requirements ?? []};
170
+ if (task.claim_validation?.version === 3 && evidence.schema !== 'ewai.task-evidence/v3') errors.push('required observed claim validation cannot use earlier evidence');
171
+ const claimValidation = ['ewai.task-evidence/v2','ewai.task-evidence/v3'].includes(evidence.schema) || task.claim_validation?.required === true;
172
+ if (task.claim_validation?.required === true && !['ewai.task-evidence/v2','ewai.task-evidence/v3'].includes(evidence.schema)) errors.push('required completion claim validation cannot use legacy evidence');
133
173
  if (evidence.task_id !== task.id) errors.push(`task_id must be ${task.id}`);
134
174
  if (text(evidence.task_branch)) {
135
175
  if (evidence.task_branch !== task.branch) errors.push(`task_branch must be ${task.branch}`);
@@ -166,6 +206,24 @@ export function validateTaskEvidence(deliveryRoot, task) {
166
206
  }
167
207
  evidenceFile(root, command.output_path, command.output_sha256, `${command.stage || 'command'} output`, errors);
168
208
  }
209
+ if (evidence.continuation) {
210
+ const history = evidence.continuation;
211
+ const enabled = task.ralph_loop?.allowed === true;
212
+ const limit = enabled ? task.ralph_loop.max_iterations : 1;
213
+ if (history.enabled !== enabled || history.max_iterations !== limit) errors.push('continuation budget must match the task contract');
214
+ const iterations = history.iterations;
215
+ if (!Array.isArray(iterations) || !iterations.length || iterations.length > limit) {
216
+ errors.push('continuation history must be non-empty and within its iteration budget');
217
+ } else {
218
+ for (const [index, iteration] of iterations.entries()) {
219
+ const check = iteration?.verification;
220
+ if (iteration?.iteration !== index + 1 || iteration?.provider_exit_code !== 0 || iteration?.execution_stopped !== true) errors.push('continuation iteration must record an ordered, successful, confirmed stopped worker');
221
+ if (!check || check.command !== task.red_green_refactor?.green_command || !Number.isInteger(check.exit_code)
222
+ || (index === iterations.length - 1 ? check.exit_code !== 0 : check.exit_code === 0)) errors.push('continuation verification must retain failures before the final passing green');
223
+ if (check) evidenceFile(root, check.output_path, check.output_sha256, 'continuation verification output', errors);
224
+ }
225
+ }
226
+ }
169
227
  const review = evidence.review ?? {};
170
228
  if (review.fresh_context !== true) errors.push('review.fresh_context must be true');
171
229
  if (review.tests_reviewed_first !== true) errors.push('review.tests_reviewed_first must be true');
@@ -174,10 +232,34 @@ export function validateTaskEvidence(deliveryRoot, task) {
174
232
  evidenceFile(root, review.evidence_path, review.evidence_sha256, 'review evidence', errors);
175
233
  const checks = Array.isArray(evidence.completion_checks) ? evidence.completion_checks : [];
176
234
  const checksByName = new Map(checks.map((check) => [text(check.name), check]));
235
+ if (claimValidation) {
236
+ if (!errors.length) {
237
+ try {
238
+ const captured = JSON.parse(readFileSync(resolve(root, review.evidence_path), 'utf8'));
239
+ if (captured.schema !== 'ewai.claim-review/v1' || !Array.isArray(captured.reviews) || !captured.reviews.length) throw new Error('captured claim reviews are missing');
240
+ const identities = String(review.reviewer).split(',').map(value => value.trim().split('/')[0]);
241
+ if (captured.reviews.length !== identities.length || new Set(captured.reviews.map(item => item.reviewer)).size !== identities.length) throw new Error('captured claim reviewer coverage differs');
242
+ let recorded;
243
+ for (const item of captured.reviews) {
244
+ if (!identities.includes(item.reviewer)) throw new Error('unknown captured claim reviewer');
245
+ recorded = parseCompletionClaimReview(item.output, list(task.completion_evidence), claimOptions);
246
+ }
247
+ if (JSON.stringify(recorded) !== JSON.stringify(review.claim_results)) throw new Error('claim results differ from captured review');
248
+ } catch (error) { errors.push(`completion claim validation differs from captured review: ${error.message}`); }
249
+ }
250
+ errors.push(...validateCompletionClaims(review.claim_results, list(task.completion_evidence), claimOptions));
251
+ if (checks.length !== list(task.completion_evidence).length || checksByName.size !== checks.length) errors.push('completion claim validation must cover exactly the planned checks');
252
+ }
177
253
  for (const planned of list(task.completion_evidence)) {
178
254
  const check = checksByName.get(planned);
179
255
  if (!check || check.status !== 'pass') errors.push(`planned completion check has no passing evidence: ${planned}`);
180
- else evidenceFile(root, check.evidence_path, check.evidence_sha256, `completion check ${planned}`, errors);
256
+ else {
257
+ evidenceFile(root, check.evidence_path, check.evidence_sha256, `completion check ${planned}`, errors);
258
+ if (claimValidation) {
259
+ const result = Array.isArray(review.claim_results) ? review.claim_results.find(item => item?.name === planned) : null;
260
+ if (!result || check.support !== result.support || observedValidation && (JSON.stringify(check.observation) !== JSON.stringify(result.observation) || JSON.stringify(check.challenge) !== JSON.stringify(result.challenge)) || check.evidence_path !== review.evidence_path || check.evidence_sha256 !== review.evidence_sha256) errors.push(`completion claim validation is not bound to its supporting review: ${planned}`);
261
+ }
262
+ }
181
263
  }
182
264
  }
183
265
  }
@@ -412,8 +494,8 @@ export function validateTaskGraph(deliveryRoot, options = {}) {
412
494
  if (!list(task.completion_evidence).length) add('afk-completion-evidence-missing', `${id} requires completion evidence.`, id);
413
495
  }
414
496
  if (task?.ralph_loop?.allowed === true) {
415
- if (!Number.isInteger(task.ralph_loop.max_iterations) || task.ralph_loop.max_iterations < 1) {
416
- add('ralph-loop-unbounded', `${id} Ralph loop requires a finite positive max_iterations.`, id);
497
+ if (!Number.isInteger(task.ralph_loop.max_iterations) || task.ralph_loop.max_iterations < 1 || task.ralph_loop.max_iterations > 100) {
498
+ add('ralph-loop-unbounded', `${id} Ralph loop requires max_iterations between 1 and 100.`, id);
417
499
  }
418
500
  if (!text(task.ralph_loop.completion_promise)) add('ralph-loop-promise-missing', `${id} Ralph loop requires a completion promise.`, id);
419
501
  }
@@ -421,6 +503,16 @@ export function validateTaskGraph(deliveryRoot, options = {}) {
421
503
  if (task?.review?.fresh_context_required !== true) add('fresh-review-required', `${id} requires fresh-context review.`, id);
422
504
  }
423
505
 
506
+ for (const task of tasks) {
507
+ if (task.completion_requirements !== undefined) {
508
+ const requirements = task.completion_requirements;
509
+ const names = Array.isArray(requirements) ? requirements.map(item => item?.name) : [];
510
+ if (!Array.isArray(requirements) || names.length !== list(task.completion_evidence).length || new Set(names).size !== names.length || names.some(name => !list(task.completion_evidence).includes(name))
511
+ || requirements.some(item => !['behaviour','configuration','command-result','review-result'].includes(item?.evidence_type))) add('completion-requirements-invalid', `${task.id} requires one valid evidence type per completion claim.`, task.id);
512
+ }
513
+ if (task.claim_validation?.version !== undefined && task.claim_validation.version !== 3) add('claim-validation-version-invalid', `${task.id} declares an unsupported claim validation version.`, task.id);
514
+ }
515
+
424
516
  const ledgerClaims = recordsFrom(claimLedger, ['implementation_claims', 'claims']);
425
517
  const actionableClaims = new Set(ledgerClaims.filter(actionableClaim).map(recordId).filter(Boolean));
426
518
  if (!ledgerClaims.length) add('claim-ledger-empty', 'The Claim Ledger must contain implementation claims.', graph.source?.claim_ledger);