@thebackstoryis/engineering-with-ai 0.3.4-beta.6 → 0.3.5-beta.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/Docs/README.md +1 -1
- package/Docs/autonomous-intent-delivery.md +20 -0
- package/Docs/explanation/delivery-workflow.md +6 -0
- package/Docs/releases/0.3.4.md +7 -0
- package/Docs/releases/0.3.5-beta.1.md +7 -0
- package/README.md +31 -66
- package/package.json +1 -1
- package/skills-src/ewai-deliver/SKILL.md +18 -2
- package/src/runtime/afk-conductor.mjs +165 -57
- package/src/runtime/provider-adapters.mjs +5 -3
- package/src/task-graph.mjs +96 -4
package/Docs/README.md
CHANGED
|
@@ -8,7 +8,7 @@ For feature delivery, **`ewai-deliver` coordinates the full fourteen-stage workf
|
|
|
8
8
|
|
|
9
9
|
If you want EWAI to act on an approved, exact pool of registered intents, read [governed autonomous intent delivery](autonomous-intent-delivery.md). It starts off, needs a named and bounded grant, and leaves human Build approval and Manual QA separate.
|
|
10
10
|
|
|
11
|
-
Read
|
|
11
|
+
Read [what's new in beta 0.3.5-beta.1](releases/0.3.5-beta.1.md), the [stable 0.3.4 release notes](releases/0.3.4.md), or the [earlier Jev and OpenAI Decisions beta notes](releases/0.3.4-beta.6.md).
|
|
12
12
|
|
|
13
13
|
## New to EWAI?
|
|
14
14
|
|
|
@@ -35,3 +35,23 @@ Answering an autonomy question records a private owner response for the run. It
|
|
|
35
35
|
Use `ewai autonomy status --project . --run RUN_ID --json` to inspect a run. `ewai autonomy pause|resume|cancel|recover --project . --run RUN_ID --expected-revision N --yes --json` applies a revision-bound control; a stale revision must be inspected again. For a human question, use the dashboard's guarded answer form or `ewai autonomy answer --project . --input PROJECT_RELATIVE_JSON --yes --json`, with a private project-local input file. Do not put private answers or credentials in chat or command arguments.
|
|
36
36
|
|
|
37
37
|
The package tests exercise a tarball-installed consumer with preseeded offline dependencies; they do not prove a fresh registry dependency bootstrap or migration from every historical version. A controlled test simulates a lost response and verifies no repeated canonical effect. A separate installed-worker test substitutes only the provider boundary with a local child process; it does not establish live provider conformance. Human Manual QA, external acceptance, deployment and npm release remain separate decisions and evidence.
|
|
38
|
+
|
|
39
|
+
## Continuing incomplete Build workers
|
|
40
|
+
|
|
41
|
+
An approved task can opt into continuation through `ralph_loop.allowed: true` and `max_iterations` between 1 and 100; a usual starting limit is 3. Put this bound in the task graph before Build approval. Each iteration is a provider invocation and consumes the supervisor's existing provider-attempt budget when autonomy owns the run. No new grant or larger budget is inferred.
|
|
42
|
+
|
|
43
|
+
If a worker exits successfully before its declared green check passes, the AFK conductor reads the actual result and starts another implementation turn in the same isolated worktree, with a reference to the captured failure output. It retains partial work and records each iteration. One task timeout covers the turns, verification, review and post-merge checks; it does not reset on each continuation. Disabled loops retain one implementation turn.
|
|
44
|
+
|
|
45
|
+
A worker summary or completion promise cannot finish the task. Tests, file scope, conductor-owned commits, fresh-context review and integration checks still apply. The loop stops on an `EWAI_BLOCKED:` worker decision, provider failure or timeout, unknown termination, scope breach, changed authority or identity, failed review, or exhausted limits. Pause prevents another implementation turn while preserving the stopped task's worktree; explicit resume follows the existing recovery/revalidation route and can create a new task attempt. Cancellation never assumes a child has stopped.
|
|
46
|
+
|
|
47
|
+
The ordinary interactive companion does not supervise individual chat-turn endings. Use the existing AFK executor for eligible approved Build work that should continue across worker turns. This capability does not autonomously complete every lifecycle phase: pre-Build drafts need review, and human approvals remain separate.
|
|
48
|
+
|
|
49
|
+
## Lightweight delivery-claim validation
|
|
50
|
+
|
|
51
|
+
The existing fresh-context reviewer assesses every named completion claim separately and supplies a specific source/test/check citation and pass/fail result. The conductor rejects omitted, duplicate, unknown, unsupported or failed results before integration; a generic verdict or green exit is insufficient. No additional model call is needed. New task evidence uses v3 and attaches each claim to hashed review evidence rather than automatically citing the same green output. New task contracts declare claim_validation.required=true so removing required claim metadata cannot satisfy them with legacy evidence. Historical v1/v2 remains readable. This is model-assisted semantic review, not proof of production deployment or human acceptance; those require their separate evidence and authority.
|
|
52
|
+
|
|
53
|
+
## Challenge, observe and recover
|
|
54
|
+
|
|
55
|
+
The existing reviewer tries to falsify every completion claim and records a concrete attempt/result, distinguishing source tracing from execution. v3 claim results include `challenge: {attempt, result}` and `observation: {evidence_type, expected, actual, evidence}`. Planned `completion_requirements` assign an evidence type to each exact completion name; behaviour is the default. A configured switch proves configuration, an executed command proves its own check, a reviewed artefact proves review, and delivered behaviour needs evidence of the claimed outcome. Missing, skipped or unavailable checks stay unverified. This remains model-assisted judgement with host-enforced structure and hashes. New contracts require claim_validation.version=3; historical v1/v2 evidence remains readable but cannot satisfy a v3 requirement.
|
|
56
|
+
|
|
57
|
+
AFK status returns an evidence-derived checkpoint: repository integration branches, declared/actual task branches and worktrees, verified completed claim labels, concrete next action and open questions. The projection is persisted atomically with run.json and recomputed on status reads. It never grants authority or certifies human acceptance. Tampered evidence removes completion from the refreshed view; uncertain worker termination directs recovery before resume. Use the existing resume preflight and canonical phase gates.
|
|
@@ -46,3 +46,9 @@ EWAI saves the intent and phase evidence in SPECS. Ask it to show what's complet
|
|
|
46
46
|
If the outcome changes materially, revisit the intent and plan. A new Build approval may be needed. If only an evidence reference needs a correction, follow the [amendment guide](../completed-phase-evidence-amendments.md), which explains the narrow permitted operations.
|
|
47
47
|
|
|
48
48
|
Try the [first-delivery tutorial](../tutorials/first-delivery.md). For exact commands and evidence contracts, use the [developer delivery guide](../developer-delivery-guide.md).
|
|
49
|
+
|
|
50
|
+
## A worker turn is not task completion
|
|
51
|
+
|
|
52
|
+
The delivery workflow governs what may happen next. For eligible approved Build tasks, the AFK conductor also keeps execution moving between worker turns. A task can approve a finite continuation loop (`ralph_loop.allowed: true`, `max_iterations: 3`, for example). A normal worker exit followed by failing green verification then leads to another turn in the same worktree, within the original deadline and scope. Host verification decides when the task is ready for review. Failed reviews and genuine blockers require attention.
|
|
53
|
+
|
|
54
|
+
This is distinct from merely instructing the interactive companion to keep going: EWAI cannot restart that host's conversation after a chat turn ends. The conductor supplies the persistent execution path for approved Build tasks, while canonical phase gates continue to govern the wider delivery.
|
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
# What's new in this version
|
|
2
|
+
|
|
3
|
+
Version `0.3.4` adds optional **Jev and OpenAI Decisions** support to rank choices, compare AI recommendations, prioritise context and suggest personas or additional checks. Both start switched off. Configure them privately through the CLI or dashboard, compare advice in shadow mode or apply supported recommendations in active mode, and track usage, estimated cost and call limits.
|
|
4
|
+
|
|
5
|
+
**Concise answers and Grok Build.** EWAI puts the result or recommendation first, with clearer choices and less repetition. Grok Build is available as a coding companion, with dashboard and CLI controls for choosing your primary assistant and independent reviewers.
|
|
6
|
+
|
|
7
|
+
**Pick up the next ready piece of work.** Approve a list of work items and EWAI can select the next ready item using your recorded priorities. Set time and attempt limits, follow progress, and pause or recover a run. Build approval, required reviews, Manual QA and release decisions stay with you.
|
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
## What's new in this version
|
|
2
|
+
|
|
3
|
+
Version `0.3.5-beta.1` adds **bounded Build continuation**. When an approved autonomous Build worker finishes before its required checks pass, EWAI can give it the failed-check feedback and continue in the same isolated worktree. Continuation is opt-in, with configured iteration and time limits, and respects pause, cancellation and blocked-work signals.
|
|
4
|
+
|
|
5
|
+
**Check the delivery claim.** The existing independent Build review now checks each declared completion claim against supporting evidence and attempts to challenge it. Reviews distinguish observed behaviour from configuration, command results and review results. Missing, failed or incomplete claim evidence prevents integration, without adding another review call.
|
|
6
|
+
|
|
7
|
+
**Recover with a clearer checkpoint.** Saved runs record verified completed work, branches and worktrees, the next action and unresolved questions. Status checks revalidate the evidence before listing work as complete. These changes support continued execution within approved Build boundaries; Manual QA, acceptance and release decisions remain human checkpoints.
|
package/README.md
CHANGED
|
@@ -1,33 +1,18 @@
|
|
|
1
1
|
# Engineering With AI
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
**EWAI now supports Jev and the OpenAI Decisions API.**
|
|
6
|
-
|
|
7
|
-
- **Rank options and recommend a choice.** Evaluate available approaches and return an ordered list with a recommended answer.
|
|
8
|
-
- **Check AI recommendations.** Compare supported options returned by a larger language model with a second recommendation.
|
|
9
|
-
- **Recommend suitable personas.** Select from personas available to your project, including core, project, personal and installed premium personas.
|
|
10
|
-
- **Reduce larger-model calls.** Evaluate supported choices through focused decision calls instead of sending the same options to a larger language model.
|
|
11
|
-
- **Prioritise context.** Help order supporting context for coding and review tasks.
|
|
12
|
-
- **Support additional checks.** Assess supplied claim evidence, classify sanitised failures and recommend supplementary impact reviews and tests.
|
|
13
|
-
- **Choose your provider.** Select Jev, OpenAI Decisions or Automatic. Automatic prefers configured and authorised OpenAI access, then Jev.
|
|
14
|
-
- **Configure through the CLI or dashboard.** Save API keys privately, enable either integration and select the uses you want.
|
|
15
|
-
- **Choose shadow or active mode.** Shadow shows recommendations for comparison. Active can apply supported recommendations to context ordering and persona selection.
|
|
16
|
-
- **Track usage and estimated cost.** View decision-call usage and set call and input-token limits.
|
|
3
|
+
Engineering With AI (EWAI) helps you plan, build and review software with an AI assistant. It gives the assistant a shared record of the project, a delivery workflow and checks against your engineering standards. You keep control of the decisions and approve implementation before it starts.
|
|
17
4
|
|
|
18
|
-
|
|
5
|
+
[](https://www.conversationalcoding.dev/engineering-with-ai-harness/?utm_source=readme&utm_medium=referral&utm_campaign=ewai)
|
|
19
6
|
|
|
20
|
-
|
|
7
|
+
**Website, guides and books:** [conversationalcoding.dev](https://www.conversationalcoding.dev/engineering-with-ai-harness/?utm_source=readme&utm_medium=referral&utm_campaign=ewai) · [Online documentation](https://www.conversationalcoding.dev/engineering-with-ai-harness/docs/?utm_source=readme&utm_medium=referral&utm_campaign=ewai) · [Persona library](https://www.conversationalcoding.dev/personas/?utm_source=readme&utm_medium=referral&utm_campaign=ewai)
|
|
21
8
|
|
|
22
|
-
|
|
23
|
-
npm install --save-dev @thebackstoryis/engineering-with-ai@beta
|
|
24
|
-
```
|
|
9
|
+
## What's new in this version
|
|
25
10
|
|
|
26
|
-
|
|
11
|
+
Version `0.3.5-beta.1` adds **bounded Build continuation**. When an approved autonomous Build worker finishes before its required checks pass, EWAI can give it the failed-check feedback and continue in the same isolated worktree. Continuation is opt-in, with configured iteration and time limits, and respects pause, cancellation and blocked-work signals.
|
|
27
12
|
|
|
28
|
-
|
|
13
|
+
**Check the delivery claim.** The existing independent Build review now checks each declared completion claim against supporting evidence and attempts to challenge it. Reviews distinguish observed behaviour from configuration, command results and review results. Missing, failed or incomplete claim evidence prevents integration, without adding another review call.
|
|
29
14
|
|
|
30
|
-
**
|
|
15
|
+
**Recover with a clearer checkpoint.** Saved runs record verified completed work, branches and worktrees, the next action and unresolved questions. Status checks revalidate the evidence before listing work as complete. These changes support continued execution within approved Build boundaries; Manual QA, acceptance and release decisions remain human checkpoints.
|
|
31
16
|
|
|
32
17
|
## Quick start
|
|
33
18
|
|
|
@@ -41,27 +26,7 @@ ewai
|
|
|
41
26
|
|
|
42
27
|
`ewai` opens your assistant and walks you through setting up the project. See [Install and start](#install-and-start) for details.
|
|
43
28
|
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
## New in 0.3.1: concise answers and Grok Build
|
|
47
|
-
|
|
48
|
-
Version `0.3.1` makes EWAI's guidance more concise and adds Grok Build as a coding provider.
|
|
49
|
-
|
|
50
|
-
**Concise answers and guided decisions.** EWAI's managed instructions now prioritise correctness and usefulness, then brevity. Expect the result or recommendation first, with less repetition and routine narration. Decision requests explain the action, options and consequences, then give a recommendation with its reason and a suggested response when useful. When a request is unclear, EWAI leads with its recommended interpretation and states the assumptions that matter. The guidance is designed to cut unnecessary output; tool-result compaction is still planned, and token or cost savings haven't been measured yet. See [concise answers and guided decisions](Docs/context-management-and-token-efficiency.md#concise-answers-and-guided-decisions).
|
|
51
|
-
|
|
52
|
-
**Grok Build and coding provider settings.** EWAI can open Grok Build as a native companion, install its skills and set up project MCP. **Configuration → Coding providers** in the dashboard, or `ewai providers` in a terminal, lets you choose a primary coding CLI, independent reviewers and an eligible pool for unattended work. Each CLI keeps its own model choice by default. Grok can run isolated coding, read-only review and restricted proposal workers. Each mode passes an offline conformance check before dispatch, and your xAI key is saved privately outside the project. Build approval, required review and Manual QA stay with you. See [provider settings](Docs/reference/cli-and-configuration.md#coding-provider-settings) and [Grok Build setup](Docs/operations/installation-updating-and-entitlements.md#grok-build).
|
|
53
|
-
|
|
54
|
-
After updating, initialisation or the next check-in refreshes EWAI's managed instructions in `AGENTS.md` and `CLAUDE.md`, preserving project-authored guidance outside that block. Start a fresh host conversation after the refresh.
|
|
55
|
-
|
|
56
|
-
## EWAI can now pick up the next ready piece of work
|
|
57
|
-
|
|
58
|
-
If you’ve prepared several work items, you can choose which ones EWAI is allowed to take on. EWAI checks what’s ready, uses the priorities you’ve recorded to pick the next item, and starts its delivery workflow. Once you’ve separately approved the Build, it can run the approved build tasks, their tests and a fresh review.
|
|
59
|
-
|
|
60
|
-
This `0.3.0` release brings the capability out of beta. It adds dashboard and command-line controls to preview the work, approve the exact list, set time and attempt limits, follow progress, and pause, cancel or recover a run. New work isn’t added to the list automatically. EWAI stops when it needs a decision from you.
|
|
61
|
-
|
|
62
|
-
The final whole-delivery test stage, Manual QA and release preparation still happen through the normal EWAI workflow. EWAI doesn’t approve or complete those steps for you.
|
|
63
|
-
|
|
64
|
-
To use the release in a project without replacing a global installation, run:
|
|
29
|
+
To use EWAI in a project without replacing a global installation, run:
|
|
65
30
|
|
|
66
31
|
```bash
|
|
67
32
|
npm install --save-dev @thebackstoryis/engineering-with-ai
|
|
@@ -70,7 +35,7 @@ npx ewai
|
|
|
70
35
|
|
|
71
36
|
To install or update EWAI globally, run `npm install --global @thebackstoryis/engineering-with-ai`.
|
|
72
37
|
|
|
73
|
-
|
|
38
|
+
### Install and start
|
|
74
39
|
|
|
75
40
|
You'll need Node.js 22.5 or newer, npm, and a supported AI command-line tool such as Codex or Claude Code, installed and signed in.
|
|
76
41
|
|
|
@@ -88,12 +53,12 @@ It first asks where to keep your project's SPECS records, then helps you describ
|
|
|
88
53
|
|
|
89
54
|
To return to the project, run `ewai` in the same folder again. It picks up the saved project state rather than starting setup from scratch. Keep your normal Git workflow for your application's code; EWAI itself is installed and updated through npm.
|
|
90
55
|
|
|
91
|
-
If npm returns `E404`, follow [installation troubleshooting](Docs/operations/installation-updating-and-entitlements.md). That's a package-acquisition failure, not a premium-persona licence error. Don't paste licence keys into npm commands or chat.
|
|
56
|
+
If npm returns `E404`, follow [installation troubleshooting](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/operations/installation-updating-and-entitlements.md). That's a package-acquisition failure, not a premium-persona licence error. Don't paste licence keys into npm commands or chat.
|
|
92
57
|
|
|
93
|
-
- [First session: set up a small project](Docs/tutorials/first-session.md)
|
|
94
|
-
- [First delivery: work through a CSV export](Docs/tutorials/first-delivery.md)
|
|
95
|
-
- [Already have a codebase?](Docs/existing-project-onboarding-guide.md) Archaeology can reconstruct missing documentation, but it's your choice whether to run it.
|
|
96
|
-
- [Installation, updates and host options](Docs/operations/installation-updating-and-entitlements.md)
|
|
58
|
+
- [First session: set up a small project](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/tutorials/first-session.md)
|
|
59
|
+
- [First delivery: work through a CSV export](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/tutorials/first-delivery.md)
|
|
60
|
+
- [Already have a codebase?](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/existing-project-onboarding-guide.md) Archaeology can reconstruct missing documentation, but it's your choice whether to run it.
|
|
61
|
+
- [Installation, updates and host options](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/operations/installation-updating-and-entitlements.md)
|
|
97
62
|
|
|
98
63
|
## Start with a conversation
|
|
99
64
|
|
|
@@ -113,7 +78,7 @@ The supporting CLI commands are available for direct control, troubleshooting an
|
|
|
113
78
|
|
|
114
79
|
**SPECS** means Scope, Purpose, Evidence, Constraints and Strategy. These readable project records hold the purpose, requirements, decisions and evidence the team has agreed. They remain useful outside an AI session.
|
|
115
80
|
|
|
116
|
-
The dashboard runs locally and lets you inspect work and make supported choices. Its loopback address isn't a shared team website. The AI host does the guided work; the runtime records progress and checks the conditions for moving on. [How these parts fit together](Docs/explanation/core-concepts.md).
|
|
81
|
+
The dashboard runs locally and lets you inspect work and make supported choices. Its loopback address isn't a shared team website. The AI host does the guided work; the runtime records progress and checks the conditions for moving on. [How these parts fit together](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/explanation/core-concepts.md).
|
|
117
82
|
|
|
118
83
|
Human judgement remains essential. A persona isn't a real stakeholder, green tests aren't human acceptance, and a delivery handoff isn't permission to deploy.
|
|
119
84
|
|
|
@@ -127,21 +92,21 @@ The source is published for transparency and review, but EWAI is not open
|
|
|
127
92
|
source. You may not modify, repackage, redistribute, rebrand or commercially
|
|
128
93
|
exploit the EWAI core without separate written permission from Backstory Group.
|
|
129
94
|
The licence does not restrict the project content or output you create by using
|
|
130
|
-
EWAI. Read the [Backstory Group Source-Available Licence](LICENSE) for the full
|
|
95
|
+
EWAI. Read the [Backstory Group Source-Available Licence](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/LICENSE) for the full
|
|
131
96
|
terms.
|
|
132
97
|
|
|
133
98
|
## Choose what you need
|
|
134
99
|
|
|
135
100
|
| When you want to… | Start here |
|
|
136
101
|
| --- | --- |
|
|
137
|
-
| Decide which work needs attention next | [Companion](Docs/context-aware-delivery-companion-user-guide.md) |
|
|
138
|
-
| Describe a feature and agree its boundaries | [Intent Studio](Docs/guided-intent-workspace-guide.md) |
|
|
139
|
-
| Understand how delivery moves through its stages | [The fourteen-stage workflow](Docs/explanation/delivery-workflow.md) |
|
|
140
|
-
| Investigate dependencies before a change | [Source Map](Docs/repository-source-map-guide.md) and [Blast Radius](Docs/blast-radius-and-impact-routing-guide.md) |
|
|
141
|
-
| Turn a meeting into reviewed project evidence | [Meeting evidence](Docs/meeting-evidence-user-guide.md) |
|
|
142
|
-
| Check an implemented feature with a person | [Manual QA](Docs/quality/manual-qa-and-acceptance.md) |
|
|
143
|
-
| Adapt the dashboard to your work | [Dashboard configuration](Docs/operations/dashboard-configuration.md) |
|
|
144
|
-
| Use shared organisational guidance | [Blueprints](Docs/designing-organisation-blueprint-packs.md) and [rollout](Docs/organisation-rollout-guide.md) |
|
|
102
|
+
| Decide which work needs attention next | [Companion](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/context-aware-delivery-companion-user-guide.md) |
|
|
103
|
+
| Describe a feature and agree its boundaries | [Intent Studio](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/guided-intent-workspace-guide.md) |
|
|
104
|
+
| Understand how delivery moves through its stages | [The fourteen-stage workflow](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/explanation/delivery-workflow.md) |
|
|
105
|
+
| Investigate dependencies before a change | [Source Map](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/repository-source-map-guide.md) and [Blast Radius](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/blast-radius-and-impact-routing-guide.md) |
|
|
106
|
+
| Turn a meeting into reviewed project evidence | [Meeting evidence](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/meeting-evidence-user-guide.md) |
|
|
107
|
+
| Check an implemented feature with a person | [Manual QA](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/quality/manual-qa-and-acceptance.md) |
|
|
108
|
+
| Adapt the dashboard to your work | [Dashboard configuration](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/operations/dashboard-configuration.md) |
|
|
109
|
+
| Use shared organisational guidance | [Blueprints](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/designing-organisation-blueprint-packs.md) and [rollout](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/organisation-rollout-guide.md) |
|
|
145
110
|
|
|
146
111
|
Start with the parts your project needs. Portfolio, Team Hub and other advanced dashboard views are optional; hiding a view doesn't disable mandatory project checks.
|
|
147
112
|
|
|
@@ -149,18 +114,18 @@ Start with the parts your project needs. Portfolio, Team Hub and other advanced
|
|
|
149
114
|
|
|
150
115
|
The included personas support the normal workflow. You can also create project-specific personas and use your own personal library.
|
|
151
116
|
|
|
152
|
-
Premium personas are optional specialist perspectives. If you have a subscription, [enter your key privately and install the pack](Docs/operations/premium-personas-setup.md) before the analysis you want it to support. Installing a persona doesn't give it authority to approve a requirement, bypass a check or speak for a real user.
|
|
117
|
+
Premium personas are optional specialist perspectives. If you have a subscription, [enter your key privately and install the pack](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/operations/premium-personas-setup.md) before the analysis you want it to support. Installing a persona doesn't give it authority to approve a requirement, bypass a check or speak for a real user.
|
|
153
118
|
|
|
154
119
|
## Go deeper
|
|
155
120
|
|
|
156
|
-
- [User guides and learning routes](Docs/README.md)
|
|
157
|
-
- [Complete guide catalogue](Docs/guide-catalogue.md)
|
|
158
|
-
- [Capabilities and project layout](Docs/reference/capabilities-and-project-layout.md)
|
|
159
|
-
- [Commands and configuration](Docs/reference/cli-and-configuration.md)
|
|
160
|
-
- [Troubleshooting and recovery](Docs/operations/troubleshooting-and-recovery.md)
|
|
121
|
+
- [User guides and learning routes](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/README.md)
|
|
122
|
+
- [Complete guide catalogue](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/guide-catalogue.md)
|
|
123
|
+
- [Capabilities and project layout](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/reference/capabilities-and-project-layout.md)
|
|
124
|
+
- [Commands and configuration](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/reference/cli-and-configuration.md)
|
|
125
|
+
- [Troubleshooting and recovery](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/operations/troubleshooting-and-recovery.md)
|
|
161
126
|
- [EWAI on the web: harness overview, online docs and changelog](https://www.conversationalcoding.dev/engineering-with-ai-harness/?utm_source=readme&utm_medium=referral&utm_campaign=ewai)
|
|
162
127
|
- [Engineering With AI, the book behind the method](https://www.conversationalcoding.dev/books/?utm_source=readme&utm_medium=referral&utm_campaign=ewai)
|
|
163
128
|
|
|
164
129
|
## Contributing to EWAI
|
|
165
130
|
|
|
166
|
-
If you're changing the harness itself, use the [contributor guide](Docs/maintainers/contributing.md) and [verification walkthroughs](Docs/maintainers/verification-walkthroughs.md). Those source-checkout and regression-test instructions aren't part of setting up your own application.
|
|
131
|
+
If you're changing the harness itself, use the [contributor guide](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/maintainers/contributing.md) and [verification walkthroughs](https://github.com/TheBackstoryIs/EngineeringWithAIHarness/blob/main/Docs/maintainers/verification-walkthroughs.md). Those source-checkout and regression-test instructions aren't part of setting up your own application.
|
package/package.json
CHANGED
|
@@ -77,6 +77,8 @@ Use guarded operations only:
|
|
|
77
77
|
|
|
78
78
|
Never mutate a raw phase or status. A phase completes only with its passing EWAI gate ledger and fresh, hashed project evidence. Record meaningful progress events for the dashboard.
|
|
79
79
|
|
|
80
|
+
New task contracts must set `claim_validation.required=true` and `claim_validation.version=3`. Declare one `completion_requirements` entry per completion name with `evidence_type`: `behaviour`, `configuration`, `command-result` or `review-result`. Default evidence type for undeclared historical tasks is behaviour. Challenge every claim: try a counterexample and record `challenge.attempt` and `challenge.result`, stating whether executed or source-traced. Record `observation.evidence_type`, `expected`, `actual` and `evidence`; configuration does not establish delivered behaviour and a green command proves only its exercised assertions. Unknown/skipped outcomes cannot pass their obligation. During the existing fresh-context review, assess each `completion_evidence` claim individually against specific source, test assertions or observed check results. Emit one `COMPLETION_CHECK:` JSON line per exact name with `status` and `support`; do not infer all claims from a generic passing test. Unsupported, missing or failed claims block integration. New evidence uses `ewai.task-evidence/v3` and binds completion claims to captured review evidence. Historical v1/v2 evidence remains readable, but cannot satisfy a task requiring claim validation. This adds no separate model invocation; it is model-assisted evidence assessment, not Manual QA or release authority.
|
|
81
|
+
|
|
80
82
|
Plan must produce `build-plan.md`, `destination.md`, the Claim Ledger, Plan Contract, and complete `task-graph.json`. The deterministic task-graph validator checks source records, claim and slice coverage, dependencies, cycles, write-set isolation, branch uniqueness, task contracts, waves, evidence paths, and review/merge ownership. Do not author its passing output manually.
|
|
81
83
|
|
|
82
84
|
## Persona-led test scenario integration
|
|
@@ -109,13 +111,15 @@ Start the intent-level active session with a stable `ownerId`. A different orche
|
|
|
109
111
|
|
|
110
112
|
## Unattended Build conductor
|
|
111
113
|
|
|
112
|
-
When the user explicitly asks EWAI to continue approved Build work while they are away, use the AFK conductor instead of inventing a host-specific background loop:
|
|
114
|
+
When the user explicitly asks EWAI to continue approved Build work to completion without repeated prompts, or while they are away, use the AFK conductor for eligible tasks instead of inventing a host-specific background loop:
|
|
113
115
|
|
|
114
116
|
1. Call `ewai_afk_preflight` and explain every blocker. Never weaken the task graph to make preflight pass.
|
|
115
117
|
2. Confirm the user has already approved Build and understands the bounded scope, provider choice, maximum parallel tasks, and timeout.
|
|
116
118
|
3. Call `ewai_afk_start`. Report the durable run ID and how to pause, resume, cancel, and inspect it; do not expose lease tokens or raw prompts.
|
|
117
119
|
4. Use `ewai_afk_status` for updates. Translate semantic run/task states into human language rather than streaming model internals.
|
|
118
|
-
5.
|
|
120
|
+
5. For tasks expected to need several implementation turns, propose an explicit `ralph_loop.allowed: true` with a finite `max_iterations` from 1 to 100 (normally 3) before Build approval. Existing disabled contracts remain single-shot; never expand a recorded task or grant to obtain more attempts.
|
|
121
|
+
6. The conductor can continue only a successfully exited, confirmed stopped implementation worker whose declared green check still fails. It keeps the same task worktree, captures each verification output, and rechecks provider/authority/control boundaries on every invocation. The task timeout is shared across its turns, verification, review and post-merge checks. A completion promise never replaces host verification.
|
|
122
|
+
7. A blocked run requires attention; diagnose its preserved log/evidence before `ewai_afk_resume`. Never silently retry a worker-reported `EWAI_BLOCKED:` decision, provider error/timeout, unknown termination, stop condition, failed review, scope breach, merge conflict, or post-merge failure.
|
|
119
123
|
|
|
120
124
|
The conductor is a local Build executor, not a replacement for the fourteen-stage harness. It detects simple and multi-repository topology from the configured `pipeline.yaml`; every task's `repo` must match a configured repository, and exactly one configured repository must contain the configured SPECS root to own durable evidence. The workspace root need not be a Git repository. It creates a missing integration branch from each selected repository's clean current branch, then creates a real task branch/worktree there, leases only ready AFK tasks, limits concurrency to the validated graph, requires a fresh-context review, merges into that repository's declared parent branch in graph order, and runs post-merge checks. It blocks detached heads and refuses to silently switch when the declared integration branch already exists elsewhere. Use `task-graph.json.repository_branches` when repositories have different parent branches. Completion means all eligible Build tasks were integrated; it does not complete the Build phase, Standards Sweep, Test Execute, external validation, Delivery, Manual QA, or Retro.
|
|
121
125
|
|
|
@@ -124,3 +128,15 @@ The conductor is a local Build executor, not a replacement for the fourteen-stag
|
|
|
124
128
|
Translate control-plane work into warm, semantic progress such as “Validating the plan against project standards” or “Checking what changed since this work was shelved.” Do not make the user learn the command surface. Ask only decisions the evidence cannot answer.
|
|
125
129
|
|
|
126
130
|
When blocked, retain the exact phase, evidence, questions, and recovery point. When complete, Retro must capture learning and route accepted improvements into project-owned SPECS assets.
|
|
131
|
+
|
|
132
|
+
## Execution continuity
|
|
133
|
+
|
|
134
|
+
After an authorised action completes, inspect the saved next action and continue all remaining permitted work. A successful tool call, completed task, progress summary or model turn is not delivery completion. Do not end the interaction just to offer to continue work the owner has already authorised.
|
|
135
|
+
|
|
136
|
+
Stop for a real blocker, a required human decision, exhausted configured execution budget, cancellation, or the actual agreed completion boundary. Preserve the exact recovery point and explain the reason. Never interpret an earlier Build instruction as Manual QA, deployment, release or acceptance approval.
|
|
137
|
+
|
|
138
|
+
For eligible approved Build tasks, the conductor provides execution continuity outside individual worker turns. These instructions improve interactive behaviour but cannot restart a native host conversation after its model ends a turn; do not advertise guaranteed unattended completion for the ordinary interactive companion. Phase preparation remains draft-only and stops for its required human review.
|
|
139
|
+
|
|
140
|
+
## Evidence-derived recovery checkpoint
|
|
141
|
+
|
|
142
|
+
Use the checkpoint returned by AFK status to present integration/task branches, evidence-validated completed claims, the exact next action and open questions. It is stored atomically in the run record and recomputed on status reads, so changed evidence is never accepted from a stale checkpoint. Reports and guarded action/preflight surfaces remain authoritative; checkpoint output never grants resume, Build, release or acceptance authority. Unknown worker termination requires recovery, and missing/tampered completion evidence requires diagnosis. Present the checkpoint at interruptions and handoffs rather than replaying raw logs.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
import { randomUUID } from 'node:crypto';
|
|
2
2
|
import {
|
|
3
|
-
copyFileSync, cpSync, existsSync, mkdirSync, readFileSync, readdirSync, realpathSync, rmSync, writeFileSync,
|
|
3
|
+
copyFileSync, cpSync, existsSync, mkdirSync, readFileSync, readdirSync, realpathSync, rmSync, statSync, writeFileSync,
|
|
4
4
|
} from 'node:fs';
|
|
5
5
|
import { dirname, relative, resolve } from 'node:path';
|
|
6
6
|
import { execFileSync, spawn } from 'node:child_process';
|
|
@@ -8,7 +8,7 @@ import { fileURLToPath } from 'node:url';
|
|
|
8
8
|
import { deliveryPaths, atomicJson, isWithin, sha256, now } from '../delivery-documents.mjs';
|
|
9
9
|
import { projectPaths } from '../paths.mjs';
|
|
10
10
|
import { loadProjectConfig, validationStatus } from '../project.mjs';
|
|
11
|
-
import { validateTaskGraph, validateTaskReport } from '../task-graph.mjs';
|
|
11
|
+
import { validateTaskGraph, validateTaskReport, parseCompletionClaimReview } from '../task-graph.mjs';
|
|
12
12
|
import { syncIntentIndex } from './intents.mjs';
|
|
13
13
|
import {
|
|
14
14
|
abandonExecutionLeasesForRun,
|
|
@@ -74,6 +74,49 @@ function requireRun(projectRoot, runId) {
|
|
|
74
74
|
return { paths, run };
|
|
75
75
|
}
|
|
76
76
|
|
|
77
|
+
export function deriveAfkCheckpoint(projectRoot, run) {
|
|
78
|
+
try { return deriveCheckpointFromEvidence(projectRoot,run); }
|
|
79
|
+
catch {
|
|
80
|
+
// A derived convenience view must never obstruct durable control writes.
|
|
81
|
+
return {schema:'ewai.afk-checkpoint/v1',runId:run.id,status:'unavailable',updatedAt:run.updatedAt,
|
|
82
|
+
repositories:[],tasks:[],done:[],openQuestions:['Recovery checkpoint unavailable: task or run evidence is malformed; inspect preserved records.'],
|
|
83
|
+
next:run.executionStopped!==true || run.unconfirmedExecution
|
|
84
|
+
? {action:'confirm-worker-termination',instruction:'Establish owned worker termination before a human recovery decision.'}
|
|
85
|
+
: {action:'inspect-evidence',instruction:'Repair or reconcile malformed saved evidence through existing guarded operations.'}};
|
|
86
|
+
}
|
|
87
|
+
}
|
|
88
|
+
|
|
89
|
+
function deriveCheckpointFromEvidence(projectRoot, run) {
|
|
90
|
+
const checkpoint = {schema:'ewai.afk-checkpoint/v1',runId:run.id,status:run.status,updatedAt:run.updatedAt,
|
|
91
|
+
repositories:(run.topology?.repositories ?? []).map(repo=>({name:repo.name,integrationBranch:repo.parentBranch})),tasks:[],done:[],openQuestions:[],next:null};
|
|
92
|
+
let graph;
|
|
93
|
+
try { graph = readGraph(projectRoot,run.slug); if (!Array.isArray(graph.tasks)) throw new Error('Malformed task graph'); }
|
|
94
|
+
catch { graph=null; checkpoint.openQuestions.push('The task graph cannot be read; inspect saved delivery evidence.'); }
|
|
95
|
+
for (const task of graph?.tasks ?? []) {
|
|
96
|
+
const state=run.tasks?.[task.id] ?? {};
|
|
97
|
+
let report;
|
|
98
|
+
try { report=validateTaskReport(deliveryPaths(projectRoot,run.slug).deliveryRoot,task); }
|
|
99
|
+
catch { report={status:'incomplete'}; }
|
|
100
|
+
checkpoint.tasks.push({id:task.id,repository:task.repo,declaredBranch:task.branch,actualBranch:state.actualBranch ?? null,
|
|
101
|
+
worktree:state.worktree ?? null,status:state.status ?? 'pending',evidenceStatus:report.status});
|
|
102
|
+
if (report.status === 'complete') checkpoint.done.push({taskId:task.id,claims:task.completion_evidence ?? []});
|
|
103
|
+
else if (state.status === 'completed') checkpoint.openQuestions.push(`${task.id} recorded completion has missing or invalid evidence; inspect it before continuing.`);
|
|
104
|
+
if (state.error) checkpoint.openQuestions.push(`${task.id}: ${state.error}`);
|
|
105
|
+
}
|
|
106
|
+
if (['blocked','cancel-unknown'].includes(run.status)) for(const message of (run.messages ?? []).slice(-1)) checkpoint.openQuestions.push(message.message);
|
|
107
|
+
const stopped=run.executionStopped===true && !run.unconfirmedExecution && !(run.activeWorkers ?? []).length;
|
|
108
|
+
const orphaned = run.status==='running' && !processAlive(run.pid);
|
|
109
|
+
if (!stopped && (!['running','starting'].includes(run.status) || orphaned)) checkpoint.next={action:'confirm-worker-termination',instruction:'Establish that owned workers stopped; inspect preserved work before a human recovery decision.'};
|
|
110
|
+
else if (checkpoint.openQuestions.some(question=>question.includes('evidence') || question.includes('task graph'))) checkpoint.next={action:'inspect-evidence',instruction:'Resolve missing or changed task evidence before claiming completion or resuming.'};
|
|
111
|
+
else if (run.status==='cancelled') checkpoint.next={action:'cancelled',instruction:'Work is preserved; starting further work needs a new permitted action.'};
|
|
112
|
+
else if (run.status==='completed' && graph?.tasks?.length && checkpoint.done.length===graph.tasks.length) checkpoint.next={action:'complete-build-gate',instruction:'All task evidence validates; use the canonical Build gate and subsequent phases. Human acceptance is still separate.'};
|
|
113
|
+
else if (orphaned) checkpoint.next={action:'guarded-resume',instruction:'The saved conductor is no longer running; inspect preserved work and use the existing guarded recovery preflight.'};
|
|
114
|
+
else if (run.status==='paused') checkpoint.next={action:'guarded-resume',instruction:'Inspect the preserved checkpoint, then use the existing resume preflight; this checkpoint grants no authority.'};
|
|
115
|
+
else if (run.status==='blocked') checkpoint.next={action:'inspect-blocker',instruction:'Diagnose the recorded blocker and preserved logs; resume only through existing guarded recovery.'};
|
|
116
|
+
else checkpoint.next={action:'continue-bounded-work',instruction:'The conductor continues only permitted tasks under its saved controls; inspect status before any manual intervention.'};
|
|
117
|
+
return checkpoint;
|
|
118
|
+
}
|
|
119
|
+
|
|
77
120
|
function saveRun(paths, run, resumeFrom = null) {
|
|
78
121
|
if (existsSync(paths.runPath)) {
|
|
79
122
|
const current = JSON.parse(readFileSync(paths.runPath, 'utf8'));
|
|
@@ -84,6 +127,7 @@ function saveRun(paths, run, resumeFrom = null) {
|
|
|
84
127
|
}
|
|
85
128
|
}
|
|
86
129
|
run.updatedAt = now();
|
|
130
|
+
run.checkpoint = deriveAfkCheckpoint(paths.projectRoot,run);
|
|
87
131
|
atomicJson(paths.runPath, run);
|
|
88
132
|
return run;
|
|
89
133
|
}
|
|
@@ -420,7 +464,7 @@ export function afkRunStatus(projectRoot, runId = '') {
|
|
|
420
464
|
const root = resolve(projectRoot);
|
|
421
465
|
if (runId) {
|
|
422
466
|
const { run } = requireRun(root, runId);
|
|
423
|
-
return { ...run, alive: processAlive(run.pid) };
|
|
467
|
+
return { ...run, checkpoint:deriveAfkCheckpoint(root,run), alive: processAlive(run.pid) };
|
|
424
468
|
}
|
|
425
469
|
const { runsRoot } = pathsFor(root);
|
|
426
470
|
if (!existsSync(runsRoot)) return [];
|
|
@@ -428,7 +472,7 @@ export function afkRunStatus(projectRoot, runId = '') {
|
|
|
428
472
|
.filter((entry) => entry.isDirectory() && existsSync(resolve(runsRoot, entry.name, 'run.json')))
|
|
429
473
|
.map((entry) => JSON.parse(readFileSync(resolve(runsRoot, entry.name, 'run.json'), 'utf8')))
|
|
430
474
|
.sort((left, right) => right.createdAt.localeCompare(left.createdAt))
|
|
431
|
-
.map((run) => ({ ...run, alive: processAlive(run.pid) }));
|
|
475
|
+
.map((run) => ({ ...run, checkpoint:deriveAfkCheckpoint(root,run), alive: processAlive(run.pid) }));
|
|
432
476
|
}
|
|
433
477
|
|
|
434
478
|
export function pauseAfkRun(projectRoot, runId) {
|
|
@@ -509,7 +553,7 @@ export function prepareAfkContext(projectRoot, task, repository, options = {}) {
|
|
|
509
553
|
const implementationCommit = String(options.implementationCommit ?? '').trim();
|
|
510
554
|
const taskContent = JSON.stringify(task);
|
|
511
555
|
let rules = review
|
|
512
|
-
? `TESTS_FIRST Read tests before implementation. Read cited standards from ${specsRoot}. Review implementation commit ${implementationCommit}. Check the diff, exact task contract, write_set, desired outcome, stop conditions, test evidence, correctness, standards, security and scope. Do not edit files. End with exactly one line: VERDICT: PASS or VERDICT: FAIL.`
|
|
556
|
+
? `TESTS_FIRST Read tests before implementation. Read cited standards from ${specsRoot}. Review implementation commit ${implementationCommit}. Check the diff, exact task contract, write_set, desired outcome, stop conditions, test evidence, correctness, standards, security and scope. Do not edit files. For EVERY completion_evidence name, assess whether the claimed result is actually delivered. Cite specific test assertions, source symbols or observed check results that support it; a generic green exit is insufficient. Never infer deployment, release or human acceptance from code/tests. Emit one line per name: COMPLETION_CHECK: followed by JSON with name (exact approved label), status (pass or fail), and support (specific evidence and reasoning), challenge (object with attempt and result), and observation (object with evidence_type, expected, actual and evidence). Try to falsify each claim using a concrete counterexample, boundary, missing input, retry or state sequence; report the attempt and what it showed, including if you could only trace source rather than execute it. Observation must describe what was actually observed, not an intention or a configured setting. Evidence type must match completion_requirements for that name; default is behaviour. Configuration proves only configuration; a command-result proves only its check; behaviour requires the claimed outcome; review-result proves only a completed review. Keep skipped or unavailable checks unverified, never invent human acceptance or deployment. Missing or unsupported delivery must be fail. End with exactly one line: VERDICT: PASS or VERDICT: FAIL.`
|
|
513
557
|
: `You are an EWAI Build worker in an isolated Git worktree. Read standards from ${specsRoot} and tests before implementation. Repository ${repository.name} (${repository.role}). Work only inside write_set. Run only allowed_commands. Stop on every stop_condition. Make the smallest code and test change. Do not create SPECS records, commit, merge, change delivery state or update phase status. AUTHORITY_NONE. Finish with files changed and commands run.`;
|
|
514
558
|
if(options.evidenceCandidates?.length)rules += ' For this isolated Grok worker, read the complete source-standard and recorded-check sections supplied in this context. Host paths and Git history are inaccessible. Use the supplied exact commit diff for review. Do not execute commands: the conductor runs approved checks and owns commits. Read tests from the declared file snapshot before implementation.';
|
|
515
559
|
const ruleMarkers = review
|
|
@@ -712,7 +756,7 @@ async function executeCommand(command, cwd, outputPath, timeoutMs, projectRoot,
|
|
|
712
756
|
const result = await invokeControlled(projectRoot, run, task, { provider: 'command', command: '/bin/sh',
|
|
713
757
|
args: ['-lc', command], cwd, prompt: '', timeoutMs, logPath: outputPath });
|
|
714
758
|
if (result.cancelled || result.timedOut || requireRun(projectRoot, run.id).run.desiredState === 'cancelled') throw new Error('AFK command stopped before completion.');
|
|
715
|
-
return { command, exitCode: result.exitCode, outputPath };
|
|
759
|
+
return { command, exitCode: result.exitCode, outputPath, outputSha256: sha256(readFileSync(outputPath)) };
|
|
716
760
|
}
|
|
717
761
|
|
|
718
762
|
function writeTaskEvidence(projectRoot, run, result, postMerge) {
|
|
@@ -722,14 +766,15 @@ function writeTaskEvidence(projectRoot, run, result, postMerge) {
|
|
|
722
766
|
const delivery = deliveryPaths(projectRoot, run.slug);
|
|
723
767
|
const evidencePath = resolve(delivery.deliveryRoot, task.evidence_path);
|
|
724
768
|
const reportPath = resolve(delivery.deliveryRoot, task.report_path);
|
|
725
|
-
const reviewRelative = `tasks/${task.id}/evidence/fresh-context-review.
|
|
769
|
+
const reviewRelative = `tasks/${task.id}/evidence/fresh-context-review.json`;
|
|
726
770
|
const reviewPath = resolve(delivery.deliveryRoot, reviewRelative);
|
|
727
771
|
mkdirSync(dirname(reviewPath), { recursive: true });
|
|
728
|
-
writeFileSync(reviewPath,
|
|
772
|
+
writeFileSync(reviewPath, JSON.stringify({schema:'ewai.claim-review/v1',reviews:review.reviewResults.map(item=>({reviewer:item.reviewer,output:item.output}))},null,2)+'\n', 'utf8');
|
|
729
773
|
const commands = verifiedCommands.map((commandResult) => {
|
|
730
774
|
const relativePath = `tasks/${task.id}/evidence/${commandResult.stage}.txt`;
|
|
731
775
|
const destination = resolve(delivery.deliveryRoot, relativePath);
|
|
732
776
|
mkdirSync(dirname(destination), { recursive: true });
|
|
777
|
+
if (sha256(readFileSync(commandResult.outputPath)) !== commandResult.outputSha256) throw new Error('Captured command output changed before acceptance.');
|
|
733
778
|
copyFileSync(commandResult.outputPath, destination);
|
|
734
779
|
return {
|
|
735
780
|
stage: commandResult.stage,
|
|
@@ -740,9 +785,22 @@ function writeTaskEvidence(projectRoot, run, result, postMerge) {
|
|
|
740
785
|
output_sha256: sha256(readFileSync(destination)),
|
|
741
786
|
};
|
|
742
787
|
});
|
|
743
|
-
const
|
|
788
|
+
const continuation = result.continuation && {
|
|
789
|
+
enabled: result.continuation.enabled,
|
|
790
|
+
max_iterations: result.continuation.maxIterations,
|
|
791
|
+
iterations: result.continuation.iterations.map(entry => {
|
|
792
|
+
const verification = entry.verification;
|
|
793
|
+
if (!verification) return { iteration: entry.iteration, provider_exit_code: entry.providerExitCode, execution_stopped: entry.executionStopped };
|
|
794
|
+
const outputPath = `tasks/${task.id}/evidence/green-iteration-${entry.iteration}.txt`;
|
|
795
|
+
const destination = resolve(delivery.deliveryRoot, outputPath);
|
|
796
|
+
if (sha256(readFileSync(verification.outputPath)) !== verification.outputSha256) throw new Error('Continuation verification output changed before acceptance.');
|
|
797
|
+
copyFileSync(verification.outputPath, destination);
|
|
798
|
+
return { iteration: entry.iteration, provider_exit_code: entry.providerExitCode, execution_stopped: entry.executionStopped,
|
|
799
|
+
verification: { command: verification.command, exit_code: verification.exitCode, output_path: outputPath, output_sha256: sha256(readFileSync(destination)) } };
|
|
800
|
+
}),
|
|
801
|
+
};
|
|
744
802
|
const evidence = {
|
|
745
|
-
schema: 'ewai.task-evidence/
|
|
803
|
+
schema: 'ewai.task-evidence/v3',
|
|
746
804
|
task_id: task.id,
|
|
747
805
|
repository: repository.name,
|
|
748
806
|
task_branch: task.branch,
|
|
@@ -751,19 +809,22 @@ function writeTaskEvidence(projectRoot, run, result, postMerge) {
|
|
|
751
809
|
commit: implementationCommit,
|
|
752
810
|
changed_files: changedFiles,
|
|
753
811
|
commands,
|
|
812
|
+
...(continuation ? { continuation } : {}),
|
|
754
813
|
review: {
|
|
755
814
|
fresh_context: true,
|
|
756
815
|
tests_reviewed_first: true,
|
|
757
816
|
status: 'pass',
|
|
758
817
|
reviewer,
|
|
818
|
+
claim_results: review.claimResults,
|
|
759
819
|
evidence_path: reviewRelative,
|
|
760
820
|
evidence_sha256: sha256(readFileSync(reviewPath)),
|
|
761
821
|
},
|
|
762
822
|
completion_checks: (task.completion_evidence ?? []).map((name) => ({
|
|
763
823
|
name,
|
|
764
824
|
status: 'pass',
|
|
765
|
-
|
|
766
|
-
|
|
825
|
+
...review.claimResults.find(result => result.name === name),
|
|
826
|
+
evidence_path: reviewRelative,
|
|
827
|
+
evidence_sha256: sha256(readFileSync(reviewPath)),
|
|
767
828
|
})),
|
|
768
829
|
post_merge_checks: postMerge,
|
|
769
830
|
};
|
|
@@ -778,7 +839,7 @@ Implemented ${task.name} in repository ${repository.name} on branch \`${actualBr
|
|
|
778
839
|
The declared red command failed as expected; the green and refactor verification commands passed. Exact outputs and hashes are recorded in the evidence sidecar.
|
|
779
840
|
|
|
780
841
|
## Feedback loops run
|
|
781
|
-
The conductor ran the task's bounded verification commands, a fresh-context review, and ${postMerge.length} post-merge check(s).
|
|
842
|
+
The conductor ran ${result.continuation?.iterations.length ?? 1} implementation turn(s), the task's bounded verification commands, a fresh-context review, and ${postMerge.length} post-merge check(s). Per-turn verification outputs are retained in the evidence sidecar.
|
|
782
843
|
|
|
783
844
|
## Files changed
|
|
784
845
|
${changedFiles.map((file) => `- ${file}`).join('\n')}
|
|
@@ -798,7 +859,24 @@ No declared stop condition occurred.
|
|
|
798
859
|
Automated task and integration evidence passed. Human QA remains governed by the later EWAI delivery gate.
|
|
799
860
|
`;
|
|
800
861
|
writeFileSync(reportPath, report, 'utf8');
|
|
801
|
-
return [reportPath, evidencePath, reviewPath, ...commands.map(command => resolve(delivery.deliveryRoot, command.output_path))
|
|
862
|
+
return [reportPath, evidencePath, reviewPath, ...commands.map(command => resolve(delivery.deliveryRoot, command.output_path)),
|
|
863
|
+
...(continuation?.iterations ?? []).filter(entry => entry.verification).map(entry => resolve(delivery.deliveryRoot, entry.verification.output_path))];
|
|
864
|
+
}
|
|
865
|
+
|
|
866
|
+
// Only assistant/result messages can signal a blocker. Tool output may quote it.
|
|
867
|
+
function workerReportedBlocker(output) {
|
|
868
|
+
const blocked = value => typeof value === 'string' && /^EWAI_BLOCKED:/m.test(value);
|
|
869
|
+
if (blocked(output)) return true;
|
|
870
|
+
for (const line of String(output ?? '').split('\n')) {
|
|
871
|
+
let event;
|
|
872
|
+
try { event = JSON.parse(line); } catch { continue; }
|
|
873
|
+
if (event?.type === 'item.completed' && event.item?.type === 'agent_message' && blocked(event.item.text)) return true;
|
|
874
|
+
if (event?.type === 'result' && blocked(event.result)) return true;
|
|
875
|
+
const message = event?.type === 'assistant' ? event.message : event?.role === 'assistant' ? event : null;
|
|
876
|
+
if (message && (blocked(message.content) || Array.isArray(message.content)
|
|
877
|
+
&& message.content.some(part => part?.type === 'text' && blocked(part.text)))) return true;
|
|
878
|
+
}
|
|
879
|
+
return false;
|
|
802
880
|
}
|
|
803
881
|
|
|
804
882
|
async function implementTask(projectRoot, run, task, provider, dependencies = {}) {
|
|
@@ -806,9 +884,20 @@ async function implementTask(projectRoot, run, task, provider, dependencies = {}
|
|
|
806
884
|
const paths = pathsFor(projectRoot, run.id);
|
|
807
885
|
const repository = repositoryForTask(run, task);
|
|
808
886
|
const attempt = Number(run.tasks?.[task.id]?.attempt ?? 0) + 1;
|
|
887
|
+
const loop = task.ralph_loop?.allowed === true;
|
|
888
|
+
const maxIterations = loop ? task.ralph_loop.max_iterations : 1;
|
|
889
|
+
const monotonicNow = dependencies.monotonicNow ?? (() => performance.now());
|
|
890
|
+
const deadline = monotonicNow() + run.timeoutMs;
|
|
891
|
+
const remainingTaskMs = () => {
|
|
892
|
+
if (!loop) return run.timeoutMs;
|
|
893
|
+
const remaining = Math.floor(deadline - monotonicNow());
|
|
894
|
+
if (remaining < 1) throw new Error(`${task.id} exhausted its task deadline.`);
|
|
895
|
+
return remaining;
|
|
896
|
+
};
|
|
897
|
+
const continuation = { enabled: loop, maxIterations, iterations: [] };
|
|
809
898
|
const worktree = resolve(paths.worktreesRoot, run.id, repository.name, `${task.id}-attempt-${attempt}`);
|
|
810
899
|
const actualBranch = taskBranch(repository.root, run, task, attempt);
|
|
811
|
-
const logRoot = resolve(paths.runRoot, 'tasks', task.id);
|
|
900
|
+
const logRoot = resolve(paths.runRoot, 'tasks', task.id, `attempt-${attempt}`);
|
|
812
901
|
let lease = null;
|
|
813
902
|
try {
|
|
814
903
|
lease = acquireExecutionLease(projectRoot, run.intentId, {
|
|
@@ -832,50 +921,65 @@ async function implementTask(projectRoot, run, task, provider, dependencies = {}
|
|
|
832
921
|
task.red_green_refactor.red_command,
|
|
833
922
|
worktree,
|
|
834
923
|
resolve(logRoot, 'red.txt'),
|
|
835
|
-
|
|
924
|
+
remainingTaskMs(), projectRoot, run, task,
|
|
836
925
|
);
|
|
837
926
|
red.stage = 'red';
|
|
838
927
|
if (red.exitCode === 0) throw new Error(`${task.id} red command already passes; the task contract is stale and must be reconciled.`);
|
|
839
|
-
|
|
840
|
-
|
|
841
|
-
|
|
842
|
-
|
|
843
|
-
|
|
844
|
-
|
|
845
|
-
|
|
846
|
-
|
|
847
|
-
task,
|
|
848
|
-
|
|
849
|
-
|
|
850
|
-
|
|
851
|
-
|
|
852
|
-
|
|
853
|
-
|
|
854
|
-
|
|
855
|
-
|
|
856
|
-
|
|
857
|
-
|
|
858
|
-
|
|
928
|
+
let green;
|
|
929
|
+
for (let iteration = 1; iteration <= maxIterations; iteration++) {
|
|
930
|
+
const control = requireRun(projectRoot, run.id).run;
|
|
931
|
+
if (control.desiredState === 'cancelled' || iteration > 1 && control.desiredState === 'paused') throw new Error(`AFK continuation stopped: ${control.desiredState}.`);
|
|
932
|
+
const timeoutMs = remainingTaskMs();
|
|
933
|
+
heartbeatExecutionLease(projectRoot, lease.id, {
|
|
934
|
+
token: lease.token, durationMs: Math.min(86_400_000, timeoutMs + 10 * 60 * 1000),
|
|
935
|
+
});
|
|
936
|
+
taskState(run, task, { status: 'implementing', continuation });
|
|
937
|
+
saveRun(paths, run);
|
|
938
|
+
event(projectRoot, run, `Agent implementing ${task.id}.`, `${provider} · iteration ${iteration}/${maxIterations}`);
|
|
939
|
+
const prior = continuation.iterations.at(-1)?.verification;
|
|
940
|
+
const feedback = prior
|
|
941
|
+
? `\nCONTINUATION ${iteration}/${maxIterations}: the previous turn ended, but the task is incomplete. Continue from the files already in this worktree. The declared green command exited ${prior.exitCode}. Read its captured output at ${prior.outputPath}, repair only within write_set and recheck the declared tests. A summary or completion promise is not completion.`
|
|
942
|
+
: '';
|
|
943
|
+
if (prior && sha256(readFileSync(prior.outputPath)) !== prior.outputSha256) throw new Error('Continuation verification output changed; inspect preserved evidence.');
|
|
944
|
+
const continuationPrompt = (loop ? '\nWork to completion within this task. If a stop condition or owner decision prevents completion, end with EWAI_BLOCKED: followed by the reason.' : '') + feedback;
|
|
945
|
+
const implementationEvidence = provider === 'grok' ? prepareGrokTaskEvidence(projectRoot, worktree, task, { mode: 'implementation', commands: [red, ...(green ? [green] : [])] }) : null;
|
|
946
|
+
const invocation = buildProviderInvocation(provider, {
|
|
947
|
+
cwd: worktree, prompt: taskPrompt(projectRoot, task, repository, implementationEvidence) + continuationPrompt,
|
|
948
|
+
grokEvidence: implementationEvidence, task, timeoutMs: remainingTaskMs(), mode: 'implementation',
|
|
949
|
+
codingPolicy: run.codingPolicy, policyDigest: run.codingPolicyDigest, policyRoot: projectRoot,
|
|
950
|
+
});
|
|
951
|
+
invocation.logPath = resolve(logRoot, loop ? `implementation-iteration-${iteration}.log` : 'implementation.log');
|
|
952
|
+
const result = await invokeControlled(projectRoot, run, task, invocation, invoke);
|
|
953
|
+
const entry = { iteration, providerExitCode: result.exitCode, executionStopped: result.executionStopped === true };
|
|
954
|
+
continuation.iterations.push(entry);
|
|
955
|
+
taskState(run, task, { continuation });
|
|
956
|
+
atomicJson(resolve(logRoot, 'continuation.json'), continuation);
|
|
957
|
+
saveRun(paths, run);
|
|
958
|
+
if (requireRun(projectRoot, run.id).run.desiredState === 'cancelled' || result.cancelled) throw new Error('AFK run was cancelled while the task agent was active.');
|
|
959
|
+
if (git(worktree, ['rev-parse', 'HEAD']) !== baseHead) throw new Error(`${task.id} created a commit; commits and merges belong to the conductor.`);
|
|
960
|
+
if (result.timedOut) throw new Error(`${provider} exceeded the task timeout.`);
|
|
961
|
+
if (result.exitCode !== 0) throw new Error(`${provider} exited ${result.exitCode}; see ${relative(projectRoot, result.logPath)}.`);
|
|
962
|
+
// Check scope before running tests or giving another worker turn authority.
|
|
963
|
+
const outside = changedPaths(worktree).filter(path => !allowedPath(path, task));
|
|
964
|
+
if (outside.length) throw new Error(`${task.id} changed files outside its write_set: ${outside.join(', ')}.`);
|
|
965
|
+
if (loop && workerReportedBlocker(result.output)) throw new Error(`${task.id} worker reported a blocker; inspect its preserved log.`);
|
|
966
|
+
green = await executeCommand(task.red_green_refactor.green_command, worktree,
|
|
967
|
+
resolve(logRoot, loop ? `green-iteration-${iteration}.txt` : 'green.txt'), remainingTaskMs(), projectRoot, run, task);
|
|
968
|
+
green.stage = 'green';
|
|
969
|
+
entry.verification = green;
|
|
970
|
+
atomicJson(resolve(logRoot, 'continuation.json'), continuation);
|
|
971
|
+
taskState(run, task, { continuation });
|
|
972
|
+
saveRun(paths, run);
|
|
973
|
+
if (green.exitCode === 0) break;
|
|
974
|
+
if (!loop) throw new Error(`${task.id} green command failed after implementation.`);
|
|
975
|
+
if (iteration === maxIterations) throw new Error(`${task.id} reached its continuation iteration limit (${maxIterations}).`);
|
|
976
|
+
event(projectRoot, run, `Continuing incomplete task ${task.id}.`, `Green verification failed on iteration ${iteration}; remaining work stays in its isolated worktree.`);
|
|
859
977
|
}
|
|
860
|
-
if (result.timedOut) throw new Error(`${provider} exceeded the ${run.timeoutMs}ms task timeout.`);
|
|
861
|
-
if (result.exitCode !== 0) throw new Error(`${provider} exited ${result.exitCode}; see ${relative(projectRoot, result.logPath)}.`);
|
|
862
|
-
heartbeatExecutionLease(projectRoot, lease.id, {
|
|
863
|
-
token: lease.token,
|
|
864
|
-
durationMs: Math.min(86_400_000, run.timeoutMs + 10 * 60 * 1000),
|
|
865
|
-
});
|
|
866
|
-
const green = await executeCommand(
|
|
867
|
-
task.red_green_refactor.green_command,
|
|
868
|
-
worktree,
|
|
869
|
-
resolve(logRoot, 'green.txt'),
|
|
870
|
-
run.timeoutMs, projectRoot, run, task,
|
|
871
|
-
);
|
|
872
|
-
green.stage = 'green';
|
|
873
|
-
if (green.exitCode !== 0) throw new Error(`${task.id} green command failed after implementation.`);
|
|
874
978
|
const refactor = await executeCommand(
|
|
875
979
|
task.red_green_refactor.green_command,
|
|
876
980
|
worktree,
|
|
877
981
|
resolve(logRoot, 'refactor.txt'),
|
|
878
|
-
|
|
982
|
+
remainingTaskMs(), projectRoot, run, task,
|
|
879
983
|
);
|
|
880
984
|
refactor.stage = 'refactor';
|
|
881
985
|
if (refactor.exitCode !== 0) throw new Error(`${task.id} refactor verification failed.`);
|
|
@@ -903,7 +1007,7 @@ async function implementTask(projectRoot, run, task, provider, dependencies = {}
|
|
|
903
1007
|
prompt: reviewPrompt(projectRoot, task, implementationCommit,reviewEvidence),
|
|
904
1008
|
grokEvidence:reviewEvidence,
|
|
905
1009
|
task,
|
|
906
|
-
timeoutMs:
|
|
1010
|
+
timeoutMs: remainingTaskMs(),
|
|
907
1011
|
mode: 'review',
|
|
908
1012
|
codingPolicy:run.codingPolicy,policyDigest:run.codingPolicyDigest,policyRoot:projectRoot,
|
|
909
1013
|
});
|
|
@@ -916,6 +1020,7 @@ async function implementTask(projectRoot, run, task, provider, dependencies = {}
|
|
|
916
1020
|
if (review.timedOut || review.exitCode !== 0 || !/^VERDICT:\s*PASS\s*$/im.test(review.output)) {
|
|
917
1021
|
throw new Error(`Fresh-context review did not pass; see ${relative(projectRoot, review.logPath)}.`);
|
|
918
1022
|
}
|
|
1023
|
+
review.claimResults = parseCompletionClaimReview(review.output, task.completion_evidence ?? [], {observed:true,requirements:task.completion_requirements ?? []});
|
|
919
1024
|
authorityChecks.get(run)?.({ stage: 'accept', mode: 'review' });
|
|
920
1025
|
if (git(worktree, ['rev-parse', 'HEAD']) !== implementationCommit || gitDirty(worktree)) {
|
|
921
1026
|
throw new Error('Fresh-context review changed the inspected implementation.');
|
|
@@ -927,6 +1032,7 @@ async function implementTask(projectRoot, run, task, provider, dependencies = {}
|
|
|
927
1032
|
token: lease.token,
|
|
928
1033
|
durationMs: Math.min(86_400_000, run.timeoutMs + 10 * 60 * 1000),
|
|
929
1034
|
});
|
|
1035
|
+
review.reviewResults = reviewResults;
|
|
930
1036
|
const branchHead = git(worktree, ['rev-parse', 'HEAD']);
|
|
931
1037
|
const reviewerIdentity = reviewers.map(reviewer=>`${reviewer}/${reviewer === provider ? 'fresh-session' : 'independent-cli'}`).join(',');
|
|
932
1038
|
taskState(run, task, { status: 'ready-to-integrate', implementationCommit, branchHead, reviewer: reviewerIdentity });
|
|
@@ -934,7 +1040,7 @@ async function implementTask(projectRoot, run, task, provider, dependencies = {}
|
|
|
934
1040
|
return {
|
|
935
1041
|
task, lease, worktree, branchHead, implementationCommit, review,
|
|
936
1042
|
reviewer: reviewerIdentity, verifiedCommands: [red, green, refactor], changedFiles: changed,
|
|
937
|
-
actualBranch, repository,
|
|
1043
|
+
actualBranch, repository, continuation, remainingTaskMs,
|
|
938
1044
|
};
|
|
939
1045
|
} catch (error) {
|
|
940
1046
|
const control = requireRun(projectRoot, run.id).run;
|
|
@@ -960,6 +1066,7 @@ async function integrateTask(projectRoot, run, result) {
|
|
|
960
1066
|
const specsRepository = run.topology.repositories.find((candidate) => candidate.name === run.topology.specsRepository);
|
|
961
1067
|
if (!specsRepository) throw new Error('AFK run has no configured SPECS repository.');
|
|
962
1068
|
const paths = pathsFor(projectRoot, run.id);
|
|
1069
|
+
result.remainingTaskMs?.();
|
|
963
1070
|
taskState(run, task, { status: 'integrating' });
|
|
964
1071
|
saveRun(paths, run);
|
|
965
1072
|
event(projectRoot, run, `Integrating ${task.id}.`, 'The orchestrator is merging sequentially and running post-merge checks.');
|
|
@@ -982,7 +1089,7 @@ async function integrateTask(projectRoot, run, result) {
|
|
|
982
1089
|
const untracked = git(root, ['ls-files', '--others', '--exclude-standard', '-z']).split('\0').filter(Boolean);
|
|
983
1090
|
if (git(root, ['rev-parse', 'HEAD']) !== baseline.head || git(root, ['branch', '--show-current']) !== baseline.branch
|
|
984
1091
|
|| stagedDelta.some(path => !staged || !owned(path) || sha256(execFileSync('git', ['show', `:${path}`],
|
|
985
|
-
{ cwd: root, stdio: ['ignore', 'pipe', 'pipe'] })) !== ownedEvidence.get(resolve(root, path)))
|
|
1092
|
+
{ cwd: root, stdio: ['ignore', 'pipe', 'pipe'], maxBuffer: Math.max(1024 * 1024, statSync(resolve(root, path)).size + 1024) })) !== ownedEvidence.get(resolve(root, path)))
|
|
986
1093
|
|| [...workingDelta, ...untracked].some(path => !owned(path))) {
|
|
987
1094
|
integrationConflict = true;
|
|
988
1095
|
throw new Error('Repository changed outside the owned integration; preserve the checkout for review.');
|
|
@@ -1008,7 +1115,7 @@ async function integrateTask(projectRoot, run, result) {
|
|
|
1008
1115
|
const postMerge = [];
|
|
1009
1116
|
for (const [index, command] of (task.merge?.post_merge_checks ?? []).entries()) {
|
|
1010
1117
|
const outputPath = resolve(delivery.deliveryRoot, 'tasks', task.id, 'evidence', `post-merge-${index + 1}.txt`);
|
|
1011
|
-
const check = await executeCommand(command, repository.root, outputPath, run.timeoutMs, projectRoot, run, task);
|
|
1118
|
+
const check = await executeCommand(command, repository.root, outputPath, result.remainingTaskMs?.() ?? run.timeoutMs, projectRoot, run, task);
|
|
1012
1119
|
const evidencePath = relative(delivery.deliveryRoot, outputPath).replaceAll('\\', '/');
|
|
1013
1120
|
postMerge.push({ command, exit_code: check.exitCode, output_path: evidencePath, output_sha256: sha256(readFileSync(outputPath)) });
|
|
1014
1121
|
ownedEvidence.set(realpathSync(outputPath), postMerge.at(-1).output_sha256);
|
|
@@ -1145,14 +1252,15 @@ export async function executeAfkRun(projectRoot, runId, dependencies = {}) {
|
|
|
1145
1252
|
} catch (error) {
|
|
1146
1253
|
const control = requireRun(paths.projectRoot, run.id).run;
|
|
1147
1254
|
run.desiredState = control.desiredState;
|
|
1148
|
-
run.status = control.desiredState === 'cancelled' ? run.unconfirmedExecution ? 'cancel-unknown' : 'cancelled'
|
|
1255
|
+
run.status = control.desiredState === 'cancelled' ? run.unconfirmedExecution ? 'cancel-unknown' : 'cancelled'
|
|
1256
|
+
: error.message === 'AFK continuation stopped: paused.' && !run.unconfirmedExecution ? 'paused' : 'blocked';
|
|
1149
1257
|
if (control.desiredState === 'cancelled') run.cancellation = { ...control.cancellation,
|
|
1150
1258
|
status: run.unconfirmedExecution ? 'unknown' : 'confirmed', settledAt: now() };
|
|
1151
1259
|
run.pid = null;
|
|
1152
1260
|
run.messages.push({ at: now(), message: error.message });
|
|
1153
1261
|
run.executionStopped = !run.unconfirmedExecution;
|
|
1154
1262
|
if (!run.unconfirmedExecution) abandonExecutionLeasesForRun(paths.projectRoot, run.id, { outcome: 'conductor-blocked' });
|
|
1155
|
-
event(paths.projectRoot, run, 'Unattended Build is blocked.', error.message, 'blocked',
|
|
1263
|
+
event(paths.projectRoot, run, run.status === 'paused' ? 'Unattended Build safely paused.' : 'Unattended Build is blocked.', error.message, run.status === 'paused' ? 'progress' : 'blocked', run.status !== 'paused');
|
|
1156
1264
|
saveRun(paths, run);
|
|
1157
1265
|
return run;
|
|
1158
1266
|
} finally { authorityChecks.delete(run); }
|
|
@@ -152,6 +152,7 @@ export function invokeProvider(invocation, options = {}) {
|
|
|
152
152
|
const maxCapture = Number(options.maxCaptureBytes ?? 1024 * 1024);
|
|
153
153
|
|
|
154
154
|
return new Promise((resolvePromise, reject) => {
|
|
155
|
+
output.once('error', reject);
|
|
155
156
|
const child = spawn(invocation.command, invocation.args, {
|
|
156
157
|
cwd: invocation.cwd,
|
|
157
158
|
env: options.env ?? process.env,
|
|
@@ -204,8 +205,9 @@ export function invokeProvider(invocation, options = {}) {
|
|
|
204
205
|
try { process.kill(-child.pid, 0); kill('SIGKILL'); }
|
|
205
206
|
catch (error) { processGroupStopped = error.code === 'ESRCH'; }
|
|
206
207
|
}
|
|
207
|
-
|
|
208
|
-
|
|
208
|
+
// Return only after the captured log has finished flushing. Callers hash
|
|
209
|
+
// this file immediately; a process exit alone does not settle its stream.
|
|
210
|
+
output.end(() => resolvePromise({
|
|
209
211
|
provider: invocation.provider,
|
|
210
212
|
exitCode: exitCode ?? -1,
|
|
211
213
|
signal: signal ?? '',
|
|
@@ -223,7 +225,7 @@ export function invokeProvider(invocation, options = {}) {
|
|
|
223
225
|
invocation.provider,
|
|
224
226
|
options.providerUsage ?? invocation.providerUsage,
|
|
225
227
|
),
|
|
226
|
-
});
|
|
228
|
+
}));
|
|
227
229
|
});
|
|
228
230
|
options.signal?.addEventListener('abort', cancel, { once: true });
|
|
229
231
|
if (options.signal?.aborted) cancel();
|
package/src/task-graph.mjs
CHANGED
|
@@ -110,6 +110,41 @@ function evidenceFile(root, configuredPath, expectedHash, label, errors) {
|
|
|
110
110
|
}
|
|
111
111
|
}
|
|
112
112
|
|
|
113
|
+
// Semantic support is assessed by the existing fresh reviewer; the host enforces
|
|
114
|
+
// complete, explicit results rather than inferring every claim from one green exit.
|
|
115
|
+
export function validateCompletionClaims(results, names, options = {}) {
|
|
116
|
+
const errors = [], expected = new Set(names), seen = new Set();
|
|
117
|
+
if (!Array.isArray(results)) return ['completion claim validation results are missing'];
|
|
118
|
+
for (const result of results) {
|
|
119
|
+
if (!result || !expected.has(result.name)) errors.push('completion claim validation contains an unknown claim');
|
|
120
|
+
if (seen.has(result?.name)) errors.push('completion claim validation contains a duplicate claim');
|
|
121
|
+
seen.add(result?.name);
|
|
122
|
+
if (result?.status !== 'pass' || !text(result?.support)) errors.push('completion claim validation requires a passing result with specific supporting evidence');
|
|
123
|
+
}
|
|
124
|
+
if (options.observed) for (const result of results) {
|
|
125
|
+
const requiredType = options.requirements?.find(item => item.name === result?.name)?.evidence_type ?? 'behaviour';
|
|
126
|
+
const observation = result?.observation, challenge = result?.challenge;
|
|
127
|
+
if (!['behaviour','configuration','command-result','review-result'].includes(requiredType) || observation?.evidence_type !== requiredType) errors.push('completion claim evidence type does not match its planned obligation');
|
|
128
|
+
if (!['expected','actual','evidence'].every(key => typeof observation?.[key] === 'string' && observation[key].trim())) errors.push('completion claim requires expected and actual observed result with evidence');
|
|
129
|
+
if (!['attempt','result'].every(key => typeof challenge?.[key] === 'string' && challenge[key].trim())) errors.push('completion claim requires a concrete challenge attempt and result');
|
|
130
|
+
}
|
|
131
|
+
for (const name of expected) if (!seen.has(name)) errors.push(`completion claim validation is missing: ${name}`);
|
|
132
|
+
return errors;
|
|
133
|
+
}
|
|
134
|
+
|
|
135
|
+
export function parseCompletionClaimReview(output, names, options = {}) {
|
|
136
|
+
const results = [];
|
|
137
|
+
for (const line of String(output).split(/\r?\n/)) {
|
|
138
|
+
const match = /^COMPLETION_CHECK:\s*(.+)$/.exec(line.trim());
|
|
139
|
+
if (!match) continue;
|
|
140
|
+
try { results.push(JSON.parse(match[1])); }
|
|
141
|
+
catch { throw new Error('Completion claim validation contains malformed JSON.'); }
|
|
142
|
+
}
|
|
143
|
+
const errors = validateCompletionClaims(results, names, options);
|
|
144
|
+
if (errors.length) throw new Error(errors.join('; '));
|
|
145
|
+
return results;
|
|
146
|
+
}
|
|
147
|
+
|
|
113
148
|
export function validateTaskEvidence(deliveryRoot, task) {
|
|
114
149
|
const root = resolve(deliveryRoot);
|
|
115
150
|
const configured = text(task?.evidence_path);
|
|
@@ -129,7 +164,12 @@ export function validateTaskEvidence(deliveryRoot, task) {
|
|
|
129
164
|
errors.push(`structured task evidence is not valid JSON: ${error.message}`);
|
|
130
165
|
}
|
|
131
166
|
if (evidence) {
|
|
132
|
-
if (evidence.
|
|
167
|
+
if (!['ewai.task-evidence/v1', 'ewai.task-evidence/v2', 'ewai.task-evidence/v3'].includes(evidence.schema)) errors.push('unsupported task evidence schema');
|
|
168
|
+
const observedValidation = evidence.schema === 'ewai.task-evidence/v3' || task.claim_validation?.version === 3;
|
|
169
|
+
const claimOptions = {observed:observedValidation,requirements:task.completion_requirements ?? []};
|
|
170
|
+
if (task.claim_validation?.version === 3 && evidence.schema !== 'ewai.task-evidence/v3') errors.push('required observed claim validation cannot use earlier evidence');
|
|
171
|
+
const claimValidation = ['ewai.task-evidence/v2','ewai.task-evidence/v3'].includes(evidence.schema) || task.claim_validation?.required === true;
|
|
172
|
+
if (task.claim_validation?.required === true && !['ewai.task-evidence/v2','ewai.task-evidence/v3'].includes(evidence.schema)) errors.push('required completion claim validation cannot use legacy evidence');
|
|
133
173
|
if (evidence.task_id !== task.id) errors.push(`task_id must be ${task.id}`);
|
|
134
174
|
if (text(evidence.task_branch)) {
|
|
135
175
|
if (evidence.task_branch !== task.branch) errors.push(`task_branch must be ${task.branch}`);
|
|
@@ -166,6 +206,24 @@ export function validateTaskEvidence(deliveryRoot, task) {
|
|
|
166
206
|
}
|
|
167
207
|
evidenceFile(root, command.output_path, command.output_sha256, `${command.stage || 'command'} output`, errors);
|
|
168
208
|
}
|
|
209
|
+
if (evidence.continuation) {
|
|
210
|
+
const history = evidence.continuation;
|
|
211
|
+
const enabled = task.ralph_loop?.allowed === true;
|
|
212
|
+
const limit = enabled ? task.ralph_loop.max_iterations : 1;
|
|
213
|
+
if (history.enabled !== enabled || history.max_iterations !== limit) errors.push('continuation budget must match the task contract');
|
|
214
|
+
const iterations = history.iterations;
|
|
215
|
+
if (!Array.isArray(iterations) || !iterations.length || iterations.length > limit) {
|
|
216
|
+
errors.push('continuation history must be non-empty and within its iteration budget');
|
|
217
|
+
} else {
|
|
218
|
+
for (const [index, iteration] of iterations.entries()) {
|
|
219
|
+
const check = iteration?.verification;
|
|
220
|
+
if (iteration?.iteration !== index + 1 || iteration?.provider_exit_code !== 0 || iteration?.execution_stopped !== true) errors.push('continuation iteration must record an ordered, successful, confirmed stopped worker');
|
|
221
|
+
if (!check || check.command !== task.red_green_refactor?.green_command || !Number.isInteger(check.exit_code)
|
|
222
|
+
|| (index === iterations.length - 1 ? check.exit_code !== 0 : check.exit_code === 0)) errors.push('continuation verification must retain failures before the final passing green');
|
|
223
|
+
if (check) evidenceFile(root, check.output_path, check.output_sha256, 'continuation verification output', errors);
|
|
224
|
+
}
|
|
225
|
+
}
|
|
226
|
+
}
|
|
169
227
|
const review = evidence.review ?? {};
|
|
170
228
|
if (review.fresh_context !== true) errors.push('review.fresh_context must be true');
|
|
171
229
|
if (review.tests_reviewed_first !== true) errors.push('review.tests_reviewed_first must be true');
|
|
@@ -174,10 +232,34 @@ export function validateTaskEvidence(deliveryRoot, task) {
|
|
|
174
232
|
evidenceFile(root, review.evidence_path, review.evidence_sha256, 'review evidence', errors);
|
|
175
233
|
const checks = Array.isArray(evidence.completion_checks) ? evidence.completion_checks : [];
|
|
176
234
|
const checksByName = new Map(checks.map((check) => [text(check.name), check]));
|
|
235
|
+
if (claimValidation) {
|
|
236
|
+
if (!errors.length) {
|
|
237
|
+
try {
|
|
238
|
+
const captured = JSON.parse(readFileSync(resolve(root, review.evidence_path), 'utf8'));
|
|
239
|
+
if (captured.schema !== 'ewai.claim-review/v1' || !Array.isArray(captured.reviews) || !captured.reviews.length) throw new Error('captured claim reviews are missing');
|
|
240
|
+
const identities = String(review.reviewer).split(',').map(value => value.trim().split('/')[0]);
|
|
241
|
+
if (captured.reviews.length !== identities.length || new Set(captured.reviews.map(item => item.reviewer)).size !== identities.length) throw new Error('captured claim reviewer coverage differs');
|
|
242
|
+
let recorded;
|
|
243
|
+
for (const item of captured.reviews) {
|
|
244
|
+
if (!identities.includes(item.reviewer)) throw new Error('unknown captured claim reviewer');
|
|
245
|
+
recorded = parseCompletionClaimReview(item.output, list(task.completion_evidence), claimOptions);
|
|
246
|
+
}
|
|
247
|
+
if (JSON.stringify(recorded) !== JSON.stringify(review.claim_results)) throw new Error('claim results differ from captured review');
|
|
248
|
+
} catch (error) { errors.push(`completion claim validation differs from captured review: ${error.message}`); }
|
|
249
|
+
}
|
|
250
|
+
errors.push(...validateCompletionClaims(review.claim_results, list(task.completion_evidence), claimOptions));
|
|
251
|
+
if (checks.length !== list(task.completion_evidence).length || checksByName.size !== checks.length) errors.push('completion claim validation must cover exactly the planned checks');
|
|
252
|
+
}
|
|
177
253
|
for (const planned of list(task.completion_evidence)) {
|
|
178
254
|
const check = checksByName.get(planned);
|
|
179
255
|
if (!check || check.status !== 'pass') errors.push(`planned completion check has no passing evidence: ${planned}`);
|
|
180
|
-
else
|
|
256
|
+
else {
|
|
257
|
+
evidenceFile(root, check.evidence_path, check.evidence_sha256, `completion check ${planned}`, errors);
|
|
258
|
+
if (claimValidation) {
|
|
259
|
+
const result = Array.isArray(review.claim_results) ? review.claim_results.find(item => item?.name === planned) : null;
|
|
260
|
+
if (!result || check.support !== result.support || observedValidation && (JSON.stringify(check.observation) !== JSON.stringify(result.observation) || JSON.stringify(check.challenge) !== JSON.stringify(result.challenge)) || check.evidence_path !== review.evidence_path || check.evidence_sha256 !== review.evidence_sha256) errors.push(`completion claim validation is not bound to its supporting review: ${planned}`);
|
|
261
|
+
}
|
|
262
|
+
}
|
|
181
263
|
}
|
|
182
264
|
}
|
|
183
265
|
}
|
|
@@ -412,8 +494,8 @@ export function validateTaskGraph(deliveryRoot, options = {}) {
|
|
|
412
494
|
if (!list(task.completion_evidence).length) add('afk-completion-evidence-missing', `${id} requires completion evidence.`, id);
|
|
413
495
|
}
|
|
414
496
|
if (task?.ralph_loop?.allowed === true) {
|
|
415
|
-
if (!Number.isInteger(task.ralph_loop.max_iterations) || task.ralph_loop.max_iterations < 1) {
|
|
416
|
-
add('ralph-loop-unbounded', `${id} Ralph loop requires
|
|
497
|
+
if (!Number.isInteger(task.ralph_loop.max_iterations) || task.ralph_loop.max_iterations < 1 || task.ralph_loop.max_iterations > 100) {
|
|
498
|
+
add('ralph-loop-unbounded', `${id} Ralph loop requires max_iterations between 1 and 100.`, id);
|
|
417
499
|
}
|
|
418
500
|
if (!text(task.ralph_loop.completion_promise)) add('ralph-loop-promise-missing', `${id} Ralph loop requires a completion promise.`, id);
|
|
419
501
|
}
|
|
@@ -421,6 +503,16 @@ export function validateTaskGraph(deliveryRoot, options = {}) {
|
|
|
421
503
|
if (task?.review?.fresh_context_required !== true) add('fresh-review-required', `${id} requires fresh-context review.`, id);
|
|
422
504
|
}
|
|
423
505
|
|
|
506
|
+
for (const task of tasks) {
|
|
507
|
+
if (task.completion_requirements !== undefined) {
|
|
508
|
+
const requirements = task.completion_requirements;
|
|
509
|
+
const names = Array.isArray(requirements) ? requirements.map(item => item?.name) : [];
|
|
510
|
+
if (!Array.isArray(requirements) || names.length !== list(task.completion_evidence).length || new Set(names).size !== names.length || names.some(name => !list(task.completion_evidence).includes(name))
|
|
511
|
+
|| requirements.some(item => !['behaviour','configuration','command-result','review-result'].includes(item?.evidence_type))) add('completion-requirements-invalid', `${task.id} requires one valid evidence type per completion claim.`, task.id);
|
|
512
|
+
}
|
|
513
|
+
if (task.claim_validation?.version !== undefined && task.claim_validation.version !== 3) add('claim-validation-version-invalid', `${task.id} declares an unsupported claim validation version.`, task.id);
|
|
514
|
+
}
|
|
515
|
+
|
|
424
516
|
const ledgerClaims = recordsFrom(claimLedger, ['implementation_claims', 'claims']);
|
|
425
517
|
const actionableClaims = new Set(ledgerClaims.filter(actionableClaim).map(recordId).filter(Boolean));
|
|
426
518
|
if (!ledgerClaims.length) add('claim-ledger-empty', 'The Claim Ledger must contain implementation claims.', graph.source?.claim_ledger);
|