devflow-kit 3.1.0 → 3.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +52 -0
- package/README.md +2 -2
- package/dist/cli/agents-view/render.js +69 -15
- package/dist/cli/agents-view/state.js +40 -14
- package/dist/cli/commands/agents.js +135 -45
- package/dist/cli/commands/init.js +128 -53
- package/dist/cli/commands/learning.js +61 -13
- package/dist/cli/commands/memory.js +35 -14
- package/dist/cli/commands/uninstall.js +163 -39
- package/dist/commands/code-review.md +1 -3
- package/dist/commands/debug.md +15 -12
- package/dist/commands/dynamic-build.md +172 -135
- package/dist/commands/dynamic-plan.md +9 -3
- package/dist/commands/explore.md +10 -4
- package/dist/commands/implement.md +149 -145
- package/dist/commands/plan.md +13 -9
- package/dist/commands/release.md +8 -2
- package/dist/commands/research.md +8 -2
- package/dist/commands/resolve.md +28 -19
- package/dist/commands/self-review.md +16 -13
- package/dist/core/agent-frontmatter.js +25 -0
- package/dist/core/agent-models.js +201 -36
- package/dist/core/agent-state.js +27 -5
- package/dist/core/assets.js +1 -1
- package/dist/core/feature-config.js +68 -10
- package/dist/core/flags.js +24 -0
- package/dist/core/learning-queue-cleanup.js +10 -11
- package/dist/core/learning-tuning-config.js +8 -0
- package/dist/core/linked-path.js +46 -0
- package/dist/core/plugins.js +16 -5
- package/dist/core/queue-drain.js +31 -0
- package/dist/hud/components/learning-counts.js +54 -8
- package/dist/skills/git/references/tracker/github/create-release.md +2 -2
- package/dist/skills/git/references/tracker/github/gather-release-evidence.md +1 -1
- package/dist/skills/git/references/tracker/jira/create-release.md +2 -2
- package/dist/skills/git/references/tracker/jira/gather-release-evidence.md +1 -1
- package/dist/skills/git/references/tracker/linear/create-release.md +2 -2
- package/dist/skills/git/references/tracker/linear/gather-release-evidence.md +1 -1
- package/dist/targets/claude-code/installer.js +36 -9
- package/dist/targets/claude-code/post-install.js +128 -38
- package/package.json +1 -1
- package/src/assets/agents/code.md +85 -35
- package/src/assets/agents/design.md +12 -0
- package/src/assets/agents/diagnose.md +18 -11
- package/src/assets/agents/evaluate.md +17 -24
- package/src/assets/agents/knowledge.md +7 -3
- package/src/assets/agents/learning.md +4 -6
- package/src/assets/agents/research.md +21 -0
- package/src/assets/agents/review.md +12 -0
- package/src/assets/agents/scrutinize.md +37 -9
- package/src/assets/agents/simplify.md +24 -0
- package/src/assets/agents/skim.md +6 -2
- package/src/assets/agents/synthesize.md +18 -0
- package/src/assets/agents/test.md +19 -11
- package/src/assets/agents/triage.md +8 -0
- package/src/assets/agents/validate.md +20 -11
- package/src/assets/commands/_partials/_engine.mds +36 -55
- package/src/assets/commands/_partials/_knowledge.mds +1 -3
- package/src/assets/commands/_partials/_plan_contract.mds +1 -1
- package/src/assets/commands/_partials/_tracker.mds +1 -1
- package/src/assets/commands/_partials/_wave.mds +8 -6
- package/src/assets/commands/code-review.mds +1 -3
- package/src/assets/commands/debug.mds +13 -8
- package/src/assets/commands/dynamic-build.mds +126 -72
- package/src/assets/commands/dynamic-plan.mds +7 -1
- package/src/assets/commands/explore.mds +9 -1
- package/src/assets/commands/implement.mds +147 -141
- package/src/assets/commands/plan.mds +12 -8
- package/src/assets/commands/release.md +8 -2
- package/src/assets/commands/research.mds +8 -2
- package/src/assets/commands/resolve.mds +27 -16
- package/src/assets/commands/self-review.mds +15 -10
- package/src/assets/mds/tracker/_common.mds +1 -1
- package/src/assets/mds/tracker/_github.mds +2 -2
- package/src/assets/mds/tracker/_jira.mds +2 -2
- package/src/assets/mds/tracker/_linear.mds +2 -2
- package/src/assets/scripts/ci-wait.cjs +636 -0
- package/src/assets/scripts/hooks/assets/orchestrator-charter.md +4 -3
- package/src/assets/scripts/hooks/background-memory-update +356 -17
- package/src/assets/scripts/hooks/capture-prompt +4 -3
- package/src/assets/scripts/hooks/capture-question +4 -3
- package/src/assets/scripts/hooks/capture-turn +4 -3
- package/src/assets/scripts/hooks/ensure-devflow-init +13 -1
- package/src/assets/scripts/hooks/ensure-root-gitignore +122 -10
- package/src/assets/scripts/hooks/git-marker +71 -0
- package/src/assets/scripts/hooks/json-helper.cjs +12 -145
- package/src/assets/scripts/hooks/json-parse +24 -129
- package/src/assets/scripts/hooks/lib/learning-store.cjs +169 -64
- package/src/assets/scripts/hooks/lib/render-decisions.cjs +1 -1
- package/src/assets/scripts/hooks/memory-worker +10 -0
- package/src/assets/scripts/hooks/pre-compact-memory +66 -14
- package/src/assets/scripts/hooks/preamble +9 -1
- package/src/assets/scripts/hooks/queue-append +53 -21
- package/src/assets/scripts/hooks/session-start-context +108 -29
- package/src/assets/scripts/hooks/session-start-memory +33 -11
- package/src/assets/scripts/release-trace.cjs +27 -10
- package/src/assets/skills/accessibility/SKILL.md +1 -1
- package/src/assets/skills/apply-decisions/SKILL.md +12 -82
- package/src/assets/skills/apply-feature-knowledge/SKILL.md +8 -42
- package/src/assets/skills/architecture/SKILL.md +1 -1
- package/src/assets/skills/boundary-validation/SKILL.md +1 -1
- package/src/assets/skills/complexity/SKILL.md +1 -1
- package/src/assets/skills/compliance/SKILL.md +1 -1
- package/src/assets/skills/consistency/SKILL.md +1 -1
- package/src/assets/skills/database/SKILL.md +1 -1
- package/src/assets/skills/dependencies/SKILL.md +1 -1
- package/src/assets/skills/dependency-research/SKILL.md +3 -6
- package/src/assets/skills/design-review/SKILL.md +1 -1
- package/src/assets/skills/docs-framework/SKILL.md +1 -1
- package/src/assets/skills/documentation/SKILL.md +1 -1
- package/src/assets/skills/gap-analysis/SKILL.md +1 -1
- package/src/assets/skills/git/SKILL.md +1 -1
- package/src/assets/skills/go/SKILL.md +1 -1
- package/src/assets/skills/java/SKILL.md +1 -1
- package/src/assets/skills/patterns/SKILL.md +1 -1
- package/src/assets/skills/performance/SKILL.md +1 -1
- package/src/assets/skills/python/SKILL.md +1 -1
- package/src/assets/skills/qa/SKILL.md +1 -3
- package/src/assets/skills/quality-gates/SKILL.md +9 -12
- package/src/assets/skills/quality-gates/references/report-template.md +20 -20
- package/src/assets/skills/react/SKILL.md +1 -1
- package/src/assets/skills/regression/SKILL.md +1 -1
- package/src/assets/skills/reliability/SKILL.md +1 -1
- package/src/assets/skills/research-codebase/SKILL.md +1 -1
- package/src/assets/skills/research-competitor/SKILL.md +1 -1
- package/src/assets/skills/research-external/SKILL.md +1 -1
- package/src/assets/skills/research-technology/SKILL.md +1 -1
- package/src/assets/skills/review-methodology/SKILL.md +1 -1
- package/src/assets/skills/rust/SKILL.md +1 -1
- package/src/assets/skills/security/SKILL.md +1 -1
- package/src/assets/skills/software-design/SKILL.md +1 -1
- package/src/assets/skills/test-driven-development/SKILL.md +15 -33
- package/src/assets/skills/testing/SKILL.md +1 -1
- package/src/assets/skills/typescript/SKILL.md +1 -1
- package/src/assets/skills/ui-design/SKILL.md +1 -1
- package/src/assets/skills/worktree-support/SKILL.md +3 -55
- package/src/assets/skills/worktree-support/references/discovery.md +48 -0
- package/src/assets/skills/worktree-support/references/roots.md +2 -2
|
@@ -2,9 +2,29 @@
|
|
|
2
2
|
name: Simplify
|
|
3
3
|
description: Simplifies and refines code for clarity, consistency, and maintainability while preserving all functionality. Focuses on recently modified code unless instructed otherwise.
|
|
4
4
|
model: sonnet
|
|
5
|
+
effort: medium
|
|
5
6
|
skills:
|
|
6
7
|
- devflow:software-design
|
|
7
8
|
- devflow:worktree-support
|
|
9
|
+
disallowedTools:
|
|
10
|
+
- Agent
|
|
11
|
+
- SendMessage
|
|
12
|
+
- NotebookEdit
|
|
13
|
+
- EnterWorktree
|
|
14
|
+
- ExitWorktree
|
|
15
|
+
- ArtifactComments
|
|
16
|
+
- ArtifactData
|
|
17
|
+
- TodoWrite
|
|
18
|
+
- AskUserQuestion
|
|
19
|
+
- TaskOutput
|
|
20
|
+
- ScheduleWakeup
|
|
21
|
+
- CronCreate
|
|
22
|
+
- CronDelete
|
|
23
|
+
- CronList
|
|
24
|
+
- RemoteTrigger
|
|
25
|
+
- PushNotification
|
|
26
|
+
- DesignSync
|
|
27
|
+
- Skill
|
|
8
28
|
---
|
|
9
29
|
|
|
10
30
|
# Simplify Agent
|
|
@@ -76,6 +96,8 @@ Your refinement process:
|
|
|
76
96
|
|
|
77
97
|
You operate autonomously and proactively, refining code immediately after it's written or modified without requiring explicit requests. Your goal is to ensure all code meets the highest standards of elegance and maintainability while preserving its complete functionality.
|
|
78
98
|
|
|
99
|
+
**Gate ownership:** Run no build, test or lint command. Reading files and git diffs is not verification. Only Validate runs the full suite.
|
|
100
|
+
|
|
79
101
|
## Output
|
|
80
102
|
|
|
81
103
|
Return structured completion status:
|
|
@@ -93,6 +115,8 @@ Return structured completion status:
|
|
|
93
115
|
- {file} ({change description})
|
|
94
116
|
```
|
|
95
117
|
|
|
118
|
+
Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file (via Bash or Write) and the message gives its path. Exempt, inline in full: the completion status, `### Files Modified` and your commits.
|
|
119
|
+
|
|
96
120
|
## Boundaries
|
|
97
121
|
|
|
98
122
|
**Escalate to orchestrator:**
|
|
@@ -1,10 +1,12 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: Skim
|
|
3
3
|
description: Codebase orientation using rskim to identify relevant files, functions, and patterns for a feature or task
|
|
4
|
-
model:
|
|
4
|
+
model: haiku
|
|
5
|
+
effort: medium
|
|
5
6
|
tools: ["Bash", "Read"]
|
|
6
7
|
skills:
|
|
7
8
|
- devflow:worktree-support
|
|
9
|
+
omitClaudeMd: true
|
|
8
10
|
---
|
|
9
11
|
|
|
10
12
|
# Skim Agent
|
|
@@ -70,7 +72,7 @@ If you already know a file needs content, go straight to Read — don't skim it
|
|
|
70
72
|
|
|
71
73
|
### Step 6: Project Knowledge
|
|
72
74
|
|
|
73
|
-
|
|
75
|
+
Read the decisions TL;DR at the repository's main worktree, where the ledger lives. Run `git -C "{start}" rev-parse --path-format=absolute --show-toplevel --git-common-dir`, `{start}` being `WORKTREE_PATH` if provided, otherwise cwd. `{ledger}` is the first that applies: the main worktree — when the output is two absolute lines and line 2 ends in `/.git`, its parent, provided that directory contains `.devflow/` and is not your home directory; else the toplevel, line 1 (on a git older than 2.31, the line after the echoed flag); else `{start}`, when the command failed. If `{ledger}/.devflow/learning/decisions.md` exists, Read its first line, `<!-- TL;DR: N decisions -->`, and report N under "### Active Decisions". Only the TL;DR — intentional for token efficiency.
|
|
74
76
|
|
|
75
77
|
### Step 7: Generate Summary
|
|
76
78
|
|
|
@@ -127,6 +129,8 @@ skim also handles prose/config files (`.md`, `.json`, `.yaml`, `.toml`) — the
|
|
|
127
129
|
{Brief recommendation based on codebase structure}
|
|
128
130
|
```
|
|
129
131
|
|
|
132
|
+
Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file via Bash and the message gives its path. Exempt, inline in full: the `### Relevant Files for Task` table.
|
|
133
|
+
|
|
130
134
|
## Principles
|
|
131
135
|
|
|
132
136
|
1. **Speed and focus** — Get oriented quickly on what's relevant; task-focused exploration only
|
|
@@ -2,10 +2,20 @@
|
|
|
2
2
|
name: Synthesize
|
|
3
3
|
description: "Combines outputs from multiple agents into actionable summaries (modes: exploration, planning, review, bug-analysis, design, research)"
|
|
4
4
|
model: haiku
|
|
5
|
+
effort: medium
|
|
5
6
|
skills:
|
|
6
7
|
- devflow:review-methodology
|
|
7
8
|
- devflow:docs-framework
|
|
8
9
|
- devflow:worktree-support
|
|
10
|
+
tools:
|
|
11
|
+
- Read
|
|
12
|
+
- Write
|
|
13
|
+
- Edit
|
|
14
|
+
- Bash
|
|
15
|
+
- Grep
|
|
16
|
+
- Glob
|
|
17
|
+
- StructuredOutput
|
|
18
|
+
omitClaudeMd: true
|
|
9
19
|
---
|
|
10
20
|
|
|
11
21
|
# Synthesize Agent
|
|
@@ -24,6 +34,8 @@ The orchestrator provides:
|
|
|
24
34
|
|
|
25
35
|
**Worktree Support**: If `WORKTREE_PATH` is provided, follow the `devflow:worktree-support` skill for path resolution. If omitted, use cwd.
|
|
26
36
|
|
|
37
|
+
**Decisions and pitfalls in words**: in any text you write, state a decision or pitfall as its rule in words, never by its ledger ID, because your summaries can be posted to a pull request.
|
|
38
|
+
|
|
27
39
|
---
|
|
28
40
|
|
|
29
41
|
## Mode: Exploration
|
|
@@ -378,6 +390,12 @@ If CYCLE_NUMBER >= 3 and prior FP ratio > 70%: append "Note: High false-positive
|
|
|
378
390
|
|
|
379
391
|
---
|
|
380
392
|
|
|
393
|
+
## Output
|
|
394
|
+
|
|
395
|
+
Each mode's template above is its output.
|
|
396
|
+
|
|
397
|
+
Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file (via Bash or Write) and the message gives its path. Exempt, inline in full: the synthesis in exploration, planning and design modes. Review, bug-analysis and research modes write the summary to disk and return its path.
|
|
398
|
+
|
|
381
399
|
## Principles
|
|
382
400
|
|
|
383
401
|
1. **No new research** - Only synthesize what agents found
|
|
@@ -2,10 +2,10 @@
|
|
|
2
2
|
name: Test
|
|
3
3
|
description: Scenario-based QA agent. Designs and executes acceptance tests from criteria and implementation. Reports pass/fail with evidence — never fixes code.
|
|
4
4
|
model: sonnet
|
|
5
|
+
effort: medium
|
|
5
6
|
tools: ["Read", "Grep", "Glob", "Bash", "mcp__claude-in-chrome__tabs_context_mcp", "mcp__claude-in-chrome__tabs_create_mcp", "mcp__claude-in-chrome__navigate", "mcp__claude-in-chrome__get_page_text", "mcp__claude-in-chrome__read_page", "mcp__claude-in-chrome__find", "mcp__claude-in-chrome__form_input", "mcp__claude-in-chrome__javascript_tool", "mcp__claude-in-chrome__read_console_messages"]
|
|
6
7
|
skills:
|
|
7
8
|
- devflow:qa
|
|
8
|
-
- devflow:testing
|
|
9
9
|
- devflow:worktree-support
|
|
10
10
|
---
|
|
11
11
|
|
|
@@ -72,18 +72,24 @@ For each scenario:
|
|
|
72
72
|
|
|
73
73
|
If a previous run failed (PREVIOUS_FAILURES provided), prioritize re-testing those scenarios first.
|
|
74
74
|
|
|
75
|
-
|
|
75
|
+
**Gate ownership:** Run the scenario commands. Never the full suite. Only Validate runs the full suite.
|
|
76
76
|
|
|
77
|
-
|
|
77
|
+
## Running commands
|
|
78
78
|
|
|
79
|
-
|
|
80
|
-
`<command> > /tmp/df-test-<slug>.log 2>&1; echo "EXIT=$?" > /tmp/df-test-<slug>.done`
|
|
81
|
-
2. Poll with the `Monitor` tool (load it via ToolSearch `select:Monitor` if it is not available): set `persistent: false`, `timeout_ms` above the expected run time (e.g. 600000), and
|
|
82
|
-
`command: until [ -f /tmp/df-test-<slug>.done ]; do echo running; sleep 25; done; echo DONE; cat /tmp/df-test-<slug>.done`
|
|
83
|
-
The 25s heartbeat (≪ 180s) is delivered as a notification that keeps you alive past the watchdog.
|
|
84
|
-
3. When the monitor reports `DONE`: the scenario's command PASSED iff the `.done` file contains `EXIT=0`. Read the `.log` for evidence.
|
|
79
|
+
Run builds, typechecks, lints and tests in the foreground, each with an explicit Bash `timeout` above its expected run time. The ceiling is 600000 ms, or `BASH_MAX_TIMEOUT_MS` when set (`echo ${BASH_MAX_TIMEOUT_MS:-600000}`).
|
|
85
80
|
|
|
86
|
-
|
|
81
|
+
- Capture, then tail, in one Bash call (shell state does not persist): `LOG=$(mktemp); echo "LOG=$LOG"; <command> >"$LOG" 2>&1; rc=$?; tail -n 40 "$LOG"; echo "EXIT=$rc"`. The printed `EXIT=` value is the result; never decide one from a grep count.
|
|
82
|
+
- Never background a command and wait on it, and never poll across turns: no `sleep` or `true` turns, no sentinel-file checks, no Monitor.
|
|
83
|
+
- Prefer the scoped command for the change (a package, a path or a test file); for the whole set, one workspace-level command over a per-package loop.
|
|
84
|
+
- A run that exceeds its timeout is BLOCKED: report its duration and log path. Do not wait on it, poll it or re-run it.
|
|
85
|
+
- A run expected to exceed the ceiling is split into parts, each under about 90% of it, run in sequence. If it cannot be split, report BLOCKED with the remedy `devflow flags --set bash-max-timeout-ms=<ms>`.
|
|
86
|
+
- Never re-run a command when nothing it reads has changed.
|
|
87
|
+
- Never wrap a build or test command in `sh -c`, `bash -c`, `python3 -c` or `node -e`: permission rules deny wrapped commands they would allow directly.
|
|
88
|
+
- The same rules hold inside a dynamic Workflow sub-agent.
|
|
89
|
+
|
|
90
|
+
## Dev server (test.md only)
|
|
91
|
+
|
|
92
|
+
When a scenario needs the dev server, follow the lifecycle in `devflow:qa/references/browser-testing.md`: start the server in the background, check readiness with one bounded loop inside a single Bash call, and kill the server before you finish. Never end the turn while it runs.
|
|
87
93
|
|
|
88
94
|
## Output
|
|
89
95
|
|
|
@@ -139,9 +145,11 @@ One row per TP line, in TP order. PASS only when every scenario covering the TP
|
|
|
139
145
|
- **Remediation**: {what Code agent should fix}
|
|
140
146
|
|
|
141
147
|
### Evidence Log
|
|
142
|
-
{
|
|
148
|
+
{Each command run's `LOG=` path, for traceability}
|
|
143
149
|
```
|
|
144
150
|
|
|
151
|
+
Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file via Bash and the message gives its path. Exempt, inline in full: `### Test Plan Evidence` (HEAD line and TP table), the Scenario Results table (`| ID | TP | Type |`) and `### Failed Scenarios`. `### Evidence Log` lists each run's `LOG=` path, never raw output.
|
|
152
|
+
|
|
145
153
|
## Principles
|
|
146
154
|
|
|
147
155
|
1. **User perspective** - Test what the user asked for, not implementation internals
|
|
@@ -2,11 +2,17 @@
|
|
|
2
2
|
name: Triage
|
|
3
3
|
description: Validates review issues against blast-radius disposition matrix. Assigns one verdict per issue. Never edits code.
|
|
4
4
|
model: opus
|
|
5
|
+
effort: high
|
|
5
6
|
skills:
|
|
6
7
|
- devflow:security
|
|
7
8
|
- devflow:worktree-support
|
|
8
9
|
- devflow:apply-decisions
|
|
9
10
|
- devflow:apply-feature-knowledge
|
|
11
|
+
tools:
|
|
12
|
+
- Read
|
|
13
|
+
- Grep
|
|
14
|
+
- Glob
|
|
15
|
+
- Bash
|
|
10
16
|
---
|
|
11
17
|
|
|
12
18
|
# Triage Agent
|
|
@@ -145,6 +151,8 @@ Return the verdict ledger grouped by disposition:
|
|
|
145
151
|
- DUPLICATE: {n}
|
|
146
152
|
```
|
|
147
153
|
|
|
154
|
+
Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file via Bash and the message gives its path. Exempt, inline in full: the whole ledger, since `/resolve` checks that every issue id appears in it.
|
|
155
|
+
|
|
148
156
|
## Boundaries
|
|
149
157
|
|
|
150
158
|
**You are TRIAGE ONLY — read and judge, never write:**
|
|
@@ -2,9 +2,14 @@
|
|
|
2
2
|
name: Validate
|
|
3
3
|
description: Dedicated agent for running validation commands (build, typecheck, lint, test). Reports pass/fail with structured failure details - never fixes.
|
|
4
4
|
model: haiku
|
|
5
|
+
effort: medium
|
|
5
6
|
skills:
|
|
6
|
-
- devflow:testing
|
|
7
7
|
- devflow:worktree-support
|
|
8
|
+
tools:
|
|
9
|
+
- Bash
|
|
10
|
+
- Read
|
|
11
|
+
- Grep
|
|
12
|
+
- Glob
|
|
8
13
|
---
|
|
9
14
|
|
|
10
15
|
# Validate Agent
|
|
@@ -38,18 +43,20 @@ Execute in this order, stopping on first failure:
|
|
|
38
43
|
| 3 | Lint | `npm run lint`, `cargo clippy`, `make lint` |
|
|
39
44
|
| 4 | Test | `npm test`, `cargo test`, `make test` |
|
|
40
45
|
|
|
41
|
-
|
|
46
|
+
**Gate ownership:** Run the full suite once per HEAD. You are the only agent that does.
|
|
42
47
|
|
|
43
|
-
|
|
48
|
+
## Running commands
|
|
44
49
|
|
|
45
|
-
|
|
46
|
-
`<command> > /tmp/df-val-<slug>.log 2>&1; echo "EXIT=$?" > /tmp/df-val-<slug>.done`
|
|
47
|
-
2. Poll with the `Monitor` tool (load it via ToolSearch `select:Monitor` if it is not available): set `persistent: false`, `timeout_ms` above the expected run time (e.g. 600000), and
|
|
48
|
-
`command: until [ -f /tmp/df-val-<slug>.done ]; do echo running; sleep 25; done; echo DONE; cat /tmp/df-val-<slug>.done`
|
|
49
|
-
The 25s heartbeat (≪ 180s) is delivered as a notification that keeps you alive past the watchdog.
|
|
50
|
-
3. When the monitor reports `DONE`: the command PASSED iff the `.done` file contains `EXIT=0`. Read the `.log` for failure details to parse.
|
|
50
|
+
Run builds, typechecks, lints and tests in the foreground, each with an explicit Bash `timeout` above its expected run time. The ceiling is 600000 ms, or `BASH_MAX_TIMEOUT_MS` when set (`echo ${BASH_MAX_TIMEOUT_MS:-600000}`).
|
|
51
51
|
|
|
52
|
-
|
|
52
|
+
- Capture, then tail, in one Bash call (shell state does not persist): `LOG=$(mktemp); echo "LOG=$LOG"; <command> >"$LOG" 2>&1; rc=$?; tail -n 40 "$LOG"; echo "EXIT=$rc"`. The printed `EXIT=` value is the result; never decide one from a grep count.
|
|
53
|
+
- Never background a command and wait on it, and never poll across turns: no `sleep` or `true` turns, no sentinel-file checks, no Monitor.
|
|
54
|
+
- Prefer the scoped command for the change (a package, a path or a test file); for the whole set, one workspace-level command over a per-package loop.
|
|
55
|
+
- A run that exceeds its timeout is BLOCKED: report its duration and log path. Do not wait on it, poll it or re-run it.
|
|
56
|
+
- A run expected to exceed the ceiling is split into parts, each under about 90% of it, run in sequence. If it cannot be split, report BLOCKED with the remedy `devflow flags --set bash-max-timeout-ms=<ms>`.
|
|
57
|
+
- Never re-run a command when nothing it reads has changed.
|
|
58
|
+
- Never wrap a build or test command in `sh -c`, `bash -c`, `python3 -c` or `node -e`: permission rules deny wrapped commands they would allow directly.
|
|
59
|
+
- The same rules hold inside a dynamic Workflow sub-agent.
|
|
53
60
|
|
|
54
61
|
## Principles
|
|
55
62
|
|
|
@@ -78,7 +85,7 @@ HEAD: {the 40-hex `git rev-parse HEAD`, read before the first command}
|
|
|
78
85
|
|
|
79
86
|
### Failures (if FAIL)
|
|
80
87
|
|
|
81
|
-
#### typecheck
|
|
88
|
+
#### typecheck (first 30 lines; full output at {log path})
|
|
82
89
|
```
|
|
83
90
|
src/auth/login.ts:42:15 - error TS2339: Property 'email' does not exist on type 'User'.
|
|
84
91
|
src/auth/login.ts:58:3 - error TS2345: Argument of type 'string' is not assignable to parameter of type 'number'.
|
|
@@ -92,6 +99,8 @@ src/auth/login.ts:58:3 - error TS2345: Argument of type 'string' is not assignab
|
|
|
92
99
|
{Description of why validation couldn't run - e.g., missing dependencies, broken config}
|
|
93
100
|
```
|
|
94
101
|
|
|
102
|
+
Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file via Bash and the message gives its path. Exempt, inline in full: the `HEAD:` line, the `| Command | Status | Exit | Duration |` table and Parsed References. Failure output: at most 30 lines per failing command, then its log path.
|
|
103
|
+
|
|
95
104
|
## Boundaries
|
|
96
105
|
|
|
97
106
|
**Escalate to orchestrator (BLOCKED):**
|
|
@@ -3,23 +3,23 @@
|
|
|
3
3
|
|
|
4
4
|
ORDER IS LOAD-BEARING. Run exactly in this sequence:
|
|
5
5
|
|
|
6
|
-
1. **
|
|
7
|
-
|
|
8
|
-
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
- If
|
|
6
|
+
1. **Simplify agent** — reduce complexity, remove duplication
|
|
7
|
+
2. **Scrutinize agent** — 9-pillar self-review (deep structural analysis); returns `{"status": "PASS" | "FIXED" | "BLOCKED"}`
|
|
8
|
+
- BLOCKED (a P0 issue cannot be fixed in scope), or a missing or unrecognised status, which counts as BLOCKED → stop the pass and escalate as `scrutiny-blocked`; Validate does not run on code Scrutinize could not accept
|
|
9
|
+
3. **Validate agent** — build / typecheck / lint / test; returns `{"verdict": "PASS" | "FAIL", "details": "..."}`. It runs LAST and UNCONDITIONALLY: one full run over the HEAD that Simplify and Scrutinize left, so it covers their commits whether or not Scrutinize changed code
|
|
10
|
+
- Build/test commands follow `build_execution_doctrine()`: the Validate agent's Running commands block, foreground under an explicit Bash timeout.
|
|
11
|
+
- FAIL → Code agent fix (max 2 retries), each followed by a Validate agent re-run
|
|
12
|
+
- If still FAIL after 2 retries → escalate as `validation-exhausted` (do not loop endlessly)
|
|
13
13
|
|
|
14
14
|
**Gate 1 contains NO Evaluate agent and NO Test agent.** Those are Gate 2 only.
|
|
15
15
|
|
|
16
16
|
Depth scales to change size + budget: a trivial one-line fix warrants a lighter pass; a multi-file refactor warrants the full depth.
|
|
17
17
|
|
|
18
18
|
**Cadence — Gate 1 runs at exactly TWO points per ticket:**
|
|
19
|
-
1. **Gate 1 #1** — immediately after the initial Code agent implementation (inside `implement_bundle()`).
|
|
20
|
-
2. **Gate 1 #2** — the FINAL gate, after ALL Gate-2 fixes AND the entire review pass have completed.
|
|
19
|
+
1. **Gate 1 #1** — immediately after the initial Code agent implementation (inside `implement_bundle()`). An escalated result — Validate exhausted, or Scrutinize BLOCKED — returns early: the ticket's engine returns `ESCALATED` with one escalation, and Gate 2 and the review pass never run on a broken or unfinished build.
|
|
20
|
+
2. **Gate 1 #2** — the FINAL gate, after ALL Gate-2 fixes AND the entire review pass have completed. Scrutinize BLOCKED here is `ESCALATED` as well.
|
|
21
21
|
|
|
22
|
-
It does NOT run inside the review pass, nor after each individual Gate-2 / review / QA fix. At those points the fixing Code agent self-verifies its OWN build compiles (see `review_pass()` and the Code agent's
|
|
22
|
+
It does NOT run inside the review pass, nor after each individual Gate-2 / review / QA fix. At those points the fixing Code agent self-verifies its OWN build compiles (see `review_pass()` and the Code agent's Running commands block). The final Gate 1 #2 is the invariant that all written code passes before merge.
|
|
23
23
|
@end
|
|
24
24
|
|
|
25
25
|
@define gate2_acceptance():
|
|
@@ -29,32 +29,29 @@ Gate 2 fires ONCE: after the implement-bundle and BEFORE the review pass. It doe
|
|
|
29
29
|
|
|
30
30
|
Gate 2 inputs are produced by `/devflow:dynamic-plan`'s plan-challenge step — the acceptance criteria and test plan written for the Evaluate agent and Test agent.
|
|
31
31
|
|
|
32
|
-
**Evaluate agent
|
|
33
|
-
- Run `evaluator_panel()` — see that block for the
|
|
34
|
-
- If
|
|
32
|
+
**Evaluate agent** (only if a plan exists):
|
|
33
|
+
- Run `evaluator_panel()` — see that block for the one spawn and the two lenses it names
|
|
34
|
+
- If the Evaluate agent returns FAIL (either lens failed): fix-and-continue — the demanded fixes are applied by a Code agent that self-verifies its own build (batched per the review-pass batching doctrine if numerous). The recorded verdict becomes `FAIL-FIXED` (issues found, fixes applied, not re-evaluated by design); Gate 2 then proceeds. In SINGLE mode the run reports it as `UNVERIFIED`, never PASS.
|
|
35
35
|
|
|
36
36
|
**Test agent** (only if acceptance criteria or a test plan exist):
|
|
37
37
|
- Scenario-based acceptance tests covering functionality, API contracts, performance
|
|
38
38
|
- FAIL → fix-and-continue — a Code agent applies the demanded fixes and self-verifies its own build. The recorded verdict becomes `FAIL-FIXED`; Gate 2 then proceeds. In SINGLE mode the run reports it as `UNVERIFIED`, never PASS.
|
|
39
39
|
|
|
40
40
|
**When Gate 2 inputs are absent:**
|
|
41
|
-
- No plan → skip Evaluate agent
|
|
41
|
+
- No plan → skip the Evaluate agent silently (note in output: "Gate 2 Evaluate agent skipped — no plan available")
|
|
42
42
|
- No acceptance criteria and no test plan → skip Test agent silently (note in output: "Gate 2 Test agent skipped — no criteria available")
|
|
43
43
|
- Build proceeds Gate-1-only. Never refuse to build; never force-generate fake criteria. Trust the user.
|
|
44
44
|
@end
|
|
45
45
|
|
|
46
46
|
@define evaluator_panel():
|
|
47
|
-
### Evaluate agent
|
|
47
|
+
### Evaluate agent spawn (one agent, two lenses)
|
|
48
48
|
|
|
49
|
-
Run
|
|
49
|
+
Run ONE Evaluate agent whose prompt names both lenses, so a ticket pays for one spawn and one read of the plan:
|
|
50
50
|
|
|
51
|
-
1. **Acceptance
|
|
52
|
-
2. **Scope / intent
|
|
53
|
-
3. **Cross-ticket-consistency evaluator** — (wave-only; skip for single-ticket runs) "Does this ticket honor the API contracts and invariants that other wave tickets depend on?"
|
|
51
|
+
1. **Acceptance criteria** — "Does the implementation satisfy each numbered acceptance criterion, INCLUDING the negative criteria (what it must NOT do)?"
|
|
52
|
+
2. **Scope / intent drift** — "Did the Code agent smuggle in unplanned changes, deviate from the plan's intent, or introduce anti-features?"
|
|
54
53
|
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
Keep the panel to 2–3 agents. Do not spawn redundant agents asking the same question — that is cost with no diversity benefit.
|
|
54
|
+
It returns `{"verdict": "PASS" | "FAIL", "rationale": "..."}`: FAIL when either lens fails, the rationale naming the failing criterion or the drift. A FAIL on either lens blocks acceptance. A wave ticket runs this same skeleton through `runSingleTicketEngine`, so it gets the same single spawn.
|
|
58
55
|
@end
|
|
59
56
|
|
|
60
57
|
@define implement_bundle():
|
|
@@ -63,12 +60,12 @@ Keep the panel to 2–3 agents. Do not spawn redundant agents asking the same qu
|
|
|
63
60
|
The standard implementation unit for one ticket. Run in order:
|
|
64
61
|
|
|
65
62
|
```
|
|
66
|
-
Code(agentType:"Code", prompt: full task + plan + DECISIONS_CONTEXT + handoff if sequential)
|
|
63
|
+
Code(agentType:"Code", prompt: "OPERATION: implement" first line, then full task + plan + DECISIONS_CONTEXT + handoff if sequential)
|
|
67
64
|
→ gate1_postcode()
|
|
68
65
|
→ gate2_acceptance() ← Gate 2 runs HERE — before the review pass, not after
|
|
69
66
|
```
|
|
70
67
|
|
|
71
|
-
|
|
68
|
+
Every Code agent prompt opens with `OPERATION: <mode>` as its first line (`implement`, `issue-fix`, `validation-fix`, `alignment-fix` or `qa-fix`). It must include: task description, implementation plan (if one exists), relevant DECISIONS_CONTEXT (the index you loaded before authoring), the compliance lens (`COMPLIANCE_FRAMEWORKS` — every Code prompt carries it, fix prompts included), and any PRIOR_PHASE_SUMMARY / HANDOFF_FILE for sequential multi-phase tickets.
|
|
72
69
|
|
|
73
70
|
Gate 2 runs at implementation acceptance — this matches devflow's deliberate placement: "evaluation is part of implementation acceptance, not post-review" (§6.1).
|
|
74
71
|
@end
|
|
@@ -115,7 +112,7 @@ Majority-survives: a finding needs >50% of verification lenses to confirm it. St
|
|
|
115
112
|
|
|
116
113
|
If no surviving findings: return early (no fixes needed). Any coverageGaps are carried in the return — they block a PASS verdict downstream, not the early exit.
|
|
117
114
|
|
|
118
|
-
If survivors remain: batch the confirmed findings for fixing: group findings by file — one file per set of sub-batches, chunked at max 5 findings per sub-batch; never mix two files in one batch. A finding with no `file` field is its own singleton batch. Sub-batches for the SAME file run sequentially (never two Code agents editing the same file concurrently — same-file edits in `parallel()` cause index contention and lost fixes); sub-batches for DISTINCT files run via `parallel()` in staggered chunks of ~5, same pacing bar as the Review spawn path (different code areas — safe per concurrency doctrine). Each Code agent's prompt pins a return contract: `{"status": "fixed"|"blocked", "commitShas": [...], "unresolved": [...]}` — a chunk is FIXED only when `result.status === "fixed"` AND `commitShas` is non-empty AND `result.unresolved` is empty; never decide disposition from status alone. A non-empty `unresolved` list means the agent named work it could not complete — carry the whole chunk into `survivingFindings` rather than guessing which findings the strings map to. `survivingFindings` = findings NOT addressed: fix Code agent dead/failed/blocked/deferred OR committed but left work named in `unresolved`. The fixing Code agent **self-verifies its own fix builds** (build/typecheck per the Code agent's
|
|
115
|
+
If survivors remain: batch the confirmed findings for fixing: group findings by file — one file per set of sub-batches, chunked at max 5 findings per sub-batch; never mix two files in one batch. A finding with no `file` field is its own singleton batch. Sub-batches for the SAME file run sequentially (never two Code agents editing the same file concurrently — same-file edits in `parallel()` cause index contention and lost fixes); sub-batches for DISTINCT files run via `parallel()` in staggered chunks of ~5, same pacing bar as the Review spawn path (different code areas — safe per concurrency doctrine). Each Code agent's prompt pins a return contract: `{"status": "fixed"|"blocked", "commitShas": [...], "unresolved": [...]}` — a chunk is FIXED only when `result.status === "fixed"` AND `commitShas` is non-empty AND `result.unresolved` is empty; never decide disposition from status alone. A non-empty `unresolved` list means the agent named work it could not complete — carry the whole chunk into `survivingFindings` rather than guessing which findings the strings map to. `survivingFindings` = findings NOT addressed: fix Code agent dead/failed/blocked/deferred OR committed but left work named in `unresolved`. The fixing Code agent **self-verifies its own fix builds** (build/typecheck per the Code agent's Running commands block, once over its whole batch). Do **NOT** run Gate 1 or Gate 2 inside the pass (no Validate agent, no Simplify agent, no Scrutinize agent, no Evaluate agent, no Test agent). The engine runs ONE final Gate 1 after the pass exits — see the `gate1_postcode()` cadence (Gate 1 #2).
|
|
119
116
|
@end
|
|
120
117
|
|
|
121
118
|
@define concurrency_doctrine():
|
|
@@ -135,42 +132,26 @@ This applies to both: multiple Code agents working on a single ticket AND multi-
|
|
|
135
132
|
@end
|
|
136
133
|
|
|
137
134
|
@define build_execution_doctrine():
|
|
138
|
-
### Build execution doctrine —
|
|
139
|
-
|
|
140
|
-
The Workflow runtime KILLS any sub-agent that emits no output for 180 seconds. A cold `cargo build`, `cargo test`, a large `tsc`, `gradle build`, `go build ./...`, etc. routinely runs silent far longer and trips this watchdog (the failure reads `agent stalled on all N attempts`). Plain foreground `Bash` also defaults to a 120s timeout.
|
|
141
|
-
|
|
142
|
-
**RULE: any agent (Validate agent, Code agent, Test agent) running a build / test / compile / install that may run silent for more than ~120s MUST run it in the BACKGROUND and POLL — never as a single silent foreground command.**
|
|
143
|
-
|
|
144
|
-
Build commands are NEVER wrapped in `sh -c`, `bash -c`, or inline interpreters (`python3 -c`, `node -e`). Invoke commands directly with the step-1 redirect form — permission systems deny wrapper-invoked commands that would be allowed directly.
|
|
135
|
+
### Build execution doctrine — the Running commands block (LOAD-BEARING)
|
|
145
136
|
|
|
146
|
-
|
|
137
|
+
The Validate, Code and Test agents each carry this `## Running commands` block in their bodies, so each prompt in this engine names the block ("your Running commands block") instead of copying it:
|
|
147
138
|
|
|
148
|
-
|
|
149
|
-
1. Choose ONE unique base path for this run and reuse it verbatim in steps 1–3, e.g. `BASE=/tmp/df-build-<ticket-slug>`. Launch the command with the Bash tool using `run_in_background: true`:
|
|
150
|
-
```
|
|
151
|
-
<build/test command> > <BASE>.log 2>&1; echo "EXIT=$?" > <BASE>.done
|
|
152
|
-
```
|
|
153
|
-
This returns immediately with a background task id — do NOT block on it.
|
|
154
|
-
2. Arm ONE Monitor that emits a heartbeat well under 180s AND exits when the job finishes:
|
|
155
|
-
- description: short, e.g. `await <build cmd>`
|
|
156
|
-
- persistent: false
|
|
157
|
-
- timeout_ms: comfortably ABOVE the expected job time (e.g. 600000)
|
|
158
|
-
- command: `until [ -f <BASE>.done ]; do echo building; sleep 25; done; echo BUILD_DONE; cat <BASE>.done`
|
|
139
|
+
Run builds, typechecks, lints and tests in the foreground, each with an explicit Bash `timeout` above its expected run time. The ceiling is 600000 ms, or `BASH_MAX_TIMEOUT_MS` when set (`echo ${BASH_MAX_TIMEOUT_MS:-600000}`).
|
|
159
140
|
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
141
|
+
- Capture, then tail, in one Bash call (shell state does not persist): `LOG=$(mktemp); echo "LOG=$LOG"; <command> >"$LOG" 2>&1; rc=$?; tail -n 40 "$LOG"; echo "EXIT=$rc"`. The printed `EXIT=` value is the result; never decide one from a grep count.
|
|
142
|
+
- Never background a command and wait on it, and never poll across turns: no `sleep` or `true` turns, no sentinel-file checks, no Monitor.
|
|
143
|
+
- Prefer the scoped command for the change (a package, a path or a test file); for the whole set, one workspace-level command over a per-package loop.
|
|
144
|
+
- A run that exceeds its timeout is BLOCKED: report its duration and log path. Do not wait on it, poll it or re-run it.
|
|
145
|
+
- A run expected to exceed the ceiling is split into parts, each under about 90% of it, run in sequence. If it cannot be split, report BLOCKED with the remedy `devflow flags --set bash-max-timeout-ms=<ms>`.
|
|
146
|
+
- Never re-run a command when nothing it reads has changed.
|
|
147
|
+
- Never wrap a build or test command in `sh -c`, `bash -c`, `python3 -c` or `node -e`: permission rules deny wrapped commands they would allow directly.
|
|
148
|
+
- The same rules hold inside a dynamic Workflow sub-agent.
|
|
166
149
|
|
|
167
150
|
**Cheapest-sufficient validation:** iterate with the fastest check that proves the change — `tsc --noEmit`, `cargo check`, `go vet`, a single test file — the ecosystem's cheapest sufficient signal. Reserve expensive full/optimized builds and whole-suite runs for the final gates only.
|
|
168
151
|
|
|
169
152
|
**One build gate per phase:** batch related fixes, validate once. A fix-pass Code agent runs ONE light check over its whole batch — never several invocations per small fix. Do NOT validate after every individual mutation; validate once after the batch is complete.
|
|
170
153
|
|
|
171
|
-
**Scope commands to stay short.** During the engine,
|
|
172
|
-
|
|
173
|
-
**Invariants:** heartbeat interval MUST stay well under 180s (25–30s is the tested value); Monitor `timeout_ms` MUST exceed the expected job duration (a too-short timeout kills the poll, not the build). Never substitute a single silent long command for this procedure.
|
|
154
|
+
**Scope commands to stay short.** During the engine, prefer the scoped command for the change over the whole workspace — `cargo build -p <crate>`, `cargo test -p <crate>`, `npm test -- <path>`, `go test ./pkg/...` — and, when the whole set is needed, one workspace-level command over a per-package loop. After the wave, the full-workspace regression is the human's job (the wave already hands the integrated branch back to the user).
|
|
174
155
|
@end
|
|
175
156
|
|
|
176
157
|
@define engine_output_schema():
|
|
@@ -210,7 +191,7 @@ Each ticket engine run returns a structured result. The Synthesize agent or the
|
|
|
210
191
|
},
|
|
211
192
|
"escalations": [
|
|
212
193
|
{
|
|
213
|
-
"type": "merge-conflict | gate2-fail | validation-exhausted | ambiguous-resolution | review-coverage-incomplete | dependency-blocked | engine-crash | ticket-link-missing | branch-missing",
|
|
194
|
+
"type": "merge-conflict | gate2-fail | validation-exhausted | scrutiny-blocked | ambiguous-resolution | review-coverage-incomplete | dependency-blocked | engine-crash | ticket-link-missing | branch-missing",
|
|
214
195
|
"description": "string"
|
|
215
196
|
}
|
|
216
197
|
],
|
|
@@ -228,7 +209,7 @@ Each ticket engine run returns a structured result. The Synthesize agent or the
|
|
|
228
209
|
|
|
229
210
|
1. **Code is written ONLY by Code agents.** No other agent type writes code — not Review agent, not Evaluate agent.
|
|
230
211
|
2. **Findings are verified before any fix is written.** The adversarial verification step is not optional; unverified findings are not passed to the Code agent.
|
|
231
|
-
3. **All written code passes Gate 1.** No code merge, commit, or handoff before
|
|
212
|
+
3. **All written code passes Gate 1.** No code merge, commit, or handoff before Simplify agent + Scrutinize agent + Validate agent (in that order — Validate last, so its one run covers the commits of the two before it).
|
|
232
213
|
4. **Gate 2 runs once, at implementation acceptance.** It does not re-run after review-fixes.
|
|
233
214
|
5. **NEVER auto-merge to main or master.** All merges target the integration branch. The user merges to main themselves.
|
|
234
215
|
6. **No unauthorized tracker or remote side-effects.** Sub-agents NEVER create issues/PRs on the tracker, comment on them, or push beyond the ticket-authorized branch unless the ticket, plan, or user explicitly authorizes that exact action. This applies to whatever tracker is resolved, not to one vendor. Proposed follow-ups go in the run report.
|
|
@@ -69,7 +69,7 @@ If the settings line says `KNOWLEDGE=off`, skip write-back entirely. The machine
|
|
|
69
69
|
**Step 2 — Evaluate whether write-back is warranted:**
|
|
70
70
|
|
|
71
71
|
Only proceed if **at least one** of these is true:
|
|
72
|
-
- This workflow changed files in a directory that is documented by an existing feature knowledge base (a documented area changed).
|
|
72
|
+
- This workflow changed files in a directory that is documented by an existing feature knowledge base (a documented area changed). Knowledge bases are written through at that point, never on a background schedule.
|
|
73
73
|
- This workflow surfaced durable, cross-cutting knowledge about a codebase area that would help future agents working in the same area — patterns, anti-patterns, integration points, gotchas not visible from a single file read.
|
|
74
74
|
|
|
75
75
|
**Never spawn unconditionally.** If neither condition is met, skip write-back silently.
|
|
@@ -86,8 +86,6 @@ DIRECTORIES: {list of primary directories touched by this workflow}
|
|
|
86
86
|
FILES_CHANGED: {list of files changed}
|
|
87
87
|
DECISIONS_CONTEXT: {DECISIONS_CONTEXT if available, else (none)}
|
|
88
88
|
|
|
89
|
-
Load the devflow:feature-knowledge skill and follow its authoring process.
|
|
90
|
-
|
|
91
89
|
Write the knowledge base to:
|
|
92
90
|
{worktree}/.devflow/features/{slug}/KNOWLEDGE.md
|
|
93
91
|
|
|
@@ -46,7 +46,7 @@ Every line of a `## Test Plan` section, or of a PR's test-plan block, follows th
|
|
|
46
46
|
|
|
47
47
|
#### Consumption by Gate 2
|
|
48
48
|
|
|
49
|
-
The Evaluate agent
|
|
49
|
+
The Evaluate agent receives: the per-ticket plan + the numbered acceptance criteria (positive and negative).
|
|
50
50
|
|
|
51
51
|
The Test agent receives: the test plan's TP lines, once `check tp` has admitted them.
|
|
52
52
|
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
@define issue_ref_grammar():
|
|
2
|
-
**Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan
|
|
2
|
+
**Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `COMMAND_INPUT` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
|
|
3
3
|
|
|
4
4
|
**A token of any other shape is neither coerced nor dropped silently — and no producer-side grammar check rejects it before the fetch.** Adjudication belongs to the operation that runs, and each one answers in its own Output block: `fetch-issue` strips a leading `#` and takes the text branch, so a non-numeric token is used as a **search term** and the operation returns the first open match or nothing; `fetch-issues-batch` resolves each token to an issue number, drops the ones it cannot resolve, and names them in `NOT_FOUND ({refs})` beside the issues it did fetch. Read the outcome from the operation that ran — a token's shape is a verdict nowhere, and there is nothing upstream holding it back.
|
|
5
5
|
|
|
@@ -40,8 +40,8 @@ For each ready ticket (sequentially by default; parallel only past the §7.1 bar
|
|
|
40
40
|
- Branch setup: the engine's setup-task creates the ticket's branch off integration HEAD at ready-time (so it already contains merged deps); every later phase, and the merge, uses the branch setup-task created, and a setup-task that reports none stops the ticket before implementing
|
|
41
41
|
- Run the single-ticket engine inside a try/catch — one ticket's crash/stall never kills the wave; catch the exception, quarantine that ticket, and continue with the remaining ready set
|
|
42
42
|
- The engine gets the ticket's own reference, the one the pre-fetch printed, as its setup-task input — never the wave's tracking issue
|
|
43
|
-
- On engine PASS or UNVERIFIED: merge to integration branch,
|
|
44
|
-
-
|
|
43
|
+
- On engine PASS or UNVERIFIED: merge to the integration branch locally through Git (no push, no build or test), then the workflow spawns the Validate agent (build + test) over the merge. The merge counts as kept only when that Validate returns PASS
|
|
44
|
+
- Validate FAIL or no PASS (build red after merge): the workflow spawns Git to undo the merge, quarantines the ticket, marks it as escalated, and continues. If the undo is refused, the wave stops taking merges and rounds
|
|
45
45
|
- On any other verdict (PARTIAL, FAIL, ESCALATED) or none: quarantine ticket, do not block independent siblings
|
|
46
46
|
|
|
47
47
|
**Cascade quarantine:** when a ticket is quarantined for any reason (Gate-1 exhausted, engine crash/stall, build-red after merge, review coverage incomplete after retry), the quarantine cascades to its direct and transitive dependents — each is marked blocked with the named reason, naming the blocker by its `{ISSUE_REF}` (e.g. "blocked: depends on {ISSUE_REF} which failed Gate-1"). Independent siblings are never affected. The quarantined list is injected into every subsequent Design agent reader prompt so the reader never schedules dependents of failed tickets.
|
|
@@ -70,7 +70,9 @@ MAX_ROUNDS = LLM judgment based on ticket count (heuristic: ticket_count * 2 + 5
|
|
|
70
70
|
|
|
71
71
|
**Parallel independent tickets:** each gets its own `git worktree add` + durable branch managed by the Git agent. Use explicit `git worktree add` — NOT the Workflow tool's ephemeral `isolation:'worktree'`. The branch must persist across implement → review → resolve → merge stages; ephemeral worktrees are gone when the agent call ends.
|
|
72
72
|
|
|
73
|
-
**Post-merge validation:** after EVERY merge into the integration branch,
|
|
73
|
+
**Post-merge validation:** after EVERY merge into the integration branch, the workflow itself spawns the Validate agent (build + test) over the integration branch; Git runs no build or test. A red build immediately after merge is the cheapest possible conflict detector. It always runs — the merge reports `treeEqual` (whether the merge commit's tree equals the ticket branch head's tree) and the wave row records it, but a true value never skips the Validate.
|
|
74
|
+
|
|
75
|
+
**Undo of a red merge:** on FAIL or no PASS verdict the workflow spawns Git to undo the merge. The undo runs only when no remote branch contains the merge commit and the integration HEAD still equals it, as `git reset --keep {mergeSha}^1` on the integration branch — never `--hard`, never a push. It returns `{"undone": true}` or `{"undone": false, "reason": "..."}`. Undone → the row reads `merged: false` and the ticket is quarantined ("post-merge build red; merge undone"). Not undone → the wave stops taking merges and rounds, and its report names the red integration HEAD.
|
|
74
76
|
|
|
75
77
|
**Commit discipline:** Git agent creates atomic commits per logical change, conventional-commit format, on the ticket branch before merge.
|
|
76
78
|
@end
|
|
@@ -83,13 +85,13 @@ Two parallel sibling tickets can produce real git conflicts. The resolution is *
|
|
|
83
85
|
**Resolution procedure:**
|
|
84
86
|
|
|
85
87
|
1. Git agent detects the conflict and reports the conflicting files + sections
|
|
86
|
-
2. Spawn a Code agent with FULL intent context:
|
|
88
|
+
2. Spawn a Code agent whose prompt opens with `OPERATION: implement` and carries `COMPLIANCE_FRAMEWORKS`, with FULL intent context:
|
|
87
89
|
- Both ticket descriptions and plans
|
|
88
90
|
- The conflicting diff sections (both sides)
|
|
89
91
|
- Relevant ADRs from DECISIONS_CONTEXT (loaded by main model before authoring)
|
|
90
92
|
3. Code agent resolves to PRESERVE BOTH INTENTS — the resolution must honor what both tickets were trying to achieve
|
|
91
93
|
4. If the correct resolution is NOT UNAMBIGUOUS from the intent context: **do NOT guess** → quarantine + escalate
|
|
92
|
-
5. After any resolution: Validate
|
|
94
|
+
5. After any resolution: the post-merge Validate covers it — the workflow spawns it after the merge, and no Validate is requested from Git
|
|
93
95
|
|
|
94
96
|
**Conservative-or-escalate is absolute.** An LLM silently guessing a wrong merge is the highest-danger failure mode in the whole design. When in doubt: quarantine + surface in report. The user re-runs (resume) with the escalated context.
|
|
95
97
|
|
|
@@ -106,7 +108,7 @@ A workflow cannot pause mid-run (F4). "Escalate" means: quarantine-and-continue
|
|
|
106
108
|
- Ticket engine FAIL after max retries (Gate 1 exhausted)
|
|
107
109
|
- Gate 2 FAIL after max retries (Evaluate agent/Test agent not satisfied)
|
|
108
110
|
- Circular dependency detected (all remaining tickets blocked on each other)
|
|
109
|
-
- Build red after merge (Validate agent fails
|
|
111
|
+
- Build red after merge (the post-merge Validate agent fails: the merge is undone and the ticket quarantined; a refused undo halts the wave)
|
|
110
112
|
- Review coverage incomplete after retry (a focus area failed to produce a live Review agent result after the retry)
|
|
111
113
|
- Ticket engine crash/stall (unrecoverable exception or watchdog kill — quarantine cascades to dependents)
|
|
112
114
|
- No ticket link while issues are required (the engine stops before implementing)
|
|
@@ -30,7 +30,7 @@ Run a comprehensive code review of the current branch by spawning parallel revie
|
|
|
30
30
|
|
|
31
31
|
1. **Discover reviewable worktrees** using the `devflow:worktree-support` skill discovery algorithm:
|
|
32
32
|
- Run `git worktree list --porcelain` → parse, filter (skip protected/detached/mid-rebase), dedup by branch, sort by recent commit
|
|
33
|
-
-
|
|
33
|
+
- Invoke the `devflow:worktree-support` skill, then Read its `references/discovery.md` from the skill's base directory for the full 7-step algorithm; the canonical protected branch list stays in the skill
|
|
34
34
|
2. **If `--path` flag provided:** use only that worktree, skip discovery
|
|
35
35
|
**`--path` validation**: Before proceeding, verify the path exists as a directory and appears in `git worktree list` output. If not: report error and stop.
|
|
36
36
|
3. **If only 1 reviewable worktree** (the common case): proceed as single-worktree flow — zero behavior change
|
|
@@ -198,8 +198,6 @@ Only `exit=0` means the skill is installed. On any other result, do NOT add that
|
|
|
198
198
|
|
|
199
199
|
**Produces:** DECISIONS_CONTEXT, FEATURE_KNOWLEDGE, COMPLIANCE_FRAMEWORKS (carried from Step 0b)
|
|
200
200
|
|
|
201
|
-
**Load Companion Skills** — Load via Skill tool: `devflow:quality-gates`, `devflow:software-design`. If a skill fails to load, continue without it.
|
|
202
|
-
|
|
203
201
|
{{decisions_load()}}
|
|
204
202
|
|
|
205
203
|
This produces a compact index of active ADR/PF entries. Pass `DECISIONS_CONTEXT` to all Review agents. Review agents use `devflow:apply-decisions` to Read full entry bodies on demand.
|