jorgex-stack 1.9.6 → 1.9.8
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +11 -11
- package/dist/cli.js +6 -6
- package/package.json +1 -1
- package/stack/scripts/post-pr-review.cjs +1 -2
- package/stack/skills/orchestrator/SKILL.md +30 -202
- package/stack/skills/orchestrator/references/standard-workflow.md +189 -0
- package/stack/skills/work-lifecycle/SKILL.md +7 -3
- package/stack/system-prompt/AGENTS.md +2 -2
package/README.md
CHANGED
|
@@ -106,21 +106,21 @@ Programmatic mode does **not** provide:
|
|
|
106
106
|
|
|
107
107
|
### Pi runtime
|
|
108
108
|
|
|
109
|
-
Pi combines the frozen **snapshot v2** package with a Stack-owned shared projection. The published Stack `1.9.
|
|
109
|
+
Pi combines the frozen **snapshot v2** package with a Stack-owned shared projection. The published Stack `1.9.7` recognizes the exact package **`jorgex-pi@0.8.4`**. Pi `0.8.5` is published independently; this checkout keeps the Stack `1.9.7` baseline with the Pi `0.8.5` candidate pin. On merge, Stack is expected to auto-bump to `1.9.8` and publish without waiting 24 hours.
|
|
110
110
|
|
|
111
|
-
The current published command set uses Stack `1.9.
|
|
111
|
+
The current published command set uses Stack `1.9.7` with Pi `0.8.4`:
|
|
112
112
|
|
|
113
113
|
```bash
|
|
114
|
-
pnpm dlx jorgex-stack@1.9.
|
|
115
|
-
pnpm dlx jorgex-stack@1.9.
|
|
116
|
-
pnpm dlx jorgex-stack@1.9.
|
|
117
|
-
pnpm dlx jorgex-stack@1.9.
|
|
118
|
-
pnpm dlx jorgex-stack@1.9.
|
|
114
|
+
pnpm dlx jorgex-stack@1.9.7 install --agents pi
|
|
115
|
+
pnpm dlx jorgex-stack@1.9.7 doctor --agents pi
|
|
116
|
+
pnpm dlx jorgex-stack@1.9.7 models --agents pi
|
|
117
|
+
pnpm dlx jorgex-stack@1.9.7 sync --agents pi
|
|
118
|
+
pnpm dlx jorgex-stack@1.9.7 uninstall --agents pi
|
|
119
119
|
```
|
|
120
120
|
|
|
121
|
-
Stack downloads the frozen registry tarball, verifies its exact size plus SHA-256/SHA-512, backs up Pi's `settings.json`, and only then asks Pi to install that local file. The historical `0.8.0` tarball was `89128340` bytes; the exact current artifact and integrity values are authoritative in `src/lib/pi-runtime.ts`. Pi's own package-manager invocation is the narrow runtime exception to the repository's pnpm-only rule; the Stack lifecycle never launches npm directly. After the package is healthy, Stack projects the shared resources into Pi: marked `jorgex:system-prompt` and `jorgex:engram-protocol` sections in `~/.pi/agent/AGENTS.md`, canonical skills under `~/.agents/skills`, and `~/.pi/agent/prompts/lean-audit.md`. When the managed Playwright preference is active, the projection also adds or removes the marked `jorgex:browser` section dynamically. The Pi-only `install --agents pi --playwright` flow installs and persists that Playwright capability just like the other harnesses. Chrome DevTools MCP and Context7 remain outside the Pi scope. The published `1.9.
|
|
121
|
+
Stack downloads the frozen registry tarball, verifies its exact size plus SHA-256/SHA-512, backs up Pi's `settings.json`, and only then asks Pi to install that local file. The historical `0.8.0` tarball was `89128340` bytes; the exact current artifact and integrity values are authoritative in `src/lib/pi-runtime.ts`. Pi's own package-manager invocation is the narrow runtime exception to the repository's pnpm-only rule; the Stack lifecycle never launches npm directly. After the package is healthy, Stack projects the shared resources into Pi: marked `jorgex:system-prompt` and `jorgex:engram-protocol` sections in `~/.pi/agent/AGENTS.md`, canonical skills under `~/.agents/skills`, and `~/.pi/agent/prompts/lean-audit.md`. When the managed Playwright preference is active, the projection also adds or removes the marked `jorgex:browser` section dynamically. The Pi-only `install --agents pi --playwright` flow installs and persists that Playwright capability just like the other harnesses. Chrome DevTools MCP and Context7 remain outside the Pi scope. The published `1.9.7` entry is `{ "source": "npm:jorgex-pi@0.8.4", "skills": [], "prompts": [] }`; the checkout's `0.8.5` entry is a candidate only. Filters are applied only after this projection exists, so the package does not duplicate shared resources. Package ownership is recorded separately in `~/.jorgex-stack/pi-receipt.json`; projection ownership is recorded in `~/.jorgex-stack/pi-projection-receipt.json`. Both receipts are scope-bound and fail closed for manual, duplicate, divergent, partial, corrupt, copied-to-another-scope, or unknown-history state.
|
|
122
122
|
|
|
123
|
-
Historically, the published Pi 0.8.0 direct-package snapshot added `work-audit`: the snapshot grew from **17 to 18 skill trees** (96 to 97 files), and the active runtime allowlist grew from **16 to 17 skills**. `playwright-cli` remains in the snapshot but inactive because browser automation is a separate opt-in integration.
|
|
123
|
+
Historically, the published Pi 0.8.0 direct-package snapshot added `work-audit`: the snapshot grew from **17 to 18 skill trees** (96 to 97 files), and the active runtime allowlist grew from **16 to 17 skills**. The Pi 0.8.5 candidate records 18 skill trees and 98 files; reference F2-A is included while the private F1 skills remain preserved. `playwright-cli` remains in the snapshot but inactive because browser automation is a separate opt-in integration.
|
|
124
124
|
|
|
125
125
|
The historical published artifact has two separate provenance anchors. The local size/SHA-256/SHA-512 checks bind the downloaded bytes to Stack's accepted artifact; they are checks within that checkout, not independent trust roots. npm's external provenance/attestation is outside Stack runtime verification, and `provenance.commit` is informative unless that external attestation is independently verified.
|
|
126
126
|
|
|
@@ -132,9 +132,9 @@ The package owns Pi's native primary-model projection: `openai-codex/gpt-5.6-sol
|
|
|
132
132
|
|
|
133
133
|
Engram remains mandatory and user-owned. An existing binary is preserved. Interactive install may offer the native `brew`/`go`/release channel with explicit confirmation; `--yes` and non-TTY installs fail with a remedy when Engram is absent. The database and memories are never updated or deleted, and uninstall never deletes the Engram binary. Under `--target-dir`, Stack accepts only `<target>/bin/engram`, isolates Pi/Home/XDG/AppData/temp/npm-cache paths inside the target, and never consults the host Engram or Pi configuration.
|
|
134
134
|
|
|
135
|
-
The published Stack `1.9.
|
|
135
|
+
The published Stack `1.9.7` recognizes the exact Pi receipt `npm:jorgex-pi@0.8.4`. Use exact versions, never `latest`, and never edit receipts or hashes or delete `HOME`, Engram, or another runtime's projection to force trust. The transition and rollback commands are in [docs/references/pi-runtime.md](docs/references/pi-runtime.md).
|
|
136
136
|
|
|
137
|
-
|
|
137
|
+
The 24-hour managed-consumption maturity rule applies only to real installation or consumption of the new Pi package; development, PR validation, merge and Stack publication may proceed immediately. Installing it on a real user scope before the maturity window requires Jorge's explicit exception.
|
|
138
138
|
|
|
139
139
|
`update --agents pi` only runs the Pi package lifecycle; it does not enter the global Stack updater. `update --check --agents pi` is a read-only Pi doctor. Uninstall runs package cleanup, backs up Pi's settings before removal, removes only the exact receipt-owned package after verifying absence, and preserves all companion/user state. Full behavior, failure states and troubleshooting are in [docs/references/pi-runtime.md](docs/references/pi-runtime.md).
|
|
140
140
|
|
package/dist/cli.js
CHANGED
|
@@ -5142,16 +5142,16 @@ import { createHash } from "crypto";
|
|
|
5142
5142
|
var PI_RUNTIME_CANDIDATE = {
|
|
5143
5143
|
package: {
|
|
5144
5144
|
name: "jorgex-pi",
|
|
5145
|
-
version: "0.8.
|
|
5146
|
-
source: "npm:jorgex-pi@0.8.
|
|
5145
|
+
version: "0.8.5",
|
|
5146
|
+
source: "npm:jorgex-pi@0.8.5"
|
|
5147
5147
|
},
|
|
5148
5148
|
provenance: {
|
|
5149
|
-
commit: "
|
|
5149
|
+
commit: "57cb15413c1dda408251ff80ff9d5658b7e91793"
|
|
5150
5150
|
},
|
|
5151
5151
|
tarball: {
|
|
5152
|
-
bytes:
|
|
5153
|
-
sha256: "
|
|
5154
|
-
sha512: "
|
|
5152
|
+
bytes: 89133857,
|
|
5153
|
+
sha256: "ea7ce0fa88c324d15756fbb1f7e222d2d87156de8a17cede8f3f317ad0d90c7e",
|
|
5154
|
+
sha512: "0f6be4e79be3b7add6949b0d3d913a79acb0dd338ea0459d500358fec46ddb1925ad2ec926d0fbe899c7e48f0a7cb9681424fc1a0d1ab104e7e25d350478b989"
|
|
5155
5155
|
},
|
|
5156
5156
|
pi: {
|
|
5157
5157
|
testedVersions: ["0.84.2"]
|
package/package.json
CHANGED
|
@@ -132,8 +132,7 @@ function isPrReadinessCommand(command) {
|
|
|
132
132
|
const message = `<pr-lifecycle-state-required>
|
|
133
133
|
A PR readiness transition was attempted through \`gh pr create\` without \`--draft\` or through \`gh pr ready\`. Do not infer success or PR state from the command text. Resolve the current PR and run \`gh pr view --json number,isDraft,headRefOid\` before the next action.
|
|
134
134
|
|
|
135
|
-
- The review boundary is the final draft diff. If the full review was not already completed, ensure the PR is draft (run \`gh pr ready --undo <number>\` if necessary), finish code, the applicable version bump, local tests, \`pnpm qa:quality\` when defined, Vercel preview review when applicable, and final diff inspection.
|
|
136
|
-
- Load and run the portable \`xreview\` skill against that exact final diff. When an orchestrator owns an active work context, it must pass the exact \`work/{name}\` to every reviewer.
|
|
135
|
+
- The review boundary is the final draft diff. If the full review was not already completed, ensure the PR is draft (run \`gh pr ready --undo <number>\` if necessary), finish code, the applicable version bump, local tests, \`pnpm qa:quality\` when defined, Vercel preview review when applicable, and final diff inspection, then load and run the portable \`xreview\` skill against that exact final diff. When an orchestrator owns an active work context, it must pass the exact \`work/{name}\` to every reviewer. If the review was already completed, retain its evidence and do not open another panel merely because readiness was attempted.
|
|
137
136
|
- After fixing findings, repeat xreview only when the fixes materially change the diff or introduce a distinct risk. For ordinary fixes, explicit evidence of the prior review plus deterministic verification is sufficient even though \`headRefOid\` changed.
|
|
138
137
|
- If the PR is actually ready, do not push. If the project has PR checks configured, wait for the complete Quality Gates, run \`gh pr checks <number>\`, and verify the checked headRefOid is the candidate SHA.
|
|
139
138
|
- If no PR checks are configured, confirm that from project configuration such as workflows, rulesets or integrations, and record it; their absence does not block the merge. An empty \`gh pr checks\` result immediately after ready is not evidence that no checks are configured.
|
|
@@ -1,241 +1,69 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: orchestrator
|
|
3
|
-
description: Main coordinator for non-trivial tasks.
|
|
3
|
+
description: Main coordinator for non-trivial tasks. Routes work through a short or standard path, then coordinates the appropriate work. Use it when work spans several layers, several files, or requires coordination.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Orchestrator
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
Route the work before starting a workflow. `short` and `standard` are routes, not new human/programmatic modes: their existing output contracts stay intact.
|
|
9
9
|
|
|
10
|
-
##
|
|
10
|
+
## Routing
|
|
11
11
|
|
|
12
|
-
|
|
13
|
-
INIT → EXPLORE → SPEC → PLAN → EXECUTE → VERIFY → SHIP → CLOSE
|
|
14
|
-
```
|
|
12
|
+
Choose **short** only when the objective is clear, the affected contract is understood, the scope is bounded, and the change has sufficient verification. Choose **standard** when scope, uncertainty, risk, or verification needs the formal workflow. A small change is not automatically safe: assess configuration, publication, security, and other affected contracts before routing.
|
|
15
13
|
|
|
16
|
-
|
|
14
|
+
Short work has one primary responsible person; involve a specialist only when it adds value. It has no mandatory analyst, PRE, or POST chain. A short standalone change does not create a PRD, plan, formal task spec, PRE, or POST merely for ceremony.
|
|
17
15
|
|
|
18
|
-
|
|
16
|
+
If short work expands in scope, risk, uncertainty, verification needs, or requires a material decision, promote it to standard **before** continuing. Do not use file counts, elapsed time, or delegation mechanics as routing rules.
|
|
19
17
|
|
|
20
|
-
|
|
18
|
+
An existing or active formal SDD work keeps its approved scope, Spec, plan row, ownership, and lifecycle even when a bounded implementation step is handled short. Do not create child formal tasks for phases or polls; if the work grows into an independent or persistent item, formalize it through `work-lifecycle` before continuing.
|
|
21
19
|
|
|
22
|
-
|
|
23
|
-
- Identify the project's constraints.
|
|
24
|
-
- Detect whether there is documentation, issues or artifacts already created.
|
|
20
|
+
The primary may implement or execute when a handoff or delegation adds no value. Otherwise, use the `agent-delegation` skill for the specialist's defined scope. Every subagent keeps that assigned scope and does not spawn subagents.
|
|
25
21
|
|
|
26
|
-
|
|
22
|
+
Apply the `lean-code` skill as a scope gate for code-bearing work: ask whether the code is needed at all, whether project or platform capabilities already solve it, and whether the smallest obvious change is enough.
|
|
27
23
|
|
|
28
|
-
|
|
24
|
+
For the **standard** route, explicitly read and follow [references/standard-workflow.md](references/standard-workflow.md). Resolve that reference relative to this skill directory, never the repository CWD; if it cannot be read, report a blocker rather than silently falling back to short. It contains the formal PRD, plan, PRE, POST, change-first, and delivery workflow. Do not load it for short work unless the work is promoted.
|
|
29
25
|
|
|
30
|
-
|
|
31
|
-
- `frontend-analyst` if it affects UI, hooks, state or rendering
|
|
32
|
-
- `security-auditor` if the area is sensitive
|
|
26
|
+
## Shared guards
|
|
33
27
|
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
- Your priority is to delegate.
|
|
37
|
-
- If a task has a clear subagent scope, delegate.
|
|
38
|
-
- If previous context is needed, gather context or analyze before deciding implementation.
|
|
39
|
-
|
|
40
|
-
### Delegation triggers
|
|
41
|
-
|
|
42
|
-
Once a task crosses any of these thresholds, delegating stops being optional:
|
|
43
|
-
|
|
44
|
-
| Trigger | Expected behavior |
|
|
45
|
-
| --- | --- |
|
|
46
|
-
| Reading 4+ files just to understand a flow | Delegate exploration to the matching analyst. |
|
|
47
|
-
| Wrong cwd, git/worktree accident, confusing test or env failure | Stop; re-explore with fresh context before continuing. |
|
|
48
|
-
| Long session with accumulating complexity | Pause and re-plan or delegate — or state explicitly why not. |
|
|
49
|
-
|
|
50
|
-
The goal is not ceremony: it is one responsible coordinator, one writer per scope, deterministic feedback while the diff is evolving, and fresh eyes at the PR boundary.
|
|
51
|
-
|
|
52
|
-
## 3. SPEC
|
|
53
|
-
|
|
54
|
-
- Synthesize findings.
|
|
55
|
-
- Propose a simple approach.
|
|
56
|
-
- Clarify only the real ambiguities.
|
|
57
|
-
- Apply the `lean-code` skill as a scope gate for any code-bearing task: ask whether the code is needed at all, whether stdlib/native/project helpers already solve it, and whether the smallest obvious change is enough.
|
|
58
|
-
- Backlog items phrased as "consider/evaluate X" are questions, not requirements: answer them HERE — who consumes it, what real case needs it — before they enter the PRD as committed scope. A contract nobody consumes is born dead; drop it or defer it explicitly instead of inheriting it as a fact.
|
|
59
|
-
- Create the PRD before moving to PLAN (see PRD rules).
|
|
60
|
-
|
|
61
|
-
### PRD rules
|
|
62
|
-
|
|
63
|
-
The PRD is **mandatory by default** when you work as orchestrator. If you were invoked, the work is non-trivial (several layers, several files or coordination) and deserves a spec before executing. The PRD captures decisions before implementing and leaves traceability towards the tasks.
|
|
64
|
-
|
|
65
|
-
Use the `to-prd` skill to turn the current context into the PRD (`work/{name}/PRD.md`) before planning execution.
|
|
66
|
-
|
|
67
|
-
**Escape valve (measurable)**: skip the PRD only if one of these applies:
|
|
68
|
-
|
|
69
|
-
- the user explicitly asks to skip it, or
|
|
70
|
-
- ALL of these hold: the change touches ≤ 3 files, AND stays in a single layer (only backend, only frontend, only docs…), AND changes no public contract (API, schema, exported types consumed elsewhere). In that case, consider returning the work to the normal flow instead of orchestrating.
|
|
71
|
-
|
|
72
|
-
If you skip it, say so explicitly and state which condition applied.
|
|
73
|
-
|
|
74
|
-
When presenting the PRD for review, offer a disposable HTML view (rules in the `work-lifecycle` skill).
|
|
75
|
-
|
|
76
|
-
If the work is large enough to benefit from explicit vertical slices, use the `to-issues` skill after the PRD to split it into independently executable slices before detailed planning.
|
|
77
|
-
|
|
78
|
-
## 4. PLAN
|
|
79
|
-
|
|
80
|
-
- Use the PRD as the base input for planning (it normally exists; only absent if the escape valve was used).
|
|
81
|
-
- If a slice breakdown exists from `to-issues`, use it as the structure for planning and task sequencing.
|
|
82
|
-
- Divide the work into clear tasks.
|
|
83
|
-
- One task = one agent = one scope.
|
|
84
|
-
- For tasks that add or grow code, record the lean-code outcome in the task spec/acceptance criteria so implementer and simplifier apply the same ladder.
|
|
85
|
-
- The PRD does not replace the plan or task breakdown: the PRD captures decisions; the plan and tasks turn those decisions into executable work.
|
|
86
|
-
- **Change-first**: for intentional material contract changes discovered in EXECUTE or VERIFY—not bugfixes that restore the approved contract—return to SPEC before further implementation. Update the PRD first, then propagate it to the plan, task specs, `SC-*` success criteria and testing decisions; rerun PRE until `clean`, obtain human approval of the delta, then resume EXECUTE and repeat VERIFY.
|
|
87
|
-
- Materialize the plan per the Work state rules: `work/{name}/plan.md` with the task table and one declared recoverable spec source per formal task — an Engram observation or canonical Markdown (templates in the `work-lifecycle` skill). Record the exact source in `Spec`; never create two active copies.
|
|
88
|
-
- Load and run the `work-audit` skill in **PRE** mode after the plan and task specs exist and before presenting the final plan. Pass the exact active `work/{name}` path and the exact PR/checkpoint scope; never infer either from the branch or scan other work folders. PRE is read-only: during audit remediation you are the only writer of active work artifacts. Route every finding to its owner artifact, correct it, and rerun PRE until it reports `clean`.
|
|
89
|
-
- An unresolved `[NEEDS CLARIFICATION: ...]` marker blocks PRE. Return to SPEC and resolve the ambiguity with the user only when existing context cannot answer it; never approve or execute a plan while PRE is not clean.
|
|
90
|
-
- When presenting the plan for human review, offer a disposable HTML view (rules in the `work-lifecycle` skill). If human review changes the PRD, plan or task specs, rerun PRE and require `clean` again before approval or EXECUTE. Requested changes go to plan.md; delete the HTML once its artifact is approved, before EXECUTE.
|
|
28
|
+
Both routes preserve mandatory memory/Engram saves, testing/TDD by risk, Git/worktree discipline, final-draft review, configured gates, and explicit user approval for merge. Security, permissions, ownership, backups, dependency consent, and the existing human/programmatic output contracts are never relaxed by routing. Product documentation remains with `docs-maintainer` when it is needed; short routing does not absorb that owner's scope.
|
|
91
29
|
|
|
92
30
|
## Work state
|
|
93
31
|
|
|
94
|
-
The `work-lifecycle` skill is the single source
|
|
32
|
+
The `work-lifecycle` skill is the single source for **formal SDD** work. Each formal task has one recoverable Spec source and one plan row. A direct or inline message is an auxiliary microassignment under a parent task, never a formal or independent task or a second Spec; if it grows into independent work, persist its formal task spec and plan row before continuing. Pass a delegated worker the exact Spec as read-only. Its distinct outcome topic_key must never overwrite that Spec; significant decisions and findings still require immediate memory saves, and checkpoint outcomes persist as required. Do not write memory for a poll or status without new information.
|
|
95
33
|
|
|
96
|
-
|
|
97
|
-
- The full spec of each formal task has one declared recoverable source: an Engram observation identified by project + topic_key `work/{name}/task/{NN}` with verified identity/access, or canonical Markdown at `work/{name}/tasks/{NN}.md`. Record it in the plan's `Spec` column; existing Engram specs remain valid and need no migration. When you delegate, pass the title and exact `Spec` reference as read-only for the worker: Engram project + topic_key and an optional ID already bound to that identity in the current store, or the accessible absolute Markdown path. Follow the lifecycle handoff: direct get only with a known current-store binding; otherwise resolve by project/topic before get or block, then recheck identity.
|
|
98
|
-
- A direct or inline message is an auxiliary, self-contained microassignment under a parent task, never a formal or independent task. If it grows or becomes independent, persist its task spec and add the plan row before continuing.
|
|
99
|
-
- Phase outcomes, decisions and PR checkpoints → Engram under `work/{name}/{phase}` and `work/{name}/pr/{NN}`; when assigning a phase outcome, give the subagent an outcome topic_key distinct from its `Spec` reference. Never direct `mem_save` or `mem_update` of a result to the Spec observation or its topic_key.
|
|
100
|
-
- Pending work → the project's single `work/backlog` topic_key, or issues (`to-issues`) if the project uses a tracker. Never a TODOs folder. For Engram, you are the **single writer**: before every change, retrieve the exact observation with `mem_get_observation`, preserve unrelated entries, send the complete content with `mem_update`, then read it again to verify. Never write it concurrently or use a blind topic-key upsert. Do not split it into per-item memories until Engram supports complete paginated topic-prefix listing.
|
|
101
|
-
- On final close: `mem_save` the outcome under `work/{name}/done`, move the PRD to the project's docs only if it has lasting documentation value, then delete `work/{name}/`. `work/{name}/done` is only for the last PR / final outcome. History is memory + git.
|
|
34
|
+
Phase outcomes, decisions and PR checkpoints → Engram under `work/{name}/{phase}` and `work/{name}/pr/{NN}`. The coordinator is the single writer for `work/backlog`: call `mem_get_observation`, preserve unrelated entries, send the complete content with `mem_update`, then read it again to verify. Never mutate it concurrently. For multi-PR work, each merge is a checkpoint; keep `work/{name}/PRD.md` and `plan.md` alive until the roadmap is finished.
|
|
102
35
|
|
|
103
36
|
## Delegation map
|
|
104
37
|
|
|
105
|
-
Load the `agent-delegation` skill: it defines the available subagents,
|
|
106
|
-
|
|
107
|
-
Every subagent follows its active result contract. Process it:
|
|
108
|
-
|
|
109
|
-
- Launch every specialist named by the active delegation field or format.
|
|
110
|
-
- A delegation is unfinished work in another scope, not a request to append a generic quality pipeline. Normal handoffs between `implementer` and `tester` do not by themselves justify reviewers or analyzers.
|
|
111
|
-
- If a subagent reports `partial`, keep the safe work and relaunch only what still needs guidance.
|
|
112
|
-
- If a subagent reports `blocked` with one concrete uncertainty question, answer it from existing context when possible; if it still cannot be resolved, ask the user only if genuinely necessary, then relaunch the original or a suitable specialist with explicit guidance.
|
|
113
|
-
- Don't declare a phase done while a delegation remains unprocessed.
|
|
114
|
-
- If the reported status is `partial` or `blocked`, resolve the cause before moving on.
|
|
115
|
-
|
|
116
|
-
## 5. EXECUTE
|
|
117
|
-
|
|
118
|
-
### Worktree
|
|
119
|
-
|
|
120
|
-
Before the first task, create a git worktree for this work and run the ENTIRE execution inside it — implementation, tests, commits and pushes happen there, never on the user's main checkout.
|
|
121
|
-
|
|
122
|
-
Canonical location is mandatory: resolve the project root with `git rev-parse --show-toplevel`, ensure `worktrees/` is ignored in the repo-local `.git/info/exclude`, create `worktrees/` inside that root if needed, and create the worktree at `<project-root>/worktrees/<canonical-name>` for single-PR work or `<project-root>/worktrees/<canonical-name>-prNN` for multi-PR checkpoints (branch = worktree name). Do not create worktrees next to the repo, in the repo root, under `work/`, or in any external temp/shared folder.
|
|
123
|
-
|
|
124
|
-
Every delegation prompt must state the worktree path as the ONLY allowed write root. After each writer subagent finishes, verify the user's main checkout is still clean (`git status` there); if the subagent wrote outside the worktree, STOP, move those changes into the worktree (patch/apply) and restore the main checkout before continuing. Subagent obedience is not a safety boundary — this check is.
|
|
125
|
-
|
|
126
|
-
### Commit cadence
|
|
127
|
-
|
|
128
|
-
Commit after each task or bounded group of tasks, with a message that reflects that task — the branch history must map to the plan. Never accumulate the whole work into one giant commit at the end.
|
|
129
|
-
|
|
130
|
-
### Draft PR cadence
|
|
131
|
-
|
|
132
|
-
- After the first coherent commit, push the branch and create the PR against its real base with `gh pr create --draft`. Do not wait until SHIP to open it.
|
|
133
|
-
- Keep every code change, commit and push inside the draft phase. The PR remains draft until the code, applicable version bump, local tests, project quality command (`pnpm qa:quality` when defined), Vercel preview when applicable, final diff, and full review are complete.
|
|
134
|
-
- Never push to a ready PR. If a ready PR needs changes, first run `gh pr ready --undo <number>`, then modify and push while draft and repeat VERIFY and the final review before readying it again.
|
|
135
|
-
|
|
136
|
-
### Handoff rule
|
|
137
|
-
|
|
138
|
-
The analyst's **Recommendation** is the implementer's input. Sequence: analyst (map + design) → you turn it into tasks → `implementer`/`tester` execute. Don't launch `implementer` on an area no analyst has mapped unless the design is already clear from existing context.
|
|
139
|
-
|
|
140
|
-
### Testing decision
|
|
141
|
-
|
|
142
|
-
Every implementation task needs a testing decision, not automatically a new test. Establish:
|
|
143
|
-
|
|
144
|
-
- the meaningful regression risk introduced by the change
|
|
145
|
-
- the existing test that already protects it, if any
|
|
146
|
-
- the new or changed behavior that needs protection
|
|
147
|
-
- the strongest seam closest to that risk
|
|
148
|
-
- the action: TDD/new test, update, reuse existing coverage, or no new test with a concrete trivial/mechanical/already-covered reason
|
|
149
|
-
|
|
150
|
-
Prefer one authoritative test per behavior. Another layer is justified only when it protects a distinct contract. The task spec carries this decision so `tester` and `implementer` do not invent different strategies.
|
|
151
|
-
|
|
152
|
-
### TDD mode
|
|
153
|
-
|
|
154
|
-
Use for business rules, bugs/regressions, public contracts, invariants, security/data boundaries, or other behavior whose risk warrants new protection.
|
|
155
|
-
|
|
156
|
-
```text
|
|
157
|
-
tester (RED) → implementer (GREEN/REFACTOR)
|
|
158
|
-
```
|
|
159
|
-
|
|
160
|
-
### Direct mode
|
|
161
|
-
|
|
162
|
-
Use for styling, wiring, generated code, mechanical refactors, trivial code, or changes already covered by an authoritative test. Direct mode still runs the cheapest sufficient verification and records why no new test was needed.
|
|
163
|
-
|
|
164
|
-
```text
|
|
165
|
-
implementer (direct change)
|
|
166
|
-
```
|
|
167
|
-
|
|
168
|
-
### Special delegations
|
|
169
|
-
|
|
170
|
-
- `translator` for translations or multilingual visible text
|
|
171
|
-
- `docs-maintainer` for documentation
|
|
172
|
-
- `security-auditor` for sensitive review
|
|
173
|
-
|
|
174
|
-
### Verification cadence
|
|
175
|
-
|
|
176
|
-
Deterministic checks are the routine feedback loop while implementation is in progress: run the relevant tests, lint and typecheck/build checks at the cheapest seam that can catch the section's regressions. Verify by bounded, coherent sections (e.g. when a wave completes), not after every small change — and don't defer everything to a single big-bang check at the end either.
|
|
177
|
-
|
|
178
|
-
Each writer verifies its own bounded area (e.g. its test file). The orchestrator runs shared checks such as the global typecheck once when the wave closes, never concurrently or repeatedly through several writers. Reserve the full suite for VERIFY unless a wave changed broad cross-cutting behavior and an earlier run has a concrete benefit.
|
|
179
|
-
|
|
180
|
-
### CI and workflow scope
|
|
181
|
-
|
|
182
|
-
CI guidance belongs here only when the task explicitly affects workflows, gates, path filters, execution frequency, or CI cost. Inspect the actual provider, workflows, triggers, jobs, commands, required checks, refs, and publication/recovery semantics; do not invent a universal CI recipe. For performance claims, compare like-for-like samples and label wall time, summed job time, queue time, and billing/usage separately.
|
|
183
|
-
|
|
184
|
-
- Use explicit base/head (or equivalent) refs for diff and path decisions; validate command errors and shared configuration. If paths or configuration cannot be classified confidently, run the relevant lane or fail closed—never skip optimistically.
|
|
185
|
-
- Draft/candidate validation may cancel obsolete validation runs when project semantics allow it, but never cancel a mutable publish/release job or leave publication halfway. Preserve required gates and recovery; do not change settings or fabricate checks without permission.
|
|
186
|
-
|
|
187
|
-
### Early-review budget
|
|
38
|
+
Load the `agent-delegation` skill: it defines the available subagents, their scopes, and when a specialist adds value. A delegation is unfinished work in another scope, not a generic quality pipeline. If a subagent reports `partial`, keep the safe work and relaunch only what still needs guidance. If a subagent reports `blocked` with one concrete uncertainty question, answer it from existing context when possible; if a material decision still cannot be resolved, suspend the work and ask the user for guidance before relaunching the original or a suitable specialist with that explicit guidance.
|
|
188
39
|
|
|
189
|
-
|
|
40
|
+
### Worktree and PR lifecycle
|
|
190
41
|
|
|
191
|
-
-
|
|
192
|
-
- Use the single most relevant specialist. Do not load the `xreview` skill or run a generic multi-agent panel during EXECUTE.
|
|
193
|
-
- Run at most one early review per bounded critical section, after that section is coherent rather than after each task inside it.
|
|
194
|
-
- Do not launch `code-reviewer`, `code-simplifier`, `test-analyzer` or `silent-failure-hunter` merely because a writer finished, a test task completed, several files changed or a commit is due.
|
|
195
|
-
- File count, writer completion, commit, push, or draft PR creation are not early-review triggers. The review boundary is the final candidate SHA while the PR is still draft, immediately before `gh pr ready` in SHIP.
|
|
42
|
+
Both routes follow the project Git/worktree rules; short is not an exception. Where those rules explicitly permit a trivial direct-main change, keep that exception; otherwise resolve the root with `git rev-parse --show-toplevel`, ensure `worktrees/` is ignored in the repo-local `.git/info/exclude`, and create/use `<project-root>/worktrees/<canonical-name>` or `<project-root>/worktrees/<canonical-name>-prNN` with the branch matching the worktree name.
|
|
196
43
|
|
|
197
|
-
|
|
44
|
+
After the first coherent commit, push the work branch and open the PR with `gh pr create --draft`. Keep it draft while it changes. Mark it ready once with `gh pr ready <number>`. If the project has PR checks configured, wait for Quality Gates, run `gh pr checks <number>`, and verify they pass for the latest commit candidate. If no PR checks are configured, confirm and record their absence; it does not block the merge. An empty `gh pr checks` result immediately after ready is not evidence that no checks are configured. Immediately before reporting or merging, compare `gh pr view --json headRefOid` with the recorded candidate SHA. If a ready PR needs a fix, run `gh pr ready --undo <number>` before editing, then repeat verification, review, ready, and configured gates.
|
|
198
45
|
|
|
199
|
-
|
|
200
|
-
- Reserve heavy suites for cases where they provide real value or the project requires them.
|
|
201
|
-
- If POST identifies an intentional material contract change, follow the PLAN's change-first procedure before further implementation.
|
|
202
|
-
- Load and run the `work-audit` skill in **POST** mode after deterministic checks. Pass the exact active `work/{name}` path and the exact current checkpoint scope. POST is read-only and must report `converged`; when it reports `gaps`, during audit remediation you are the only writer of active work artifacts: add normal plan tasks and persist their one declared spec source when needed, return to the phase that owns each gap, and rerun POST after the fixes.
|
|
203
|
-
- Only after POST reports `converged`, validate against the plan's **Success criteria** and mark the success criteria complete. Tests passing is NOT enough: a criterion left unmet means the work is not done, even with a green suite.
|
|
204
|
-
- Before SHIP, ensure all applicable preflight work is complete: code, version bump, local tests, the project's quality command (`pnpm qa:quality` when defined), and Vercel preview review when the project uses Vercel. React Doctor is manual/local, never assumed to be a GitHub Actions gate.
|
|
205
|
-
- If something fails, go back to EXECUTE with fix tasks.
|
|
206
|
-
- **Anti-thrashing**: max 3 attempts per failing task or criterion. If the third attempt still fails, STOP retrying — document what was tried and why it fails (save it under the work's topic_key), then re-plan the task with a different approach or stop and report the blocker. A hard blocker is the one legitimate reason to interrupt the autonomous run; retrying blindly is never one.
|
|
46
|
+
After each intermediate merge: persist the checkpoint to `work/{name}/pr/{NN}`, update `plan.md`, and keep `work/{name}/` alive. Merge always requires explicit user approval.
|
|
207
47
|
|
|
208
|
-
|
|
48
|
+
### Deterministic verification
|
|
209
49
|
|
|
210
|
-
|
|
50
|
+
Deterministic checks are the routine feedback loop while implementation is in progress: run the relevant tests, lint, and typecheck/build checks at the cheapest seam that can catch the section's regressions. Verify coherent sections rather than every small edit or one big-bang run. Normal handoffs between `implementer` and `tester` do not by themselves justify reviewers or analyzers.
|
|
211
51
|
|
|
212
|
-
|
|
213
|
-
2. Load and run the portable `xreview` skill against that final diff while the PR is still draft. Use the exact active `work/{name}` already established for this work and include it verbatim as the work context in every review subagent prompt; never infer it from the branch or scan other `work/*` folders. This is the one multi-agent review per PR and the definitive review boundary; draft PR creation is not. Process the report by its three levels:
|
|
214
|
-
- **Critical Issues (must fix)**: apply ALL of them — the PR must not reach merge with these open.
|
|
215
|
-
- **Important Improvements (should fix)**: apply the ones worth doing now, at your judgment.
|
|
216
|
-
- **Suggestions (nice to have)**: apply only if trivial and safe.
|
|
217
|
-
3. Every finding you decide NOT to apply now goes to the project's `work/backlog` single topic_key — one line each: what + why deferred. Apply the safe serialized backlog protocol above; subagents only return candidate lines.
|
|
218
|
-
4. For what you DO apply: add the new tasks to plan.md and persist one declared recoverable spec source per formal task, execute them as in EXECUTE, re-verify, and push the fixes while the PR remains draft. Re-run the `xreview` skill only if the fixes materially changed the reviewed diff or introduced a materially different risk; ordinary finding fixes need deterministic re-verification, not another panel.
|
|
219
|
-
5. Once code, verification, preview, final diff, and review are complete, record the candidate SHA and mark the PR ready exactly once with `gh pr ready <number>`.
|
|
220
|
-
6. Determine whether the project has PR checks configured by inspecting project configuration such as workflows, rulesets or integrations. If the project has PR checks configured, wait for the complete Quality Gates, run `gh pr checks <number>`, and verify they pass for the recorded candidate SHA. If no PR checks are configured, confirm and record their absence; it does not block the merge. An empty `gh pr checks` result immediately after ready is not evidence that no checks are configured. In either case, do not push while the PR is ready. Immediately before reporting or merging, compare `gh pr view --json headRefOid` with the recorded candidate SHA.
|
|
221
|
-
7. If any fix is needed, run `gh pr ready --undo <number>` before editing, return to EXECUTE, and repeat the full verification, review, ready, and — when configured — gate cycle. Never treat checks from an older SHA as merge evidence.
|
|
52
|
+
### Early review
|
|
222
53
|
|
|
223
|
-
|
|
54
|
+
An early review during EXECUTE is an exception, not a default phase. Use it only for a concrete risk that deterministic checks cannot cover and whose feedback can change the remaining implementation. State the bounded risk, use the single most relevant specialist, and run at most one early review per bounded critical section. Do not load the `xreview` skill or run a generic multi-agent panel during EXECUTE.
|
|
224
55
|
|
|
225
|
-
|
|
226
|
-
- NEVER merge the PR yourself — merge only on an explicit user order. After each intermediate merge: persist the checkpoint to `work/{name}/pr/{NN}`, update `plan.md`, and keep `work/{name}/` alive. After the final merge: persist the final outcome to memory, clean up `work/{name}/` and remove the worktree (see Work state).
|
|
227
|
-
- If the repo has its own skill for the closing steps (release, deploy, git, cleanup), that skill takes precedence over the default behavior.
|
|
56
|
+
Do not launch `code-reviewer`, `code-simplifier`, `test-analyzer`, or `silent-failure-hunter` merely because a writer finished, a test task completed, several files changed, or a commit is due. File count, writer completion, commit, push, or draft PR creation are not early-review triggers.
|
|
228
57
|
|
|
229
|
-
|
|
58
|
+
### Bounded retries
|
|
230
59
|
|
|
231
|
-
|
|
60
|
+
Both routes cap a failing task or criterion at three attempts. After the third failure, stop retrying, save the meaningful failure and what was tried through the required outcome or memory path, then re-plan with a different approach or report the blocker; never loop blindly.
|
|
61
|
+
### Final review and PR lifecycle
|
|
232
62
|
|
|
233
|
-
|
|
63
|
+
The review boundary is the final candidate SHA while the PR is still draft, immediately before `gh pr ready`. It is one final review, not a mandatory panel or multiple reviewers for short work. Load and run the portable `xreview` skill only when a full review has not already covered that final draft and its multi-agent review adds value; in that case, pass the exact active `work/{name}` to every review subagent prompt and do not infer it from another `work/*` folder. Re-run it only if fixes materially change the diff or introduce a materially different risk; otherwise retain the prior review evidence and run deterministic verification.
|
|
234
64
|
|
|
235
|
-
|
|
236
|
-
- Read-only agents can run in parallel.
|
|
237
|
-
- Write agents only run in parallel if they don't touch the same files.
|
|
65
|
+
Both routes keep the PR draft while it changes, use the canonical Git worktree, complete review before ready, wait for configured gates, compare the candidate SHA before reporting or merging, and never merge without explicit user approval.
|
|
238
66
|
|
|
239
67
|
## Closing rule
|
|
240
68
|
|
|
241
|
-
|
|
69
|
+
Do not declare work finished after analysis or planning alone: complete the routed execution or report the concrete blocker.
|
|
@@ -0,0 +1,189 @@
|
|
|
1
|
+
# Standard workflow
|
|
2
|
+
|
|
3
|
+
Read this reference only after the orchestrator routes work to **standard**. The standard route uses a formal SDD PRD and plan, with PRE before approval and POST before SHIP; it preserves change-first and the shared guards in SKILL.md.
|
|
4
|
+
|
|
5
|
+
## Phases
|
|
6
|
+
|
|
7
|
+
```text
|
|
8
|
+
INIT → EXPLORE → SPEC → PLAN → EXECUTE → VERIFY → SHIP → CLOSE
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
### Autonomy
|
|
12
|
+
|
|
13
|
+
The human drives the flow UP TO the plan: the idea, the PRD review and the plan review are interactive. Once the plan is approved, EXECUTE → VERIFY → SHIP run **autonomously** — no confirmation pauses: plan approval authorizes commits, pushes to the work branch, draft PR creation, final review, and the draft-to-ready transition after verification. Task-critical uncertainty from a subagent is an operational blocker, not a pause in autonomy: answer from existing context first; only if the decision genuinely cannot be made from available context may you ask the user, then relaunch with explicit guidance. Control returns to the user at CLOSE. Merging the PR is NEVER yours: it always requires an explicit user order. For multi-PR work, each merge is a checkpoint; keep `work/{name}/PRD.md` and `plan.md` alive until the roadmap is finished. Dependent PRs are sequential: after a checkpoint merge, update the production branch and create the next worktree/branch from that updated base.
|
|
14
|
+
|
|
15
|
+
## 1. INIT
|
|
16
|
+
|
|
17
|
+
- Load previous context from Engram memory: for non-trivial reads, delegate to the `engram` subagent (`mem_context` / `mem_search` filtered to the task).
|
|
18
|
+
- Identify the project's constraints.
|
|
19
|
+
- Detect whether there is documentation, issues or artifacts already created.
|
|
20
|
+
|
|
21
|
+
## 2. EXPLORE
|
|
22
|
+
|
|
23
|
+
Launch analysts according to scope:
|
|
24
|
+
|
|
25
|
+
- `backend-analyst` if it affects backend, DB, APIs or server functions
|
|
26
|
+
- `frontend-analyst` if it affects UI, hooks, state or rendering
|
|
27
|
+
- `security-auditor` if the area is sensitive
|
|
28
|
+
|
|
29
|
+
## 3. SPEC
|
|
30
|
+
|
|
31
|
+
- Synthesize findings.
|
|
32
|
+
- Propose a simple approach.
|
|
33
|
+
- Clarify only the real ambiguities.
|
|
34
|
+
- Apply the `lean-code` skill as a scope gate for any code-bearing task: ask whether the code is needed at all, whether stdlib/native/project helpers already solve it, and whether the smallest obvious change is enough.
|
|
35
|
+
- Backlog items phrased as "consider/evaluate X" are questions, not requirements: answer them HERE — who consumes it, what real case needs it — before they enter the PRD as committed scope. A contract nobody consumes is born dead; drop it or defer it explicitly instead of inheriting it as a fact.
|
|
36
|
+
- Create the PRD before moving to PLAN (see PRD rules).
|
|
37
|
+
|
|
38
|
+
### PRD rules
|
|
39
|
+
|
|
40
|
+
The PRD is **mandatory by default** when you work as orchestrator. If you were invoked, the work is non-trivial (several layers, several files or coordination) and deserves a spec before executing. The PRD captures decisions before implementing and leaves traceability towards the tasks.
|
|
41
|
+
|
|
42
|
+
Use the `to-prd` skill to turn the current context into the PRD (`work/{name}/PRD.md`) before planning execution.
|
|
43
|
+
|
|
44
|
+
|
|
45
|
+
When presenting the PRD for review, offer a disposable HTML view (rules in the `work-lifecycle` skill).
|
|
46
|
+
|
|
47
|
+
If the work is large enough to benefit from explicit vertical slices, use the `to-issues` skill after the PRD to split it into independently executable slices before detailed planning.
|
|
48
|
+
|
|
49
|
+
## 4. PLAN
|
|
50
|
+
|
|
51
|
+
- Use the PRD as the base input for planning.
|
|
52
|
+
- If a slice breakdown exists from `to-issues`, use it as the structure for planning and task sequencing.
|
|
53
|
+
- Divide the work into clear tasks.
|
|
54
|
+
- One task = one agent = one scope.
|
|
55
|
+
- For tasks that add or grow code, record the lean-code outcome in the task spec/acceptance criteria so implementer and simplifier apply the same ladder.
|
|
56
|
+
- The PRD does not replace the plan or task breakdown: the PRD captures decisions; the plan and tasks turn those decisions into executable work.
|
|
57
|
+
- **Change-first**: for intentional material contract changes discovered in EXECUTE or VERIFY—not bugfixes that restore the approved contract—return to SPEC before further implementation. Update the PRD first, then propagate it to the plan, task specs, `SC-*` success criteria and testing decisions; rerun PRE until `clean`, obtain human approval of the delta, then resume EXECUTE and repeat VERIFY.
|
|
58
|
+
- Materialize the plan per the Work state rules: `work/{name}/plan.md` with the task table and one declared recoverable spec source per formal task — an Engram observation or canonical Markdown (templates in the `work-lifecycle` skill). Record the exact source in `Spec`; never create two active copies.
|
|
59
|
+
- Load and run the `work-audit` skill in **PRE** mode after the plan and task specs exist and before presenting the final plan. Pass the exact active `work/{name}` path and the exact PR/checkpoint scope; never infer either from the branch or scan other work folders. PRE is read-only: during audit remediation you are the only writer of active work artifacts. Route every finding to its owner artifact, correct it, and rerun PRE until it reports `clean`.
|
|
60
|
+
- An unresolved `[NEEDS CLARIFICATION: ...]` marker blocks PRE. Return to SPEC and resolve the ambiguity with the user only when existing context cannot answer it; never approve or execute a plan while PRE is not clean.
|
|
61
|
+
- When presenting the plan for human review, offer a disposable HTML view (rules in the `work-lifecycle` skill). If human review changes the PRD, plan or task specs, rerun PRE and require `clean` again before approval or EXECUTE. Requested changes go to plan.md; delete the HTML once its artifact is approved, before EXECUTE.
|
|
62
|
+
|
|
63
|
+
|
|
64
|
+
## 5. EXECUTE
|
|
65
|
+
|
|
66
|
+
### Worktree
|
|
67
|
+
|
|
68
|
+
Follow the project Git/worktree rules before the first task. Standard work uses the canonical git worktree unless those rules explicitly permit a trivial direct-main exception; short never creates its own exception.
|
|
69
|
+
|
|
70
|
+
Canonical location is mandatory: resolve the project root with `git rev-parse --show-toplevel`, ensure `worktrees/` is ignored in the repo-local `.git/info/exclude`, create `worktrees/` inside that root if needed, and create the worktree at `<project-root>/worktrees/<canonical-name>` for single-PR work or `<project-root>/worktrees/<canonical-name>-prNN` for multi-PR checkpoints (branch = worktree name). Do not create worktrees next to the repo, in the repo root, under `work/`, or in any external temp/shared folder.
|
|
71
|
+
|
|
72
|
+
Every delegation prompt must state the worktree path as the ONLY allowed write root. After each writer subagent finishes, verify the user's main checkout is still clean (`git status` there); if the subagent wrote outside the worktree, STOP, move those changes into the worktree (patch/apply) and restore the main checkout before continuing. Subagent obedience is not a safety boundary — this check is.
|
|
73
|
+
|
|
74
|
+
### Commit cadence
|
|
75
|
+
|
|
76
|
+
Commit after each task or bounded group of tasks, with a message that reflects that task — the branch history must map to the plan. Never accumulate the whole work into one giant commit at the end.
|
|
77
|
+
|
|
78
|
+
### Draft PR cadence
|
|
79
|
+
|
|
80
|
+
- After the first coherent commit, push the branch and create the PR against its real base with `gh pr create --draft`. Do not wait until SHIP to open it.
|
|
81
|
+
- Keep every code change, commit and push inside the draft phase. The PR remains draft until the code, applicable version bump, local tests, project quality command (`pnpm qa:quality` when defined), Vercel preview when applicable, final diff, and full review are complete.
|
|
82
|
+
- Never push to a ready PR. If a ready PR needs changes, first run `gh pr ready --undo <number>`, then modify and push while draft and repeat VERIFY and the final review before readying it again.
|
|
83
|
+
|
|
84
|
+
### Handoff rule
|
|
85
|
+
|
|
86
|
+
The analyst's **Recommendation** is the implementer's input. Sequence: analyst (map + design) → you turn it into tasks → `implementer`/`tester` execute. Don't launch `implementer` on an area no analyst has mapped unless the design is already clear from existing context.
|
|
87
|
+
|
|
88
|
+
### Testing decision
|
|
89
|
+
|
|
90
|
+
Every implementation task needs a testing decision, not automatically a new test. Establish:
|
|
91
|
+
|
|
92
|
+
- the meaningful regression risk introduced by the change
|
|
93
|
+
- the existing test that already protects it, if any
|
|
94
|
+
- the new or changed behavior that needs protection
|
|
95
|
+
- the strongest seam closest to that risk
|
|
96
|
+
- the action: TDD/new test, update, reuse existing coverage, or no new test with a concrete trivial/mechanical/already-covered reason
|
|
97
|
+
|
|
98
|
+
Prefer one authoritative test per behavior. Another layer is justified only when it protects a distinct contract. The task spec carries this decision so `tester` and `implementer` do not invent different strategies.
|
|
99
|
+
|
|
100
|
+
### TDD path
|
|
101
|
+
|
|
102
|
+
Use for business rules, bugs/regressions, public contracts, invariants, security/data boundaries, or other behavior whose risk warrants new protection.
|
|
103
|
+
|
|
104
|
+
```text
|
|
105
|
+
tester (RED) → implementer (GREEN/REFACTOR)
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
### Direct path
|
|
109
|
+
|
|
110
|
+
Use for styling, wiring, generated code, mechanical refactors, trivial code, or changes already covered by an authoritative test. Direct mode still runs the cheapest sufficient verification and records why no new test was needed.
|
|
111
|
+
|
|
112
|
+
```text
|
|
113
|
+
implementer (direct change)
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
### Special delegations
|
|
117
|
+
|
|
118
|
+
- `translator` for translations or multilingual visible text
|
|
119
|
+
- `docs-maintainer` for documentation
|
|
120
|
+
- `security-auditor` for sensitive review
|
|
121
|
+
|
|
122
|
+
### Verification cadence
|
|
123
|
+
|
|
124
|
+
Deterministic checks are the routine feedback loop while implementation is in progress: run the relevant tests, lint and typecheck/build checks at the cheapest seam that can catch the section's regressions. Verify by bounded, coherent sections (e.g. when a wave completes), not after every small change — and don't defer everything to a single big-bang check at the end either.
|
|
125
|
+
|
|
126
|
+
Each writer verifies its own bounded area (e.g. its test file). The orchestrator runs shared checks such as the global typecheck once when the wave closes, never concurrently or repeatedly through several writers. Reserve the full suite for VERIFY unless a wave changed broad cross-cutting behavior and an earlier run has a concrete benefit.
|
|
127
|
+
|
|
128
|
+
### CI and workflow scope
|
|
129
|
+
|
|
130
|
+
CI guidance belongs here only when the task explicitly affects workflows, gates, path filters, execution frequency, or CI cost. Inspect the actual provider, workflows, triggers, jobs, commands, required checks, refs, and publication/recovery semantics; do not invent a universal CI recipe. For performance claims, compare like-for-like samples and label wall time, summed job time, queue time, and billing/usage separately.
|
|
131
|
+
|
|
132
|
+
- Use explicit base/head (or equivalent) refs for diff and path decisions; validate command errors and shared configuration. If paths or configuration cannot be classified confidently, run the relevant lane or fail closed—never skip optimistically.
|
|
133
|
+
- Draft/candidate validation may cancel obsolete validation runs when project semantics allow it, but never cancel a mutable publish/release job or leave publication halfway. Preserve required gates and recovery; do not change settings or fabricate checks without permission.
|
|
134
|
+
|
|
135
|
+
### Early-review budget
|
|
136
|
+
|
|
137
|
+
An early review during EXECUTE is an **exception**, not a default phase. Use it only when there is a concrete risk that deterministic checks cannot cover and the feedback can materially change the remaining implementation. Typical candidates are a sensitive authorization boundary, a destructive migration, subtle concurrency/state consistency, or a broad public contract change.
|
|
138
|
+
|
|
139
|
+
- State the exact risk and the bounded diff section to inspect before launching anyone.
|
|
140
|
+
- Use the single most relevant specialist. Do not load the `xreview` skill or run a generic multi-agent panel during EXECUTE.
|
|
141
|
+
- Run at most one early review per bounded critical section, after that section is coherent rather than after each task inside it.
|
|
142
|
+
- Do not launch `code-reviewer`, `code-simplifier`, `test-analyzer` or `silent-failure-hunter` merely because a writer finished, a test task completed, several files changed or a commit is due.
|
|
143
|
+
- File count, writer completion, commit, push, or draft PR creation are not early-review triggers. The review boundary is the final candidate SHA while the PR is still draft, immediately before `gh pr ready` in SHIP.
|
|
144
|
+
|
|
145
|
+
## 6. VERIFY
|
|
146
|
+
|
|
147
|
+
- Run the minimum verification that is sufficient.
|
|
148
|
+
- Reserve heavy suites for cases where they provide real value or the project requires them.
|
|
149
|
+
- If POST identifies an intentional material contract change, follow the PLAN's change-first procedure before further implementation.
|
|
150
|
+
- Load and run the `work-audit` skill in **POST** mode after deterministic checks. Pass the exact active `work/{name}` path and the exact current checkpoint scope. POST is read-only and must report `converged`; when it reports `gaps`, during audit remediation you are the only writer of active work artifacts: add normal plan tasks and persist their one declared spec source when needed, return to the phase that owns each gap, and rerun POST after the fixes.
|
|
151
|
+
- Only after POST reports `converged`, validate against the plan's **Success criteria** and mark the success criteria complete. Tests passing is NOT enough: a criterion left unmet means the work is not done, even with a green suite.
|
|
152
|
+
- Before SHIP, ensure all applicable preflight work is complete: code, version bump, local tests, the project's quality command (`pnpm qa:quality` when defined), and Vercel preview review when the project uses Vercel. React Doctor is manual/local, never assumed to be a GitHub Actions gate.
|
|
153
|
+
- If something fails, go back to EXECUTE with fix tasks.
|
|
154
|
+
- Apply the shared **Bounded retries** guard in `SKILL.md`.
|
|
155
|
+
|
|
156
|
+
## 7. SHIP (automatic)
|
|
157
|
+
|
|
158
|
+
When the plan is fully applied and VERIFY passes:
|
|
159
|
+
|
|
160
|
+
1. Confirm the draft PR exists, the worktree is clean, and the draft head matches the local HEAD. Inspect the final diff against the PR's real base.
|
|
161
|
+
2. Apply **Final review and PR lifecycle** in the entry [SKILL.md](../SKILL.md) and the project's review requirements. Reuse valid prior review evidence; choosing standard does not require another panel. Process the review findings by their three levels:
|
|
162
|
+
- **Critical Issues (must fix)**: apply ALL of them — the PR must not reach merge with these open.
|
|
163
|
+
- **Important Improvements (should fix)**: apply the ones worth doing now, at your judgment.
|
|
164
|
+
- **Suggestions (nice to have)**: apply only if trivial and safe.
|
|
165
|
+
3. Every finding you decide NOT to apply now goes to the project's `work/backlog` single topic_key — one line each: what + why deferred. Apply the safe serialized backlog protocol above; subagents only return candidate lines.
|
|
166
|
+
4. For what you DO apply: add the new tasks to plan.md and persist one declared recoverable spec source per formal task, execute them as in EXECUTE, re-verify, and push the fixes while the PR remains draft. Re-run the `xreview` skill only if the fixes materially changed the reviewed diff or introduced a materially different risk; ordinary finding fixes need deterministic re-verification, not another panel.
|
|
167
|
+
5. Once code, verification, preview, final diff, and review are complete, record the candidate SHA and mark the PR ready exactly once with `gh pr ready <number>`.
|
|
168
|
+
6. Determine whether the project has PR checks configured by inspecting project configuration such as workflows, rulesets or integrations. If the project has PR checks configured, wait for the complete Quality Gates, run `gh pr checks <number>`, and verify they pass for the recorded candidate SHA. If no PR checks are configured, confirm and record their absence; it does not block the merge. An empty `gh pr checks` result immediately after ready is not evidence that no checks are configured. In either case, do not push while the PR is ready. Immediately before reporting or merging, compare `gh pr view --json headRefOid` with the recorded candidate SHA.
|
|
169
|
+
7. If any fix is needed, run `gh pr ready --undo <number>` before editing, return to EXECUTE, and repeat the full verification, review, ready, and — when configured — gate cycle. Never treat checks from an older SHA as merge evidence.
|
|
170
|
+
|
|
171
|
+
## 8. CLOSE
|
|
172
|
+
|
|
173
|
+
- STOP here and hand control back to the user only after configured Quality Gates pass for the latest commit, or after confirming that the project has no PR checks configured: report the candidate SHA, check result or confirmed absence, review findings applied vs deferred to `work/backlog`, and whether manual testing is advisable (recommend it for big or user-facing changes; small well-tested changes may not need it).
|
|
174
|
+
- NEVER merge the PR yourself — merge only on an explicit user order. After each intermediate merge: persist the checkpoint to `work/{name}/pr/{NN}`, update `plan.md`, and keep `work/{name}/` alive. After the final merge: persist the final outcome to memory, clean up `work/{name}/` and remove the worktree (see Work state).
|
|
175
|
+
- If the repo has its own skill for the closing steps (release, deploy, git, cleanup), that skill takes precedence over the default behavior.
|
|
176
|
+
|
|
177
|
+
## Task rule
|
|
178
|
+
|
|
179
|
+
A task must correspond to a single agent and a single scope. Don't mix production, tests, docs and translations in the same task.
|
|
180
|
+
|
|
181
|
+
## Operational rules
|
|
182
|
+
|
|
183
|
+
- The coordinator must not mix scopes in a single task.
|
|
184
|
+
- Read-only agents can run in parallel.
|
|
185
|
+
- Write agents only run in parallel if they don't touch the same files.
|
|
186
|
+
|
|
187
|
+
## Closing rule
|
|
188
|
+
|
|
189
|
+
Don't declare the task finished if you have only analyzed or planned. There must be real execution by the subagents or a concrete blocker.
|
|
@@ -7,16 +7,18 @@ description: Single source for how a piece of work is tracked and advances — P
|
|
|
7
7
|
|
|
8
8
|
One rule kills duplication: **every piece of information has exactly ONE home**. Files hold human-reviewed artifacts and task specs deliberately chosen as Markdown; memory holds Engram-backed task specs and history. Nothing is ever stored in two places.
|
|
9
9
|
|
|
10
|
+
This lifecycle applies after the canonical routing in the `orchestrator` skill selects **formal SDD** work. A short standalone change stays outside this lifecycle and does not create PRD, plan, formal task specs, PRE, or POST for ceremony. A bounded short step inside active formal SDD work preserves its approved Spec, plan row, ownership, and tracking. Routing criteria live only in `orchestrator`; do not copy them here.
|
|
11
|
+
|
|
10
12
|
## Identity
|
|
11
13
|
|
|
12
|
-
Every piece of work gets a **canonical kebab-case name** when it starts (e.g. `checkout-refactor`), shared with the single-PR branch/worktree name and the base name for multi-PR checkpoints. The name stays the same across its whole life — it is the key to everything else.
|
|
14
|
+
Every formal SDD piece of work gets a **canonical kebab-case name** when it starts (e.g. `checkout-refactor`), shared with the single-PR branch/worktree name and the base name for multi-PR checkpoints. The name stays the same across its whole life — it is the key to everything else.
|
|
13
15
|
|
|
14
16
|
## Where everything lives
|
|
15
17
|
|
|
16
18
|
| Piece | Single home | Why there |
|
|
17
19
|
|---|---|---|
|
|
18
|
-
| PRD | `work/{name}/PRD.md` | Written once, reviewed by the human |
|
|
19
|
-
|
|
|
20
|
+
| Formal SDD PRD | `work/{name}/PRD.md` | Written once, reviewed by the human |
|
|
21
|
+
| Formal SDD plan (goal, approach, task board) | `work/{name}/plan.md` | The status board: humans glance at it; statuses flip with surgical edits |
|
|
20
22
|
| Full spec of each formal task | One recoverable source: Engram project + topic_key `work/{name}/task/{NN}` **or** `work/{name}/tasks/{NN}.md` | Choose for durable access; record the exact reference in the plan |
|
|
21
23
|
| PR checkpoint outcome | Engram `work/{name}/pr/{NN}` | Intermediate PR merge record |
|
|
22
24
|
| Phase outcomes, decisions, findings | Engram `work/{name}/{phase}` | History — must survive the folder and compactions |
|
|
@@ -27,6 +29,8 @@ Every piece of work gets a **canonical kebab-case name** when it starts (e.g. `c
|
|
|
27
29
|
|
|
28
30
|
## Starting
|
|
29
31
|
|
|
32
|
+
Use this section only after routing selects formal SDD work, including work promoted from short before it expands. Do not use it to formalize a standalone short change by ceremony.
|
|
33
|
+
|
|
30
34
|
1. Pick the canonical name and create `work/{name}/`. If the item came from the backlog, remove it from `work/backlog` in the same step.
|
|
31
35
|
2. Produce the PRD with the `to-prd` skill → `work/{name}/PRD.md`. The human reviews it there.
|
|
32
36
|
3. Write `work/{name}/plan.md` (structure in `references/plan-template.md`): goal, chosen approach, `SC-*` success criteria, PR roadmap, and the task table — number, PR, agent, scope, **Spec**, title, one-line description, SC coverage, status, wave, deps.
|
|
@@ -89,9 +89,9 @@ docs/
|
|
|
89
89
|
|
|
90
90
|
## Work State
|
|
91
91
|
|
|
92
|
-
Every piece of information about a piece of work has exactly ONE home — never two. The `work-lifecycle` skill is the single source of this flow.
|
|
92
|
+
Every piece of information about a piece of work has exactly ONE home — never two. The `work-lifecycle` skill is the single source of this flow for formal SDD work. The canonical `orchestrator` routing decides whether work is short or standard; a standalone short route does not create PRD, plan, formal Specs, PRE, or POST by ceremony, while a short step inside active formal SDD work preserves its approved tracking and ownership.
|
|
93
93
|
|
|
94
|
-
- In-progress work lives in `work/{name}/` (gitignored): `PRD.md` + `plan.md`. They stay there across intermediate PR merges; `plan.md` is the ONLY task status board — update statuses with surgical edits. An empty `work/` means nothing is half-done.
|
|
94
|
+
- In-progress formal SDD work lives in `work/{name}/` (gitignored): `PRD.md` + `plan.md`. They stay there across intermediate PR merges; `plan.md` is the ONLY task status board — update statuses with surgical edits. An empty `work/` means nothing is half-done.
|
|
95
95
|
- Execution worktrees and their branches always use the same name. Resolve the root with `git rev-parse --show-toplevel`, ensure `worktrees/` is ignored in the repo-local `.git/info/exclude`, then create/use `<project-root>/worktrees/<canonical-name>` for single-PR work or `<project-root>/worktrees/<canonical-name>-prNN` for multi-PR checkpoints; never create worktrees next to the repo, in the repo root, under `work/`, or in external temp/shared folders.
|
|
96
96
|
- Every formal task has one declared recoverable spec source: an Engram observation identified by project + topic_key `work/{name}/task/{NN}` with verified identity/access (optional local ID bound in the current store; resolve per the lifecycle handoff before get), or canonical Markdown at `work/{name}/tasks/{NN}.md`; its plan `Spec` column is the reference. Direct messages are only auxiliary microassignments under a parent task; if one becomes independent, persist its spec and add its plan row before continuing. Phase outcomes, PR checkpoints and history remain in Engram under `work/{name}/{phase}`, `work/{name}/pr/{NN}` and `work/{name}/done`.
|
|
97
97
|
- Pending work: the project's single `work/backlog` topic_key, or issues (`to-issues`) if the project uses a tracker. Never a TODOs folder. The coordinator/orchestrator is its **single writer**: retrieve the exact observation with `mem_get_observation`, preserve unrelated entries, send the complete content with `mem_update`, then read it again to verify; never mutate it concurrently or use a blind topic-key upsert. Do not split it into one memory per item until Engram supports complete paginated topic-prefix listing.
|