autonomous-sdlc-harness 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +201 -0
- package/NOTICE +7 -0
- package/README.md +24 -0
- package/dist/cli.js +194 -0
- package/dist/cli.js.map +1 -0
- package/dist/commands/config.js +561 -0
- package/dist/commands/config.js.map +1 -0
- package/dist/commands/daemon.js +791 -0
- package/dist/commands/daemon.js.map +1 -0
- package/dist/commands/doctor.js +336 -0
- package/dist/commands/doctor.js.map +1 -0
- package/dist/commands/init.js +2023 -0
- package/dist/commands/init.js.map +1 -0
- package/dist/commands/registry.js +42 -0
- package/dist/commands/registry.js.map +1 -0
- package/dist/config/check.js +505 -0
- package/dist/config/check.js.map +1 -0
- package/dist/config/io.js +177 -0
- package/dist/config/io.js.map +1 -0
- package/dist/config/model.js +406 -0
- package/dist/config/model.js.map +1 -0
- package/dist/core/errors.js +71 -0
- package/dist/core/errors.js.map +1 -0
- package/dist/core/git.js +537 -0
- package/dist/core/git.js.map +1 -0
- package/dist/core/json.js +125 -0
- package/dist/core/json.js.map +1 -0
- package/dist/core/layerCoverage.js +141 -0
- package/dist/core/layerCoverage.js.map +1 -0
- package/dist/core/layerGapRemedy.js +62 -0
- package/dist/core/layerGapRemedy.js.map +1 -0
- package/dist/core/nameList.js +23 -0
- package/dist/core/nameList.js.map +1 -0
- package/dist/core/paths.js +153 -0
- package/dist/core/paths.js.map +1 -0
- package/dist/core/prompt.js +206 -0
- package/dist/core/prompt.js.map +1 -0
- package/dist/core/repoPaths.js +55 -0
- package/dist/core/repoPaths.js.map +1 -0
- package/dist/core/report.js +150 -0
- package/dist/core/report.js.map +1 -0
- package/dist/core/templating.js +88 -0
- package/dist/core/templating.js.map +1 -0
- package/dist/core/writer.js +479 -0
- package/dist/core/writer.js.map +1 -0
- package/dist/daemon/backend.js +180 -0
- package/dist/daemon/backend.js.map +1 -0
- package/dist/daemon/units.js +380 -0
- package/dist/daemon/units.js.map +1 -0
- package/dist/detect/nestedApplication.js +79 -0
- package/dist/detect/nestedApplication.js.map +1 -0
- package/dist/detect/presets.js +2033 -0
- package/dist/detect/presets.js.map +1 -0
- package/dist/detect/signals.js +1368 -0
- package/dist/detect/signals.js.map +1 -0
- package/dist/doctor/checks.js +3530 -0
- package/dist/doctor/checks.js.map +1 -0
- package/dist/generators/claudeContext.js +588 -0
- package/dist/generators/claudeContext.js.map +1 -0
- package/dist/generators/githooks.js +446 -0
- package/dist/generators/githooks.js.map +1 -0
- package/dist/generators/harnessConfig.js +632 -0
- package/dist/generators/harnessConfig.js.map +1 -0
- package/dist/generators/notifications.js +191 -0
- package/dist/generators/notifications.js.map +1 -0
- package/dist/generators/outerLoopScripts.js +165 -0
- package/dist/generators/outerLoopScripts.js.map +1 -0
- package/dist/generators/permissionProfile.js +1172 -0
- package/dist/generators/permissionProfile.js.map +1 -0
- package/dist/generators/projectSettings.js +322 -0
- package/dist/generators/projectSettings.js.map +1 -0
- package/dist/generators/repoRoot.js +417 -0
- package/dist/generators/repoRoot.js.map +1 -0
- package/dist/generators/scripts.js +557 -0
- package/dist/generators/scripts.js.map +1 -0
- package/dist/generators/stateDir.js +221 -0
- package/dist/generators/stateDir.js.map +1 -0
- package/dist/machine/paths.js +111 -0
- package/dist/machine/paths.js.map +1 -0
- package/dist/machine/plugins.js +224 -0
- package/dist/machine/plugins.js.map +1 -0
- package/dist/machine/registry.js +330 -0
- package/dist/machine/registry.js.map +1 -0
- package/package.json +23 -0
- package/scripts/README.md +13 -0
- package/scripts/daemon/launchd.plist.template +59 -0
- package/scripts/daemon/systemd.service.template +58 -0
- package/templates/README.md +15 -0
- package/templates/claude/CLAUDE.md +54 -0
- package/templates/claude/README.md +5 -0
- package/templates/claude/context/api.md +29 -0
- package/templates/claude/context/conventions.md +23 -0
- package/templates/claude/context/data-layer.md +28 -0
- package/templates/claude/context/data-storage.md +29 -0
- package/templates/claude/context/docs-catalog.md +29 -0
- package/templates/claude/context/domain.md +28 -0
- package/templates/claude/context/layer.md +20 -0
- package/templates/claude/context/module.md +30 -0
- package/templates/claude/context/package.md +29 -0
- package/templates/claude/context/presentation.md +32 -0
- package/templates/claude/context/state-slices.md +28 -0
- package/templates/claude/context/tests.md +28 -0
- package/templates/claude/harness-task-offer.md +58 -0
- package/templates/claude/push-notify.env.example +21 -0
- package/templates/claude/qa-accounts.env.example +38 -0
- package/templates/claude/qa_test_scenarios.md +110 -0
- package/templates/claude/settings.autonomous.json +93 -0
- package/templates/claude/settings.autonomous.qa.json +36 -0
- package/templates/githooks/README.md +3 -0
- package/templates/githooks/pre-push +72 -0
- package/templates/repo/README.md +3 -0
- package/templates/repo/gitattributes +16 -0
- package/templates/repo/gitignore +61 -0
- package/templates/repo/gitignore.qa +25 -0
- package/templates/repo/mcp.json +17 -0
- package/templates/scripts/README.md +5 -0
- package/templates/scripts/autonomous-format-stream.sh +95 -0
- package/templates/scripts/autonomous-notify.sh +337 -0
- package/templates/scripts/autonomous-watcher.sh +3087 -0
- package/templates/scripts/cleanup-merged-worktrees.sh +327 -0
- package/templates/scripts/commit-on-branch.sh +288 -0
- package/templates/scripts/create-worktree.sh +360 -0
- package/templates/scripts/deploy.sh +47 -0
- package/templates/scripts/lib/harness-run-lib.sh +1481 -0
- package/templates/scripts/push-branch.sh +140 -0
- package/templates/scripts/refresh-branch.sh +244 -0
- package/templates/scripts/restart-watcher.sh +401 -0
- package/templates/scripts/scratch-run.sh +302 -0
- package/templates/scripts/setup-worktree.sh +262 -0
- package/templates/scripts/start-dev-server.sh +99 -0
- package/templates/scripts/test.sh +50 -0
- package/templates/scripts/typecheck.sh +50 -0
- package/templates/state-dir/README-root.md +13 -0
- package/templates/state-dir/README.md +9 -0
- package/templates/state-dir/architecture_branch_review_point_reviews/README.md +9 -0
- package/templates/state-dir/architecture_branch_reviews/README.md +9 -0
- package/templates/state-dir/architecture_reviews/README.md +9 -0
- package/templates/state-dir/architecture_user_review_reviews/README.md +9 -0
- package/templates/state-dir/autonomous_inbox/README.md +9 -0
- package/templates/state-dir/autonomous_logs/README.md +9 -0
- package/templates/state-dir/branch_statistics/README.md +9 -0
- package/templates/state-dir/business_parity_branch_review_point_reviews/README.md +9 -0
- package/templates/state-dir/business_parity_branch_reviews/README.md +9 -0
- package/templates/state-dir/business_parity_reviews/README.md +9 -0
- package/templates/state-dir/business_parity_user_review_reviews/README.md +9 -0
- package/templates/state-dir/clarification_digests/README.md +9 -0
- package/templates/state-dir/clarifications/README.md +9 -0
- package/templates/state-dir/code_reviews/README.md +9 -0
- package/templates/state-dir/dispatch_additions/README.md +19 -0
- package/templates/state-dir/docs_catalog/README.md +9 -0
- package/templates/state-dir/flow_progress/README.md +9 -0
- package/templates/state-dir/improvement_observations/README.md +19 -0
- package/templates/state-dir/improvement_suggestions.md +29 -0
- package/templates/state-dir/lessons.md +23 -0
- package/templates/state-dir/qa_review_point_reviews/README.md +9 -0
- package/templates/state-dir/qa_reviews/README.md +9 -0
- package/templates/state-dir/review_plan_point_reviews/README.md +9 -0
- package/templates/state-dir/review_plan_reviews/README.md +9 -0
- package/templates/state-dir/scratch/README.md +11 -0
- package/templates/state-dir/skeptic_review_plan_reviews/README.md +9 -0
- package/templates/state-dir/skeptic_review_point_reviews/README.md +9 -0
- package/templates/state-dir/skeptic_reviews/README.md +9 -0
- package/templates/state-dir/story_plans/README.md +9 -0
- package/templates/state-dir/task_plan_point_reviews/README.md +9 -0
- package/templates/state-dir/task_plan_reviews/README.md +9 -0
- package/templates/state-dir/task_plans/README.md +9 -0
- package/templates/state-dir/task_prompts/README.md +9 -0
- package/templates/state-dir/ui_test_plan_reviews/README.md +9 -0
- package/templates/state-dir/ui_test_plans/README.md +9 -0
- package/templates/state-dir/user_review_fix_plan_point_reviews/README.md +9 -0
- package/templates/state-dir/user_reviews/README.md +9 -0
|
@@ -0,0 +1,50 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# Written by `autonomous-sdlc-harness init`, and yours from there on: edit it freely, a re-run
|
|
3
|
+
# keeps your copy. It is what `commands.test` invokes: it runs the raw test command line below
|
|
4
|
+
# from the repository root, forwards whatever arguments it was given to it — what that line makes
|
|
5
|
+
# of them is the command's own property, stated at the line — and prints one verdict line.
|
|
6
|
+
|
|
7
|
+
# Deliberately no `-e`: this script has to outlive its own command's failure long enough to
|
|
8
|
+
# print the verdict line below.
|
|
9
|
+
set -uo pipefail
|
|
10
|
+
|
|
11
|
+
# Exit with a diagnostic and no verdict line: nothing ran, so there is no result to report. A
|
|
12
|
+
# function rather than an inline `echo … && exit 1`, because the command line below is handed the
|
|
13
|
+
# caller's arguments and `exit` refuses extra ones — `init` writes a call to it as that line when
|
|
14
|
+
# no command line resolved.
|
|
15
|
+
harness_fail() {
|
|
16
|
+
echo "$1" >&2
|
|
17
|
+
exit 1
|
|
18
|
+
}
|
|
19
|
+
|
|
20
|
+
# Anchor to the repository root, derived from this script's own location and never from the
|
|
21
|
+
# caller's directory: `commands.*` are command lines run from the repository root and every path
|
|
22
|
+
# in `harness.config.json` is repo-relative. `git -C` on the script's own directory answers with
|
|
23
|
+
# the checkout this copy belongs to, so all three forms the permission profile emits — the
|
|
24
|
+
# repo-relative one, its absolute twin and a sibling worktree's own copy — run at the root of the
|
|
25
|
+
# checkout they were invoked out of, whatever directory the caller stood in.
|
|
26
|
+
script_dir="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
27
|
+
repo_root="$(git -C "$script_dir" rev-parse --show-toplevel 2>/dev/null)"
|
|
28
|
+
if [ -z "$repo_root" ] || ! cd "$repo_root"; then
|
|
29
|
+
harness_fail "{{name}}: could not resolve a repository root from '${script_dir}' (is git on PATH?)"
|
|
30
|
+
fi
|
|
31
|
+
|
|
32
|
+
# Forward the caller's arguments to the command; with none the whole line runs, which every
|
|
33
|
+
# command takes. `"$@"` is safe under `set -u` even when empty. They attach to the *last* command
|
|
34
|
+
# of the line below, so keep that line one command rather than a compound — and whether that
|
|
35
|
+
# command accepts a bare path or name filter is a property of the command, not of this wrapper: a
|
|
36
|
+
# test runner usually does, a line ending in a sub-command's own flags does not (`ctest` needs
|
|
37
|
+
# `-R <regex>`, `cmake --build` needs `--target <name>`, a `cargo clippy … -- <flags>` line
|
|
38
|
+
# forwards into the compiler's arguments). Read the line below before passing any.
|
|
39
|
+
{{command}} "$@"
|
|
40
|
+
status=$?
|
|
41
|
+
|
|
42
|
+
# Exactly one verdict line, so no caller ever appends an exit-code probe to this script. That
|
|
43
|
+
# compound form is what stalls an unattended run on a permission prompt.
|
|
44
|
+
if [ "$status" -eq 0 ]; then
|
|
45
|
+
echo "PASS: {{name}}"
|
|
46
|
+
else
|
|
47
|
+
echo "FAIL: {{name}} (exit $status)"
|
|
48
|
+
fi
|
|
49
|
+
|
|
50
|
+
exit "$status"
|
|
@@ -0,0 +1,50 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# Written by `autonomous-sdlc-harness init`, and yours from there on: edit it freely, a re-run
|
|
3
|
+
# keeps your copy. It is what `commands.typecheck` invokes: it runs the raw type-check command
|
|
4
|
+
# line below from the repository root, forwards whatever arguments it was given to it, and
|
|
5
|
+
# prints one verdict line.
|
|
6
|
+
|
|
7
|
+
# Deliberately no `-e`: this script has to outlive its own command's failure long enough to
|
|
8
|
+
# print the verdict line below.
|
|
9
|
+
set -uo pipefail
|
|
10
|
+
|
|
11
|
+
# Exit with a diagnostic and no verdict line: nothing ran, so there is no result to report. A
|
|
12
|
+
# function rather than an inline `echo … && exit 1`, because the command line below is handed the
|
|
13
|
+
# caller's arguments and `exit` refuses extra ones — `init` writes a call to it as that line when
|
|
14
|
+
# no command line resolved.
|
|
15
|
+
harness_fail() {
|
|
16
|
+
echo "$1" >&2
|
|
17
|
+
exit 1
|
|
18
|
+
}
|
|
19
|
+
|
|
20
|
+
# Anchor to the repository root, derived from this script's own location and never from the
|
|
21
|
+
# caller's directory: `commands.*` are command lines run from the repository root and every path
|
|
22
|
+
# in `harness.config.json` is repo-relative. `git -C` on the script's own directory answers with
|
|
23
|
+
# the checkout this copy belongs to, so all three forms the permission profile emits — the
|
|
24
|
+
# repo-relative one, its absolute twin and a sibling worktree's own copy — run at the root of the
|
|
25
|
+
# checkout they were invoked out of, whatever directory the caller stood in.
|
|
26
|
+
script_dir="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
27
|
+
repo_root="$(git -C "$script_dir" rev-parse --show-toplevel 2>/dev/null)"
|
|
28
|
+
if [ -z "$repo_root" ] || ! cd "$repo_root"; then
|
|
29
|
+
harness_fail "{{name}}: could not resolve a repository root from '${script_dir}' (is git on PATH?)"
|
|
30
|
+
fi
|
|
31
|
+
|
|
32
|
+
# Forward the caller's arguments to the command; with none the whole line runs, which every
|
|
33
|
+
# command takes. `"$@"` is safe under `set -u` even when empty. They attach to the *last* command
|
|
34
|
+
# of the line below, so keep that line one command rather than a compound — and whether that
|
|
35
|
+
# command accepts a bare path or name filter is a property of the command, not of this wrapper: a
|
|
36
|
+
# test runner usually does, a line ending in a sub-command's own flags does not (`ctest` needs
|
|
37
|
+
# `-R <regex>`, `cmake --build` needs `--target <name>`, a `cargo clippy … -- <flags>` line
|
|
38
|
+
# forwards into the compiler's arguments). Read the line below before passing any.
|
|
39
|
+
{{command}} "$@"
|
|
40
|
+
status=$?
|
|
41
|
+
|
|
42
|
+
# Exactly one verdict line, so no caller ever appends an exit-code probe to this script. That
|
|
43
|
+
# compound form is what stalls an unattended run on a permission prompt.
|
|
44
|
+
if [ "$status" -eq 0 ]; then
|
|
45
|
+
echo "PASS: {{name}}"
|
|
46
|
+
else
|
|
47
|
+
echo "FAIL: {{name}} (exit $status)"
|
|
48
|
+
fi
|
|
49
|
+
|
|
50
|
+
exit "$status"
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
# Harness run artifacts
|
|
2
|
+
|
|
3
|
+
Everything the delivery flow writes and reads lives here: the task prompt a branch starts from, the plans made for it, the findings of every review round, the progress ledger a resumed run reads back, the end-of-run statistics, and the two long-lived ledgers at the root of this directory. **The tree is committed** — these artifacts are the record of how each branch was planned, reviewed and fixed, and a reviewer or a resumed run reads them out of the repository rather than out of one machine's scratch space.
|
|
4
|
+
|
|
5
|
+
What git ignores is the machine-local part of it, and it is a short list: the run daemon's per-run transcripts, event logs and run registry under `<state_dir>/autonomous_logs/`; the prompts dropped into `<state_dir>/autonomous_inbox/`; the questions a parked run asks with the answers it is given, under `<state_dir>/clarifications/`; and the throwaway files an agent runs a probe or a mutation check from, under `<state_dir>/scratch/`. All four are ignored **by their contents**, so the committed `README.md` in each survives the rule and the directory keeps its contract. Ignored with them are the stop, pause and dispatch-count files a run leaves flat at the root of this directory while it is in flight — a committed `STOP` being the one that would halt every run for everyone who clones, at a step whose own instruction forbids deleting the file.
|
|
6
|
+
|
|
7
|
+
**Do not rename this directory to a dot-name.** That is this tree's own hard constraint, not a preference. It is tempting to tidy the tree out of sight, and it is the one change that risks unattended operation: this tree has to be writable by an unattended run, and a dot-path is where a host reserves directories an unattended run may not write to — the measured case is the host's own `.claude/**`, where an unattended run completes reporting success with nothing written. The name is configured as `stateDir` in `harness.config.json`, which refuses a dot at the start of any of its path segments rather than trusting which dot-paths a given host reserves — so a value that only reaches a dot-directory by traversal is refused as well as a plainly dot-named one.
|
|
8
|
+
|
|
9
|
+
Some of the harness's state deliberately lives **outside** this tree, so looking for it here is time wasted. The usage assessment unattended runs publish for one another, and the advisory lock that keeps a single repository spending the account's rate-limit window at a time, belong to the **account** rather than to any one repository, so they sit in the platform state directory under the harness's own name; the notification credentials and the watcher's tunable overrides sit likewise in the platform config directory. The harness's own watcher documentation — which ships with the harness rather than with this repository — is the format of record for all four; this note states only where the boundary is.
|
|
10
|
+
|
|
11
|
+
Each directory here carries its own `README.md` naming what a file in it is and who writes it, so the contract lives beside the artifacts rather than in a central index that drifts from them. That README is also what settles whether a directory's ignore rule needs a negation: a directory that has one is ignored by its **contents** plus a negation for that single file, never as a directory, because git cannot re-include a file whose parent directory is excluded. The rule belongs to the generator that writes this repository's ignore file, in the managed block `init` maintains there — not to the tree.
|
|
12
|
+
|
|
13
|
+
_Written by `autonomous-sdlc-harness init`, and yours from there on: edit it freely, a re-run keeps your copy._
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# state-dir/
|
|
2
|
+
|
|
3
|
+
The run-artifact tree `init` materializes at the configured `stateDir`: the task prompts, story plans and per-task plans a run reads and writes, the per-discipline review directories and the per-item findings directories that re-check them, the park-and-ask clarification channel, the flow-progress records, the branch statistics, the unattended inbox and its logs, the scratch directory an agent writes a probe file into, the observations intake, and the committed digest of the clarification channel's resolved exchanges — together with the two long-lived ledgers, the recurring-escape lessons ledger and the improvement-suggestions ledger. The phase-gated families sit here too and are written only when their phase is on: the interactive-test directories, the documentation catalog, and the reference-implementation-parity reviews with their per-item findings root.
|
|
4
|
+
|
|
5
|
+
**Naming inside this subdirectory.** One child directory per artifact directory of the tree, each holding the `README.md` that becomes that directory's contract; the two ledgers sit at this level under their adopter-side names; and the tree's own top-level README is `README-root.md`, because `README.md` here already describes this template directory to a reader of this repository and is not written to an adopter.
|
|
6
|
+
|
|
7
|
+
**The child directories mirror `STATE_DIR_ENTRIES` one for one, and that table — not this listing — is where their order lives:** the always-on families first, then the phase-gated ones grouped by phase. The table in `cli/src/generators/stateDir.ts` is the single declaration of which directories the tree has, and this subdirectory is its material on disk: the writer reads each selected row's contract from the child directory the row names. A row added there without a child directory here fails the write outright — the read happens while the write plan is assembled, so `init` aborts on a missing template before anything reaches the adopter's disk, and a phase-gated row fails only for the adopters who have that phase on, which is the later and worse way to find out. Edit the two together.
|
|
8
|
+
|
|
9
|
+
**Hard constraint, carried here so a later tidy-up does not "fix" it: `stateDir` must not be a dot-directory.** The tree has to be writable by an unattended run, and a dot-path is where a host reserves directories an unattended run may not write to — measured for `.claude/**`, where such a run completes having written nothing. The configuration schema rejects a dot at the start of any path segment of this key rather than trusting which dot-paths a given host reserves, so a value that reaches a dot-directory by traversal is refused as well as a dot-named one.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# architecture_branch_review_point_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>_arch_review/item_<N>/` folder per architecture fix item, holding the findings from re-checking that item's fix: written by the item's layer reviewer and read by the orchestrator. `<N>` is the item's position in the architecture review's ordered fix list — the first entry is `item_1`, the second `item_2`. Inside the folder sits one `review_<iteration>.md` per re-check, numbered from `review_0.md` and incremented on each failed one; a later number is added beside the earlier files, never written over them.
|
|
4
|
+
|
|
5
|
+
The layer reviewer the item's `_(layer: …)_` tag routed to writes the file. The orchestrator reads its verdict to decide whether the item closes, and on a failure the layer implementer is dispatched again and handed the folder's most recent file as the thing to fix. The architecture reviewer never writes here: its own findings against the branch sit in `<state_dir>/architecture_branch_reviews/`, and this directory holds only the re-checks of the fixes for them.
|
|
6
|
+
|
|
7
|
+
A file is written only when a re-check has something to report, and the reviewer creates the folder itself at that moment — so an absent folder is not a statement about the fix. It equally covers an item that passed its first re-check, an item the loop never reached, and a mode running with the per-unit review step off; where an item's outcome is recorded is the readiness checkbox in the architecture-review index, which the committer flips as the fix lands. Nothing supersedes an earlier file — the numbered set is that item's convergence history — and the whole directory is committed with the branch, since no ignore rule reaches it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is expecting this directory's twin to be beside it. The reference-parity fix loop keeps its own re-checks at `<state_dir>/business_parity_branch_review_point_reviews/`, in the identical `item_<N>/review_<iteration>.md` shape — but this directory is always part of the tree while that one exists only while `phases.parity` is `true` in `harness.config.json`. The two families look symmetrical and are not, so a reader who finds this one and not the other has found a phase that is off, not a tree that was written incompletely.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# architecture_branch_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>_arch_review.md` index plus one `<branch>_arch_review/finding_<N>.md` per finding, holding the architecture review of a finished branch: written by the architecture reviewer and read by the orchestrator. The index is thin — a `## Context` paragraph, the `## Phase 2 Readiness — Ordered Fix List`, and one `### N. Title` pointer per finding under its severity section — while the full problem description and the concrete relocation or refactor live in the finding's own file, self-contained enough to implement from alone. The readiness heading is byte-identical to the code review's because the same machinery walks it, and the format both halves are written to is the code review's too: `${CLAUDE_PLUGIN_ROOT}/samples/sample_code_review.md` for the index and `${CLAUDE_PLUGIN_ROOT}/samples/sample_code_review/finding_1.md` for a detail file, from which only the names differ.
|
|
4
|
+
|
|
5
|
+
This is the architecture reviewer's other mode: instead of a plan it judges the branch diff against the default branch, so the findings are about where code actually landed rather than where a plan said it would. The orchestrator reads the readiness list to drive its fix loop and routes each item off that entry's `_(layer: …)_` tag without opening the finding; the layer implementer is handed one `finding_<N>.md`; the layer reviewer re-checks the fix into `<state_dir>/architecture_branch_review_point_reviews/`; and the committer flips the item's `[ ]` to `[x]` as the fix lands.
|
|
6
|
+
|
|
7
|
+
The index appears when the end-of-branch architecture phase has something to report — a clean review writes no file at all and the flow moves straight on. It is committed once, index and per-finding folder together, before the first of its fixes is implemented, so the review that prompted the fixes is on record even if the loop is interrupted. Nothing supersedes it afterwards; a later round is written beside it under the next suffix, and the whole directory is committed with the branch, since no ignore rule reaches it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is renaming the index to whatever a branch calls its own review. The name is load-bearing twice over: the per-finding folder is that name without its extension, so a rename orphans every pointer the index carries, and a re-review round appends its suffix to this exact base — the same rule the code-review index follows, and the reason both are shaped `<branch>_<what>_review`. The folder always mirrors the suffix its index carries.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# architecture_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>/review_{iteration}.md` per planning iteration, holding the plan-gate architecture review of the task plan: written by the architecture reviewer and read by the plan writer on the next iteration. One `<branch>` subdirectory holds one branch's set. `{iteration}` is this directory's own index, and a new file takes the **next free index in that branch's directory** rather than restarting at zero. The numbers in one directory are consecutive: each plan gate resolves its index in its own directory, so this directory holds its own series — a gap is a file that went missing, not a round another gate owned.
|
|
4
|
+
|
|
5
|
+
What is under review is the **task plan** — the story index's `## Context` plus every per-task file — and the lens is layer placement and dependency direction only: where each described file would land, which layer owns each responsibility, which way imports cross a boundary, judged against the conventions documents the configured layers name. The reader is the plan writer, which applies the Must Fix items to the story index or to the per-task files the findings name. The reviewer creates the branch directory itself and only when it has findings to write: a clean review touches no disk, so a plan whose architecture passed on the first gate leaves this directory empty.
|
|
6
|
+
|
|
7
|
+
Files appear one per failed gate and stop when the loop ends — at convergence, or at the fifth iteration, where the gate escalates and reports the latest path here. Nothing supersedes an earlier file; the accumulated set is the record of how the plan's architecture converged, and the whole directory is committed with the branch, since no ignore rule reaches it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is justifying a deliberate architecture decision anywhere but in the plan itself. This gate returns FAIL — and so drives another writer iteration — whenever it has **anything** to write, and it verifies a claimed exception rather than accepting it: it opens the authorization the justification cites and confirms that authorization is real and actually licenses the deviation. A decision defended only in the writer's head, or in a note the plan does not carry, therefore comes back every round and the loop ping-pongs to the cap over a call that was fine. The justification has to be in the plan for the gate to close.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# architecture_user_review_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>/review_{iteration}.md` per iteration, holding the architecture reviewer's plan-gate review of a hands-on-review fix plan rather than of a task plan: read by the fix-plan writer on the next iteration. One `<branch>` subdirectory holds one branch's set. The fix-plan index plays the role the story index plays at the task-plan gate, and its per-finding folder the role of the per-task files. `{iteration}` is this gate's own counter — each gate of the fix-plan flow is a separate loop carrying its own — and it is a **plain counter, not a free-index resolution**: `${CLAUDE_PLUGIN_ROOT}/instructions/user_review_fix_plan_writing_instructions_core.md` → `## Flow` step 8 opens with `Set arch_iteration = 0` and increments it once per FAIL, and the reviewer writes to that index without listing the directory. A **re-entered** gate therefore restarts at `0` and **writes over** the earlier round's file at that index.
|
|
4
|
+
|
|
5
|
+
The writer is the architecture reviewer, the same agent and the same layering lens that grades a task plan into `<state_dir>/architecture_reviews/`; it is told which of the two it is holding by the paths it is passed, never by a flag. The reader is the fix-plan writer, re-dispatched in revision mode with the path to the newest file and expected to apply its Must Fix items to the fix-plan index or to the per-finding files they name. The reviewer creates the branch directory itself and only when it has findings to write, so a fix plan whose architecture was clean leaves this directory empty.
|
|
6
|
+
|
|
7
|
+
A file appears only for a failed gate, and the gate sits early: the drafted fix plan is graded **before** anyone approves it and before a single one of its fixes is implemented. Files stop at convergence or at the fifth iteration, where the flow escalates with the latest path here. No file supersedes an earlier one in meaning, but the numbered set is only the **latest** pass through this gate rather than the fix plan's whole convergence history: step 7's FAIL path sends the revised plan back through step 6 and into this gate again, and that re-entry rewrites `review_0.md` in place — copy a round aside first if you need it. The whole directory is committed with the branch, since no ignore rule reaches it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is writing for the orchestrator. The flow deliberately does **not** read this file to judge the review's quality; it parses the verdict line, and on FAIL it hands the path straight to the fix-plan writer's revision dispatch. So the only consumer of the prose is that writer, and a finding stated as an observation rather than as a concrete change to a named plan file gives it nothing to apply — the gate fails again on the same point, and nothing in the loop notices why.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# autonomous_inbox/
|
|
2
|
+
|
|
3
|
+
The unattended run loop's drop point: one file per run request, and its **name** is the whole request — it selects the engine command, the working-copy strategy and where the dropped file lands, all three at once. `<branch>_task_prompt.md` starts a delivery run in a fresh working copy and lands the prompt in `<state_dir>/task_prompts/`; `<branch>_review[_<n>].md` starts a fix round for a hands-on review in that branch's existing working copy and lands the review in `<state_dir>/user_reviews/` with its round suffix intact; `<branch>_docs.md` starts a documentation run and lands the checklist in `<state_dir>/docs_catalog/`. The three patterns are anchored suffixes and mutually exclusive, so a branch whose own name contains `review` or `task_prompt` is not mis-routed, and only the branch is read out of the name — never the round, which the engine resolves inside the working copy.
|
|
4
|
+
|
|
5
|
+
A drop is written by the harness's prompt command, by its hands-on-review command, or by hand for a documentation checklist, which no shipped command writes. The only reader is the run daemon, on its next poll pass: there is no queue server and no webhook in between. Once it has consumed a file the daemon moves it aside into `.processed/` under a timestamp; a file matching **none** of the three patterns is archived into the same place under a `rejected_` prefix, with one log line and no registry record. A drop the daemon defers instead — the kill switch is up, the run cap is full, the usage hold is on — is left exactly where it is, so a later pass picks it up with nothing to re-drop by hand.
|
|
6
|
+
|
|
7
|
+
A file therefore lives here for at most one poll interval, and this directory's steady state is empty. Its contents are **machine-local and gitignored**: the directory is ignored by its contents rather than as a directory, and this README survives by an explicit negation, which makes it the only committed file here.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is renaming this file, because its own name is a runtime contract. The watcher scans this directory for `*.md` and archives anything that does not match one of its drop patterns, exempting the basename `README.md` exactly — rename it and the first tick moves it into `.processed/`, leaving a deletion of a tracked file in this repository's `git status`. The exemption is there because that archiving behaviour was once observed taking this very file; the harness's verification record for the outer loop, which ships with its documentation, keeps the account of the defect and of the fix.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# autonomous_logs/
|
|
2
|
+
|
|
3
|
+
One readable transcript `<branch>.log` and one raw event log `<branch>.stream.jsonl` per dispatched run, plus the daemon's own `watcher.log`, the service manager's capture of its standard output and error as `watcher.out.log` and `watcher.err.log`, and `registry.json` — the run registry that indexes every run's status. All of it resolves to `<state_dir>/autonomous_logs/` **in the main checkout**, whichever sibling working copy a run actually executes in, so one repository has one set of these files however many runs it has in flight.
|
|
4
|
+
|
|
5
|
+
Everything here is written by the run daemon and read by the operator watching or diagnosing a run: the transcript is formatted live and tailable while the run is going, and the raw stream beside it is the deep-debug copy and the one thing the usage gate can parse, since a run cannot read the stream it is itself producing. The registry has readers beyond the operator, and that is what makes it different in kind from its neighbours: it is the single record of which runs exist and what state each is in, and the harness's status command, its pause and resume commands, `<scripts_dir>/restart-watcher.sh` and `<scripts_dir>/cleanup-merged-worktrees.sh` all read it before they act.
|
|
6
|
+
|
|
7
|
+
A run's two files appear when it launches and are **appended** to, not replaced — a resumed run, a later review round or a re-launch on the same branch adds to the same pair, so one transcript may hold several sessions end to end. The registry likewise holds one record per branch, re-used across a branch's runs rather than accumulating one per launch. None of it is committed: the contents are **machine-local and gitignored**, the directory is ignored by its contents rather than as a directory, and this README survives by an explicit negation, the same way the unattended loop's inbox does.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is treating the registry as a log. `registry.json` is **state** — hand-editing it, or deleting it to clear some clutter, is how a live run becomes invisible to every command that would otherwise pause, resume or clean up after it, and none of them will re-derive it from the transcripts sitting beside it. The transcripts are the disposable part of this directory; the registry is not.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# branch_statistics/
|
|
2
|
+
|
|
3
|
+
One `<branch>/statistics.md` per branch, holding that run's end-of-branch metrics: the story-point denominator summed off the story index, the severity-weighted cost of the issues a hands-on review raised, the success rate derived from those two, and the raw counts they came from, retained so every number is re-derivable by hand from the same files. Note the per-branch **subdirectory** — the report is `<state_dir>/branch_statistics/<branch>/statistics.md`, and the writer creates that subdirectory itself if it is absent. The exact format is not this file's to describe: headings, section order, the formula and the edge-case rules are fixed by `${CLAUDE_PLUGIN_ROOT}/samples/sample_statistics.md`, whose numbers are illustrative and never copied.
|
|
4
|
+
|
|
5
|
+
It is written by the statistics writer and read by the operator. Nothing in the flow keys off it: it carries no readiness list and no per-finding files, no committer flips a checkbox for it, and the caller that dispatches the writer reports the headline rate from the writer's own return block without opening the file.
|
|
6
|
+
|
|
7
|
+
A branch is written twice over its life. The first write lands at the end of the delivery run, before any hands-on review exists — no issues are counted yet, so the rate is `100%` — and the second lands once a fix round's items are all `[x]`, recounted from scratch and **overwriting** the first rather than sitting beside it. `status:` is what tells the two apart, `pre-user-review` or `post-user-review`, so there is one report per branch and never a round-suffixed series. The directory is committed with the branch, since no ignore rule reaches it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is reading this directory as a scoreboard. Each report is a record of one branch, and the tree deliberately holds no cross-branch roll-up to rank them against each other — so an absent file means that branch's run never reached its final phase, not that it scored nothing.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# business_parity_branch_review_point_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>_parity_review/item_<N>/` folder per reference-parity fix item, holding the findings from re-checking that item's fix. `<N>` is the item's position in the parity review's ordered fix list — the first entry is `item_1`, the second `item_2`. Inside the folder sits one `review_<iteration>.md` per re-check, numbered from `review_0.md` and incremented on each failed one; a later number is added beside the earlier files, never written over them.
|
|
4
|
+
|
|
5
|
+
A file here is written by the layer reviewer the fix item was routed to, and read by the orchestrator to decide whether the item closes — and, when it did not, by the layer implementer dispatched to fix it, which is handed the folder's most recent file. The parity reviewer itself never writes here: its own findings against the branch go to `<state_dir>/business_parity_branch_reviews/`.
|
|
6
|
+
|
|
7
|
+
A folder appears the first time one of its item's re-checks has something to report; a re-check with nothing to report writes no file at all, so an item that converged immediately leaves no folder behind. Nothing supersedes an earlier file — the numbered files accumulate into that item's convergence history — and the whole directory is committed with the branch it judged, since no ignore rule reaches it. It exists only while `phases.parity` is on in `harness.config.json`: with no reference implementation to compare against there is no parity fix loop to write into it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is reading this directory as `<state_dir>/business_parity_branch_reviews/`. That one holds the **original** parity findings against the finished branch; this one holds the **re-check of a fix** for one of them. The two have different writers, different readers and different lifetimes, which is why the tree keeps them apart instead of nesting the re-checks under the review that prompted them.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# business_parity_branch_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>_parity_review.md` index plus one `<branch>_parity_review/finding_<N>.md` per finding, holding the reference-implementation-parity review of a finished branch: written by the parity reviewer and read by the orchestrator. The index is thin — a `## Context` paragraph, the `## Phase 2 Readiness — Ordered Fix List`, and one `### N. Title` pointer per finding under its severity section — while each finding's file carries the full deviation, the reference source it was compared against, and the exact rename, payload field or predicate change that fixes it. The split, the byte-identical readiness heading and the format are the code review's: `${CLAUDE_PLUGIN_ROOT}/samples/sample_code_review.md` for the index and `${CLAUDE_PLUGIN_ROOT}/samples/sample_code_review/finding_1.md` for a detail file.
|
|
4
|
+
|
|
5
|
+
This is the parity reviewer's other mode: instead of a plan's citations it judges the branch diff against the reference implementation itself, stating every finding in the vocabulary `parity.referenceName` gives that implementation. The orchestrator reads the readiness list to drive its fix loop, routing each item off its `_(layer: …)_` tag without opening the finding; the layer implementer is handed one `finding_<N>.md`; the layer reviewer re-checks the fix into `<state_dir>/business_parity_branch_review_point_reviews/`; and the committer flips the item's `[ ]` to `[x]` as the fix lands. This phase runs ahead of the architecture review of the same branch, so a parity deviation is caught before structural effort is spent on the code carrying it.
|
|
6
|
+
|
|
7
|
+
The directory exists only while `phases.parity` is `true` in `harness.config.json`: a project with no reference implementation to compare against will not have it, and turning the phase on and re-running `init` creates it — an absent directory here is a configuration answer, not a missing artifact. While the phase is on, the index appears when the review has something to report and is committed once, index and per-finding folder together, before the first of its fixes is implemented. Nothing supersedes it; a later round is written beside it under the next suffix, and the whole directory is committed with the branch, since no ignore rule reaches it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is treating the index filename as cosmetic. Two end-of-branch split indices feed the same fix loop through the same readiness heading, and `_parity_review` versus `_arch_review` is the whole of what tells them apart — while the per-finding folder is that filename without its extension, and a re-review round appends its suffix to that same base. Renaming the index to a branch's own convention therefore orphans its pointers and puts two unrelated fix loops on one name.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# business_parity_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>/review_{iteration}.md` per planning iteration, holding the plan-gate review of a task plan against the reference implementation: written by the parity reviewer and read by the plan writer on the next iteration. One `<branch>` subdirectory holds one branch's set. `{iteration}` is this directory's own index, and a new file takes the **next free index in that branch's directory** rather than restarting at zero; because each plan gate resolves its index in its own directory, this directory holds its own consecutive series and a gap is a file that went missing.
|
|
4
|
+
|
|
5
|
+
The reviewer reads the story index's `## Context` plus every per-task file and judges one lens only — business-logic parity: external call names and payload shapes, stored-document shapes and queries and serialized field names, threshold constants, gating predicates, side-effect ordering, strict-versus-loose comparisons. It states each finding in the vocabulary `parity.referenceName` gives that implementation and cites the reference source it compared against. The reader is the plan writer, which applies the Must Fix items to the story index or to the per-task files the findings name; the reviewer creates the branch directory itself and only when it has findings, so a plan that was clean on parity leaves this directory empty.
|
|
6
|
+
|
|
7
|
+
The directory exists only while `phases.parity` is `true` in `harness.config.json`: a project with no reference implementation to compare against has no parity gate and will not have it, and turning the phase on and re-running `init` creates it — an absent directory here is a configuration answer, not a missing artifact. While the phase is on, files appear one per failed gate and stop at convergence or at the fifth iteration, where the gate escalates with the latest path. Nothing supersedes an earlier file, and the whole directory is committed with the branch, since no ignore rule reaches it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is expecting a plan to pass here on plausible prose. This gate is graded on **citations**: every constant, call name, wire field and boundary comparison the plan describes has to trace to a source line in the reference implementation, and a behaviour the reference has that no per-task file covers is a Must Fix of the same grade as a wrong field. A plan that describes the right behaviour without saying where it came from is the ordinary way a round is lost here.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# business_parity_user_review_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>/review_{iteration}.md` per iteration, holding the parity reviewer's review of a hands-on-review fix plan rather than of a task plan: read by the fix-plan writer on the next iteration. One `<branch>` subdirectory holds one branch's set. The fix-plan index plays the role the story index plays at the task-plan gate, and its per-finding folder the role of the per-task files. `{iteration}` is this gate's own counter, separate from the architecture gate that runs after it, and it is a **plain counter, not a free-index resolution**: `${CLAUDE_PLUGIN_ROOT}/instructions/user_review_fix_plan_writing_instructions_core.md` → `## Flow` step 7 opens with `Set parity_iteration = 0` and increments it once per FAIL, and the reviewer writes to that index without listing the directory — so an index an earlier round already used is **overwritten**, not skipped.
|
|
4
|
+
|
|
5
|
+
The writer is the parity reviewer, the same agent and the same lens that grades a task plan into `<state_dir>/business_parity_reviews/`; it is told which of the two it is holding by the paths it is passed. It asks of each drafted fix what it asks of a planned task: does it cite a real source in the reference implementation for every constant, call name and field, do the described payloads match field by field, do the described predicates match the reference's boundary comparisons. The reader is the fix-plan writer, re-dispatched in revision mode with the newest path and expected to apply its Must Fix items to the fix-plan index or to the per-finding files they name.
|
|
6
|
+
|
|
7
|
+
The directory exists only while `phases.parity` is `true` in `harness.config.json`: a project with no reference implementation to compare against will not have it, and turning the phase on and re-running `init` creates it — an absent directory here is a configuration answer, not a missing artifact. While the phase is on, a file appears only for a failed gate, written before the fix plan is approved and before any of its fixes is implemented; a clean review touches no disk. Files stop at convergence or at the fifth iteration, where the flow escalates with the latest path. No file supersedes an earlier one in meaning, though a re-entered gate can overwrite one on disk (below), and the whole directory is committed with the branch, since no ignore rule reaches it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is reading the first PASS as this gate's last word. Parity is graded ahead of architecture, and an architecture failure sends the revised plan back through **both** gates — so this gate re-runs rounds after it had already passed once, against a plan the writer changed in between. Because step 7 resets `parity_iteration` to `0` on that re-entry, the new round lands **on top of `review_0.md`** and the round that first caught the deviation is gone, with no gap or renumbering to mark the loss. Read the numbered set as the latest pass through the loop, not as a full history, and copy a round aside before re-entering the gate if you need it.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# clarification_digests/
|
|
2
|
+
|
|
3
|
+
One `<state_dir>/clarification_digests/<branch>.md` per branch, holding one block per clarification the runs on that branch parked on and had answered: what was asked, the operator's ruling in the words it was given, and whatever that ruling settles beyond the question that prompted it. One file per **branch**, appended to and never rewritten — each block is headed by the `question_<n>` index it was derived from, and a later round or a resumed session adds only the blocks whose key is not in it already.
|
|
4
|
+
|
|
5
|
+
Each block is written by the run's own orchestrator at the end of the run, from the question and answer files in `<state_dir>/clarifications/<branch>/` and its `answered/` archive. Those files are machine-local and go with the working copy — a working copy removed once its branch has merged takes every question and every answer with it, and **this file is the copy that survives**. Its readers are whoever comes to the branch afterwards: a reviewer asking why a decision was taken the way it was, and the human folding a ruling that outlives the branch into the rules the flow reads.
|
|
6
|
+
|
|
7
|
+
The file is committed and pushed by the run itself, best-effort and never a gate: a run that parked on nothing writes nothing at all, and a write, commit or push that fails is logged while the run finishes unchanged. So an **absent** file means one of two things, and they are worth telling apart: the branch's runs reached the end and parked on nothing, or a run parked and never got that far — a question left unanswered, a halt, an abandoned pause. The digest is written at the end of a run, so a park that is never answered is never digested and its question goes with the working copy. Answering a question you intend to abandon is what preserves it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is editing a recorded ruling. A block is the operator's words as they were given, and a later disagreement is a new question — parked and answered like the first. Rewriting the record instead leaves the branch describing a decision nobody took, and the run that acted on the original ruling looking like it acted against it.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# clarifications/
|
|
2
|
+
|
|
3
|
+
The park-and-ask channel: one question per file at `<state_dir>/clarifications/<branch>/question_<n>.md`, with the answer it waits for beside it as `answer_<n>.md`. Pairing is by the index `<n>`, which is 1-based and unique across the branch's directory **including its `answered/` archive** — a second question takes the next index above the highest `question_<n>.md` in either location, never merely the next free name at the top level and never overwriting the first, so an archived pair does not free its index for reuse and `answer_3.md` answers `question_3.md` and nothing else. One `<branch>` subdirectory holds one branch's exchange and is created with that branch's first question.
|
|
4
|
+
|
|
5
|
+
A question is written by the agent that reached a decision it must not take alone: writing the file and ending the session is what parks the run. The answer is written by the harness's answer command on the operator's behalf, verbatim and beside the question. The run daemon reads the pair and re-launches the same engine in the same working copy for the lowest answered index, and the re-entering run is what consumes the answer — the daemon never answers a question itself.
|
|
6
|
+
|
|
7
|
+
The pair stays at the top level of the branch's directory while that resumed session runs, and moves into `<branch>/answered/` only once it has exited, so the answer is still where the resumed run looks for it and an already-consumed one is never read a second time. Everything the runs write here is machine-local and gitignored; this README is the only committed file in **this** directory. What survives the working copy is the record: at the end of each run the branch's resolved questions are digested into `<state_dir>/clarification_digests/<branch>.md`, one committed block per question, keyed by that same index — what was asked and the operator's ruling in the words it was given. Reusing an index would make two parks indistinguishable there, which is the other reason it is never reused. That sibling directory's own README states the format. The harness's watcher documentation, which ships with it, states the same protocol from the daemon's side.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is tidying up by hand. Deleting a question, or archiving it into `answered/` before the run that parked on it has been resumed, does not unpark that run — an unanswered question is the sole signal that classifies a run as parked, so removing it strands the run instead, waiting on an answer nothing will now deliver.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# code_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>_code_review.md` index plus one `<branch>_code_review/finding_<K>.md` per finding, each carrying its own severity and fix state: written by the branch reviewer, and worked through by the implementers, the fix reviewers and the committer. The index is thin — a Context paragraph, the `## Phase 2 Readiness — Ordered Fix List`, and one `### K. Title` pointer per finding under its severity section — while the full problem description and the concrete fix live in the finding's own file, self-contained enough to implement from alone. **The index filename is load-bearing**: a re-review of the same branch becomes `<branch>_code_review_2.md`, then `_3.md`, with its detail files under `<branch>_code_review_2/`, so the round-suffix logic depends on the base name being exactly this shape, and the folder always mirrors the suffix its index carries. The canonical format both halves are written to is `${CLAUDE_PLUGIN_ROOT}/samples/sample_code_review.md` for the index and `${CLAUDE_PLUGIN_ROOT}/samples/sample_code_review/finding_1.md` for a detail file.
|
|
4
|
+
|
|
5
|
+
The branch reviewer writes the set, and on a failed meta-review adjusts the files that meta-review names rather than regenerating it. Everything after that reads it: the review-plan reviewer meta-reviews the index against its finding files before any fix is implemented; the fix loop walks the readiness list and hands one implementer one `finding_<K>.md`; the layer reviewer re-checks that fix; the committer flips the item's readiness entry as the fix lands; and the skeptic reviewer reads the whole set as the set its own findings must not duplicate.
|
|
6
|
+
|
|
7
|
+
The index appears when the end-of-branch review runs. It is committed once — index, per-finding folder and the meta-review findings written about them, in a single commit — before the first fix is implemented, so the review that prompted a branch's fixes is on record even if the fix loop is interrupted. Nothing supersedes it afterwards: a re-review is written beside it under the next round suffix, and the accumulated rounds are the branch's review history. The whole directory is committed with the branch, since no ignore rule reaches it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is reading a checkbox inside a per-finding file as that fix's status. The index's readiness list is the fix loop's single source of truth — the list the loop walks, and the only place the committer flips anything — while the `- [ ]` sub-steps in a finding body are informational markers the implementer ticks off to track its own progress, and nothing ever flips them.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# dispatch_additions/
|
|
2
|
+
|
|
3
|
+
One `<state_dir>/dispatch_additions/<branch>.md` per branch, holding what the orchestrating agent added to a sub-agent's dispatch prompt beyond the block that agent's governing instruction defines. One file per **branch**, not per run or per phase: every phase in which an addition was made appends its blocks to the same file, and a resumed session skips any block whose `## <dispatch key>` heading is already there rather than rewriting or renumbering one.
|
|
4
|
+
|
|
5
|
+
Each block is written by the run's own orchestrator — the only participant that composes those prompts, and the added text otherwise exists nowhere but the session transcript — once per phase in which at least one addition was made, before the flow leaves that phase. **No agent reads it back**, the one exception being the session checking its own file so it does not re-append a block it already wrote. Its reader is a human auditing a finding, a verdict or a round count after the fact: whether the agent that produced it was told anything beyond its instruction, and if so what. That question cannot be answered from the review artifacts alone, which is the whole reason this directory exists.
|
|
6
|
+
|
|
7
|
+
The file is committed by the run itself, on the run's own branch, as one commit per phase that wrote to it. The step is best-effort and never a gate: a phase with no additions writes nothing — no empty block and no empty commit — and a write, commit or push that fails is logged while the run finishes unchanged. So an **absent** file, or a run whose hand-off says none were recorded, means nothing was added, not that the step failed.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is reading the per-branch split as an accident and consolidating it into one shared target. Every parallel run appending to one file is among the most reliable merge-conflict generators there is, and a record every run had to touch would re-serialise exactly the parallel runs this tree exists to allow. Nothing collates these files; each is read on its own, against the branch it belongs to.
|
|
10
|
+
|
|
11
|
+
One block, illustrative rather than real, in the shape every entry takes — as the file for the branch `feat/saved_filters_panel` would carry it. The format is fixed by `${CLAUDE_PLUGIN_ROOT}/instructions/dispatch_discipline_instructions.md` → `## The entry format`, which also settles what may be added at all; the `verbatim:` line is quoted, never paraphrased, because the wording is what decides which side of that boundary an addition fell on:
|
|
12
|
+
|
|
13
|
+
```markdown
|
|
14
|
+
## [C2 · item 4 · review · iter 1] → skeptic-reviewer (#37)
|
|
15
|
+
- **added:** `context_notes:`
|
|
16
|
+
- **verbatim:** "the type-check gate ran at sha 4f1c9ab over the whole diff and reported no error"
|
|
17
|
+
- **why the agent could not derive it:** the gate's output is not committed anywhere in its read set
|
|
18
|
+
- **reported back by the agent:** (none)
|
|
19
|
+
```
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# docs_catalog/
|
|
2
|
+
|
|
3
|
+
The documentation phase's own bookkeeping — the per-branch checklist of documents to write, the gap and needs-review registers, and one review file per document round: written and read by that phase, and never the documents themselves, which live in the configured documentation root. The checklist is `<branch>_docs.md`, an ordered list whose `[ ]`/`[x]` box is the phase's **only** iteration state; each entry names one document by slug, type, title, output path and a blob of research starting-point hints. The registers are `needs_review.md`, appended when a document exhausts its fix rounds, and — only while `phases.parity` is `true` in `harness.config.json` — `parity_gaps.md`, appended when a document turns up a gap against the reference implementation; with that phase off the writer reports no gaps and the file is never created, so its absence is a configuration answer rather than a missing register. A review file is `reviews/<slug>/review_<i>.md`, and `<i>` is the next free index **in that slug's directory** — a slug already carrying `review_0.md` and `review_1.md` starts at `2`.
|
|
4
|
+
|
|
5
|
+
Everything here is written and read inside the phase: the orchestrator parses the checklist, flips its boxes and appends the registers; the document reviewer writes its own findings file; the document writer reads that findings file to revise. The documents this bookkeeping is about are **never** in this directory — they are written into the repository-relative directory `docs.root` names in `harness.config.json`, and what stays here is only the record of which ones were written, checked or deferred.
|
|
6
|
+
|
|
7
|
+
On the unattended path the checklist is committed **before** the run launches, because it is a precondition of the phase's first step; everything else accumulates entry by entry, each document, its checkbox flip and any register or findings change landing in one commit. Nothing supersedes an earlier file — a later round is a higher `<i>` beside it — and the whole directory is committed with the branch, since no ignore rule reaches it. It exists only while `phases.docs` is `true` in `harness.config.json`.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is hand-editing a checkbox mid-run. The committed boxes are the phase's entire resume state: it re-reads the checklist on every re-entry and continues at the first `[ ]`, with no other record of where it got to. Ticking an entry skips a document that was never written; clearing one re-runs work that already landed. An entry is atomic — written, reviewed and committed together — so a crash needs no repair by hand, and there is never a good reason to make one.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# flow_progress/
|
|
2
|
+
|
|
3
|
+
One append-only `<branch>_progress.md` per branch, carrying a line for each phase boundary a run reaches: written by the run's driving fork as it advances and read back when a paused or interrupted run resumes, which is what makes resuming possible at all. Its entries are **phase-level and fixed per engine** — a delivery run's planning, implementation and review boundaries, a hands-on-review round's five fix entries — so the ledger answers *which phase* a run is in and nothing finer. Which item *within* a phase is the detail index's answer: the story index's readiness list, a review's findings index, the interactive-test index.
|
|
4
|
+
|
|
5
|
+
The writer is always the run's **driving fork**, never a sub-agent: the fork that plans creates the ledger and flips the planning entries, the orchestrator flips the implementation and review entries, and on a hands-on-review round the fix-plan and fixes forks flip theirs. Each flips only its own. An entry goes `[x]` only once its phase reached one of the recorded outcomes that entry names — its underlying detail checklist fully `[x]` **and** that phase's artifact committed where the phase produces one, or the clean-pass outcome where it correctly produces none — so the ledger is never ahead of what is on the branch. The reader is the fork that re-enters the run: it resumes at the **first `[ ]` entry** and skips every entry already resolved — `[x]` for done, `[-]` for a phase excluded before the run began, by the run's authored run mode or by a `phases.*` flag that is `false` in `harness.config.json` — with no reviewer re-dispatch, no review regeneration and no re-commit.
|
|
6
|
+
|
|
7
|
+
The file is created once per branch by the fork that starts the run, idempotently: a resume reads the existing ledger rather than resetting it. It is **committed and pushed** on creation and on every flip — like the rest of this tree, and for a reason of its own, since a resumed run may execute in a different working copy from the one that wrote it and would otherwise resume from nothing. Nothing supersedes it within a run. A fresh hands-on-review round re-seeds it with that round's own checklist, which is the one sanctioned rewrite and is what stops the previous round's all-`[x]` ledger from marking the new round already done.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is tidying it. It is append-only and is never rewritten, compacted or cleared by hand — a compacted ledger is a resume that redoes work the run already did, and the loss is silent because redone work looks exactly like new work.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# improvement_observations/
|
|
2
|
+
|
|
3
|
+
One `<state_dir>/improvement_observations/<branch>.md` per branch, holding the harness and workflow problems the runs on that branch actually hit. One file per **branch**, not per run: a later hands-on-review fix round on the same branch appends beneath its own `#` round label rather than overwriting, and a resumed session appends only what that session itself observed, so a branch that went through a fix cycle carries several blocks in one file. The first, unlabelled block is always the delivery run's.
|
|
4
|
+
|
|
5
|
+
Each block is written by the run's own orchestrator — the only participant that sees a run end to end, which is why the observations would otherwise die with the session — once, near the end of the run, after that run's statistics land and before it reports itself done. **No agent reads it back**, the one exception being a resumed session checking its own file so it does not re-append a block it already wrote. A human folds each intake into the triage ledger at the root of this tree, `<state_dir>/improvement_suggestions.md`, on the default branch; that fold is this directory's only reader, and nothing in the flow keys off the file.
|
|
6
|
+
|
|
7
|
+
The file is committed and pushed by the run itself, on the run's own branch, near the end of the run's closing phase, with only that run's clarification digest (`<state_dir>/clarification_digests/<branch>.md`) written after it. The step is best-effort and never a gate: a run that observed nothing writes nothing at all, and a write, commit or push that fails is logged while the run finishes unchanged. So an **absent** file means nothing was observed, not that the step failed.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is reading the per-branch split as an accident and consolidating it into one shared file. Every parallel run appending to one target at the end of its work is among the most reliable merge-conflict generators there is, and it would re-serialise exactly the parallel runs this tree exists to allow. The intakes are meant to come together at triage — shared, human, and on the default branch — and nowhere earlier.
|
|
10
|
+
|
|
11
|
+
One block, illustrative rather than real, in the shape every entry takes — as the intake file for the branch `feat/recent_searches_panel` would carry it. The format and the seven category names are fixed by `${CLAUDE_PLUGIN_ROOT}/instructions/improvement_observations_instructions.md` → `## The entry format`; what belongs here is a harness or workflow problem the run hit, never a defect in the work the run was shipping:
|
|
12
|
+
|
|
13
|
+
```markdown
|
|
14
|
+
## The type-check wrapper is gated in a sibling working copy, so two tasks shipped unverified
|
|
15
|
+
- **category:** tooling-gap
|
|
16
|
+
- **evidence:** the profile entry for the wrapper matches the main checkout's path only, so both dispatches that ran it were denied; both tasks were committed on a review reading alone.
|
|
17
|
+
- **cost this run:** two of five tasks carry no type-check result.
|
|
18
|
+
- **hypothesis:** (guess) the entry predates runs executing outside the main checkout.
|
|
19
|
+
```
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
# Improvement suggestions
|
|
2
|
+
|
|
3
|
+
**What this is.** A human-triaged checklist of the harness and workflow problems that runs actually hit, kept so a recurring defect is diagnosed from how often it recurs across runs rather than rediscovered from scratch each time. **No agent reads this file.** Unlike the lessons ledger beside it, whose rules bind planning and review work, these are maintenance items with no value to an agent implementing a change, so this ledger is on no agent's reading list.
|
|
4
|
+
|
|
5
|
+
**How entries arrive.** Each run writes exactly one intake file at `improvement_observations/<branch>.md`, once, at the end of the run; a human folds those intakes into this file. One intake file per branch is the design: a single shared append target across parallel runs is a reliable generator of merge conflicts, which is the cost this split avoids.
|
|
6
|
+
|
|
7
|
+
**This ledger is appended to and never rewritten.** A closed entry stays in place with its outcome appended rather than being deleted, and an entry is cross-referenced by quoting its problem statement rather than by position, because entries are inserted as intake is folded in. The file accumulates over the whole life of the project and is the only copy of that history.
|
|
8
|
+
|
|
9
|
+
**Entry shape.** One line per item plus its attribution line, in one of three states — `[ ]` open, `[x]` done, `[~]` won't do:
|
|
10
|
+
|
|
11
|
+
```markdown
|
|
12
|
+
- [ ] **[the problem, in one line]** — [the suggested direction, optional]
|
|
13
|
+
_(observed by: [branch], [branch] · category: [category])_
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
**Categories.** A folded entry keeps the category its intake assigned it. The vocabulary is the intake module's rather than this file's, and is one of `tooling-gap`, `silent-failure`, `flow-efficiency`, `optimization`, `agent-contract`, `shared-state`, `convention`. What each one covers is defined in `${CLAUDE_PLUGIN_ROOT}/instructions/improvement_observations_instructions.md` → `## The entry format`, which is the definition of record: re-glossing the names here would give them a second owner to drift from.
|
|
17
|
+
|
|
18
|
+
---
|
|
19
|
+
|
|
20
|
+
## Resuming a paused run
|
|
21
|
+
|
|
22
|
+
_The entry below is a worked example, not an observation from this project — delete it, or replace it with the project's own first entry._
|
|
23
|
+
|
|
24
|
+
- [ ] **A phase's artifacts are regenerated from scratch after a resume, because the progress ledger is the only resume state** — commit each phase's artifacts as that phase finishes, so a resumed run finds them instead of redoing them
|
|
25
|
+
_(observed by: feat/search_result_pagination · category: flow-efficiency)_
|
|
26
|
+
|
|
27
|
+
_[Add headings of your own as themes emerge — one per recurring theme — and file each entry under the heading it belongs to.]_
|
|
28
|
+
|
|
29
|
+
_Written by `autonomous-sdlc-harness init`, and yours from there on: a re-run never touches a ledger that already exists._
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
# Lessons ledger
|
|
2
|
+
|
|
3
|
+
**What this is.** One-line rules distilled from defects that got past every automated gate on a branch and were caught only by a human's hands-on review of the finished work. An agent that plans work or grades it reads this file **before it starts** — before a plan is drafted, before a review is graded — and an entry here carries the same force as a rule in one of the project's own conventions documents, so a plan that re-plans a lesson fails its review and the agent writing it checks its own draft against the ledger before handing it over. The other end of the loop is the fix step: the agent that turns a hands-on review into a fix plan is the ledger's only writer, and it appends there once per review round. The file lives in the run-artifact tree rather than beside the agent configuration so that an unattended headless run can append to it, and it is resolved against the checkout the run is executing in — never against a fixed path, or a run in a second working copy appends its lesson to the wrong branch.
|
|
4
|
+
|
|
5
|
+
**This ledger is appended to and never rewritten.** A new lesson is added as one line under the section it belongs to; a rule taught again by a second branch gains that branch in its attribution instead of being duplicated, so grep before adding; and a rule later codified into a conventions document keeps its line here with a pointer to that document appended. Nothing is edited away, reordered or truncated: the file accumulates over the whole life of the project, it is the only copy of that history, and an overwrite destroys it.
|
|
6
|
+
|
|
7
|
+
**Entry shape.** One line per lesson — the rule in the imperative, then the branches that taught it:
|
|
8
|
+
|
|
9
|
+
```markdown
|
|
10
|
+
- **[the rule, in one line]** _(taught by: [branch], [branch])_ — codified in [conventions document] §[section]
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
---
|
|
14
|
+
|
|
15
|
+
## Layer ownership
|
|
16
|
+
|
|
17
|
+
_The entry below is a worked example, not a lesson from this project — delete it, or replace it with the project's own first lesson._
|
|
18
|
+
|
|
19
|
+
- **A unit in the outermost layer never reaches the innermost one directly: it calls through the layer between, so the logic has one place to live and one place to be tested.** _(taught by: feat/recent_searches_panel, feat/search_result_pagination)_ — codified in the layering conventions document §Call flow
|
|
20
|
+
|
|
21
|
+
_[Add headings of your own as themes emerge — one per recurring theme — and file each new lesson under the heading it belongs to.]_
|
|
22
|
+
|
|
23
|
+
_Written by `autonomous-sdlc-harness init`, and yours from there on: a re-run never touches a ledger that already exists._
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# qa_review_point_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>_qa_review[_<round>]/item_<N>/` folder per interactive-test fix item, holding the findings from re-checking that fix: written by the item's layer reviewer and read by the orchestrator. An item is one **finding** from that round's merged QA-review index, not one failed test — a single failing test that produced three findings gets three folders — and `<N>` is the item's position in that index's ordered fix list. Inside sits one `review_<iteration>.md` per re-check, numbered from `review_0.md` and incremented on each failed one, a later number added beside the earlier files rather than over them. The root mirrors the round suffix of the QA-review folder it re-checks, so resolve it each round instead of writing it literally: a hardcoded unsuffixed root silently drops a re-test round's re-checks into round 1's folder.
|
|
4
|
+
|
|
5
|
+
The layer reviewer the item's `_(layer: …)_` tag routed to writes the file — the tag being the one the tester wrote on the finding and the merge carried through verbatim. The orchestrator reads the verdict to decide whether the item closes, and on a failure the layer implementer is dispatched again and handed the folder's most recent file. The tester never writes here: its own findings against the running application are in `<state_dir>/qa_reviews/`, and this directory holds only the re-checks of the fixes for them.
|
|
6
|
+
|
|
7
|
+
A file is written only when a re-check has something to report, and the reviewer creates the folder itself at that moment, so an absent folder covers an item that passed first time as readily as one a mode with the per-unit review step off never reviewed. Nothing supersedes an earlier file — the numbered set is that item's convergence history — and the whole directory is committed with the branch, since no ignore rule reaches it. It exists only while `phases.qa` is `true` in `harness.config.json`: with no interactive tests there is no QA fix loop to write into it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is reading `item_<N>` as a test number. These items are numbered off the merged QA-review index's readiness list, which is ordered for fixing rather than for running, so `item_2/` is generally not Test 2. The way back to the test is the finding file the readiness entry points at — `finding_t<T>_<N>.md`, whose `<T>` is the test id that produced it.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# qa_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>_qa_review[_<round>]/qa_review_<iteration>.md` per interactive-test round, recording each test as **passed, failed or blocked** together with the behaviour observed: written by the tester agent and read by the orchestrator to drive the fix loop. The round suffix lives on the **folder** — round 1 unsuffixed, `<branch>_qa_review_2/` next — and never on a sibling file; `<iteration>` numbers the index inside that folder. How the round is chosen differs by flow: the delivery flow takes the unsuffixed folder for its first round and increments from its own in-session counter on each re-test, while the flows started by hand — the hands-on-review fix flow's QA pass and a standalone test session — list this directory and resolve the round to the **next free** folder, precisely so they never collide with one the delivery flow already produced. The folder also carries the tester's per-dispatch artefacts, namespaced by test id so parallel dispatches cannot overwrite each other: `qa_review_<iteration>_t<T>.md` per test, with its findings in `finding_t<T>_<N>.md`.
|
|
4
|
+
|
|
5
|
+
The tester agent is dispatched **once per test case**, so each test runs in a fresh agent session that establishes its own starting state per the tester's `## Preconditions` — the browser is a session-level service shared across dispatches, so the tester clears client storage itself rather than inheriting a fresh context — which is what buys the clean attribution this directory's artefacts depend on, and each dispatch writes only its own namespaced files. The **orchestrator** owns the round-level index and assembles it from those per-dispatch ones — the tester never writes the merged file. That merged index carries a single `## Phase 2 Readiness — Ordered Fix List` enumerating every failing test's findings exactly once, and that one list, not the per-dispatch indices, is what the fix loop iterates over.
|
|
6
|
+
|
|
7
|
+
An index appears only on a round with something to report: a clean round writes no file at all, and the tester creates the folder itself when it first has findings. On a failing round the index and the folder's per-dispatch artefacts are committed together, before the first QA fix is implemented, so the round that prompted the fixes is on record even if the loop is interrupted. Nothing supersedes it — a re-test round writes into its own folder — and the whole directory is committed with the branch, since no ignore rule reaches it. It exists only while `phases.qa` is `true` in `harness.config.json`; a project that runs no interactive tests will not have it, and that absence is a configuration answer rather than a missing artifact.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is collapsing **blocked** into failed. A test the environment prevented driving — no credentials for an auth-gated surface, an account another session holds, a backend failure that cannot be induced without request mocking — is recorded `blocked`, and that is a first-class outcome, not a weak result. Read as a failure it sends the fix loop after a defect that is not there; rolled into a pass it hides a test that never ran at all.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# review_plan_point_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>_review/item_<N>/` folder per code-review fix item, holding the findings from re-checking that item's fix: written by the item's layer reviewer and read by the orchestrator to decide whether the item closes. `<N>` is the item's position in the code-review index's ordered fix list — the first entry is `item_1`, the second `item_2`. Inside the folder sits one `review_<iteration>.md` per re-check, numbered from `review_0.md` and incremented on each failed one; a later number is added beside the earlier files, never written over them.
|
|
4
|
+
|
|
5
|
+
The layer reviewer the item's `_(layer: …)_` tag routed to writes the file. The orchestrator reads its verdict to decide whether the item closes, and on a failure the layer implementer is dispatched again and handed the folder's most recent file as the thing to fix. The reviewer that raised the finding never writes here: its own findings sit in `<state_dir>/code_reviews/`, and this directory holds only the re-checks of the fixes for them.
|
|
6
|
+
|
|
7
|
+
A file is written only when a re-check has something to report, and the reviewer creates the folder itself at that moment — so an absent folder is not a statement about the fix. It equally covers an item that passed its first re-check, an item the loop never reached, and a mode running with the per-unit review step off, which writes nothing here at all and is not an omission; where an item's outcome is recorded is the readiness checkbox in the code-review index, which the committer flips as the fix lands. Nothing supersedes an earlier file — the numbered set is that item's convergence history — and the whole directory is committed with the branch, since no ignore rule reaches it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is assuming the two per-item roots could share one folder. This one is keyed on `<branch>_review/` and its adversarial counterpart on `<branch>_skeptic_review/`, and both restart numbering at `item_1`, so a branch that ran both fix loops would otherwise write two unrelated `item_1/` folders into the same place. The distinct index names are what keeps each loop's re-checks readable as one sequence.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# review_plan_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>/review_{iteration}.md` per iteration, holding the meta-review of the code-review fix plan — does it cover every finding, is each item scoped to one layer: written by the review-plan reviewer and read by the orchestrator before the fix loop starts. One `<branch>` subdirectory holds one branch's set. `{iteration}` is that meta-review loop's own counter — `0` on the first pass and incremented once per FAIL — and it is a **plain counter, not a free-index resolution**: `${CLAUDE_PLUGIN_ROOT}/instructions/plan_orchestration_instructions_core.md` → `### B.2 Meta-review the review plan` opens with `iteration = 0` and the reviewer writes to that index without listing the directory; at index 5 the loop stops and escalates rather than iterating further. Within one loop the series is consecutive and replaces nothing. Across **re-review rounds** it is not: this directory is keyed by branch alone, while the code-review index it grades carries the round in its filename (`<branch>_code_review_2.md`), so a second round's B.2 starts at `0` again and **overwrites** the first round's `review_0.md`.
|
|
4
|
+
|
|
5
|
+
The review-plan reviewer writes here. It is the same agent that meta-reviews the adversarial review into `<state_dir>/skeptic_review_plan_reviews/`, and it is told which of the two it is holding by the paths it is passed rather than by a flag — which is why the two families are kept in separate directories instead of one. Two consumers read what it writes: the orchestrator, which parses the verdict to decide whether to loop again or move on to committing the review, and the branch reviewer, which is handed the findings file and adjusts the index or the named `finding_<K>.md` accordingly.
|
|
6
|
+
|
|
7
|
+
A file appears only for a failed meta-review. A verdict of PASS writes nothing and creates no directory, so a code review that was clean on the first pass leaves this directory empty. Nothing supersedes an earlier file **within a round** — the numbered set records how that round's review converged, and a later round rewrites it from `0` (above) — and the branch's directory is committed in the same commit as the code-review index it graded, since no ignore rule reaches it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is expecting a verdict about the branch here. A file in this directory grades the review **document**: whether its findings are real and correctly graded, whether every index pointer resolves to a detail file, whether the readiness list carries the byte-identical heading the fix loop keys off. It is written before a single fix exists, so it says nothing about whether the code is now right — that question is answered per item, later, under `<state_dir>/review_plan_point_reviews/`.
|
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
# scratch/
|
|
2
|
+
|
|
3
|
+
A throwaway file an agent writes so that it can **execute** something: a language probe, before the code that would answer the question exists, or a mutation check that breaks an implementation on purpose to prove a new test fails. It is a file in the project's own language, and the interpreter follows from its extension — `probe.py` is run as Python, `probe.mjs` as Node. **Nothing depends on its name.** This is the one directory in the tree where a file's name carries no contract, and the only reason to choose it with any care is that two dispatches in flight at once must not land on the same one.
|
|
4
|
+
|
|
5
|
+
A file here is written by the dispatched implementer or reviewer that needs the answer. **Nothing reads it**, and the harness's scratch runner is the route the flow **provides** for executing it — the route to use, not the only one that exists. It lives for the length of one dispatch — the agent that wrote it is the last participant with a reason to care about it — so this directory's steady state is empty. Nothing rotates or archives it, which makes deleting it once the answer is in hand the whole of the housekeeping.
|
|
6
|
+
|
|
7
|
+
Its contents are **machine-local and gitignored**: the directory is ignored by its contents rather than as a directory, and this README survives by an explicit negation, which makes it the only committed file here.
|
|
8
|
+
|
|
9
|
+
Two boundaries an agent does not cross. They are rules for the agent rather than a containment claim: the runner's own `WHAT THIS DOES NOT CONTAIN` paragraph records that its fence is over **which file runs** and over nothing that file does. A probe file lives **here and nowhere else** — the runner refuses any path that does not resolve inside this directory, so putting one somewhere more convenient does not get it run, it gets it refused. And a mutation check is **reverted before the task's own verification runs**: the point of one is a suite that fails, and the committer has to see that suite clean, so the revert belongs with the check rather than at the end of the task. What a file run from here reaches follows from that fence: it executes with the session's own privileges and reaches whatever the permission surfaces withhold from a command string, and the `pre-push` git hook is the layer that still holds — driven rather than asserted in the harness's `docs/outer-loop-verification.md` §3, under *What that `allow` reaches*.
|
|
10
|
+
|
|
11
|
+
The mistake worth naming is treating a file here as evidence. It is gitignored by construction, so it never reaches a commit and no later reader can open it: a finding, a verdict or a plan sentence that rests on a probe records **what the probe showed** and the command that produced it, never a pointer to the file that showed it. That holds whether or not the file is still on disk when the task ends.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# skeptic_review_plan_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>/review_{iteration}.md` per iteration, holding the coverage check of the skeptic review's own fix plan: written by the review-plan reviewer and read by the orchestrator before that fix loop starts. One `<branch>` subdirectory holds one branch's set. `{iteration}` is that meta-review loop's own counter — `0` on the first pass and incremented once per FAIL — and it is a **plain counter, not a free-index resolution**: `${CLAUDE_PLUGIN_ROOT}/instructions/plan_orchestration_instructions_core.md` → `### C2.2 Meta-review the skeptic review` opens with `iteration = 0` and the reviewer writes to that index without listing the directory; at index 5 the loop stops and escalates. Within one loop the series is consecutive and replaces nothing. Across **re-review rounds** it is not: this directory is keyed by branch alone, while the skeptic index it grades carries the round in its filename (`<branch>_skeptic_review_2.md`), so a second round's C2.2 starts at `0` again and **overwrites** the first round's `review_0.md`.
|
|
4
|
+
|
|
5
|
+
The writer is the same review-plan reviewer that meta-reviews the code review into `<state_dir>/review_plan_reviews/`, holding the adversarial index instead; it is told which of the two it has by the paths it is passed, never by a flag. The orchestrator parses its verdict to decide whether to loop again or move on to committing the skeptic review, and the skeptic reviewer is handed the findings file to adjust the index or the per-finding files it names. The meta-review exists for a reason particular to this family: the adversarial review is the one most prone to false positives and its fixes are applied with no one in between, so a bogus finding is dropped here instead of being implemented.
|
|
6
|
+
|
|
7
|
+
A file appears only for a failed meta-review. A PASS writes nothing and creates no directory, so a skeptic review that was clean on the first pass leaves this directory empty — as does a branch whose skeptic pass found nothing to review in the first place. Nothing supersedes an earlier file **within a round** — a later round rewrites the series from `0` (above) — and the branch's directory is committed in the same commit as the skeptic-review index it graded, since no ignore rule reaches it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is expecting a verdict about the branch here. A file in this directory grades the review **document**: whether its findings are real and correctly graded, whether every index pointer resolves to a detail file, whether the readiness list carries the byte-identical heading the fix loop keys off. It is written before a single fix exists — whether the code is now right is answered per item, later, under `<state_dir>/skeptic_review_point_reviews/`.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# skeptic_review_point_reviews/
|
|
2
|
+
|
|
3
|
+
One `<branch>_skeptic_review/item_<N>/` folder per skeptic fix item, holding the findings from re-checking that item's fix: written by the item's layer reviewer and read by the orchestrator. `<N>` is the item's position in the skeptic-review index's ordered fix list — the first entry is `item_1`, the second `item_2`. Inside the folder sits one `review_<iteration>.md` per re-check, numbered from `review_0.md` and incremented on each failed one; a later number is added beside the earlier files, never written over them.
|
|
4
|
+
|
|
5
|
+
The layer reviewer the item's `_(layer: …)_` tag routed to writes the file. The orchestrator reads its verdict to decide whether the item closes, and on a failure the layer implementer is dispatched again and handed the folder's most recent file as the thing to fix. The skeptic reviewer never writes here: its own findings sit in `<state_dir>/skeptic_reviews/`, and this directory holds only the re-checks of the fixes for them.
|
|
6
|
+
|
|
7
|
+
A file is written only when a re-check has something to report, and the reviewer creates the folder itself at that moment — so an absent folder is not a statement about the fix. It equally covers an item that passed its first re-check, an item the loop never reached, and a mode running with the per-unit review step off; where an item's outcome is recorded is the readiness checkbox in the skeptic-review index, which the committer flips as the fix lands. A whole branch with nothing here is the common case, because a skeptic pass that found nothing writes no index and so starts no fix loop. Nothing supersedes an earlier file, and the whole directory is committed with the branch, since no ignore rule reaches it.
|
|
8
|
+
|
|
9
|
+
The mistake worth naming is assuming the two per-item roots could share one folder. This one is keyed on `<branch>_skeptic_review/` and the code review's counterpart on `<branch>_review/`, and both restart numbering at `item_1`, so a branch that ran both fix loops would otherwise write two unrelated `item_1/` folders into the same place. The distinct index names are what keeps each loop's re-checks readable as one sequence.
|