agent-bios 0.9.8 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. package/DEPENDENCIES.md +19 -19
  2. package/README.md +43 -12
  3. package/claude/CLAUDE.md +5 -41
  4. package/claude/guides/claude-prompting.md +1 -1
  5. package/claude/guides/cli-multi-model-workflow.md +22 -4
  6. package/claude/guides/coding-staged-workflow.md +49 -15
  7. package/claude/guides/concept-economy.md +187 -0
  8. package/claude/guides/documentation-hygiene.md +112 -0
  9. package/claude/guides/gpt-prompting.md +1 -1
  10. package/claude/guides/learning-flow.md +5 -5
  11. package/claude/guides/llm-capability-boundary.md +6 -1
  12. package/claude/guides/review-request.md +9 -7
  13. package/claude/guides/session-distill-workflow.md +19 -9
  14. package/claude/guides/tooling-gotchas.md +26 -0
  15. package/claude/guides/verification-discipline.md +166 -0
  16. package/claude/hooks/tooling-gotchas-hook.py +329 -12
  17. package/codex/AGENTS.md +5 -41
  18. package/codex/guides/claude-prompting.md +1 -1
  19. package/codex/guides/cli-multi-model-workflow.md +22 -4
  20. package/codex/guides/coding-staged-workflow.md +49 -15
  21. package/codex/guides/concept-economy.md +187 -0
  22. package/codex/guides/documentation-hygiene.md +112 -0
  23. package/codex/guides/gpt-prompting.md +1 -1
  24. package/codex/guides/learning-flow.md +5 -5
  25. package/codex/guides/llm-capability-boundary.md +6 -1
  26. package/codex/guides/review-request.md +9 -7
  27. package/codex/guides/session-distill-workflow.md +19 -9
  28. package/codex/guides/tooling-gotchas.md +26 -0
  29. package/codex/guides/verification-discipline.md +166 -0
  30. package/{scripts → compose}/assemble.py +194 -17
  31. package/{scripts → compose}/canary.sh +14 -5
  32. package/compose/check-domains.py +1178 -0
  33. package/{config → compose}/domains.json +11 -44
  34. package/{scripts → compose}/pkgid.py +8 -1
  35. package/compose/prune-backups.py +204 -0
  36. package/{scripts → compose}/register-hooks.py +3 -3
  37. package/install.sh +1233 -0
  38. package/launch/agent-launch.py +5294 -0
  39. package/launch/agent-launch.toml +376 -0
  40. package/{scripts → launch}/check-prompting-targets.sh +1 -1
  41. package/{scripts → launch}/provision-venv.sh +1 -1
  42. package/{scripts → learn}/check-learning.py +7 -7
  43. package/{scripts → learn}/collect-learning.py +10 -10
  44. package/{config → learn}/learning.schema.json +3 -3
  45. package/{scripts → learn}/migrate-learnings.py +95 -54
  46. package/{scripts → learn}/redact.py +4 -4
  47. package/package.json +32 -27
  48. package/provenance.json +1 -0
  49. package/wrappers/claude-run.sh +162 -0
  50. package/{scripts → wrappers}/codex-run.sh +62 -6
  51. package/config/agent-launch.toml +0 -143
  52. package/scripts/agent-launch.py +0 -2350
  53. package/scripts/check-domains.py +0 -296
  54. package/scripts/check-parity.sh +0 -2003
  55. package/scripts/install.sh +0 -819
  56. /package/{shell → launch}/agent-launch.zsh +0 -0
  57. /package/{config → learn}/promotions.json +0 -0
  58. /package/{scripts/session-cost.py → session-cost.py} +0 -0
  59. /package/{scripts → wrappers}/codex-helm.sh +0 -0
package/DEPENDENCIES.md CHANGED
@@ -8,12 +8,12 @@ Korean: [`ko/DEPENDENCIES.md`](ko/DEPENDENCIES.md). Dates = verification time; u
8
8
 
9
9
  | Tool | Required by | Required capability | Verified |
10
10
  | --- | --- | --- | --- |
11
- | `codex` (codex-cli) | `scripts/codex-run.sh`, `scripts/codex-helm.sh`, `scripts/agent-launch.py`; guide "Codex direct-drive" + "Session relocation" bindings | existing `codex exec` contract plus `agents.<name>.description/config_file` config projection; final output on stdout and progress on stderr; honors `CODEX_HOME`; `codex resume` cwd-filtered | 0.144.1 · 2026-07-13 |
12
- | `bash` | `scripts/*.sh` | POSIX + arrays; runs on macOS system bash | 3.2.57 · 2026-07 |
13
- | `python3` | `scripts/session-cost.py`, `scripts/agent-launch.py` | stdlib for every direct / non-TTY / numbered path; Python 3.11+ (`tomllib`). The interactive preflight additionally needs `textual` (next row) | 3.14.5 · 2026-07-13 |
14
- | `textual` (managed venv) | `scripts/agent-launch.py` interactive preflight; provisioned by `scripts/provision-venv.sh` | Textual TUI framework in `~/.local/share/agent-launch/venv` (override `AGENT_LAUNCH_VENV`); the launcher re-execs into it on the interactive TTY path only. Absent/broken venv, non-TTY, or `TERM` `dumb`/unset falls back to numbered prompts and never blocks | 8.2.8 · py 3.14.5 · 2026-07-13 |
15
- | `jsonschema` (system python) | `scripts/check-learning.py` (learning record gate; chained from `scripts/check-parity.sh`) | JSON Schema Draft 2020-12 validator executing `config/learning.schema.json` as the SSOT | 4.26.0 · 2026-07-20 |
16
- | `zsh` | `shell/agent-launch.zsh` | functions, TTY tests, argument-preserving dispatch | 5.9 · 2026-07-13 |
11
+ | `codex` (codex-cli) | `wrappers/codex-run.sh`, `wrappers/codex-helm.sh`, `launch/agent-launch.py`; guide "Codex direct-drive" + "Session relocation" bindings | existing `codex exec` contract plus `agents.<name>.description/config_file` config projection; final output on stdout and progress on stderr; honors `CODEX_HOME`; `codex resume` cwd-filtered | 0.144.1 · 2026-07-13 |
12
+ | `bash` | `install.sh`, `*/*.sh` | POSIX + arrays; runs on macOS system bash | 3.2.57 · 2026-07 |
13
+ | `python3` | `session-cost.py`, `launch/agent-launch.py` | stdlib for every direct / non-TTY / numbered path; Python 3.11+ (`tomllib`). The interactive preflight additionally needs `textual` (next row) | 3.14.5 · 2026-07-13 |
14
+ | `textual` (managed venv) | `launch/agent-launch.py` interactive preflight; provisioned by `launch/provision-venv.sh` | Textual TUI framework in `~/.local/share/agent-launch/venv` (override `AGENT_LAUNCH_VENV`); the launcher re-execs into it on the interactive TTY path only. Absent/broken venv, non-TTY, or `TERM` `dumb`/unset falls back to numbered prompts and never blocks | 8.2.8 · py 3.14.5 · 2026-07-13 |
15
+ | `jsonschema` (system python) | `learn/check-learning.py` (learning record gate; chained from `gates/check-parity.sh`) | JSON Schema Draft 2020-12 validator executing `learn/learning.schema.json` as the SSOT | 4.26.0 · 2026-07-20 |
16
+ | `zsh` | `launch/agent-launch.zsh` | functions, TTY tests, argument-preserving dispatch | 5.9 · 2026-07-13 |
17
17
  | `git` | scripts, workflow (`origin/<base>..HEAD`, worktrees) | modern git; worktree support | 2.50.1 · 2026-07 |
18
18
  | coreutils (`mktemp`, `cp`) | `codex-run.sh` hermetic home; `codex-helm.sh` managed home | BSD or GNU | 2026-07 |
19
19
 
@@ -21,7 +21,7 @@ Korean: [`ko/DEPENDENCIES.md`](ko/DEPENDENCIES.md). Dates = verification time; u
21
21
 
22
22
  | CLI | Role | Required capability | Version owner |
23
23
  | --- | --- | --- | --- |
24
- | Claude Code | primary host; loads `CLAUDE.md` + `guides/`; `agent-launch` backend | `--model`; effort `low/medium/high/xhigh/max`; `--agents` JSON with per-agent `model`/`effort`; `--append-system-prompt`; `--mcp-config`; permission modes `acceptEdits/auto/bypassPermissions/manual/dontAsk/plan` | Environment Binding (v2.1.207) |
24
+ | Claude Code | primary host; loads `CLAUDE.md` + `guides/`; `agent-launch` backend | `--model`; effort `low/medium/high/xhigh/max`; `--agents` JSON with per-agent `model`/`effort`; `--append-system-prompt`; `--mcp-config`; permission modes `acceptEdits/auto/bypassPermissions/manual/dontAsk/plan` | Environment Binding (v2.1.220 — the version whose behaviour is verified here) |
25
25
  | Codex CLI | mirror host; loads `AGENTS.md` + `guides/`; worker/reviewer runtime | see codex row above | Environment Binding + this file |
26
26
 
27
27
  ## LLM models & providers — owned by `Environment Binding`
@@ -33,19 +33,19 @@ Concrete role-slot→model bindings live only in each guide's `Environment Bindi
33
33
 
34
34
  ## Deployed Codex assets
35
35
 
36
- - **Codex custom agents** (`codex/agents/*.toml`) — installed role templates for `frontier`, `workhorse`, `sweep`, and `reviewer`; optional at runtime, but part of the Restore contract and activated by `scripts/codex-helm.sh` when Codex subagent fan-out is explicitly authorized.
36
+ - **Codex custom agents** (`codex/agents/*.toml`) — installed role templates for `frontier`, `workhorse`, `sweep`, and `reviewer`; optional at runtime, but part of the Restore contract and activated by `wrappers/codex-helm.sh` when Codex subagent fan-out is explicitly authorized.
37
37
 
38
38
  ## Referenced / optional — not required by the core repo
39
39
 
40
- - **ultracode-for-codex** (0.5.0) — the `$ultracode-for-codex` Codex skill / CLI (Codex-backed, gpt). In cross-family review it is the **ultracode** route a **Claude** main dispatches (gpt review); a Codex main instead uses `claude --effort ultracode -p` (Claude Code's headless `/workflows` ultracode mode). Required only when the ultracode/hybrid route is selected and the main is Claude.
41
- - **Cross-family review reviewers** — with `review_family=cross` (default), each main routes review to the opposite family. A Claude main dispatches gpt review via `$CODEX_HOME/bin/codex-run --profile hermetic` (and `codex-helm --mode review` for hybrid fan-out); a Codex main dispatches Claude review via the `claude` CLI (`claude -p --permission-mode plan` for native/onto, and `claude --effort ultracode -p` for the ultracode workflow-orchestration route the `ultracode` effort value is accepted by claude 2.1.210 though not listed in `--help`). The reviewer command, resolved path, and opposite-family tier bindings are named in the launch contract; an absent or unauthenticated route degrades to same-family native (PROPOSED). `review_family=same` restores same-family review.
42
- - **codex-plugin-cc** (1.0.6; re-evaluated 2026-07-16) — spawns `codex app-server` with inherited env and no `--ignore-user-config`/`--profile`, so every run reads the real `~/.codex` (config.toml, auth, its MCP servers); it has no per-invocation hermetic reach, which is what makes it unfit as a **review** route: the reviewer would inherit the same config and AGENTS.md as the main, undercutting the independent lens `review_family=cross` exists to provide. The model *is* selectable (`--model`/`--effort`); what is dated is the bundled `gpt-5-4-prompting` skill, so passing a current model does not resolve it. **Not adopted**; `scripts/codex-run.sh` is preferred for controlled reach. It does not touch Claude Code's `/code-review` (no `code-review.md`; it adds namespaced `/codex:*`), so it never made that route cross-family. Capability we lack and may still want independently: its opt-in `Stop` hook review gate.
40
+ - **ultracode-for-codex** (0.5.0) — the `$ultracode-for-codex` Codex skill / CLI (Codex-backed, gpt). In cross-family review it is the **ultracode** route a **Claude** main dispatches (gpt review); a Codex main instead uses the **ultracode** capability, which is the `claude` backend itself run headless with the keyword `ultracode` in the prompt — that keyword is what opens the Workflow tool for the turn (`workflowKeywordTriggerEnabled`, default true, read in the installed 2.1.220 bundle). Required only when the ultracode/hybrid route is selected and the main is Claude.
41
+ - **Cross-family review reviewers** — with `review_family=cross` (default), each main routes review to the opposite family. A Claude main dispatches gpt review via `$CODEX_HOME/bin/codex-run --profile hermetic` (and `codex-helm --mode review` for hybrid fan-out); a Codex main dispatches Claude review via the `claude` CLI (`claude -p --permission-mode plan` for native/onto, and, for the workflow-orchestration route, the same `claude` CLI headless with the keyword `ultracode` in the prompt — the keyword trigger is what the injected contract names, so this is the mechanism to follow). The reviewer command, resolved path, and opposite-family tier bindings are named in the launch contract; an absent or unauthenticated route degrades to same-family native (PROPOSED). `review_family=same` restores same-family review.
42
+ - **codex-plugin-cc** (1.0.6; re-evaluated 2026-07-16) — spawns `codex app-server` with inherited env and no `--ignore-user-config`/`--profile`, so every run reads the real `~/.codex` (config.toml, auth, its MCP servers); it has no per-invocation hermetic reach, which is what makes it unfit as a **review** route: the reviewer would inherit the same config and AGENTS.md as the main, undercutting the independent lens `review_family=cross` exists to provide. The model *is* selectable (`--model`/`--effort`); what is dated is the bundled `gpt-5-4-prompting` skill, so passing a current model does not resolve it. **Not adopted**; `wrappers/codex-run.sh` is preferred for controlled reach. It does not touch Claude Code's `/code-review` (no `code-review.md`; it adds namespaced `/codex:*`), so it never made that route cross-family. Capability we lack and may still want independently: its opt-in `Stop` hook review gate.
43
43
  - **MCP servers** (onto, clickhouse, node_repl, …) — environment-specific; referenced by Environment Binding (VERIFIER-A; coding-staged guide's structured multi-lens review slot), not a core dependency. For cross-family review, agent-launch mounts `onto` and instructs the main to call `onto_review` with `llmOverride={provider,model}` (from `[hosts.*].onto_review`, an onto review-role registered pair) so onto runs the opposite family; onto's own model seats are not launcher-controllable, so the family is set per call.
44
44
  - **spreadsheet-processing** (skill) — referenced by the global spreadsheet rule; present in the author's Claude Code and Codex environments. If absent, the rule's inline fallback (plain tools/code + real Excel-engine validation) applies.
45
45
 
46
46
  ## Untracked — dependencies, but excluded by design
47
47
 
48
- Host `config.toml` and `settings.json` — machine-specific trust lists, hook paths, and MCP secrets. The tracked `config/agent-launch.toml` contains launch bindings but no secrets. See README Scope.
48
+ Host `config.toml` and `settings.json` — machine-specific trust lists, hook paths, and MCP secrets. The tracked `launch/agent-launch.toml` contains launch bindings but no secrets. See README Scope.
49
49
 
50
50
  ## Re-verify
51
51
 
@@ -61,13 +61,13 @@ printf '%s\n' "$codex_help" | grep -Eq '(^|[[:space:]])-p([,[:space:]]|$)' || {
61
61
  printf '%s\n' "$codex_help" | grep -Eq '(^|[[:space:]])-s([,[:space:]]|$)' || { echo "missing codex flag: -s"; exit 1; }
62
62
  claude --version; claude --help | grep -E -- '--model|--effort|--agents|--append-system-prompt|--mcp-config'
63
63
  bash --version | head -1; zsh --version; python3 --version; git --version
64
- AGENT_LAUNCH_VENV="${AGENT_LAUNCH_VENV:-$HOME/.local/share/agent-launch/venv}" bash scripts/provision-venv.sh
64
+ AGENT_LAUNCH_VENV="${AGENT_LAUNCH_VENV:-$HOME/.local/share/agent-launch/venv}" bash launch/provision-venv.sh
65
65
  "${AGENT_LAUNCH_VENV:-$HOME/.local/share/agent-launch/venv}/bin/python" -c 'import textual, sys; print("textual", textual.__version__, "py", sys.version.split()[0])'
66
- bash -n scripts/codex-run.sh scripts/codex-helm.sh scripts/check-parity.sh scripts/check-prompting-targets.sh scripts/provision-venv.sh scripts/install.sh
67
- zsh -n shell/agent-launch.zsh
68
- python3 -c 'compile(open("scripts/agent-launch.py").read(), "scripts/agent-launch.py", "exec")'
69
- ./scripts/check-parity.sh
70
- ./scripts/check-prompting-targets.sh
66
+ bash -n wrappers/codex-run.sh wrappers/codex-helm.sh gates/check-parity.sh launch/check-prompting-targets.sh launch/provision-venv.sh install.sh
67
+ zsh -n launch/agent-launch.zsh
68
+ python3 -c 'compile(open("launch/agent-launch.py").read(), "launch/agent-launch.py", "exec")'
69
+ ./gates/check-parity.sh
70
+ ./launch/check-prompting-targets.sh
71
71
  python3 - <<'PY'
72
72
  import os, pathlib, tomllib
73
73
  roots = [pathlib.Path("codex/agents"), pathlib.Path(os.environ.get("CODEX_HOME", pathlib.Path.home() / ".codex")) / "agents"]
@@ -79,7 +79,7 @@ for root in roots:
79
79
  tomllib.loads(path.read_text())
80
80
  print(f"agent TOML ok: {root}")
81
81
  PY
82
- scripts/codex-helm.sh --dry-run --mode review "probe"
82
+ wrappers/codex-helm.sh --dry-run --mode review "probe"
83
83
  ```
84
84
 
85
85
  ## Ownership
package/README.md CHANGED
@@ -2,6 +2,10 @@
2
2
 
3
3
  Single source of truth for the global instructions and scoped guides that drive multiple LLM CLI agents (Claude Code, Codex CLI) under one working discipline. Edit once; it applies to every agent in every environment.
4
4
 
5
+ A deployable instruction corpus for coding agents, plus the CLI that
6
+ installs, verifies, and evolves it. The npm package ships the corpus; `install.sh` is both the
7
+ `agent-bios` CLI entry and the deployer.
8
+
5
9
  Korean reference: [`ko/`](ko/) mirrors every doc below — reference only, never installed or loaded.
6
10
 
7
11
  ## Model
@@ -14,9 +18,13 @@ Two layers:
14
18
  ## Principles
15
19
 
16
20
  - **Rule bodies never name concrete models or tools** — only role slots and tiers. Bindings live in each guide's `Environment Binding` (dated; expire ~8 weeks or on a newer model). New model → update that row + date; leave rules alone. Declared exceptions: sections whose subject is a concrete tool surface (the cli guide's Codex direct-drive section) and optional-capability names inventoried in `DEPENDENCIES.md` (e.g. the `spreadsheet-processing` skill) — the Adopting checklist below covers swapping both.
17
- - **The two CLIs' files are mirrors** — differing only in title and config-home variable (`$CLAUDE_CONFIG_DIR` ↔ `$CODEX_HOME`), plus one declared Codex-only standing-dispatch authorization required by Codex's trigger contract. `scripts/check-parity.sh` enforces that exact exception, mirror parity, pointer resolvability, frontmatter, shared anchor phrases on the global↔guide restatement pairs that remain, and Codex role-binding / wrapper-default projections. Parity enforces content synchronization only; it does not guarantee both harnesses respond to the same wording with the same strength.
21
+ - **The Codex tree is generated, not mirrored by hand** — `gates/emit-mirrors.py` projects `claude/` → `codex/` and `ko/claude/` → `ko/codex/`, differing only in title and config-home variable (`$CLAUDE_CONFIG_DIR` ↔ `$CODEX_HOME`), plus one declared Codex-only standing-dispatch authorization required by Codex's trigger contract, inserted at a pinned position. It owns the projection rule; `gates/check-parity.sh` runs its `--check` and adds pointer resolvability, frontmatter, shared anchor phrases on the global↔guide restatement pairs that remain, and Codex role-binding / wrapper-default projections. Parity enforces content synchronization only; it does not guarantee both harnesses respond to the same wording with the same strength.
18
22
  - **The global is a per-session token budget** — it is re-sent every session and to every subagent, and each added bullet dilutes every other rule. A new global bullet must name the bullet it displaces (or why none does); procedures, tables, numbers, and worked examples belong in guides.
19
- - **`Evidence Base` (per guide) is the single owner of numbers.** Measure with `scripts/session-cost.py`.
23
+ - **Deploying is not loading** file presence cannot detect a declined import approval or a broken entry import line, so `onboard` ends with an activation canary and `verify` checks the entry import line. A corpus that landed but is not read has changed nothing.
24
+ - **Every span we write into a file we do not own has a remover** — the Codex config block, the settings hook registrations, the `AGENTS.md` central region and the zsh hook line each sit in a span identifiable as ours (a marker pair, a tagged line, or a manifest name), and `uninstall` removes exactly those and nothing around them. The entry `CLAUDE.md`/`AGENTS.md` are seeded once and then yours: the corpus is assembled under `central/` and the entry file only imports it, so your own additions are never mixed with ours and never removed with them. `uninstall` is a security operation — everything of ours goes, and what it took leaves as one archive you can hand off or delete.
25
+ - **Metadata selects, observation authorizes** — a record of what was promoted can propose an irreversible act but never authorize one; it has never seen the machine it will run on. Clearing a user's personal copy of a promoted learning asks the deployed corpus whether the replacement is really there, excludes the copy being deleted from that evidence, and keeps the copy on any uncertainty: a kept duplicate is redundant, a wrong removal is data loss.
26
+ - **The deployment outlives the act that made it** — a command's success says nothing about which version landed; a cached package can serve the previous one at exit 0. So the deployed state carries its own version marker and `status` reads that marker rather than the source it was launched from. With no marker the answer is `unknown`, never a guess.
27
+ - **`Evidence Base` (per guide) is the single owner of numbers.** Measure with `session-cost.py`.
20
28
  - **English is canonical and installed; Korean (`ko/`) is reference only.** The harness loads only fixed-name English files; the installer deploys English only.
21
29
 
22
30
  ## Layout
@@ -27,8 +35,21 @@ Two layers:
27
35
  | `claude/guides/*.md`, `codex/guides/*.md` | scoped guides (en) — installed |
28
36
  | `codex/agents/*.toml` | Codex custom subagent role templates — installed |
29
37
  | `ko/**` | Korean mirror of every doc above + this README + DEPENDENCIES (reference only) |
30
- | `config/agent-launch.toml`, `scripts/agent-launch.py`, `shell/agent-launch.zsh` | shared Codex/Claude launch profile, preflight TUI, and zero-argument shell interception |
31
- | `scripts/` | parity/cost tooling plus internal Codex wrappers deployed to `$CODEX_HOME/bin/` |
38
+ | `launch/` | launch profile, preflight TUI, zero-argument shell interception, the managed Textual venv, and the prompting-target check that guards the profile's model bindings |
39
+ | `compose/` | corpus classification and per-selection assembly domain manifest and its gate, assembler, package identity, hook registration, activation canary, deployed-corpus state |
40
+ | `learn/` | the collection loop — capture, record schema and its validator, curation intake, promotion manifest, redistribution, and the secret-redaction floor |
41
+ | `session-distill/` | the heavy curator pipeline that mines many sessions into corpus-grade items |
42
+ | `wrappers/` | internal Codex wrappers, deployed to `$CODEX_HOME/bin/` |
43
+ | `gates/` | author-side verification (mirror generation, parity, lexicon, payload, assembler scenarios) — reachable only from a repo checkout, and `check-package.sh` fails if any of it enters the npm payload |
44
+ | `ontology/` | what a change obliges elsewhere — entities, obligation edges, and the service's routes, held against real source by `check-ontology.py`. `instances/graph.json` is canonical; `LEXICON.md`, the RDF views, the HTML map, and the competency/extension docs are generated from it |
45
+ | `install.sh`, `session-cost.py` | the CLI and the cost meter — the two things you run directly |
46
+ | `decisions/` | the decision record for developing this repo — what was decided and which alternative it closed; author-side, never shipped |
47
+ | `packages/` | authored corpus packages, organized by package identity rather than by concept home |
48
+ | `.githooks/` | the pre-commit hook that runs the gates against the index, enabled per clone with `core.hooksPath` |
49
+ | `SURFACES.md` | where knowledge and tools reach a model, and what each place admits — every entry names the code that realizes it |
50
+ | `FINDINGS.md` | open implementation defects, live; closing one deletes its entry |
51
+ | `design/`, `benchmarks/` | design records and the instruction-behavior benchmark |
52
+ | `research/` | corpus research (the 12,749-file AGENTS.md/CLAUDE.md classification): reports, scripts, labeling record; bulk data stays local by `.gitignore` rule |
32
53
  | `DEPENDENCIES.md` | external tools / host CLIs / model providers + verified versions |
33
54
 
34
55
  `config.toml` and `settings.json` are machine-specific (trust lists, hook paths, secrets) and intentionally untracked.
@@ -38,18 +59,24 @@ Two layers:
38
59
  | Guide | Scope |
39
60
  | --- | --- |
40
61
  | `cli-multi-model-workflow` | multi-model CLI workflow: Default Frame, role slots/tiers, delegation mechanics, driving Codex CLI directly, cache economy, unattended-batch safety, halt/resume, handoff contract, Environment Binding |
41
- | `coding-staged-workflow` | staged development: design → process → implement, review loop, severity contract, verification menus, stop conditions |
62
+ | `coding-staged-workflow` | staged development: design → process → implement, lightweight path, review loop, severity contract, stop conditions |
63
+ | `verification-discipline` | verification depth and per-domain mix (owns the Verification Menus), case space, what a green result is worth |
64
+ | `concept-economy` | concept-surface economy: reuse / extend / rename / split, split triggers, migration compatibility |
65
+ | `documentation-hygiene` | where comments, history, and handoffs belong; how to phrase rules others follow |
42
66
  | `llm-capability-boundary` (+ `-patterns`, `-examples`) | LLM/tool/code authority boundary: field authority, accepted output channels, structural enforcement, worked examples |
43
67
  | `mock-realization-boundary` | mock/fixture realization vs product semantic path |
44
68
  | `svg-visualization-guide` | SVG diagram / service-blueprint spec |
45
69
  | `implementation-map` | IMPLEMENTATION_MAP.html current-state dashboard |
46
70
 
71
+ The table names the load-bearing guides; the full inventory and its per-domain
72
+ classification live in `compose/domains.json`, which the domains gate holds against the
73
+ tree.
74
+
47
75
  ## Edit workflow
48
76
 
49
- 1. Edit the English canonical (`claude/`). Regenerate the Codex mirror, then preserve the single `Codex-only standing authorization` bullet under `Multi-Model Workflow`:
50
- `sed -e '1s/# CLAUDE.md/# AGENTS.md/' -e 's|${CLAUDE_CONFIG_DIR:-$HOME/.claude}|${CODEX_HOME:-$HOME/.codex}|g' claude/CLAUDE.md > codex/AGENTS.md`, and copy guides to `codex/guides/`.
51
- 2. Update the Korean mirror under `ko/` the same way (globals under `ko/claude`→`ko/codex`, guides `ko/claude/guides`→`ko/codex/guides`, config-home swap for the codex side).
52
- 3. `./scripts/check-parity.sh` must pass.
77
+ 1. Edit the English canonical (`claude/`), and the Korean canonical (`ko/claude/`) when the change is user-facing.
78
+ 2. `python3 gates/emit-mirrors.py` regenerates `codex/` and `ko/codex/` from those two canonicals. Never hand-edit the Codex side: it is a generated projection, and `--check` (which the parity gate runs) fails on any file that is not exactly what the generator emits.
79
+ 3. `./gates/check-parity.sh` must pass.
53
80
  4. Commit, then deploy: `agent-bios install` (or `agent-bios update` from a clone).
54
81
 
55
82
  ## Install (deploy)
@@ -59,16 +86,16 @@ The `agent-bios` CLI deploys this SSOT into your environment by copy — idempot
59
86
  ```bash
60
87
  npm install -g agent-bios
61
88
  agent-bios install # deploy, back up replaced files, then verify
62
- agent-bios onboard # pick domain packages, packaged install, activation canary
89
+ agent-bios onboard # pick domain packages, install, activation canary
63
90
  agent-bios verify # re-check the deployed state matches the source
64
91
  agent-bios status # show what is installed and where
65
92
  agent-bios update # git pull + reinstall (clone), or print the npm update line
66
93
  agent-bios uninstall # remove deployed files and the zsh hook
67
94
  ```
68
95
 
69
- From a git clone, run `./scripts/install.sh install` directly (the same CLI). `install` respects `CLAUDE_CONFIG_DIR`, `CODEX_HOME`, `AGENT_LAUNCH_VENV`, and `ZDOTDIR`; `--dry-run` prints actions without changing anything. Deploy to **every active environment in one sitting** — a partial deploy leaves a shared global pointing at a guide some environment lacks; globals are English only. Replaced files are backed up under `~/.local/share/agent-bios/backups/<timestamp>/`, and the installed set is recorded in a manifest that `uninstall` consumes. The published npm package ships only the deploy set (never `settings.json`, `config.toml`, `ko/`, or `benchmarks/`).
96
+ From a git clone, run `./install.sh install` directly (the same CLI). `install` respects `CLAUDE_CONFIG_DIR`, `CODEX_HOME`, `AGENT_LAUNCH_VENV`, and `ZDOTDIR`; `--dry-run` prints actions without changing anything. Deploy to **every active environment in one sitting** — a partial deploy leaves a shared global pointing at a guide some environment lacks; globals are English only. Replaced files are backed up under `~/.local/share/agent-bios/backups/<timestamp>/`, and the installed set is recorded in a manifest that `uninstall` consumes. Those copies are **a manual escape hatch, not a restore mechanism**: nothing reads them back, so recovering from one means copying files yourself. The supported paths are the ones `agent-bios help` prints — a clone rolls corpus content back to a registered version, an npm install rolls the whole package back by version — and `uninstall` emits one archive of everything it removed. The backups exist for the case those three do not cover: the exact bytes that were on disk before a particular install. The published npm package ships only the deploy set (never `settings.json`, `config.toml`, `ko/`, or `benchmarks/`).
70
97
 
71
- In a TTY, zero-argument `codex` or `claude` opens the launch preflight. Every arrow-key TUI selection screen keeps the complete current setup in a fixed top panel, followed by the highlighted option's description and then the option list. Move with Up/Down, select with Enter, use Esc to return to the previous menu, and use `q` to cancel; Esc at the mode root also cancels. Each tier's model is chosen from the host's configured catalog, or via **Other** to type any model id the backend accepts; that text input preserves values that start with `q`, so Esc or Ctrl-C cancels immediately there, while submitting `q` cancels after Enter. The Textual preflight reflows to the terminal size, so there is no fixed minimum geometry. The root menu picks a mode — **Software Engineer** (repo-scoped work that defers to the project's own AGENTS.md/CLAUDE.md: **Vanilla**, the bare CLI with no launch contract, tier bindings, or applied permissions, plus **Custom**; SE-specific review defaults arrive later), **Builder** (the tier presets: Balanced, Deep review, Fast batch, Solo with delegation off, plus Custom), **Session distill** — then a preset within it. Select **Custom** (in Software Engineer or Builder) to open a persistent settings hub for the main tier, review setup, host policy, and each tier binding. Every edit returns to that hub; **Start with these settings** is the final launch confirmation, **Save these settings globally and start** additionally persists the setup as a named preset (with host-scoped tier overrides) in your user config for reuse elsewhere, and **Exit without launching** cancels the launch. The rich preflight runs from a managed virtualenv (`scripts/provision-venv.sh`, at `~/.local/share/agent-launch/venv`) that the launcher re-execs into on the interactive path; when that venv is unavailable, or the call is non-interactive, or `TERM` is `dumb`/unset, the launcher falls back to numbered prompts, where `b` is the back command. Backend command names are resolved from the calling environment's `PATH`; shell functions are not re-entered. Codex child bindings are materialized as session-selected agent configs under the user cache, while Claude receives model and effort in `--agents` JSON when delegation is enabled. Review runs cross-family by default (`review_family`, default `cross`; `same` restores today's same-family projection): because the main's tiers are one model family, every dispatchable review route — native, onto, and ultracode — runs on the opposite family. The exception is `slash-review`, the host's own built-in review command (`/code-review` on Claude, with `ultra` for its deep multi-agent pass; `/review` on Codex): it needs no dependency and always resolves, but being the main's own command it cannot be dispatched cross-family, so under `cross` it runs as the same-family floor and its verdicts are labeled PROPOSED. A Claude main dispatches gpt/codex review (native via the `codex-run` reviewer wrapper resolved under `$CODEX_HOME/bin`, onto via an `llmOverride` to the configured openai seat, ultracode via the `$ultracode-for-codex` Codex skill); a Codex main dispatches Anthropic/Claude review (native via `claude -p --permission-mode plan`, onto via an `llmOverride` to the anthropic seat, ultracode via `claude --effort ultracode -p` Claude Code's headless `/workflows` ultracode mode, verified accepted on claude 2.1.210). The concrete reviewer command, resolved absolute path, `llmOverride`, and opposite-family tier bindings are named in the injected session-start contract; cross-family reviewers are dispatched as read-only subprocesses, not CLI-native subagents, since neither CLI hosts the other family as a native subagent. When a cross-family route is unavailable at launch or unauthenticated at use time it degrades to same-family native subagent review labeled PROPOSED (family collapse) rather than blocking; a requested non-none review with no cross-family route and no same-family fallback (delegation off) stays fail-closed. Review setup means configured/requested; this launcher does not claim that review completed, and unavailable runtimes such as Ultrawork are not offered until integrated.
98
+ In a TTY, zero-argument `codex` or `claude` opens the launch preflight. Every arrow-key TUI selection screen keeps the complete current setup in a fixed top panel, followed by the highlighted option's description and then the option list. Move with Up/Down, select with Enter, use Esc to return to the previous menu, and use `q` to cancel; Esc at the mode root also cancels. Each tier's model is chosen from the host's configured catalog, or via **Other** to type any model id the backend accepts; that text input preserves values that start with `q`, so Esc or Ctrl-C cancels immediately there, while submitting `q` cancels after Enter. The Textual preflight reflows to the terminal size, so there is no fixed minimum geometry. The root menu picks a mode — **Software Engineer** (repo-scoped work that defers to the project's own AGENTS.md/CLAUDE.md: **Vanilla**, the bare CLI with no launch contract, tier bindings, or applied permissions, plus **Custom**; SE-specific review defaults arrive later), **Builder** (the tier presets: Balanced, Deep review, Fast batch, Solo with delegation off, plus Custom), **Session distill** — then a preset within it. Select **Custom** (in Software Engineer or Builder) to open a persistent settings hub for the main tier, review setup, host policy, and each tier binding. Every edit returns to that hub; **Start with these settings** is the final launch confirmation, **Save these settings globally and start** additionally persists the setup as a named preset (with host-scoped tier overrides) in your user config for reuse elsewhere, and **Exit without launching** cancels the launch. The rich preflight runs from a managed virtualenv (`launch/provision-venv.sh`, at `~/.local/share/agent-launch/venv`) that the launcher re-execs into on the interactive path; when that venv is unavailable, or the call is non-interactive, or `TERM` is `dumb`/unset, the launcher falls back to numbered prompts, where `b` is the back command. Backend command names are resolved from the calling environment's `PATH`; shell functions are not re-entered. Codex child bindings are materialized as session-selected agent configs under the user cache, while Claude receives model and effort in `--agents` JSON when delegation is enabled. Review runs cross-family by default (`review_family`, default `cross`; `same` restores today's same-family projection): because the main's tiers are one model family, every dispatchable review route — native, onto, and ultracode — runs on the opposite family. The exception is `slash-review`, the host's own built-in review command (`/code-review` on Claude, with `ultra` for its deep multi-agent pass; `/review` on Codex): it needs no dependency and always resolves, but being the main's own command it cannot be dispatched cross-family, so under `cross` it runs as the same-family floor and its verdicts are labeled PROPOSED. A Claude main dispatches gpt/codex review (native via the `codex-run` reviewer wrapper resolved under `$CODEX_HOME/bin`, onto via an `llmOverride` to the configured openai seat, ultracode via the `$ultracode-for-codex` Codex skill); a Codex main dispatches Anthropic/Claude review (native via `claude -p --permission-mode plan`, onto via an `llmOverride` to the anthropic seat, ultracode via the `claude` CLI headless with the keyword `ultracode` in the prompt, which is what opens Claude Code's dynamic workflow for that turn). The concrete reviewer command, resolved absolute path, `llmOverride`, and opposite-family tier bindings are named in the injected session-start contract; cross-family reviewers are dispatched as read-only subprocesses, not CLI-native subagents, since neither CLI hosts the other family as a native subagent. When a cross-family route is unavailable at launch or unauthenticated at use time it degrades to same-family native subagent review labeled PROPOSED (family collapse) rather than blocking; a requested non-none review with no cross-family route and no same-family fallback (delegation off) stays fail-closed. Review setup means configured/requested; this launcher does not claim that review completed, and unavailable runtimes such as Ultrawork are not offered until integrated.
72
99
 
73
100
  At the shell-wrapper boundary, every argument-bearing command (`codex exec ...`, `claude -p ...`) and every non-TTY invocation skips launch-profile projection and preserves caller arguments. The Claude direct path intentionally retains its wrapper default, `--dangerously-skip-permissions`. `codex --no-tui ...` / `claude --no-tui ...` explicitly take that direct path, and `AGENT_LAUNCH_TUI=0` disables zero-argument TUI interception for a process tree.
74
101
 
@@ -76,6 +103,10 @@ Direct `agent-launch` calls still require a valid profile to resolve the backend
76
103
 
77
104
  Add `$CODEX_DIR/bin` to `PATH` or invoke the wrappers by absolute path. `codex-helm` follows the local CLI default and launches the HELM main with `--dangerously-bypass-approvals-and-sandbox`; an explicit `--sandbox MODE` disables bypass for that run regardless of flag order. `AGENTS.md` gives root/main local Codex sessions standing ordinary-subagent authorization when the delegation gates fire. A non-Ultra HELM main sets native multi-agent off by default and instructs HELM to send tiered dispatch through the internal `codex-run` adapter, where the selected model, effort, and sandbox are pinned; native multi-agent defaults on only when the HELM main itself is explicitly Ultra. FRONTIER is instructed to run as a separate `gpt-5.6-sol` root that is always read-only, at max by default, Ultra for genuinely divisible complex work, or a lower supported effort when cost or latency dominates. Because the HELM main has bypass authority and arbitrary expert `-c` by design, this dispatch route is an instruction-backed, live-E2E-verified default rather than a security boundary. Keep `codex-run` as the low-level internal adapter, not as a user-facing policy boundary. Both wrappers accept `-c key=value` as an expert override, and that override may intentionally change wrapper defaults for a single run.
78
105
 
106
+ `claude-run` is the Claude-side adapter, deployed to `$CLAUDE_DIR/bin`, and it takes `--model` and `--effort` to pin the seat. Omitting either warns and dispatches anyway, matching `codex-run`: refusing outright turned "the review ran unpinned" into "the review did not run", which is the worse of the two. The honest signal is downstream instead — an unpinned dispatch can name no seat, so it emits no receipt and the method adjudicates to UNKNOWN rather than to a clean pass. Its default denies the mutating tools, which is not the OS-level sandbox its Codex twin gets — do not read the two defaults as equivalent guarantees.
107
+
108
+ **Review receipts.** A launch reports what it *projected*, because at launch no review has run — so a clean verdict without a receipt is PROPOSED, never ACHIEVED. Given `REVIEW_RECEIPT_DIR`, both adapters record what they observed of the dispatch they just performed: exit status, a hash of the packet fed in, a hash of the bytes returned, and the seat actually sent. Unset, they behave exactly as they would otherwise and write nothing. `agent-launch --fold-receipts DIR PACKET MAIN_DISPATCH_ID` folds a run into a `ReviewReceipts/v1` bundle — several passes of one method become the one record it is judged on — and `agent-launch --verify-receipts PLAN BUNDLE` adjudicates it, exiting non-zero unless every selected method verified. Adapting another tool needs no change here: call `agent-launch --emit-receipt` from your adapter and prove it conforms with `agent-launch --check-adapter SEAT -- CMD`, which is adjudicated by the same code that credits a real review. A receipt is still written by whoever ran the review, so this buys drift rather than honesty: what it stops is a reviewer that quietly never ran, returned nothing, or exited non-zero reading as a clean pass.
109
+
79
110
  `agent-bios verify` runs the post-deploy gate (also run at the end of `install`): every `guides/*.md` referenced by the deployed global exists in that environment's `guides/`; required agent files `frontier.toml`, `workhorse.toml`, `sweep.toml`, and `reviewer.toml` exist under `$CODEX_DIR/agents/` and parse as TOML; `frontier.toml` deliberately omits `model_reasoning_effort` for native surfaces that accept per-spawn effort; each required `codex exec --help` flag in `DEPENDENCIES.md` is present; and `$CODEX_DIR/bin/codex-helm --dry-run --mode review "probe"` succeeds as a credential-free assembly check (not a live Codex call). `install` overwrites each guide's `Environment Binding`; keep per-environment binding edits in the repo copy or an untracked file.
80
111
 
81
112
  ## Adopting elsewhere
package/claude/CLAUDE.md CHANGED
@@ -52,57 +52,25 @@
52
52
 
53
53
  ## Concept Economy
54
54
 
55
- - Keep the concept graph compact by reusing existing concepts that clearly cover the behavior.
56
- - Treat lasting or shared names as concept candidates: features, entities, variables, types, helper modules, artifacts, config keys, CLI flags, MCP/tool fields, public response fields, artifact fields, enum values, failure kinds, retry/recovery tokens, process names, and documentation terms.
57
- - Before adding or changing a concept, find the nearest existing concept and choose one path explicitly: reuse, extend, rename, or split.
58
- - Prefer broad, stable concepts with precise properties over narrow near-duplicates.
55
+ - When adding, changing, renaming, splitting, or exposing anything lasting or shared a feature, entity, type, field, config key, CLI flag, enum value, failure kind, artifact, or documentation term — read and use `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/concept-economy.md` as a scoped extension of this section.
59
56
  - Before fixing a review finding or test failure, classify the fix as reducing, preserving, or increasing the active concept surface.
60
- - Split or promote a concept when it changes runtime behavior, ownership, lifecycle, validation, failure mode, user-visible behavior, audit/replay requirements, authority, persistence, user control, or failure handling.
61
- - Keep derived values as properties or projections of their source concept when tools/code can derive them from the source authority.
62
- - Keep internal projections and helper outputs internal unless public exposure is required for user behavior, product contract, or artifact truth.
63
- - Distinguish authority from visibility: public responses may expose bounded views, while the source concept or artifact remains the truth location.
64
- - Reuse existing enum values, failure kinds, retry/recovery tokens, and result/failure surfaces before introducing new vocabulary.
65
- - Use fallback paths, compatibility shims, and deprecated alias normalization when explicit migration compatibility is required.
66
- - Keep comments and active docs aligned with runtime behavior, failure semantics, retry policy, ownership, and authority.
67
- - When a split is necessary, name the parent concept, explain the reason for the split, and map aliases or variants back to the canonical concept.
68
- - In ontology work, check existing entities and relations first, then keep the concept graph compact.
69
- - In code work, follow existing naming patterns and consolidate variations introduced by the current change.
70
- - Let the repository's shape mirror its concept graph: keep each shared, lasting concept's canonical name traceable across the layers it appears in — path, module, type/interface, field, and public API — so the structure is navigable by name (grep-findable, path-guessable) without a translation table. This binds shared concepts only; transient locals, generic containers, and framework- or tooling-imposed layout may diverge.
71
57
 
72
58
  ## Coding Guidelines
73
59
 
74
60
  - For `.xlsx` editing, generation, reconciliation, validation, or connected spreadsheet processing, use the installed `spreadsheet-processing` skill when present — with plain tools/code as the fallback — and validate formula-dependent Excel results with the real Microsoft Excel engine.
75
- - For meaningful development work, read and use `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/coding-staged-workflow.md` as a scoped extension of these Coding Guidelines.
61
+ - For development work, read and use `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/coding-staged-workflow.md` as a scoped extension of these Coding Guidelines — a change too narrow to need it is what its lightweight path decides, not a reason to skip the read.
76
62
  - For mock, fixture, fake, stub, simulated-provider, or test-realization design, read and use `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/mock-realization-boundary.md` as a scoped extension of these Coding Guidelines.
77
- - When the user asks to "설계" or design, read the coding-staged-workflow guide and focus on high-level design and implementation-process design; move to implementation only after the user asks to implement or approves the plan.
78
- - Think before coding: state key assumptions and surface ambiguity early.
79
- - Build the smallest viable functional path that satisfies the qualitative completion criteria. Minimum limits surface area, configuration, abstractions, optional scope, and implementation spread; it must not reduce required behavior, runtime authority, evidence quality, or verification depth.
80
- - Treat viability as real behavior against real inputs, real authority, and the intended runtime path. Use mocks only for tests, fixtures, or explicitly requested simulations; mock-backed paths support verification but do not count as product completion.
81
- - Make surgical changes. Touch only what the request requires, preserve existing style, and avoid casual adjacent refactors.
82
- - Clean up issues introduced by the current change. Mention unrelated dead code separately.
83
63
  - Own the full lifecycle of what you create — spawned processes and handles through teardown, artifacts out of tool-managed temp locations into a durable home — and keep differently-owned state separate: never colocate deploy-managed and user-owned data in one overwrite-managed file.
84
- - Define success criteria before multi-step coding work, then verify against them.
85
- - For bugs, prefer a reproducing test before the fix when practical.
86
- - Every changed line should trace back to the user's request.
87
- - Fix the root cause at its authority rather than the visible symptom: when downstream patches keep compensating for bad inputs, fix upstream at the source; when each fix only exposes another instance of the same defect, single-source the value and fix the whole class instead of patching instances.
88
64
  - Land risky or behavior-changing work behind a default-off path that preserves current behavior when off (proven by diff) and is enabled by an explicit opt-in, so the change stays reversible and the on/off difference is isolated. When a request would weaken a security or authority posture — removing or loosening an authentication/authorization check or access scope, or lowering a protective value such as session/token lifetime, password/crypto strength, rate limit, lockout threshold, or audit retention — treat it as a decision, not a rote edit, even when it is a one-line change and nothing in the code labels the value as security-relevant: state the consequence and at least one safer path to the real goal, and do not apply the weakening in the same turn — proceed only after the user confirms they accept the tradeoff.
89
65
 
90
66
  ## Verification Discipline
91
67
 
92
68
  - For composing a review request, packet, or reviewer role — the evidence bar, the verdict shape, and why a review returned noise, nothing, or a clean bill of health — read and use `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/review-request.md` as a scoped extension of this section.
93
69
  - After every meaningful code, ontology, config, data, spreadsheet, or documentation change, run a verification loop regardless of commit or handoff status.
94
- - Use static checks broadly: typecheck, lint, build, format, schema/config validation, graph validation, workbook structure checks, import boundaries, and security checks when available.
95
- - Add the narrowest reliable runtime or semantic test that proves the changed behavior, meaning, or contract.
96
- - Pick each domain's verification mix (code, ontology, config/data, spreadsheets, docs) from the Verification Menus in the coding-staged-workflow guide.
97
- - Let the LLM derive scenarios from the diff, user impact, concept impact, and failure modes; let tools/code execute and verify them.
98
- - Keep E2E stable with deterministic data, resilient selectors, isolated external dependencies, and explicit waits.
70
+ - For choosing verification depth, the per-domain mix, the case space, what makes a completion criterion falsifiable, how to keep an E2E stable, or what a green result is worth, read and use `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/verification-discipline.md` as a scoped extension of this section — its Verification Menus carry the per-domain mixes.
99
71
  - Report the checks run, results, and any unverified risk before calling the work done.
100
72
  - Trust a green check only when it traversed the actual changed code through the real dispatch and real calls (not a mock, dry-run, or bypass), and remember that "it ran" is not "quality met" — a fallback, floor, or mock run is not done; treat a zero-findings verdict as suspect until you confirm the harness ran rather than silently crashed, and make PASS mean concrete assertions on real output from the real path.
101
- - Make completion criteria falsifiable: prefer signals that fail when the mechanism is wrong (negative or contrast controls), and if no existing gate can judge a criterion, build the executable judge or do not claim the criterion met.
102
73
  - Before comparing two of anything (cost, performance, quality, frequency), fix a common basis — units, denominators, population, measurement surface — compare on equivalent output, and exclude or flag non-representative data (promotions, outages, smoke slices).
103
- - For non-trivial designs or high-risk changes, run independent adversarial review across distinct lenses, ideally on the design before implementation, and re-verify each finding against real code before acting on it. Apply the convergence heuristic by reviewer kind (detailed in the multi-model guide): same-kind convergence is high confidence but same-kind reviewers share blind spots — their shared "clean" is not verification; different-kind divergence is the expected signal — act on the union. Never accept an orchestrated workflow's self-reported all-green as sufficient; independently re-run the diff inspection and verification suite yourself.
104
- - Proportion verification to cost, risk, and information gain: before expensive or slow live runs, diagnose in code and replay the changed deterministic logic over persisted real artifacts, probe at N=1 with inputs precondition-checked, and reserve full design-review-plus-live verification for first-of-kind or authority-changing work; proportion assurance to the deployment context — a single-user, own-data tool does not warrant production-grade assurance; prefer delivery.
105
- - Trust a green / zero-findings verdict only if the check could have failed over a real, non-empty subject: assert the entity-under-test set has cardinality > 0 before any "no bad X" or "all X satisfy P" claim (an empty subject set passes vacuously and proves nothing), and for any test touching a branch you add or delete, confirm its inputs satisfy the live branch's entry guard — a copied fixture that fails the new guard silently routes into the about-to-be-deleted dead branch and stays green even after the real behavior breaks. When a check goes green unexpectedly fast or empty, dump what it actually ran over. Checker code itself must assert the expected shape and fail loud — a permissive fallback (`a || b`) inside a gate absorbs wrong assumptions and keeps passing.
106
74
 
107
75
  ## Tooling and Operational Safety
108
76
 
@@ -113,7 +81,6 @@
113
81
  - Never accept secrets through transcript- or history-logged channels.
114
82
  - When a secret must be supplied, provide a gitignored env slot, read the value only from the environment, verify its presence and format without echoing it, and advise rotating anything already pasted; assume a resource-creating call may echo the secret back in its success output — suppress or discard the response body, and treat an echoed secret as pasted (rotate).
115
83
  - Treat a coarse runtime signal — a failure label, a `ps`/process-inspection result, idle CPU with no output — as a hypothesis, and confirm the cause against the authoritative low-level evidence the mechanism emits before attributing blame or intervening: read the raw provider/skill log payload (e.g. `input_tokens:0` proves a pre-dispatch rejection that exonerates your content and your change), and confirm a config/env toggle reached a subprocess via a cheap artifact the gated branch emits rather than an unreliable `ps` env read. A multi-minute LLM or subprocess call at ~0% CPU with an output gap is the normal signature of I/O wait, not a hang — check process state and the call trace's in-flight duration before acting, so you do not abort healthy long-running work.
116
- - Before reasoning about what a branch contains or opening a PR, run `git fetch` and compute the range as `origin/<base>..HEAD`, never `<base>..HEAD` against the local tracking ref — on a shared repo the local base drifts behind the remote until you pull, silently inflating the diff with already-merged work; if the range is surprisingly large, suspect a stale base before suspecting the branch. Platform "mergeable" flags are computed against the base only — sibling PRs can each look clean yet conflict; before picking a merge order, diff their changed-file sets and simulate the sequence.
117
84
 
118
85
  ## Multi-Model Workflow
119
86
 
@@ -122,6 +89,7 @@
122
89
  - For work spanning multiple models or CLI agents, context resets and handoffs, unattended LLM batches (including orchestrated subagent fleets), or parallel worktree branches, read and use `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/cli-multi-model-workflow.md` as a scoped extension of this section.
123
90
  - For composing a prompt, packet, or tool description aimed at a specific model family — including cross-family review dispatch, porting a prompt written for an older model, or choosing a reasoning-effort level for a model family — read and use `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/gpt-prompting.md` for gpt-family targets and `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/claude-prompting.md` for claude-family targets as scoped extensions of this section.
124
91
  - Allocate models by difficulty × blast radius, not phase name; when implementation ran on a cheaper tier, compensate by raising reviewer effort or adding a reviewer kind — never economize on implementation and verification at once.
92
+ - Judge a review by how much independence it actually bought, per reviewer and in this order: different provider, then different model, then strictly higher effort, then the two-perspective floor. A lower effort earns nothing — cheaper is not another perspective. Isolation is a gate rather than a rung: a reviewer you cannot show ran in a fresh context is not a weak review but no review, so exclude it instead of grading it low. Several ready methods are coverage, not proof the perspectives differed; and a clean verdict is PROPOSED until a receipt evidences a fresh dispatch of the declared packet on the exact seat, since a model echo is not evidence.
125
93
  - When the user asks for design AND two or more providers are reachable at frontier tier, run dual-provider frontier design drafts: two independent drafts from the same blind packet, one per provider, compared and synthesized into the working draft. The consent gate is about metered spend, not the fan-out: a provider reached via an OAuth session (subscription-covered, no marginal cost) proceeds WITHOUT asking — if a non-main-context OAuth frontier provider exists, just run the dual-provider design; do not ask. Explicit per-request approval (never standing) is required ONLY before dispatching to a provider reachable solely via a metered API key, and it approves that spend. If withholding un-approved API spend leaves fewer than two providers, run single-provider rather than blocking the design on approval. Inject the corpus design principles (concept economy, LLM/capability boundary, staged workflow) into every dispatched design packet — an external model does not load this corpus.
126
94
  - Never retry-storm a live rate limit: give unattended batches you author a code-level circuit breaker with per-item completion tracking (thresholds, backoff, and dead-letter rules in the guide); for third-party dispatchers, confirm equivalent protection exists or attend the run.
127
95
  - On any resumed, cleared, or relocated session, re-verify where you are (pwd; in a repo, branch and HEAD) before acting on prior-session assumptions — against the pinned handoff state when one exists.
@@ -129,11 +97,7 @@
129
97
  ## Documentation Hygiene
130
98
 
131
99
  - Keep runtime code, active docs, and execution-facing docs focused on current behavior, current decisions, current contracts, current authority, and current failure handling.
132
- - Use comments for non-obvious current behavior, invariants, constraints, or risks that still apply.
133
- - Put backward-compatibility notes, deprecated behavior, migration rationale, historical alternatives, change narratives, and handoff logs in isolated documentation paths such as `docs/`, `design/`, `archive/`, or `deprecated/`.
134
- - Link from active docs or code to isolated notes only when the current task needs that history or the reference helps future maintainers.
135
- - Phrase guidelines as desired behavior and preferred patterns.
136
- - Prefer established docs such as `CHANGELOG.md`, `IMPLEMENTATION_MAP.html`, or handoff notes for change history and implementation context.
100
+ - For where a comment, a compatibility note, deprecated behavior, a rejected alternative, a migration rationale, a change narrative, or a handoff log belongs — how to phrase a rule others will follow, and whether active docs should link to history — read and use `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/documentation-hygiene.md` as a scoped extension of this section.
137
101
 
138
102
  ## Visual Explanations
139
103
 
@@ -137,5 +137,5 @@ pick the path.
137
137
  Derived from the vendor's published guidance for the `targets` models above.
138
138
  When a `targets` model changes, re-derive this guide from current vendor
139
139
  guidance rather than editing around the old rules — prompting guidance is
140
- version-bound. `scripts/check-prompting-targets.sh` fails when the launch config
140
+ version-bound. `launch/check-prompting-targets.sh` fails when the launch config
141
141
  binds a model this guide does not list.
@@ -57,7 +57,7 @@ Delegate execution, not decisions. A unit is delegable only when it is decision-
57
57
  - Use a resident teammate only for dependent slices in one burst. Verify that the CLI preserves its model and context; resume-after-completion may silently change both. Retire after the burst or cache TTL, and persist durable knowledge in files.
58
58
  - After a discard or direction change, respawn once a routine round costs about as much as a fresh slice. Recover unique in-flight state to files first.
59
59
  - Redirects to busy workers may queue rather than preempt. Check artifacts before destructive redirects, phrase them conditionally, and stop an actively harmful worker by scoped PID/worktree authority.
60
- - Idle/progress notifications are hypotheses; verify repo artifacts before re-dispatch. Cross-reset state belongs in files, not task boards or transcripts. When polling concurrent async jobs, pin the exact id/handle received at dispatch — a "latest" convenience selector can silently point at a sibling job and return plausible-but-wrong results.
60
+ - Idle/progress notifications are hypotheses; verify repo artifacts before re-dispatch. An idle signal is liveness decoupled from the report: a subagent can go idle without ever delivering its result, so idle-without-report is not done — request the report explicitly rather than waiting. Cross-reset state belongs in files, not task boards or transcripts. When polling concurrent async jobs, pin the exact id/handle received at dispatch — a "latest" convenience selector can silently point at a sibling job and return plausible-but-wrong results.
61
61
  - Give reviewers/subagents a read-only diff, snapshot, or isolated worktree — not the live tree the main is editing — and forbid destructive git ops (checkout --, reset --hard, stash, clean) on any tree with uncommitted work; re-verify tree integrity before trusting results produced mid-edit.
62
62
  - Review cost scales with the diff, so layered review preserves delegation savings. Lower reviewer tier before dropping a review kind.
63
63
 
@@ -108,11 +108,28 @@ Instruction/config reach is per invocation. A rule in AGENTS.md cannot bind a he
108
108
  - A silent/dead lens is incomplete, never clean. Confirm liveness from usage/error/report evidence; rerun, swap provider, or report PROPOSED.
109
109
  - Kind labels do not guarantee distinct backends: wrappers and rate-limit fallbacks can silently route two "different-kind" verifiers to the same model/provider. Before trusting diversity on a high-stakes verdict, confirm each verifier's actual backing model from live process or usage evidence; on collapse, treat the pair as one kind and label PROPOSED.
110
110
 
111
+ ### Review Independence
112
+
113
+ How much independence a review actually bought, as an ordinal grade per reviewer rather than a global cross/same flag. Given the main seat `M` and the reviewer seat `R`:
114
+
115
+ | Grade | When |
116
+ |---|---|
117
+ | `provider_difference` | `R.provider != M.provider` |
118
+ | `model_difference` | same provider, `R.model != M.model` |
119
+ | `higher_effort` | same provider and model, `R.effort` strictly above `M.effort` |
120
+ | `perspective_floor` | otherwise — still a real review |
121
+
122
+ - Only upward counts. A different-but-**lower** effort earns nothing and lands on the floor: cheaper is not another perspective.
123
+ - **Isolation is a gate, not a rung.** A reviewer that cannot be shown to run in a fresh context is excluded entirely (`NOT_REVIEW`), never graded low — an in-context "review" is the failure this ladder exists to make visible, so it must not appear as a weak pass. Isolation is realised per mechanism: a fresh read-only subprocess, a hermetic profile, a stdio tool call in a fresh session, a headless host workflow, or a stateless API call. If none of these can deliver the required seat, the review did not happen.
124
+ - The floor still requires **at least two distinct perspectives**; one pass on the main's own seat is self-review with extra steps.
125
+ - Multiple ready methods are **coverage, not diversity**. Distinct labels do not prove the perspectives differed.
126
+ - **Achieved is not available.** What can be projected before a review runs is `projected`; a clean verdict without a receipt evidencing a fresh dispatch, the declared packet, a non-empty result and the exact seat is `PROPOSED`, never ACHIEVED. A model echo is not a receipt.
127
+
111
128
  ## Dual-Provider Design Drafts
112
129
 
113
- - Trigger: the task is design (the staged-workflow guide's design stages) AND two or more providers are reachable at frontier tier. Reachability via an OAuth session is subscription-covered — no marginal spend, so no approval and no question: if a non-main-context OAuth frontier provider exists, proceed with the dual-provider design directly. The consent gate applies ONLY to a provider reachable solely via a metered API key: dispatching to it needs the user's explicit per-request approval of that spend (per-request, not standing — an old approval does not carry to the next design). If the only way to reach a second provider is un-approved metered API spend, stay single-provider rather than blocking the design.
130
+ - Trigger: the task is design high-level shape and implementation process, before any code — AND two or more providers are reachable at frontier tier. Reachability via an OAuth session is subscription-covered — no marginal spend, so no approval and no question: if a non-main-context OAuth frontier provider exists, proceed with the dual-provider design directly. The consent gate applies ONLY to a provider reachable solely via a metered API key: dispatching to it needs the user's explicit per-request approval of that spend (per-request, not standing — an old approval does not carry to the next design). If the only way to reach a second provider is un-approved metered API spend, stay single-provider rather than blocking the design.
114
131
  - Mechanics: compose ONE blind packet (evidence, constraints, rubric, neutral alternatives — the escalation-gate packet shape) and dispatch it unchanged to one frontier-tier model per provider; drafts stay independent — neither sees the other's output. Then adjudicate: compare the two dual-provider frontier design drafts against the rubric, take the winner as the skeleton, graft the loser's superior parts, and record what differed and why the synthesis chose as it did (FRONTIER disposition line).
115
- - Packet injection: a dispatched designer is hermetic — it reads only its packet and never loads this corpus. Inject the design principles the corpus would have supplied: concept economy (reuse/extend/rename/split, compact concept graph), the LLM/tools-code capability boundary, the staged-workflow design rules (smallest viable path, falsifiable done-when), and any domain-specific principles the design touches. A draft produced without the principles is not comparable to one produced with them.
132
+ - Packet injection: a dispatched designer is hermetic — it reads only its packet and never loads this corpus. Inject the design principles the corpus would have supplied: concept economy (reuse/extend/rename/split, compact concept graph), the LLM/tools-code capability boundary, the staged design rules (smallest viable path, falsifiable done-when), and any domain-specific principles the design touches. A draft produced without the principles is not comparable to one produced with them.
116
133
 
117
134
  ## Unattended Batch Safety
118
135
 
@@ -154,7 +171,7 @@ Write for the next agent and re-verification, not narrative. Required content:
154
171
 
155
172
  ## Environment Binding (edit per environment)
156
173
 
157
- This is the human-readable projection of concrete models/tools; `config/agent-launch.toml` is the machine launch authority and parity checks keep them aligned. Re-probe when the binding is older than ~8 weeks or a newer observable model/tool changes the surface. `agent-bios install` overwrites deployed bindings, so edit the repo copy.
174
+ This is the human-readable projection of concrete models/tools; `launch/agent-launch.toml` is the machine launch authority and parity checks keep them aligned. Re-probe when the binding is older than ~8 weeks or a newer observable model/tool changes the surface. `agent-bios install` overwrites deployed bindings, so edit the repo copy.
158
175
 
159
176
  Binding (2026-07-25):
160
177
 
@@ -178,6 +195,7 @@ Codex direct-drive (verified 0.144.1, 2026-07-12):
178
195
  - HELM is instructed to dispatch tiers through internal `codex-run`, which pins model/effort/sandbox. FRONTIER uses a separate `gpt-5.6-sol`, read-only root: max by default, Ultra for divisible work, lower effort when cost/latency dominates. Nested multi-agent is enabled only for Ultra. Native `codex exec` spawn cannot pin role/effort.
179
196
  - This is an instruction-backed, live-E2E-verified default, not a security boundary: main bypass and arbitrary expert `-c` remain available by design. `frontier.toml` omits fixed effort for native surfaces that accept overrides.
180
197
  - `codex-run` owns reach, stdin, schema, profiles, expert `-c`, channel preservation, and exit status. Keep it internal.
198
+ - `claude-run` is its Claude-side twin and the command a composable review contract names for a panel dispatch on that host. Same shape: prompt on stdin, final message on stdout, exit status mirrored, `--model`/`--effort` pinning the seat, everything it does not recognise forwarded to `claude`. It denies the mutating tools by default, which is not the OS-level sandbox `codex-run` gets — do not read the two defaults as equivalent guarantees. Dispatch whatever command the contract names rather than the bare CLI: only the adapter can report what the dispatch actually did, and a review with no receipt stays PROPOSED.
181
199
 
182
200
  Dispatch packets:
183
201
 
@@ -7,7 +7,7 @@ use_when:
7
7
  - architecture changes, new features, cross-module or ontology changes, review-driven fixes
8
8
  - the user asks to design ("설계") before implementation
9
9
  - judging materiality of review findings and deciding when to stop or redesign
10
- - choosing the per-domain verification mix (code, ontology, config/data, spreadsheets, docs)
10
+ - deciding where a stage's verification points and review gates belong in the work plan
11
11
  ---
12
12
 
13
13
  # Coding Guidelines: Staged Workflow
@@ -16,7 +16,12 @@ This guide is a scoped extension of the global Coding Guidelines. Use it for mea
16
16
 
17
17
  It operates inside the existing global rules for requested scope, concept economy, LLM/tools/code boundary, verification discipline, and documentation hygiene.
18
18
 
19
- For trivial edits, use the lightweight inspect-edit-verify path from the global Coding Guidelines.
19
+ For trivial edits, use the lightweight inspect-edit-verify path, defined here in full:
20
+ read the surface you are about to touch before changing it; if the request admits more
21
+ than one reading, state the reading you act on; make the surgical edit — every changed
22
+ line traceable to the request, adjacent code left alone; verify with the narrowest
23
+ reliable check that would fail if the edit were wrong; clean up only what your own
24
+ change introduced. The stages below are for work that outgrows that sentence.
20
25
 
21
26
  When the user asks to "설계" or design, stay in design mode. Focus on high-level design and implementation-process design, then present the plan, tradeoffs, review gates, and implementation trigger. Move to implementation after the user asks to implement or approves the plan.
22
27
 
@@ -32,8 +37,45 @@ When the user asks to "설계" or design, stay in design mode. Focus on high-lev
32
37
  2. Implementation-process design: turn the design into an ordered work plan with dependencies, verification points, review gates, and redesign triggers.
33
38
  3. Implementation: make the smallest viable functional changes that satisfy the approved design and process plan.
34
39
 
40
+ **Minimum** limits surface area, configuration, abstractions, optional scope, and implementation spread. It must not reduce required behavior, runtime authority, evidence quality, or verification depth — a change that ships less of those is not smaller, it is less finished, and "smallest viable" becomes the excuse rather than the discipline.
41
+
42
+ **Viable** means real behavior: the change runs against real inputs, real authority, and the intended runtime path. Mocks belong to tests, fixtures, and explicitly requested simulations — a mock-backed path supports verification but does not count as product completion, even when nothing labels it a mock; the full boundary lives in `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/mock-realization-boundary.md`.
43
+
44
+ Define the success criteria before multi-step work starts, then verify against those criteria rather than against what you ended up building. Criteria written afterwards describe the implementation, so they cannot fail it.
45
+
35
46
  When simplifying a pipeline, moving processing downstream and dropping captured source fields are separate decisions: relocation is free simplification, but reducing captured information is riskier and needs explicit confirmation — "no current consumer" is not evidence of no future value.
36
47
 
48
+ ## Making the change
49
+
50
+ Deciding what to build and building it are different disciplines. This one is about leaving a
51
+ change that reads as the change that was asked for.
52
+
53
+ **State the assumption before you act on it.** Most wasted implementation is not a wrong answer to
54
+ the question; it is a right answer to a question nobody asked. Where the request admits more than
55
+ one reading, say which one you took and keep going — surfacing it early costs a sentence, and
56
+ surfacing it after the work costs the work.
57
+
58
+ **Surgical means legible, not minimal.** Touch what the request requires, follow the file's
59
+ existing style rather than your preferred one, and leave adjacent code alone even when it is
60
+ worse than what you are adding. A diff carrying an unrequested refactor forces the reviewer to
61
+ separate the two by hand, and the improvement is the part that gets dropped when they run out of
62
+ patience.
63
+
64
+ **Clean up what this change introduced, and only that.** Dead code, unused imports, and debris
65
+ your own edit created belong in the same change. Debris you found belongs in a sentence: name it
66
+ so it is visible, and leave it where the person who owns it can decide.
67
+
68
+ **For a bug, reproduce before you fix, when that is practical.** A test written after the fix
69
+ proves the code does what it now does. A test written before proves you understood the failure —
70
+ and it is the only version that can tell you the fix was unnecessary, or that it addressed a
71
+ different bug than the one reported.
72
+
73
+ **Fix the cause at its authority, not the symptom where it shows.** Two signals say you are
74
+ patching downstream: compensating code keeps accumulating around bad inputs, and each fix reveals
75
+ another instance of the same defect. The first says go upstream to where the value is produced.
76
+ The second says the instances are a class — single-source the value and fix the class, because
77
+ patching them one at a time is a queue that refills.
78
+
37
79
  ## Review Loop
38
80
 
39
81
  - At each stage, run review loops as appropriate: self review, subagent review when available, and structured multi-lens review when the repository or domain supports one (concrete tool: Environment Binding below).
@@ -45,25 +87,17 @@ When simplifying a pipeline, moving processing downstream and dropping captured
45
87
  - Treat low and info as non-blocking unless requested or promoted by new evidence.
46
88
  - When a document declares sections co-authoritative for a rule (fixture blocks, conformance appendices), treat every occurrence as one replicated value: propagate edits to all declared locations in the same pass and check propagation completeness explicitly in review.
47
89
 
48
- ## Verification Menus
49
-
50
- Per-domain menus for the global Verification Discipline loop; pick the narrowest reliable mix that proves the changed behavior, meaning, or contract.
90
+ ## Verification
51
91
 
52
- - Code: a layered mix of unit tests, integration tests for E2E segments, targeted E2E for changed flows, and full E2E for release or high-risk changes.
53
- - Ontology: static graph checks, concept economy gates, changed-path integration checks, and competency-question E2E checks.
54
- - Config or data: real parsers, schema checks, fixture validation, and sample transformations.
55
- - Spreadsheets: static workbook checks, fixture-based output checks, cross-sheet flow checks, visual/layout checks, and real Microsoft Excel engine recalculation for formula-dependent results.
56
- - Docs: links, terminology, current behavior alignment, and references to isolated historical notes.
57
- - Release or distribution: after publishing to multiple independently writable channels (signed manifest, object storage, release host, embedded updater), digest-verify every referenced object against the staging original per channel — publish success and upload order are not evidence — and run the real installer/updater through its default path.
58
- - A/B or on/off measurements: before accepting a null result, verify the arms actually received different treatment in the mechanism under test — a shared default or unconditional upstream step can silently apply the treatment to both arms.
59
- - Model-behavior guardrails: verify by changed behavior, not recitation — a staged battery from named-trigger cases through disguised, deconfounded, category-wide, and single-variable framings; a clean pass means "no known defect", so re-run the battery when the model changes.
60
- - Branch/version test builds against real data: explicitly separate every state sink the app touches (files, DB, OS-level stores that ignore env overrides), confirm the launch path propagates the isolation to child processes, and back up live data before the first run — a mismatched schema that drops unknown fields on write is data loss, not a no-op.
61
- - Irreversible capture switches: when activation itself has unreproducible cost (a capture window that cannot be replayed), prove the downstream consumption path against existing samples before enabling — reversibility of the code path alone is not enough.
92
+ Verification is a subject of its own depth, per-domain menus, deriving the case space, and what
93
+ a green is worth. It lives in `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/verification-discipline.md`.
94
+ Run it at each stage's verification points, and use its Verification Menus to pick the mix.
62
95
 
63
96
  ## Stop Conditions
64
97
 
65
98
  - If the issue boundary expands compared with the previous review, stop and ask the user to choose redesign/rework or continuing the current iteration.
66
99
  - Consider the boundary expanded when review reveals a broader affected purpose, failure condition, impact area, concept boundary, architecture boundary, or severity class.
100
+ - When review rounds keep producing material findings, classify each before fixing: a regression the previous round's own fix introduced, or a fresh instance of one pre-existing root cause. Regressions say tighten the increment and continue; recurring instances with zero regressions say instance-patching is a refilling queue — trigger the redesign-versus-continue stop above.
67
101
  - Before calling the work done, report the current stage, review results, remaining material issues, verification results, and any stop reason.
68
102
 
69
103
  ## Environment Binding (edit per environment)