model-orchestrator 0.1.35 → 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (108) hide show
  1. package/AGENTS.md +31 -21
  2. package/CHANGELOG.md +43 -1
  3. package/README.md +127 -110
  4. package/bin/README.md +57 -6
  5. package/bin/aunx.js +7 -0
  6. package/bin/cli-run.mjs +21 -15
  7. package/bin/cli.js +376 -257
  8. package/docs/README.md +15 -18
  9. package/docs/catalog.md +228 -38
  10. package/docs/companions.md +28 -10
  11. package/docs/guarantees.md +21 -12
  12. package/docs/how-it-routes.md +49 -42
  13. package/docs/install.md +135 -33
  14. package/docs/part-1-beginner.md +37 -45
  15. package/docs/part-2-intermediate.md +34 -52
  16. package/docs/part-3-advanced.md +36 -26
  17. package/docs/security-review-history.md +38 -0
  18. package/llms.txt +24 -25
  19. package/package.json +15 -8
  20. package/proof/README.md +100 -0
  21. package/proof/gate-demo.cast +9 -0
  22. package/proof/gate-demo.gif +0 -0
  23. package/proof/results.json +198 -0
  24. package/proof/scripts/check-gate.js +26 -0
  25. package/proof/scripts/install-time.js +16 -0
  26. package/proof/scripts/lib.js +73 -0
  27. package/proof/scripts/measure.js +15 -0
  28. package/proof/scripts/missing-results.js +30 -0
  29. package/proof/scripts/record-gate.js +38 -0
  30. package/proof/scripts/render.js +18 -0
  31. package/proof/scripts/runner-overhead.js +21 -0
  32. package/src/README.md +9 -3
  33. package/src/activation-ownership.js +19 -0
  34. package/src/apply-companions.js +104 -0
  35. package/src/apply-snippets.js +60 -28
  36. package/src/aunx.js +262 -0
  37. package/src/catalog.js +253 -117
  38. package/src/install.js +478 -209
  39. package/src/plugin.js +13 -4
  40. package/src/postinstall.js +57 -0
  41. package/src/roles.js +184 -0
  42. package/src/uninstall.js +125 -8
  43. package/templates/README.md +19 -2
  44. package/templates/advanced/README.md +2 -2
  45. package/templates/advanced/vm/PRIVACY_GATES.md +17 -19
  46. package/templates/advanced/vm/README.md +25 -20
  47. package/templates/advanced/vm/box-CLAUDE.md +19 -18
  48. package/templates/advanced/vm/jobs/README.md +3 -1
  49. package/templates/advanced/vm/jobs/weekly-audit.service +3 -0
  50. package/templates/advanced/vm/jobs/weekly-audit.sh +2 -2
  51. package/templates/advanced/vm/setup-vm.sh +49 -2
  52. package/templates/agents/README.md +2 -2
  53. package/templates/agents/agy/README.md +20 -3
  54. package/templates/agents/agy/builder.md +11 -7
  55. package/templates/agents/agy/bulk-worker.md +9 -7
  56. package/templates/agents/agy/code-reviewer.md +13 -7
  57. package/templates/agents/agy/deep-planner.md +10 -7
  58. package/templates/agents/agy/done-verifier.md +13 -22
  59. package/templates/agents/agy/finding-verifier.md +14 -22
  60. package/templates/agents/agy/live-researcher.md +10 -7
  61. package/templates/agents/agy/reader.md +10 -12
  62. package/templates/agents/claude-code/README.md +18 -14
  63. package/templates/agents/claude-code/builder.md +10 -15
  64. package/templates/agents/claude-code/bulk-worker.md +8 -10
  65. package/templates/agents/claude-code/code-reviewer.md +11 -17
  66. package/templates/agents/claude-code/deep-planner.md +9 -11
  67. package/templates/agents/claude-code/done-verifier.md +12 -33
  68. package/templates/agents/claude-code/finding-verifier.md +13 -39
  69. package/templates/agents/claude-code/live-researcher.md +9 -11
  70. package/templates/agents/claude-code/reader.md +9 -18
  71. package/templates/agents/snippets/chat.md +9 -10
  72. package/templates/agents/snippets/claude-code.md +17 -18
  73. package/templates/agents/snippets/generic.md +9 -11
  74. package/templates/agents/snippets/route-gate.mjs +2 -2
  75. package/templates/agents/snippets/route-metrics.mjs +1 -1
  76. package/templates/agents/snippets/subagent-context.mjs +4 -4
  77. package/templates/beginner/ORCHESTRATOR.md +31 -36
  78. package/templates/beginner/README.md +1 -1
  79. package/templates/common/ACCEPTANCE_CHECKS.json +12 -0
  80. package/templates/common/CONTEXT.md +37 -0
  81. package/templates/common/DECISIONS.md +11 -0
  82. package/templates/common/README.md +24 -11
  83. package/templates/common/TASK_BRIEF.md +84 -0
  84. package/templates/common/protocols/README.md +14 -11
  85. package/templates/common/protocols/acceptance-checks.md +14 -0
  86. package/templates/common/protocols/build-protocol.md +91 -106
  87. package/templates/common/protocols/context-file.md +10 -0
  88. package/templates/common/protocols/decision-log.md +9 -0
  89. package/templates/common/protocols/deep-research.md +20 -34
  90. package/templates/common/protocols/docs-then-prove.md +13 -18
  91. package/templates/common/protocols/gap-analysis.md +15 -21
  92. package/templates/common/protocols/memory-and-record.md +21 -20
  93. package/templates/common/protocols/numbers-and-logic.md +20 -26
  94. package/templates/common/protocols/propagate.md +18 -27
  95. package/templates/intermediate/CLI-RUN.md +83 -113
  96. package/templates/intermediate/DELEGATION_MATRIX.md +9 -3
  97. package/templates/intermediate/README.md +3 -3
  98. package/templates/intermediate/RESEARCH_TRIAGE.md +23 -15
  99. package/templates/intermediate/ROUTING.md +54 -51
  100. package/templates/intermediate/TIERS.md +37 -76
  101. package/templates/tools/README.md +1 -1
  102. package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +1 -1
  103. package/docs/audit-brief.md +0 -148
  104. package/scripts/README.md +0 -7
  105. package/scripts/gen-catalog.js +0 -81
  106. package/scripts/gen-plugin.js +0 -16
  107. package/scripts/record-demo.sh +0 -45
  108. package/templates/common/TASK_BUNDLE.md +0 -56
@@ -1,143 +1,128 @@
1
- # Build Protocol
1
+ # Build protocol: from the request to a change in use
2
2
 
3
- **Three phases, eight stages, and every gate is a question that can be answered wrong.**
3
+ When a task builds, codes, implements, migrates or deploys, follow this procedure. For a lookup, prose edit, bulk classification or one-line configuration change, use the relevant routing rule and verify the result directly.
4
4
 
5
- Fires on any task that builds, codes, implements, migrates or deploys. Rough test: if it would earn a second-opinion audit or a tracker issue, it runs this.
5
+ When a check controls a decision, first demonstrate an input that makes it fail. When a tool reports a finding, reproduce it before treating it as evidence. Keep every step proportional to the approved scope.
6
6
 
7
- > **The one rule underneath:** a gate you cannot fail is not a gate. If a stage's exit reads like "confirm it looks good", it is written wrong and it will pass every time, including the times it should not.
7
+ ## Pre-build
8
8
 
9
- Three corollaries:
10
- 1. A check nobody has watched fail is not known to work. Prove a gate can go red before trusting green.
11
- 2. A tool's output is a claim, not a fact. Scanner findings, audit reports and exit codes get read and reproduced before they are repeated.
12
- 3. A gate that fires on unrelated things gets bypassed, and a bypassed gate certifies what it never checked.
9
+ ### 1. Frame the request and probe the runtime
13
10
 
14
- | Phase | Master question | Stages |
15
- |---|---|---|
16
- | 1 Pre-build | What exactly are we building, what do we need first, and what does this touch or break? | 0 Route · 1 Map · 2 Judge |
17
- | 2 Build | Is it secure, built on current code, and correct without hidden flaws? | 3 Build · 4 Scan · 5 Challenge · 5b Ship gate |
18
- | 3 Post-build | Did it land everywhere, is it proven against the real thing, and is it recorded? | 6 Verify · 7 Record |
11
+ - Quote each acceptance-critical clause from the user's request and write the observable property that must hold in the final artifact.
12
+ - Give each property an **acceptance check**: a command that exits zero when the property holds, or an explicit manual procedure. Store checks using `ACCEPTANCE_CHECKS.json` and `protocols/acceptance-checks.md`.
13
+ - Probe the tools, models, access and services available now. Use the tool's own command or capability listing and record which surface it covers.
14
+ - Freeze acceptance-critical availability facts as checks too. When a dependency or permission changes, rerun the affected probe before work that assumes it.
15
+ - When the approach depends on a risky assumption, spike it on the actual runtime before building. Carry the measured value, method and date into the task brief.
16
+ - When requirements conflict, mark the affected check unresolved and obtain a decision before dependent work proceeds.
17
+ - Record decisions as they are made with **Did / Why / Serves / Rejected** in `DECISIONS.md`; see `protocols/decision-log.md`.
19
18
 
20
- The two seams are the point. Pre-build to Build: nothing is written yet, changing your mind costs a conversation. Build to Post-build: the ship, the only irreversible step, the only one that needs an explicit human yes.
19
+ **Exit:** the ask, scope, checks, availability evidence and risky assumptions have named records and verifiers.
21
20
 
22
- ## Phase 1 · Pre-build
21
+ ### 2. Apply the ordering function
23
22
 
24
- ### Stage 0 · Route
25
- 1. Do we have everything needed to start? Every key, access path and asset **verified present by a live probe**, not assumed from a doc.
26
- 2. Is this actually a build, or a quick fix, doc edit, or question that needs no plan?
23
+ When choosing the next step or resource, apply this ordering:
27
24
 
28
- **Gate:** classified, and every required input confirmed to exist and work. Verify access here, never at ship time. A missing credential found at Stage 0 costs a message; found at Stage 5b it costs the session.
25
+ 1. Satisfy dependencies before their consumers.
26
+ 2. Check invalidators earliest: availability, permission and whether the build is needed.
27
+ 3. Use cheap deterministic probes before model judgment.
28
+ 4. Run independent units together within the available concurrency and scope.
29
+ 5. Put irreversible actions last.
29
30
 
30
- ### Stage 1 · Map
31
- 1. What does this touch, improve, scale, or replace? Docs, indexes, tracker issues, tool servers, devices, scheduled jobs, hooks, other repos.
32
- 2. What already-working thing could this break, and does something already do this?
31
+ When a step is unnecessary, record the skipped step and the evidence for skipping it. Select resources from the live probe when their capabilities serve the task.
33
32
 
34
- Four bounded questions, not four exhaustive scans. **The builder maps; the judgment tier does not.** Retrieval is mechanical and the builder holds the local tooling and the map stays in its context for the build. Paying top-tier rates for a file list is the most expensive routing mistake available.
33
+ ### 3. Research bounded questions
35
34
 
36
- **Gate:** a written map naming affected files, systems and issues.
35
+ - Write the questions, search ceiling and stop condition before research.
36
+ - Read current official documentation for changing external interfaces and inspect relevant open-source approaches.
37
+ - Vet a repository before reading its implementation: reputation, usage, age, maintenance, author and answered issues.
38
+ - When signals are strong in a hard domain, surface adoption versus reimplementation to the user before adding the dependency.
39
+ - When signals are adequate, extract the technique and its limits. When signals are weak, exclude the source and record why.
40
+ - Return the technique, what it does, its implementation location and licence. Treat documentary claims as assumptions until runtime or source verifies them.
41
+ - When research could remove the need for the build, complete it before mapping. Otherwise run it beside independent mapping work.
37
42
 
38
- ### Stage 2 · Judge (Checkpoint 1)
39
- Ask the judgment tier, on the finished map:
40
- 1. Is this the simplest way to build it, or are we overcomplicating?
41
- 2. What is most likely to go wrong, and what did the request miss?
43
+ ### 4. Write one context file
42
44
 
43
- **Gate:** **one named weak spot** and **one gap in the request**. Approval alone is not an exit; an advisor asked only to approve will approve. If a consult comes back mostly restating the map, the brief asked it to retrieve when it should have asked it to decide.
45
+ - Read the project rules, relevant records and current task state within the authorized scope.
46
+ - Map the affected files, interfaces, users and connected surfaces with bounded searches.
47
+ - List necessary steps, recorded skips and the facts already answered by existing sources.
48
+ - Batch remaining questions that would change the build before dependent implementation.
49
+ - Write one `CONTEXT.md` per run. Every task brief points to that **context file**; verify a sample of its claims against source. See `protocols/context-file.md`.
44
50
 
45
- ## Phase 2 · Build
51
+ **Exit:** a reader can reconstruct the approved scope and current evidence from one context file.
46
52
 
47
- ### Stage 3 · Build
48
- "Up to date" means two things. Ask both.
49
- 1. Is our repo clean and current? No stray uncommitted work, on a branch, base ref recorded.
50
- 2. Am I writing against an external API or SDK, and have I read its actual source this session? Never from recall. Read the installed dependency, which is what actually runs.
51
- 3. Does the new code follow existing conventions and pass typecheck, tests, or dry-run?
53
+ ### 5. Assign by fit
52
54
 
53
- **Gate:** clean baseline recorded, every external-API claim traced to source read this session, build green. Check a reference clone's date before trusting it: a stale clone read as current is worse than no clone.
55
+ - Before each build dispatch, check the lane's live model roster and compare configured pins with its current default. Flag a pin that has fallen behind, then choose model and effort for this job.
56
+ - Select by reasoning depth, tool reach, context window and remaining capacity. Assign sections separately when different resources fit them better.
57
+ - Record the lane, model, effort and reason, including why a cheaper eligible route would be insufficient.
58
+ - Give each builder the whole scope, its own section, the context file, acceptance checks, holds/lacks inventory, measured assumptions, non-goals, interfaces to preserve, order of work and a coverage-table report contract. Use `TASK_BRIEF.md` (`aunx brief`).
59
+ - Run the availability probe and frozen checks before starting work that depends on them.
54
60
 
55
- ### Stage 4 · Scan (automatic, no judgment)
56
- 1. Any secret, key or token in the new code?
57
- 2. Any vulnerability or vulnerable dependency in the lines we added?
61
+ {{BUILDER_HANDOFF_NOTE}}
58
62
 
59
- Secret detection, static analysis and dependency scanning, filtered to lines this diff added. Refuses by default: a missing or erroring scanner exits non-zero, never a silent green.
63
+ ## Build
60
64
 
61
- **Gate:** zero flags on added lines. Pre-existing flags are reported, never inherited as blockers, and never waved through unread. A scanner finding is a claim; read the code before calling it anything.
65
+ ### 6. Implement and verify
62
66
 
63
- ### Stage 5 · Challenge (Checkpoint 2, one pass, never two)
64
- 1. Can bad input or a bad actor break it, and what happens when a dependency fails?
65
- 2. Did the build stick to the approved plan, or did unintended changes sneak in?
67
+ - Record the starting branch, base revision and existing changes before editing.
68
+ - Use the selected build model at high effort; use xhigh where supported for architecture, security or irreversible work. Choose the model from the current roster.
69
+ - Use the lane's available tools and authorized capabilities, including independent subagents when useful.
70
+ - When an external interface may have changed, consult current official sources and trace each API claim to a source read during this run.
71
+ - When a sandbox or permission boundary refuses a write, hand the required patch and evidence to an authorized writer and continue independent work. A refusal is a handoff, never a reason to widen permission silently.
72
+ - For every background task, arm a **heartbeat** at launch: check it every five minutes for liveness and output growth. After two consecutive checks with no growth, diagnose the stall, stop blind waiting and report the measured cause.
73
+ - Run checks appropriate to the change, including relevant secret, static and dependency scans. Read each finding and report scanner errors as unverified coverage.
66
74
 
67
- Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to a second-opinion reviewer, ideally a **different model family**. Architecture-shaped diffs go to the judgment tier reviewing build against plan. Never both on one diff.
75
+ ### 7. Split and merge
68
76
 
69
- Findings do not go straight to a repair. Hand them to the **finding-verifier**, whose job is to DISPROVE each one: read the cited line, state what would trigger it, then hunt for the guard, caller or test that makes it impossible. It returns one of three verdicts per finding, and rounding between them is the failure mode to watch for.
77
+ - When splitting the build, put the whole scope and all section owners in every brief.
78
+ - Give each lane its own write boundaries and shared interface contracts.
79
+ - Have the assigning agent merge the sections into one artifact and name every conflict and its resolution.
80
+ - Replay acceptance checks on the merged artifact. Use that artifact for the audit.
70
81
 
71
- | Verdict | Meaning | What happens next |
72
- |---|---|---|
73
- | CONFIRMED | reproduced, or a concrete path nothing blocks | it earns a repair |
74
- | NOT_REPRODUCED | something prevents it, named and located | dropped, and not narrated |
75
- | INCONCLUSIVE | not settleable read-only | say what it would take; never round it to either side |
82
+ ## Post-build
76
83
 
77
- Use a different model family from the one that produced the finding where you have one: a family asked to check its own claim tends to agree with itself.
84
+ ### 8. Audit once, with a companion consult
78
85
 
79
- **Gate:** every finding **verified** before it reaches a human, and only CONFIRMED findings trigger a change. Hard cap one re-audit. `CLEAN` is a valid success state; an auditor that is not allowed to say so manufactures something, and so does a verifier that is expected to confirm.
86
+ Run one audit pass against the final diff: build versus approved scope, correctness, and the final acceptance checks. The auditor is a different author and a different model family from the builder. When a suitable independent reviewer is unavailable, report the audit as pending and obtain the required review before shipping.
80
87
 
81
- ### Stage 5b · Ship gate
82
- 1. What is the rollback target? Record it before shipping.
83
- 2. Has the human authorised this specific change going live?
88
+ Skip this pass only when commands prove **all three** conditions: the diff is docs-only with no executable file touched; it changes no auth, secret, migration, route, mcp or policy path; and it is below the project's recorded line threshold. When a threshold or proof is absent, run the audit.
84
89
 
85
- **Gate:** rollback identifier written down, and an explicit yes. Authorisation is per change and does not carry over.
90
+ In the **same step**, run a companion consult asking a different question: scope versus the user's ask, including omissions and unnecessary work. This companion is a second reviewer, independent of optional companion software. If the assigning agent authored a section, have the companion cover that section's build against its scope as well.
86
91
 
87
- ## Phase 3 · Post-build
92
+ - Raise review effort with the measured diff scope; use xhigh where supported for auth, tokens, OAuth, middleware, routes, MCP, untrusted input, deletion or bulk mutation.
93
+ - Reproduce every finding or drop it. Preserve inconclusive findings with the evidence needed to settle them. `CLEAN` is a valid result.
94
+ - Assign confirmed fixes to a non-author and add one regression per fix, first shown failing against the unfixed behavior.
95
+ - Record dated ROUND 1 dispositions. Verify fixes with their regressions and acceptance checks; **no second audit pass** on the same diff.
88
96
 
89
- ### Stage 6 · Verify
90
- 1. Did it land across ALL connected surfaces? Re-grep the OLD identifier everywhere and expect zero except named historical records.
91
- 2. Can we prove it works against the real thing, including when a dependency fails, and has each check been *seen* to go red?
97
+ ### 9. Ship into use
92
98
 
93
- **Gate:** real-world test passes, the negative test behaves, every gate proven capable of failing. Prefer a local reproduction of the real fault over inducing it in production.
99
+ - Record the rollback identifier before shipping.
100
+ - Confirm the user's authorization covers the concrete change and destination. When it does, continue; when authorization is missing, present the checked result and request it.
101
+ - Replay all active acceptance checks against the final artifact immediately before the irreversible action.
102
+ - Wire the change into every approved surface that makes it usable, then verify live behavior on each surface.
103
+ - Treat shipped as a state: **in use**, with the carrying surfaces named and evidence that a real invocation reaches them.
104
+ - When rendering, packaging or deployment can change a checked property, replay the checks against the materialized artifact too.
94
105
 
95
- ### Stage 7 · Record
96
- 1. Where is the ONE doc that traces this end to end? Name the path.
97
- 2. What watches this thing? Name it, or write "nothing".
98
- 3. Are the docs, indexes, memory and tracker updated with evidence rather than claims?
99
- 4. Is the plan doc deleted?
106
+ ### 10. Record and report
100
107
 
101
- **Gate:** the end-to-end doc exists at a named path in the folder that owns the domain, with that folder's index corrected in the same pass (`protocols/memory-and-record.md`); the watcher is named or its absence is written down; the tracker is Done with evidence and read back; then the plan doc is deleted, not archived. "Nothing watches it" is a valid answer and usually the valuable one: writing it down turns an invisible gap into a tracked one.
108
+ - Write one end-to-end operations document for anything that runs independently or supports another system: purpose, trigger, invocation chain, dependencies, reads, writes, feedback loop, failure modes, manual verification and source of truth.
109
+ - Name what watches it; write `nothing` when there is no watcher.
110
+ - Update documentation and indexes, then attach artifact evidence to any authorized tracker update and read the final state back.
111
+ - Retire a temporary plan only after its operational knowledge is preserved under the project's retention policy.
112
+ - Return a coverage table with one row per requirement and evidence. State what was built, where it is in use, how future sessions reach it, and what remains unverified.
102
113
 
103
- ## Roles, as capabilities
114
+ ## Tools and fallbacks
104
115
 
105
- | Role | Does | Does not |
106
- |---|---|---|
107
- {{ROLES_BUILDER_ROW}}
108
- | Judgment tier | Stage 2 and the architectural arm of Stage 5. Argues with a finished map | Perform the retrieval |
109
- | Second-opinion reviewer | The security arm of Stage 5. Reviews the diff | Fix anything |
110
- | Mechanical gates | Stage 4 and any always-on guard | Be overridden without reading |
111
- | Cheap workers | Bounded sub-parts: bulk passes, wide searches, long loops | Own a stage |
112
- | Human | Stage 5b, and any irreversible or architectural call | Be the first line of review |
116
+ When optional companion software is selected, use codecalc for computation, Context7 for documentation, and obsidian-tc for the record store. When one is absent, use the local runtime or test suite, official documentation or installed source, and repository files with search and version control respectively. These substitutions preserve the checks. A missing capability needed by a check remains explicitly unverified until an authorized tool or person can verify it.
113
117
 
114
- {{BUILDER_HANDOFF_NOTE}}
118
+ ## Roles
115
119
 
116
- ## Checklist
117
-
118
- ```
119
- PRE-BUILD
120
- [ ] 0 Inputs and access verified by live probe, not assumed
121
- [ ] 0 Confirmed this is a build and not a quick fix
122
- [ ] 1 Everything it touches written: files, systems, issues
123
- [ ] 1 Asked what could break, and whether this already exists
124
- [ ] 2 Judgment tier named a weak spot AND a gap in the request
125
-
126
- BUILD
127
- [ ] 3 Repo clean, on a branch, base ref recorded
128
- [ ] 3 External API claims traced to source read this session
129
- [ ] 3 Typecheck / tests / dry-run green
130
- [ ] 4 Scan clean on ADDED lines; pre-existing flags read, not inherited
131
- [ ] 5 One audit pass, findings reproduced, plan drift reviewed
132
- [ ] 5b Rollback target recorded
133
- [ ] 5b Human authorised this specific ship
134
-
135
- POST-BUILD
136
- [ ] 6 Old identifier re-grepped everywhere, zero hits
137
- [ ] 6 Real-world test passes; negative test goes red
138
- [ ] 6 Every gate proven capable of failing
139
- [ ] 7 End-to-end doc exists at ONE named path
140
- [ ] 7 "What watches it" answered, even if the answer is "nothing"
141
- [ ] 7 Tracker Done with evidence, state read back
142
- [ ] 7 Plan doc deleted (only after the end-to-end doc exists)
143
- ```
120
+ | Role | Responsibility | Boundary |
121
+ |---|---|---|
122
+ | Assigning agent | Scope, resource selection, section ownership, merge and reporting | Granted scope |
123
+ {{ROLES_BUILDER_ROW}}
124
+ | Planning model | Resolve ambiguity using prepared context and evidence | Read-only planning |
125
+ | Independent reviewer | Audit the merged build against scope and acceptance checks | Review and report |
126
+ | Companion reviewer | Check scope against the ask in the same audit step | Independent question |
127
+ | Cheap workers | Bounded retrieval, classification and mechanical work | Assigned section |
128
+ | Human | Resolve missing mandates and authorize irreversible actions | User decisions |
@@ -0,0 +1,10 @@
1
+ # Context file: one shared source for a run
2
+
3
+ - When a build begins, scaffold `CONTEXT.md` with `aunx context` and fill it from the approved request and bounded source reads.
4
+ - Record the live state, affected files, acceptance checks, resource inventory, measurements, dependencies and resolved decisions once.
5
+ - When sending a task brief, point it at this context file and give the receiving worker access to the content.
6
+ - When a source fact changes, update the file and rerun the affected verifier before dependent work continues.
7
+ - Verify a sample of the context file's claims against original source paths or runtime output.
8
+ - When the task ends, record final evidence and unresolved limits where the next reader can find them.
9
+
10
+ When a shared record tool is available, use it within the granted scope. Otherwise keep the file with the project and use version control to preserve edits.
@@ -0,0 +1,9 @@
1
+ # Decision log: the choice and its evidence
2
+
3
+ - When deciding an approach, record **Did / Why / Serves / Rejected** in `DECISIONS.md` while the choice is current.
4
+ - Name what was chosen, the evidence supporting it, the user requirement it serves, and the alternatives set aside.
5
+ - Keep the explanation at the level a reviewer needs to assess the choice. Private chain-of-thought stays out of the record.
6
+ - When a decision changes, append its replacement with the new evidence and identify what it supersedes.
7
+ - When a workflow step is skipped, name the step and the reason in the same record.
8
+
9
+ When a record-store companion is absent, use a project Markdown file and version control. One writer owns the shared record; other lanes return proposals with evidence.
@@ -1,44 +1,30 @@
1
- # Deep research: parallel engines, then triage
1
+ # Deep research: bounded questions and verified sources
2
2
 
3
- Two research lanes, not three. The old "quick fact / known source / deep" split was ceremony.
3
+ When one lookup or a known source answers the question, retrieve it and answer directly. When the source set is unknown, several sources need reconciliation and the result will be cited later, use this procedure.
4
4
 
5
- | Lane | Entry condition | Output |
6
- |---|---|---|
7
- | **Search** | You can name the source, or one lookup answers it | Inline answer, no artifact |
8
- | **Deep** | All three: the source set is unknown, several sources must be reconciled, and the output must survive being cited later | A dated, cited artifact |
5
+ ## Prepare and run
9
6
 
10
- ## The shape of a deep run
7
+ 1. Write the research questions, source standard, authorized scope, search ceiling and stop condition.
8
+ 2. Break the question into independent sub-questions and check that each serves the user's ask.
9
+ 3. Vet sources before relying on them. For implementation sources, inspect reputation, usage, maintenance, authorship and licence.
10
+ 4. Run independent questions on available authorized engines, then open primary sources supporting material claims.
11
+ 5. Search existing records before adding a new durable report.
12
+ 6. Write one dated synthesis with citations, claim status and unresolved questions.
13
+ 7. Name the practical next step supported by the evidence: use, prototype, monitor, reject or take no action.
11
14
 
12
- ```
13
- 0 CHARTER what topics, what counts as a source, what is worth interrupting a human for
14
- 1 PLAN decompose into sub-questions <- highest leverage stage; a mis-scoped question
15
- produces a confident report about the wrong thing
16
- 2 RUN fan out to independent engines
17
- 3 TRIAGE reconcile disagreement against primary sources you open yourself
18
- 4 DEDUPE check what you already have BEFORE writing (semantic search; see memory-and-record.md)
19
- 5 BRIEF one dated artifact with marks (below)
20
- 6 ROUTE adopt / prototype / watch / pass / no action
21
- ```
15
+ ## Mark each claim
22
16
 
23
- ## Marks every claim carries
17
+ - **CONFIRMED:** verified against a primary source, with independent corroboration where required by the question.
18
+ - **DISAGREEMENT:** keep competing readings and name the evidence that would settle them.
19
+ - **REPORTED:** attribute a source's statement to that source.
20
+ - **UNVERIFIED:** name the absent evidence or access.
24
21
 
25
- - **CONFIRMED**: at least two independent engines agreed AND you opened the primary source.
26
- - **DISAGREEMENT**: engines conflicted. Record the verdict and the rejected reading. Never average.
27
- - **REPORTED**: a named person's post, a forum thread, a tool's self-report. Quoted, not trusted.
28
- - **UNVERIFIED**: plausible, single-source, or unsourced precision. Do not cite as fact.
22
+ When engines agree, inspect their shared premise. When they disagree, keep the difference visible until source evidence resolves it. When a figure affects a decision, compute it or verify its source.
29
23
 
30
- Agreement is weak evidence. Disagreement is the signal.
24
+ ## Select available research tools
31
25
 
32
- ## Level 1: one agent
26
+ {{RESEARCH_SELECTION_ADVICE}} Have one writer reconcile the outputs and maintain the durable report. Give every engine the context file, task brief and stopping condition.
33
27
 
34
- You still get the shape. Run PLAN as its own turn and inspect it before spending anything. Run the sweep. Then run a **fresh-context second-opinion turn** with a brief that says "question the premise; list what this report would get wrong if its sources were stale". Plant one deliberately wrong figure in the brief and see whether it corrects it: if it does not, its confirmations are worth less than they look. Mark every claim.
28
+ When only one engine is available, use a fresh context to question the sources and coverage, and report that limit. When Context7 is absent, use official docs or source. When a computing companion is absent, use the local runtime for numerical checks. When a record-store companion is absent, use project files and version control.
35
29
 
36
- ## Level 2 and up: selected engines, one triager
37
-
38
- {{RESEARCH_SELECTION_ADVICE}} The orchestrator triages: it opens the primary sources itself, marks each claim, and writes the brief. Only the orchestrator writes the durable record; every other engine proposes.
39
-
40
- Known failure shape: one engine will return confident unsourced numerics and claim full coverage. Downgrade those to hypothesis. The engines that report their own gaps honestly are the ones to weight.
41
-
42
- ## Measure
43
-
44
- Count **dispositions**, not briefs. A week that produced seven briefs and zero decisions is a failure.
30
+ When evaluating a research method itself, use a labelled false-premise fixture and confirm that the method rejects it. Keep fixture claims separate from the factual report.
@@ -1,26 +1,21 @@
1
- # Docs, then prove: documentation is a lead, never a verdict
1
+ # Docs, then prove: verify interfaces on the runtime
2
2
 
3
- **Why this is not `numbers-and-logic.md`.** That protocol is scoped to arithmetic, comparisons and complexity claims. This one is scoped to a different failure: trusting what a library, SDK, API or CLI is documented to do instead of checking what it actually does on this version, in this codebase. Two different mistakes, two different tools, one rule underneath both: a model that feels finished is not the same thing as a model that checked.
3
+ When coding against a changing library, SDK, API or CLI, read current documentation for the version in use, then run a check against the actual runtime.
4
4
 
5
- Companion tool for this rule: **Context7** (Upstash), {{CONTEXT7_STATUS}}. It pairs with **codecalc**, {{CODECALC_STATUS}}: Context7 tells the agent what the library is SUPPOSED to do (current, version-aware docs); codecalc runs the code and proves what it actually does. Docs never stand as proof on their own, and where a doc and a run disagree, the run wins and the source settles it.
5
+ Optional companions: **Context7** (Upstash), {{CONTEXT7_STATUS}}, for documentation; **codecalc**, {{CODECALC_STATUS}}, for execution checks.
6
6
 
7
- ## When calling is mandatory
7
+ ## Check an interface
8
8
 
9
- | You are about to | Use |
10
- |---|---|
11
- | write code against a library, SDK, API or CLI you have not confirmed the current signature for | Context7 (`resolve-library-id` then `query-docs`, or `ctx7 library` / `ctx7 docs`), then write the call |
12
- | claim a documented behaviour is what the code actually does | run it (`execute_code`, the project's own test suite, or a REPL), never the doc alone |
13
- | a doc and a run disagree | the run wins. Say so, and say what the doc got wrong (a stale version, a changed default, a deprecated flag) |
14
- | training-data recall of a library's API surface, with no source open | treat it as a guess until a doc or a run confirms it; version drift and renamed APIs are the common failure, not a rare one |
9
+ 1. Identify the library and installed or requested version.
10
+ 2. Read its current official documentation or source. With Context7, resolve the library ID before querying its docs unless a verified ID is already available.
11
+ 3. Write the smallest representative call and run it through the local runtime, test suite or codecalc.
12
+ 4. When docs and runtime disagree, inspect the implementation and version, record the difference, and base the build on verified behavior.
13
+ 5. Name the source, version, runtime check and result in the task evidence.
15
14
 
16
- Trivial, unversioned standard-library calls you would bet the build on are exempt. Anything with a version number attached to its behaviour is not.
15
+ ## Official-source and local-runtime workflow
17
16
 
18
- ## How to report a documentation-derived claim
17
+ When Context7 is absent, use the vendor's documentation, README, changelog or installed source. When codecalc is absent, use the project's own interpreter or tests. When neither path can verify the needed behavior, mark that assumption UNVERIFIED and route the probe to a lane with authorized access.
19
18
 
20
- Name the library, the version if Context7 returned one, and that it came from docs, not a run: "per Context7, `/vercel/next.js@14.0.0`'s middleware API takes X" reads differently from "Next.js middleware takes X", because the first can be checked against a version and the second cannot. If the claim was then verified by running it, say that too, and which one actually settled the question.
19
+ When reporting a documentation-only claim, label it as such. When a run verifies it, cite the check separately. Keep consequential arithmetic under `numbers-and-logic.md`.
21
20
 
22
- ## Without Context7
23
-
24
- The rule still binds. Read the vendor's own README, changelog or source before writing code against it, the same way this repository's own `AGENTS.md` asks: read the actual API, not a recollection of it. What is not allowed is code written against a remembered shape of a library that was never opened this session.
25
-
26
- Source: https://github.com/upstash/context7
21
+ Companion source: https://github.com/upstash/context7
@@ -1,28 +1,22 @@
1
- # Gap analysis: the second pass
1
+ # Gap analysis: compare scope with the request
2
2
 
3
- **On any comprehensive task, run a second pass that hunts for what is MISSING, not just verifies what is there.** Comprehensive means research, audits, plans, builds, and multi-file work.
3
+ When a task covers several requirements or surfaces, compare the result with the user's ask and the prepared map. For a build, run this as the companion consult in the same audit step; it asks scope versus ask while the auditor checks build versus scope.
4
4
 
5
- Verification asks "is what I did correct?". Gap analysis asks "what did I not do?". They are different questions and the second one is the one a first pass cannot answer about itself.
5
+ ## Check coverage
6
6
 
7
- ## The shape
7
+ 1. Enumerate the actual files, tests, configured lanes, jobs and sources within scope.
8
+ 2. Map every clause of the request to evidence in the result.
9
+ 3. Identify missing requirements, unintended additions and changes omitted from the map.
10
+ 4. Check the behavior expected when a dependency is absent or unavailable.
11
+ 5. Return one coverage row per requirement with IMPLEMENTED, PARTIAL, MISSING or OUT-OF-SCOPE and a source path or check.
8
12
 
9
- 1. **Enumerate what exists.** The live state, not the plan: files written, tests present, lanes configured, jobs scheduled, sources consulted.
10
- 2. **Diff it against the ask and the map.** Every clause of the original request maps to at least one thing you did. Every item on the Stage 1 map maps to a change. A clause with no step is dropped scope; a step with no clause is invented scope. Say both.
11
- 3. **Hunt the absences.** For each category: what would a reader expect to find here that is not here? What does the source say that the output does not? What fails if a dependency is down?
12
- 4. **Report the gaps as findings**, not as apologies. Each one: what is missing, where it should be, and whether you are closing it now or naming it as open.
13
+ ## Choose the reviewer
13
14
 
14
- ## Who runs it
15
+ - When an independent reviewer is available, give it the whole ask and scope with the final artifact.
16
+ - At level 2 and above, consider this selected lane: {{GAP_ANALYSIS_LANE}}
17
+ - When only one agent is available, use a fresh context for the coverage question and identify the independence limit. A build's required independent audit remains pending until an eligible reviewer can perform it.
18
+ - For a recurring capability check, enumerate live state and compare it with the documented configuration on the configured schedule.
15
19
 
16
- - **Level 1 (one agent):** the same agent, in a fresh turn, with a brief that says "you are looking for what is missing; do not re-verify what is present". Fresh context matters more than a different model.
17
- - **Level 2 and up:** {{GAP_ANALYSIS_LANE}}
18
- - **Level 3:** make it recurring. A weekly audit job enumerates live state (lanes, jobs, services, model lists), diffs it against the plan, and files a report. It catches the dead lane and the silently renamed model nobody noticed.
20
+ ## Resolve reported gaps
19
21
 
20
- ## The second half: analyze, compare, suggest
21
-
22
- Once the gaps are named, for each lane or component: what does it give, is there a cheaper, better or faster alternative, and what is the best paid option beside the free default. Parked is fine; unshown is not.
23
-
24
- ## Anti-patterns
25
-
26
- - Treating a green test suite as a gap analysis. Tests verify presence; they cannot see absence.
27
- - Asking the model that wrote the thing whether it is complete, in the same context. It will say yes.
28
- - Reporting gaps you did not verify. A gap is a finding; it needs the same evidence a bug does.
22
+ When a gap is reported, verify the claim against the artifact before assigning a change. Keep unverified observations separate. When a gap changes the approved scope, return the decision to the user; when it is already in scope, complete it and rerun its acceptance check.
@@ -1,30 +1,31 @@
1
- # Memory and record: the store is part of the change
1
+ # Memory and record: keep durable information findable
2
2
 
3
- Every protocol here ends in a write: the end-to-end doc, the rename that lands everywhere, the research brief, the gap report. A write that nothing indexes is a note in a drawer. This protocol says how the store is kept honest, whatever the store is.
3
+ When writing durable information, search existing records, update their index and use one writer for the shared record.
4
4
 
5
- Companion tool for this rule: **obsidian-tc**, {{OBSIDIAN_TC_STATUS}}.
5
+ Optional companion: **obsidian-tc**, {{OBSIDIAN_TC_STATUS}}.
6
6
 
7
- ## Rules
7
+ ## Record a change
8
8
 
9
- 1. **Search before you write.** A research brief, a decision, a rule: check whether it already exists (`semantic_search` for the concept, `search_text` for the exact phrase). Duplicates are how a store starts lying: two notes, two answers, and a reader picks one.
10
- 2. **The folder index is part of the change, not a follow-up.** Every folder has one index file that says what is in it and what state it is in. Any write, edit or delete reopens that index in the same pass and corrects whatever the change made untrue. A stale index is worse than a missing one because agents believe it.
11
- 3. **One writer per run.** Several agents may propose; one records. If you are not the writer, produce the file and name it in your report.
12
- 4. **Machine output stays out of the index.** Scan dumps, logs, traces embed well and outrank the thing they describe. Keep them outside the searchable store, or in a folder the index excludes.
13
- 5. **A record is not present state.** A note, a ticket, a checkbox is a dated observation. Re-read the live thing before you act on it.
14
- 6. **Inferred content is marked as inferred.** A conclusion an agent reached, rather than copied from a source, carries `source: agent-synthesis` (and, with obsidian-tc, goes through its poison scan before it lands). A reader must be able to tell a quote from a guess.
15
- 7. **Compare-and-swap on overwrite.** Read, then write with the hash you read. A blind overwrite of a note someone else changed is a lost update nobody notices.
9
+ 1. Search for the concept and exact phrase before creating a new record. When a matching record exists, update it within the granted scope.
10
+ 2. Read the owning folder's index and correct any description, status or link changed by the work.
11
+ 3. Assign one writer; have other lanes return proposed edits with evidence.
12
+ 4. Keep machine logs and raw traces outside the curated record index, or in an explicitly excluded folder.
13
+ 5. Before acting on a dated record, probe the current state it describes.
14
+ 6. Mark synthesis and inference distinctly from source quotations.
15
+ 7. Before overwriting, verify the current version or content hash. If it changed since the read, reconcile the changes first.
16
16
 
17
- ## Where the other protocols touch the store
17
+ ## Connect records to work
18
18
 
19
- | Protocol | Store call |
19
+ | Workflow | Record action |
20
20
  |---|---|
21
- | Propagate, step 1 | `get_backlinks` on the thing being renamed; `search_text` for the literal old term; after the change, `find_unresolved_links` |
22
- | Build, Stage 7 | the end-to-end doc goes in the folder that owns the domain; its index is updated in the same pass |
23
- | Deep research, step 4 | `semantic_search` before the brief is written; a hit means append to the existing note, not a second note |
24
- | Gap analysis | the "what exists" enumeration starts from the store, then diffs against the live state |
21
+ | Shared rename | Search references and backlinks, then check for unresolved links |
22
+ | Build report | Store the operations document in its owning folder and update the index |
23
+ | Research | Search existing coverage and attach new verified evidence |
24
+ | Coverage check | Compare recorded scope against the current artifact |
25
+ | Decision | Record Did / Why / Serves / Rejected with evidence |
25
26
 
26
- ## Without obsidian-tc
27
+ ## Local record workflow
27
28
 
28
- The rules still bind. `grep -rn` is your literal search, a folder README is your index, `git` is your compare-and-swap. What is not allowed is a write nobody can find again.
29
+ When obsidian-tc is absent, use scoped file search, a folder README and version control. Before editing, compare the file with the version you read; version control alone does not prevent a concurrent lost update. Keep writes within the task's authorized paths.
29
30
 
30
- Source: https://github.com/The-40-Thieves/obsidian-tc
31
+ Companion source: https://github.com/The-40-Thieves/obsidian-tc
@@ -1,35 +1,29 @@
1
- # Numbers and logic: compute, never guess
1
+ # Numbers and logic: compute decision inputs
2
2
 
3
- **Never do arithmetic in your head when being wrong would matter.** Money, rates, percentages, margins, budgets, token and cost counts, any comparison you are about to state, any complexity or equivalence or speedup claim, anything past 2^53. A wrong number that looks right is worse than no number, because it gets acted on.
3
+ When a number or logical claim affects a decision, compute it with an appropriate tool. This includes money, rates, percentages, budgets, token counts, comparisons, complexity, equivalence and performance claims. Trivial single-digit sums can be stated directly.
4
4
 
5
- A model that is confident about `0.1 + 0.2` does not feel uncertain. It feels finished. The tool exists so the feeling is not the check.
5
+ Optional companion: **codecalc**, {{CODECALC_STATUS}}.
6
6
 
7
- Companion tool for this rule: **codecalc**, {{CODECALC_STATUS}}.
7
+ ## Select the check
8
8
 
9
- ## When calling is mandatory
10
-
11
- | You are about to | Use |
9
+ | Claim | Verification |
12
10
  |---|---|
13
- | state a number someone will act on | `evaluate_expression` (exact rationals, no float drift) |
14
- | say A is bigger, cheaper, faster than B | compute both, then compare; never eyeball |
15
- | claim two programs behave the same (a port, a rewrite) | `verify_translation` |
16
- | claim an optimization preserved behaviour | `verify_optimization` |
17
- | state a Big-O, or that something scales | `analyze_complexity` (static) or `benchmark` (measured) |
18
- | assert a logical property holds, or that a set of constraints is satisfiable | `z3_check` (SMT) or `truth_table` |
19
- | run a snippet to see what it actually does | `execute_code` (31 languages, sandboxed) |
20
-
21
- Trivial single-digit sums are exempt. Everything else is not.
22
-
23
- ## How to report a computed figure
24
-
25
- Give the exact form and the decimal, and say which tool produced it. `37/210 = 17.62%` reads differently from `about 18%`; the first can be checked, the second cannot. If a tool returned `unenforced` or a grade below its top, say so beside the number.
26
-
27
- ## Logic flow
11
+ | A figure someone will act on | Exact arithmetic with a calculator, interpreter or spreadsheet |
12
+ | One option is cheaper, larger or faster | Compute both values and their comparison |
13
+ | A port preserves behavior | Relevant tests or `verify_translation` |
14
+ | An optimization preserves behavior | Regression checks or `verify_optimization` |
15
+ | Complexity or scaling | Source analysis or `analyze_complexity`; benchmark measured performance |
16
+ | A logical property or constraint holds | An appropriate solver such as `z3_check` or a truth table |
17
+ | A snippet behaves a certain way | Run it with `execute_code` or the project's local runtime |
28
18
 
29
- Reasoning scaffolds help models that have no native reasoning mode and add nothing to models that already think before answering. So: do not narrate a chain of thought as evidence. **The authorising evidence is the computed outcome, not the thought log.** When a decision hangs on a logical claim, encode the claim and check it (`z3_check`); when it hangs on behaviour, run it (`execute_code`); when it hangs on a figure, compute it. A trace that was never checked is a story.
19
+ When using a companion command, probe that tool's current schema before calling it. When codecalc is absent, use the local runtime, test suite, spreadsheet or a suitable calculator.
30
20
 
31
- ## Without codecalc
21
+ ## Report the evidence
32
22
 
33
- The rule still binds. Use whatever computes: a shell `python3 -c`, a spreadsheet, the vendor's built-in interpreter. What is not allowed is the number that came from nowhere.
23
+ - Give the input, exact result and useful decimal representation, with method and source.
24
+ - For measurements, name the script, sample size, environment and date.
25
+ - When a tool reports a limit or an unenforced property, state it beside the result.
26
+ - When a claim cannot be computed or checked from available evidence, mark it UNVERIFIED.
27
+ - When a decision depends on logic, report the checked property and result. Keep private chain-of-thought out of the evidence record.
34
28
 
35
- Source: https://github.com/The-40-Thieves/codecalc
29
+ Companion source: https://github.com/The-40-Thieves/codecalc