model-orchestrator 0.1.16 → 0.1.17
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +8 -1
- package/README.md +13 -9
- package/docs/README.md +1 -1
- package/docs/catalog.md +1 -1
- package/docs/part-1-beginner.md +4 -4
- package/docs/part-2-intermediate.md +4 -4
- package/llms.txt +1 -1
- package/package.json +1 -1
- package/src/catalog.js +1 -1
- package/src/install.js +7 -7
- package/templates/agents/agy/README.md +1 -1
- package/templates/agents/agy/builder.md +1 -1
- package/templates/agents/agy/finding-verifier.md +1 -1
- package/templates/agents/claude-code/finding-verifier.md +1 -1
- package/templates/agents/snippets/chat.md +2 -2
- package/templates/agents/snippets/claude-code.md +1 -1
- package/templates/agents/snippets/generic.md +1 -1
- package/templates/agents/snippets/route-gate.mjs +1 -1
- package/templates/beginner/ORCHESTRATOR.md +2 -2
- package/templates/common/README.md +1 -1
- package/templates/common/protocols/build-protocol.md +10 -10
- package/templates/common/protocols/deep-research.md +2 -2
- package/templates/common/protocols/gap-analysis.md +1 -1
- package/templates/common/protocols/propagate.md +2 -2
- package/templates/intermediate/CLI-RUN.md +3 -3
- package/templates/intermediate/DELEGATION_MATRIX.md +1 -1
- package/templates/intermediate/ROUTING.md +4 -4
- package/templates/intermediate/TIERS.md +12 -8
- package/templates/tools/codecalc/CODECALC.md +1 -1
- package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +3 -3
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,12 @@ All notable changes to this project are documented here. The format follows [Kee
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [0.1.17] - 2026-09-11
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
- **User-facing text now uses plain language instead of security-audit jargon.** Words like "risk", "attack lane", "adversarial", "blast radius", "fail closed" and "threat model" read as alarming to someone deciding whether to try the tool, so they scared off exactly the readers this project needs. No rule any of them described changed, only the wording: "risk" is now "stakes" everywhere it names a routing input (with a one-line definition added to `README.md` and `TIERS.md`), "attack lane" / "Stage 5 Attack" / "attack pass" are now "challenge lane" / "Stage 5 Challenge" / "challenge pass", "adversarial" (auditor, read, critique, turn, pass) is now "second-opinion", "blast radius" is now "everything it touches", "fail(s) closed" is now "refuses by default", and "threat model" is now "security notes" in the files that link to it. `test/prose.test.js` gained a permanent check (`no alarming security wording in user-facing text`) over the purely-prose, user-facing surface (`docs/`, `templates/`, `README.md`, `llms.txt`, `CONTRIBUTING.md`, the PR template) so the old wording cannot silently creep back in. `docs/audit-brief.md`, `SECURITY.md`, `CODE_OF_CONDUCT.md` and code identifiers/comments (for example the `ATTACK_LANE` render var) are unchanged, since these words are expected or load-bearing there.
|
|
12
|
+
|
|
7
13
|
## [0.1.16] - 2026-09-11
|
|
8
14
|
|
|
9
15
|
### Added
|
|
@@ -256,7 +262,8 @@ First release.
|
|
|
256
262
|
- Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
|
|
257
263
|
- Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
|
|
258
264
|
|
|
259
|
-
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.
|
|
265
|
+
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.17...HEAD
|
|
266
|
+
[0.1.17]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.16...v0.1.17
|
|
260
267
|
[0.1.16]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.15...v0.1.16
|
|
261
268
|
[0.1.15]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...v0.1.15
|
|
262
269
|
[0.1.14]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...v0.1.14
|
package/README.md
CHANGED
|
@@ -44,7 +44,7 @@ Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/pa
|
|
|
44
44
|
| Id | What | Level |
|
|
45
45
|
|---|---|---|
|
|
46
46
|
| `claude-code` | Claude Code CLI, the default orchestrator | 1+ |
|
|
47
|
-
| `codex` | Codex CLI on a ChatGPT plan: second coder
|
|
47
|
+
| `codex` | Codex CLI on a ChatGPT plan: second coder and second-opinion reviewer (a different model family reading your diff) | 1+ |
|
|
48
48
|
| `agy` | Antigravity CLI on a Google AI plan: research sweeps, concurrent fan-out | 1+ |
|
|
49
49
|
| `grok` | Grok CLI on X Premium: live X and web reads at $0 | 1+ |
|
|
50
50
|
| `hermes` | Hermes Agent: the free tier | 2+ |
|
|
@@ -173,7 +173,7 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
|
|
|
173
173
|
6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, only the brief carries what this task needs. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
|
|
174
174
|
7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
|
|
175
175
|
|
|
176
|
-
## Routing by role, complexity and
|
|
176
|
+
## Routing by role, complexity and stakes
|
|
177
177
|
|
|
178
178
|
Role picks the agent. Two more inputs move the choice, and they move it in
|
|
179
179
|
different directions, so `TIERS.md` states them separately rather than folding
|
|
@@ -182,10 +182,14 @@ them into the role:
|
|
|
182
182
|
- **Complexity moves the effort.** A worker executing a finished plan needs less
|
|
183
183
|
reasoning than the reviewer judging its output. When the plan is airtight the
|
|
184
184
|
spec is carrying the thinking.
|
|
185
|
-
- **
|
|
186
|
-
irreversible changes buy the
|
|
187
|
-
human yes. A one-line change to an auth check is simple and high-
|
|
188
|
-
same time, and it is the
|
|
185
|
+
- **Stakes move the tier and the reader.** Security, privacy, data loss and
|
|
186
|
+
irreversible changes buy the challenge lane, a named check, a rollback path or
|
|
187
|
+
a human yes. A one-line change to an auth check is simple and high-stakes at
|
|
188
|
+
the same time, and it is the stakes that decide.
|
|
189
|
+
|
|
190
|
+
Stakes means what a mistake would cost: a security hole, leaked personal data,
|
|
191
|
+
lost data, or something you can't undo. Most tasks are low-stakes and route
|
|
192
|
+
normally.
|
|
189
193
|
|
|
190
194
|
The top of the ladder is bought with evidence: a reproduced failure, an
|
|
191
195
|
unresolved checkpoint, an irreversible change. A task that merely feels hard is
|
|
@@ -221,7 +225,7 @@ The report prints turns, **route-marker coverage** (the percentage of turns whos
|
|
|
221
225
|
A lane with no `--model`, no `--effort` and no `defaults` entry in
|
|
222
226
|
`bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
|
|
223
227
|
CLI configured months ago at a low reasoning effort keeps auditing at that
|
|
224
|
-
effort while your routing docs describe
|
|
228
|
+
effort while your routing docs describe a second-opinion pass.
|
|
225
229
|
|
|
226
230
|
```bash
|
|
227
231
|
node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
|
|
@@ -244,7 +248,7 @@ Install for the tools you have, then let the generated `ROUTING.md` decide the t
|
|
|
244
248
|
|
|
245
249
|
### How do I route tasks to cheaper models?
|
|
246
250
|
|
|
247
|
-
The rules route by role, complexity and
|
|
251
|
+
The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly never gets a frontier token.
|
|
248
252
|
|
|
249
253
|
### Is this an LLM router or an AI gateway?
|
|
250
254
|
|
|
@@ -272,7 +276,7 @@ Add an AI to `src/catalog.js` and every prompt, table, config and doc picks it u
|
|
|
272
276
|
## Credits
|
|
273
277
|
|
|
274
278
|
- [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for
|
|
275
|
-
separating role, complexity and
|
|
279
|
+
separating role, complexity and stakes instead of compressing them into one
|
|
276
280
|
scale, for recording the model and effort a lane was actually asked for, and
|
|
277
281
|
for verifying findings before they trigger repairs. All three shipped in
|
|
278
282
|
0.1.14.
|
package/docs/README.md
CHANGED
|
@@ -7,7 +7,7 @@ The three parts, as reading. The installer writes the working files; these expla
|
|
|
7
7
|
| 1 Beginner | you use one LLM or one agent and want it to route well | [part-1-beginner.md](part-1-beginner.md) |
|
|
8
8
|
| 2 Intermediate | you have several AIs and want to call them through their CLIs from one orchestrator | [part-2-intermediate.md](part-2-intermediate.md) |
|
|
9
9
|
| 3 Advanced | you want the whole thing running unattended on a virtual machine | [part-3-advanced.md](part-3-advanced.md) |
|
|
10
|
-
| Audit brief | the
|
|
10
|
+
| Audit brief | the security notes and the two second-opinion audit rounds this shipped with | [audit-brief.md](audit-brief.md) |
|
|
11
11
|
| Catalog | what each AI and companion tool in the installer is for, how it installs, how it signs in | [catalog.md](catalog.md) |
|
|
12
12
|
|
|
13
13
|
Each part ends with "what the installer gives you at this level" so the doc and the files agree.
|
package/docs/catalog.md
CHANGED
|
@@ -24,7 +24,7 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
|
|
|
24
24
|
### `codex` · Codex CLI (OpenAI, ChatGPT plan)
|
|
25
25
|
|
|
26
26
|
- **Kind:** agent-cli · **Access:** subscription · **Lane:** A · **Level:** 1+
|
|
27
|
-
- **Wins at:** second coder and
|
|
27
|
+
- **Wins at:** second coder and second-opinion reviewer (a different model family reading your diff)
|
|
28
28
|
- **Install:** `npm install -g @openai/codex@0.153.4`
|
|
29
29
|
- **Sign in:** `codex login` (add `--device-auth` on a machine with no browser)
|
|
30
30
|
- **Reads rules from:** `AGENTS.md`
|
package/docs/part-1-beginner.md
CHANGED
|
@@ -14,7 +14,7 @@ If your agent exposes model choice (Claude Code, Codex, Antigravity), map the ti
|
|
|
14
14
|
|
|
15
15
|
Three cost levers, always together: tier (price per token), token discipline (how many tokens: read only what you will touch, never re-read, deliverables not narration), effort (how hard each call thinks).
|
|
16
16
|
|
|
17
|
-
And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **
|
|
17
|
+
And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **Stakes** move the tier and who reads the result: security, privacy, data loss and irreversible changes are the four worth naming, because none of their failures can be fixed by editing the code afterwards. A one-line change to an auth check is simple and high-stakes at once, and it is the stakes that decide.
|
|
18
18
|
|
|
19
19
|
Robustness first, cost second. You split tiers because the split produces better work.
|
|
20
20
|
|
|
@@ -32,9 +32,9 @@ Modifiers: plan big, execute small · never silently retry a failed attempt at t
|
|
|
32
32
|
|
|
33
33
|
> A gate you cannot fail is not a gate.
|
|
34
34
|
|
|
35
|
-
"Does this look good?" passes every time. "Name
|
|
35
|
+
"Does this look good?" passes every time. "Name what is most likely to go wrong, and what the request as filed missed" can come back empty, which is how you know it worked.
|
|
36
36
|
|
|
37
|
-
Every build gets two checkpoints. **Before writing:** you map what it touches and what could break, then ask the deep tier on the finished map for
|
|
37
|
+
Every build gets two checkpoints. **Before writing:** you map what it touches and what could break, then ask the deep tier on the finished map for one named weak spot and one gap in the request. **After it is green:** a fresh context challenges it (bad input, failing dependency, drift from the plan), allowed to answer CLEAN, every finding reproduced before it reaches you. Then the ship, with a rollback named and an explicit yes. Then the loud negative: re-check the old name everywhere and expect zero.
|
|
38
38
|
|
|
39
39
|
## 4. Every hand-off carries a brief
|
|
40
40
|
|
|
@@ -46,7 +46,7 @@ After anything comprehensive, a fresh turn that hunts for what is **missing**, n
|
|
|
46
46
|
|
|
47
47
|
## 6. Deep research, single agent
|
|
48
48
|
|
|
49
|
-
Plan the sub-questions as their own turn and inspect them before spending anything. Sweep. Then a fresh
|
|
49
|
+
Plan the sub-questions as their own turn and inspect them before spending anything. Sweep. Then a fresh second-opinion turn told to question the premise. Plant one deliberately wrong figure and see whether it corrects it. Mark every claim CONFIRMED / DISAGREEMENT / REPORTED / UNVERIFIED. Agreement is weak evidence; disagreement is the signal.
|
|
50
50
|
|
|
51
51
|
## 7. Numbers and logic are computed, never guessed
|
|
52
52
|
|
|
@@ -13,7 +13,7 @@ Rule: never spend a frontier token on a task a cheap tier finishes correctly. Es
|
|
|
13
13
|
| Lane | Wins at |
|
|
14
14
|
|---|---|
|
|
15
15
|
| the orchestrator (Claude Code, or whichever you chose) | routes, maps, builds, verifies, records; drives the others as CLIs |
|
|
16
|
-
| Codex | second coder and
|
|
16
|
+
| Codex | second coder and second-opinion reviewer: a different model family reading your diff |
|
|
17
17
|
| Antigravity `agy` | deep research sweeps; concurrent fan-out (its subagent call takes an array) |
|
|
18
18
|
| Grok CLI | X and live web reads at $0 (the same search on the API bills per call) |
|
|
19
19
|
| Hermes | the free tier: rough drafts, first-pass summaries, divergent reads, cron jobs |
|
|
@@ -28,7 +28,7 @@ Every agent CLI can report success and deliver nothing. `bin/cli-run.mjs` builds
|
|
|
28
28
|
|
|
29
29
|
Every call goes through it. "This lane is flaky" becomes a query over its log instead of an argument. `node bin/cli-run.mjs --doctor` is the first thing to run after install: enabled lanes, binaries on PATH, the route each lane is pinned to, and with `--run` a one-word canary per lane.
|
|
30
30
|
|
|
31
|
-
There is a second thing a lane can be quietly wrong about. Left unpinned, it runs on **its own config file**, which the runner cannot see: a CLI set up months ago at a low reasoning effort keeps auditing at that effort while your routing docs describe
|
|
31
|
+
There is a second thing a lane can be quietly wrong about. Left unpinned, it runs on **its own config file**, which the runner cannot see: a CLI set up months ago at a low reasoning effort keeps auditing at that effort while your routing docs describe a second-opinion pass, and no error is ever raised. `--model` and `--effort` pin it per call, `defaults` in `bin/lanes.json` pins it per lane, and every run records the value requested and where it came from (`flag`, `lanes.json`, `lane_default`). The log claims no actual: grok reports a model id in its output, the other four lanes report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string the durable log never holds.
|
|
32
32
|
|
|
33
33
|
## 4. Every delegation carries a task bundle, on both surfaces
|
|
34
34
|
|
|
@@ -36,7 +36,7 @@ Subagents and CLI lanes are close to the same problem: something that may hold n
|
|
|
36
36
|
|
|
37
37
|
## 5. Research: three engines, one triager
|
|
38
38
|
|
|
39
|
-
Fan the same plan to three model families (web sweep,
|
|
39
|
+
Fan the same plan to three model families (web sweep, second-opinion read, live data), each as one `cli-run` call. The orchestrator opens the primary sources itself, marks every claim, and writes the only durable record. Expect one engine to return confident unsourced numerics; downgrade it. Weight the engines that report their own gaps. Count dispositions, not briefs.
|
|
40
40
|
|
|
41
41
|
## 5a. A finding is a claim, not a fact
|
|
42
42
|
|
|
@@ -48,7 +48,7 @@ The second pass is now a different model reading the same artifact, in read-only
|
|
|
48
48
|
|
|
49
49
|
## 7. The build protocol, bound to lanes
|
|
50
50
|
|
|
51
|
-
Stage 1 Map: the orchestrator sweeps; CLI lanes critique the map at $0. Stage 2: deep tier,
|
|
51
|
+
Stage 1 Map: the orchestrator sweeps; CLI lanes critique the map at $0. Stage 2: deep tier, one named weak spot and one gap in the request. Stage 4: scanners on the added lines, refuses by default. Stage 5: security-shaped diff → the second coder in read-only audit mode; architecture-shaped → deep tier reviewing build against plan; never both. Two deep checkpoints per build; CLI lanes are uncapped.
|
|
52
52
|
|
|
53
53
|
## 8. Privacy gate
|
|
54
54
|
|
package/llms.txt
CHANGED
|
@@ -24,4 +24,4 @@ Levels: 1 beginner (one agent or chat app), 2 intermediate (several agent CLIs,
|
|
|
24
24
|
## Optional
|
|
25
25
|
|
|
26
26
|
- [Security policy](https://github.com/aunysillyme/model-orchestrator/blob/main/SECURITY.md)
|
|
27
|
-
- [Audit brief](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/audit-brief.md): the
|
|
27
|
+
- [Audit brief](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/audit-brief.md): the security notes and what has already been security-reviewed
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "model-orchestrator",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.17",
|
|
4
4
|
"description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
package/src/catalog.js
CHANGED
|
@@ -87,7 +87,7 @@ export const AIS = [
|
|
|
87
87
|
bin: 'codex',
|
|
88
88
|
access: 'subscription',
|
|
89
89
|
lane: 'A',
|
|
90
|
-
role: 'second coder and
|
|
90
|
+
role: 'second coder and second-opinion reviewer (a different model family reading your diff)',
|
|
91
91
|
minLevel: 1,
|
|
92
92
|
install: { npm: '@openai/codex', pin: '0.153.4' },
|
|
93
93
|
builtAgainst: '0.153.4',
|
package/src/install.js
CHANGED
|
@@ -155,21 +155,21 @@ export function laneVars(selected) {
|
|
|
155
155
|
if (has('hermes')) step0.push(`${cr('hermes')} (the free tier) for rough drafts and divergent reads`);
|
|
156
156
|
if (has('qwen')) step0.push(`${cr('qwen')} (the cheapest metered lane) for structured bulk, never for anything citing a line, number or source`);
|
|
157
157
|
if (has('grok')) step0.push(`${cr('grok')} for X and live web reads at $0`);
|
|
158
|
-
if (has('codex')) step0.push(`${cr('codex --audit')} for
|
|
158
|
+
if (has('codex')) step0.push(`${cr('codex --audit')} for a second-opinion read by a second model family`);
|
|
159
159
|
if (has('agy')) step0.push(`${cr('agy')} for research sweeps and concurrent fan-out`);
|
|
160
160
|
const stage1 = [];
|
|
161
|
-
if (has('codex')) stage1.push(`${cr('codex')} for
|
|
161
|
+
if (has('codex')) stage1.push(`${cr('codex')} for a second-opinion critique of the map`);
|
|
162
162
|
if (has('grok')) stage1.push(`${cr('grok')} to verify current API behaviour instead of trusting recall`);
|
|
163
163
|
if (has('hermes')) stage1.push(`${cr('hermes')} for a divergent read`);
|
|
164
164
|
if (has('agy')) stage1.push(`${cr('agy')} for a wide sweep of prior art`);
|
|
165
165
|
const examples = [];
|
|
166
166
|
examples.push(has('grok') ? `| "What is trending on X today" | ${cr('grok')} |` : '| "What is trending on X today" | live-researcher (standard tier with web tools) |');
|
|
167
|
-
examples.push(has('codex') ? `| "Audit this auth diff" | ${cr('codex --audit')} |` : '| "Audit this auth diff" | code-reviewer at deep tier, in a fresh context told to
|
|
167
|
+
examples.push(has('codex') ? `| "Audit this auth diff" | ${cr('codex --audit')} |` : '| "Audit this auth diff" | code-reviewer at deep tier, in a fresh context told to challenge |');
|
|
168
168
|
examples.push(has('qwen') ? `| "Classify these 200 items" | bulk-worker, or ${cr('qwen')} if the items may leave the machine |` : '| "Classify these 200 items" | bulk-worker |');
|
|
169
|
-
examples.push(enabled.length >= 2 ? '| "Research this topic properly" | several engines in parallel, see `RESEARCH_TRIAGE.md` |' : '| "Research this topic properly" | deep tier plans, standard tier sweeps, a fresh context
|
|
169
|
+
examples.push(enabled.length >= 2 ? '| "Research this topic properly" | several engines in parallel, see `RESEARCH_TRIAGE.md` |' : '| "Research this topic properly" | deep tier plans, standard tier sweeps, a fresh context challenges; see `RESEARCH_TRIAGE.md` |');
|
|
170
170
|
const roles = [];
|
|
171
171
|
if (has('agy')) roles.push('| Web sweep | `cli-run agy` | widest landscape pass |');
|
|
172
|
-
if (has('codex')) roles.push('|
|
|
172
|
+
if (has('codex')) roles.push('| Second-opinion read | `cli-run codex --audit` | question the premise, hunt for what the others would get wrong |');
|
|
173
173
|
if (has('grok')) roles.push('| Live data | `cli-run grok` | dated primary sources, real-time reads |');
|
|
174
174
|
if (has('hermes')) roles.push('| Cheap divergent read | `cli-run hermes` | another opinion at $0 |');
|
|
175
175
|
if (has('qwen')) roles.push('| Structured extraction | `cli-run qwen` | pull the facts into a table; never trust its citations without a check |');
|
|
@@ -183,12 +183,12 @@ export function laneVars(selected) {
|
|
|
183
183
|
return {
|
|
184
184
|
LANE_STEP0: step0.length ? step0.map((l) => ' - ' + l).join('\n') : ' - none selected yet: every task stays on your primary agent\'s tiers until you add a lane (re-run the installer with more AIs)',
|
|
185
185
|
STAGE1_LANES: stage1.length ? '; ' + stage1.join(', ') : '',
|
|
186
|
-
ATTACK_LANE: has('codex') ? '`cli-run codex --audit` (a second model family in a read-only sandbox)' : 'code-reviewer at deep tier, in a fresh context told to
|
|
186
|
+
ATTACK_LANE: has('codex') ? '`cli-run codex --audit` (a second model family in a read-only sandbox)' : 'code-reviewer at deep tier, in a fresh context told to challenge and allowed to answer CLEAN',
|
|
187
187
|
LIVE_LANE: has('grok') ? '`cli-run grok` first ($0), then' : '',
|
|
188
188
|
BULK_LANE: has('qwen') ? ', or `cli-run qwen` if the data may leave your machine' : has('hermes') ? ', or `cli-run hermes` for a free rough pass' : '',
|
|
189
189
|
LANE_EXAMPLES: examples.join('\n'),
|
|
190
190
|
RESEARCH_ROLES: roles.join('\n'),
|
|
191
|
-
RESEARCH_RUN: run.length ? run.join('\n') : '# no cli-run lane selected: run the sweep on your primary agent, then a fresh
|
|
191
|
+
RESEARCH_RUN: run.length ? run.join('\n') : '# no cli-run lane selected: run the sweep on your primary agent, then a fresh second-opinion turn (protocols/deep-research.md, level 1 shape)',
|
|
192
192
|
RESEARCH_ENGINES: String(run.length)
|
|
193
193
|
};
|
|
194
194
|
}
|
|
@@ -2,4 +2,4 @@
|
|
|
2
2
|
|
|
3
3
|
Antigravity CLI custom agents, one per tier plus three checks (`finding-verifier`, `done-verifier`, `reader`), in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
|
|
4
4
|
|
|
5
|
-
`commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests;
|
|
5
|
+
`commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; deletes and other destructive commands still ask before running) and `off` for the read-only agents: `code-reviewer`, `finding-verifier`, `live-researcher`, `done-verifier`, `reader`. `model` is a tier: `pro` for deep-planner, `flash` for the rest. `done-verifier` and `reader` never write and, with `commandExecutionPolicy: off`, cannot execute any command at all here, mutating or not: unlike its claude-code counterpart, which does carry an unrestricted `Bash` and stays read-only by its prompt rather than by the tool grant, agy's `done-verifier` is mechanically blocked from shelling out and probes artifacts through whatever read or fetch capability it has instead. Neither is `bulk-worker`, which classifies, tags and transforms items and does write.
|
|
@@ -4,7 +4,7 @@ description: Well-specified execution of a bounded sub-part of a build.
|
|
|
4
4
|
model: flash
|
|
5
5
|
subagent: true
|
|
6
6
|
mainAgent: true
|
|
7
|
-
commandExecutionPolicy: auto # standard build/test commands run unattended;
|
|
7
|
+
commandExecutionPolicy: auto # standard build/test commands run unattended; destructive commands, like deletes, still ask before running
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# builder
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: finding-verifier
|
|
3
|
-
description:
|
|
3
|
+
description: Second-opinion verification of review findings; tries to disprove each one and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE. Read-only, never repairs.
|
|
4
4
|
model: flash
|
|
5
5
|
subagent: true
|
|
6
6
|
mainAgent: true
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: finding-verifier
|
|
3
|
-
description:
|
|
3
|
+
description: Second-opinion verification of review findings. Use after a review or audit returns findings and before any of them trigger a repair. No file-editing tools; Bash is for read-only checks, bound by the prompt below, not by the tool grant. Tries to DISPROVE each finding and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Do not use to find new problems, and do not use to fix anything.
|
|
4
4
|
tools: Read, Glob, Grep, Bash
|
|
5
5
|
model: sonnet
|
|
6
6
|
effort: high
|
|
@@ -4,8 +4,8 @@
|
|
|
4
4
|
|
|
5
5
|
```
|
|
6
6
|
You follow a model-orchestrator workflow inside this chat. Tiers describe effort, not automatic model switching or cost savings.
|
|
7
|
-
Route first: bulk/formatting -> fast; live data -> standard with tools; review -> standard, read-only; ambiguous/high-
|
|
8
|
-
For builds: map affected parts; identify
|
|
7
|
+
Route first: bulk/formatting -> fast; live data -> standard with tools; review -> standard, read-only; ambiguous/high-stakes -> deep; otherwise standard. State the tier. Escalate on failure instead of silently retrying.
|
|
8
|
+
For builds: map affected parts; identify what is most likely to go wrong and any gap in the request (none is valid with reasons); build and verify; use a fresh turn to challenge the result. Before irreversible actions, explain rollback and ask for approval.
|
|
9
9
|
For hand-offs: include purpose, scope, allowed and denied actions, required output, and stopping conditions. A fresh context has none of these instructions.
|
|
10
10
|
After comprehensive work, check for omissions. Compute consequential numbers and comparisons with a tool; report what was checked and what remains unverified.
|
|
11
11
|
Before durable writes, search existing records, update their index, use one writer, and label inferences.
|
|
@@ -17,7 +17,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
|
|
|
17
17
|
|
|
18
18
|
A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Measure yours once: spawn a subagent with a one-line task and read its token count. Work smaller than that stays inline.
|
|
19
19
|
|
|
20
|
-
Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one
|
|
20
|
+
Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one challenge pass, an explicit human yes before anything irreversible, then the loud negative.
|
|
21
21
|
|
|
22
22
|
Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A Claude Code subagent loads this CLAUDE.md hierarchy, so it holds the standing rules already, just not this task's scope; a second CLI or a fresh chat window may hold none of them. Absence is denial either way.
|
|
23
23
|
|
|
@@ -9,7 +9,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
|
|
|
9
9
|
|
|
10
10
|
Route by capability tier, first match wins: bulk and mechanical -> fast tier · needs live data -> standard tier with tools · review without changing -> standard, read-only · ambiguous or expensive to get wrong -> deep tier, then hand the plan down · everything else -> build it directly at standard tier.
|
|
11
11
|
|
|
12
|
-
Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: map
|
|
12
|
+
Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: map everything it touches yourself, ask the deep tier for one named weak spot and one gap in the request, build green, scan the added lines, one challenge pass with every finding reproduced, an explicit human yes before anything irreversible, then re-grep the old identifier and expect zero.
|
|
13
13
|
|
|
14
14
|
Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A fresh context holds none of these rules; absence is denial.
|
|
15
15
|
|
|
@@ -10,7 +10,7 @@
|
|
|
10
10
|
// This script always exits 0, never blocks on stdin past a short bound,
|
|
11
11
|
// reads at most 64 KB of the rules file through a fixed-size buffer (never
|
|
12
12
|
// a full read of an arbitrarily large or non-regular file), and never
|
|
13
|
-
// executes anything it reads. See docs/audit-brief.md for the
|
|
13
|
+
// executes anything it reads. See docs/audit-brief.md for the security notes.
|
|
14
14
|
import { statSync, openSync, readSync, closeSync, realpathSync } from 'node:fs';
|
|
15
15
|
import { join, isAbsolute } from 'node:path';
|
|
16
16
|
|
|
@@ -29,8 +29,8 @@ Modifiers:
|
|
|
29
29
|
|
|
30
30
|
## The two checkpoints (every build)
|
|
31
31
|
|
|
32
|
-
- **Checkpoint 1, before writing anything.** You map
|
|
33
|
-
- **Checkpoint 2, after the build is green.** Security-shaped diffs get
|
|
32
|
+
- **Checkpoint 1, before writing anything.** You map everything it touches yourself (files, systems, docs, tickets). Then ask the deep tier, on the finished map: *is this the simplest way, what is most likely to go wrong, what did the request miss?* It must return one named weak spot and one gap in the request. Approval alone is not an answer.
|
|
33
|
+
- **Checkpoint 2, after the build is green.** Security-shaped diffs get a second-opinion read (in a fresh context, told to challenge, allowed to answer CLEAN). Architecture-shaped diffs get the deep tier reviewing build against plan. Never both on one diff. Every finding reproduced before it reaches a human.
|
|
34
34
|
|
|
35
35
|
Cap: two deep-tier consults per build. The full procedure is `protocols/build-protocol.md`.
|
|
36
36
|
|
|
@@ -47,7 +47,7 @@ Level 2 adds `ROUTING.md`, `TIERS.md`, `DELEGATION_MATRIX.md`, `RESEARCH_TRIAGE.
|
|
|
47
47
|
## The three rules that carry everything
|
|
48
48
|
|
|
49
49
|
1. **Route by capability tier, not by model name.** deep = ambiguous or expensive to get wrong · standard = well-specified execution and review · fast = bulk and mechanical. Default down, escalate on evidence.
|
|
50
|
-
2. **A gate you cannot fail is not a gate.** "Does it look good?" passes every time. "Name
|
|
50
|
+
2. **A gate you cannot fail is not a gate.** "Does it look good?" passes every time. "Name what is most likely to go wrong, and what the request missed" can come back empty, which is how you know it worked.
|
|
51
51
|
3. **Exit 0 is not a deliverable.** Any tool, CLI or subagent can report success and hand back nothing. Check for the artifact, not the status line.
|
|
52
52
|
|
|
53
53
|
## Where things went
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
**Three phases, eight stages, and every gate is a question that can be answered wrong.**
|
|
4
4
|
|
|
5
|
-
Fires on any task that builds, codes, implements, migrates or deploys. Rough test: if it would earn
|
|
5
|
+
Fires on any task that builds, codes, implements, migrates or deploys. Rough test: if it would earn a second-opinion audit or a tracker issue, it runs this.
|
|
6
6
|
|
|
7
7
|
> **The one rule underneath:** a gate you cannot fail is not a gate. If a stage's exit reads like "confirm it looks good", it is written wrong and it will pass every time, including the times it should not.
|
|
8
8
|
|
|
@@ -14,7 +14,7 @@ Three corollaries:
|
|
|
14
14
|
| Phase | Master question | Stages |
|
|
15
15
|
|---|---|---|
|
|
16
16
|
| 1 Pre-build | What exactly are we building, what do we need first, and what does this touch or break? | 0 Route · 1 Map · 2 Judge |
|
|
17
|
-
| 2 Build | Is it secure, built on current code, and correct without hidden flaws? | 3 Build · 4 Scan · 5
|
|
17
|
+
| 2 Build | Is it secure, built on current code, and correct without hidden flaws? | 3 Build · 4 Scan · 5 Challenge · 5b Ship gate |
|
|
18
18
|
| 3 Post-build | Did it land everywhere, is it proven against the real thing, and is it recorded? | 6 Verify · 7 Record |
|
|
19
19
|
|
|
20
20
|
The two seams are the point. Pre-build to Build: nothing is written yet, changing your mind costs a conversation. Build to Post-build: the ship, the only irreversible step, the only one that needs an explicit human yes.
|
|
@@ -38,9 +38,9 @@ Four bounded questions, not four exhaustive scans. **The builder maps; the judgm
|
|
|
38
38
|
### Stage 2 · Judge (Checkpoint 1)
|
|
39
39
|
Ask the judgment tier, on the finished map:
|
|
40
40
|
1. Is this the simplest way to build it, or are we overcomplicating?
|
|
41
|
-
2. What is
|
|
41
|
+
2. What is most likely to go wrong, and what did the request miss?
|
|
42
42
|
|
|
43
|
-
**Gate:**
|
|
43
|
+
**Gate:** **one named weak spot** and **one gap in the request**. Approval alone is not an exit; an advisor asked only to approve will approve. If a consult comes back mostly restating the map, the brief asked it to retrieve when it should have asked it to decide.
|
|
44
44
|
|
|
45
45
|
## Phase 2 · Build
|
|
46
46
|
|
|
@@ -56,15 +56,15 @@ Ask the judgment tier, on the finished map:
|
|
|
56
56
|
1. Any secret, key or token in the new code?
|
|
57
57
|
2. Any vulnerability or vulnerable dependency in the lines we added?
|
|
58
58
|
|
|
59
|
-
Secret detection, static analysis and dependency scanning, filtered to lines this diff added.
|
|
59
|
+
Secret detection, static analysis and dependency scanning, filtered to lines this diff added. Refuses by default: a missing or erroring scanner exits non-zero, never a silent green.
|
|
60
60
|
|
|
61
61
|
**Gate:** zero flags on added lines. Pre-existing flags are reported, never inherited as blockers, and never waved through unread. A scanner finding is a claim; read the code before calling it anything.
|
|
62
62
|
|
|
63
|
-
### Stage 5 ·
|
|
63
|
+
### Stage 5 · Challenge (Checkpoint 2, one pass, never two)
|
|
64
64
|
1. Can bad input or a bad actor break it, and what happens when a dependency fails?
|
|
65
65
|
2. Did the build stick to the approved plan, or did unintended changes sneak in?
|
|
66
66
|
|
|
67
|
-
Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to
|
|
67
|
+
Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to a second-opinion reviewer, ideally a **different model family**. Architecture-shaped diffs go to the judgment tier reviewing build against plan. Never both on one diff.
|
|
68
68
|
|
|
69
69
|
Findings do not go straight to a repair. Hand them to the **finding-verifier**, whose job is to DISPROVE each one: read the cited line, state what would trigger it, then hunt for the guard, caller or test that makes it impossible. It returns one of three verdicts per finding, and rounding between them is the failure mode to watch for.
|
|
70
70
|
|
|
@@ -106,7 +106,7 @@ Use a different model family from the one that produced the finding where you ha
|
|
|
106
106
|
|---|---|---|
|
|
107
107
|
{{ROLES_BUILDER_ROW}}
|
|
108
108
|
| Judgment tier | Stage 2 and the architectural arm of Stage 5. Argues with a finished map | Perform the retrieval |
|
|
109
|
-
|
|
|
109
|
+
| Second-opinion reviewer | The security arm of Stage 5. Reviews the diff | Fix anything |
|
|
110
110
|
| Mechanical gates | Stage 4 and any always-on guard | Be overridden without reading |
|
|
111
111
|
| Cheap workers | Bounded sub-parts: bulk passes, wide searches, long loops | Own a stage |
|
|
112
112
|
| Human | Stage 5b, and any irreversible or architectural call | Be the first line of review |
|
|
@@ -119,9 +119,9 @@ Use a different model family from the one that produced the finding where you ha
|
|
|
119
119
|
PRE-BUILD
|
|
120
120
|
[ ] 0 Inputs and access verified by live probe, not assumed
|
|
121
121
|
[ ] 0 Confirmed this is a build and not a quick fix
|
|
122
|
-
[ ] 1
|
|
122
|
+
[ ] 1 Everything it touches written: files, systems, issues
|
|
123
123
|
[ ] 1 Asked what could break, and whether this already exists
|
|
124
|
-
[ ] 2 Judgment tier named a
|
|
124
|
+
[ ] 2 Judgment tier named a weak spot AND a gap in the request
|
|
125
125
|
|
|
126
126
|
BUILD
|
|
127
127
|
[ ] 3 Repo clean, on a branch, base ref recorded
|
|
@@ -31,11 +31,11 @@ Agreement is weak evidence. Disagreement is the signal.
|
|
|
31
31
|
|
|
32
32
|
## Level 1: one agent
|
|
33
33
|
|
|
34
|
-
You still get the shape. Run PLAN as its own turn and inspect it before spending anything. Run the sweep. Then run a **fresh-context
|
|
34
|
+
You still get the shape. Run PLAN as its own turn and inspect it before spending anything. Run the sweep. Then run a **fresh-context second-opinion turn** with a brief that says "question the premise; list what this report would get wrong if its sources were stale". Plant one deliberately wrong figure in the brief and see whether it corrects it: if it does not, its confirmations are worth less than they look. Mark every claim.
|
|
35
35
|
|
|
36
36
|
## Level 2 and up: three engines, one triager
|
|
37
37
|
|
|
38
|
-
Fan out the same PLAN to three different model families through their CLIs (a web-sweep lane,
|
|
38
|
+
Fan out the same PLAN to three different model families through their CLIs (a web-sweep lane, a second-opinion-read lane, a live-data lane). Run them through `cli-run` so a run that produced nothing is caught as `rc=10` rather than read as an empty finding. The orchestrator triages: it opens the primary sources itself, marks each claim, and writes the brief. Only the orchestrator writes the durable record; every other engine proposes.
|
|
39
39
|
|
|
40
40
|
Known failure shape: one engine will return confident unsourced numerics and claim full coverage. Downgrade those to hypothesis. The engines that report their own gaps honestly are the ones to weight.
|
|
41
41
|
|
|
@@ -14,7 +14,7 @@ Verification asks "is what I did correct?". Gap analysis asks "what did I not do
|
|
|
14
14
|
## Who runs it
|
|
15
15
|
|
|
16
16
|
- **Level 1 (one agent):** the same agent, in a fresh turn, with a brief that says "you are looking for what is missing; do not re-verify what is present". Fresh context matters more than a different model.
|
|
17
|
-
- **Level 2 and up:** a **different model family** reading the same artifact. Disagreement between two families is the cheapest available signal that something is soft. The
|
|
17
|
+
- **Level 2 and up:** a **different model family** reading the same artifact. Disagreement between two families is the cheapest available signal that something is soft. The second-opinion coder lane (read-only mode) is the natural fit.
|
|
18
18
|
- **Level 3:** make it recurring. A weekly audit job enumerates live state (lanes, jobs, services, model lists), diffs it against the plan, and files a report. It catches the dead lane and the silently renamed model nobody noticed.
|
|
19
19
|
|
|
20
20
|
## The second half: analyze, compare, suggest
|
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
# Propagate: change completeness
|
|
2
2
|
|
|
3
|
-
**A rename is a refactor, not a single-file edit.** Any change to a name, term, path, slug, schema field, routing rule or shared convention
|
|
3
|
+
**A rename is a refactor, not a single-file edit.** Any change to a name, term, path, slug, schema field, routing rule or shared convention reaches everything that uses it, and the goal is zero silent strays.
|
|
4
4
|
|
|
5
5
|
This is retrieval work. It stays with the orchestrator (or a cheap worker for the grep sweep). It never goes to the deep tier: a judgment model re-deriving a file list is the most expensive routing mistake there is.
|
|
6
6
|
|
|
7
|
-
## 1. Map
|
|
7
|
+
## 1. Map everything it touches (before editing anything)
|
|
8
8
|
|
|
9
9
|
- **Docs and notes:** backlinks to the thing being renamed; literal search for the old term and its link forms. With obsidian-tc: `get_backlinks`, `search_text`, then `find_unresolved_links` after the change (`protocols/memory-and-record.md`).
|
|
10
10
|
- **Memory / instructions:** grep every instructions file your agents read (`CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, `QWEN.md`, custom instructions) and any memory store.
|
|
@@ -68,7 +68,7 @@ qwen is the lane whose own success flags lie: an upstream 400 comes back as exit
|
|
|
68
68
|
|
|
69
69
|
## The route: which model, and how hard it thinks
|
|
70
70
|
|
|
71
|
-
A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe
|
|
71
|
+
A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe a second-opinion pass, and nothing anywhere says so.
|
|
72
72
|
|
|
73
73
|
Pin it per call, or per lane:
|
|
74
74
|
|
|
@@ -114,9 +114,9 @@ The log records what was **requested**, on every record including a run refused
|
|
|
114
114
|
|
|
115
115
|
That is each vendor's documented headless shape (`-p`, `exec`). Two consequences: argv is visible to other processes on the machine, so a prompt is never the place for a key; and argv is bounded by the OS (`ARG_MAX`), so a very large brief should be referenced by path inside the prompt rather than pasted whole.
|
|
116
116
|
|
|
117
|
-
## lanes.json
|
|
117
|
+
## lanes.json refuses by default
|
|
118
118
|
|
|
119
|
-
Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all
|
|
119
|
+
Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all make the whole file refuse by default rather than being skipped quietly.
|
|
120
120
|
|
|
121
121
|
## A killed lane is not a deliverable
|
|
122
122
|
|
|
@@ -15,7 +15,7 @@ Generated {{DATE}} from the AIs you said you have: `{{AI_IDS}}`.
|
|
|
15
15
|
| Many independent items each needing its own agent turn | a concurrent fan-out lane | one call, N children, on a subscription |
|
|
16
16
|
| Live web or social reads | the live-data CLI | subscription-covered; the same search on the API bills per call |
|
|
17
17
|
| Code review, no changes | standard tier, or the second-coder CLI | a different model family catches what one misses |
|
|
18
|
-
|
|
|
18
|
+
| Second-opinion audit of a security-shaped diff | the second-coder CLI in read-only audit mode | Claude writes, a second family challenges, the orchestrator reproduces |
|
|
19
19
|
| Deep architecture / planning | deep tier | expensive to get wrong |
|
|
20
20
|
| Well-specified execution | the orchestrator | execution does not need the top tier |
|
|
21
21
|
| Long-document analysis | the largest-context lane, or caching on the primary | window size vs re-query cost |
|
|
@@ -31,10 +31,10 @@ Rule of thumb: never spend a frontier token on a task a cheap tier finishes corr
|
|
|
31
31
|
|---|---|
|
|
32
32
|
| 0 Route | live probe for access; `cli-run` lanes are $0 and uncapped |
|
|
33
33
|
| 1 Map | the orchestrator sweeps{{STAGE1_LANES}} |
|
|
34
|
-
| 2 Judge | deep tier, on the finished map:
|
|
34
|
+
| 2 Judge | deep tier, on the finished map: one named weak spot and one gap in the request |
|
|
35
35
|
| 3 Build | the orchestrator, against the installed dependency's source |
|
|
36
|
-
| 4 Scan | secret + static + dependency scanners, diff-scoped,
|
|
37
|
-
| 5
|
|
36
|
+
| 4 Scan | secret + static + dependency scanners, diff-scoped, refuses by default |
|
|
37
|
+
| 5 Challenge | security-shaped diff → {{ATTACK_LANE}}. Architecture-shaped → deep tier, build against plan. Never both |
|
|
38
38
|
| 5a Verify findings | finding-verifier, a different model family where you have one: CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a repair |
|
|
39
39
|
| 5b Ship | rollback id recorded, explicit human yes |
|
|
40
40
|
| 6 Verify | real test, negative test seen red, old identifier re-grepped to zero |
|
|
@@ -58,7 +58,7 @@ One writer per run; every other lane proposes. Search before writing, index in t
|
|
|
58
58
|
- **Long context:** mechanical digestion → fast tier in chunks; judgment over a long input → standard tier.
|
|
59
59
|
- **Token discipline on every delegation:** pass only the context the delegate needs, never the conversation.
|
|
60
60
|
- **Effort per agent:** deep xhigh, review, verification and build high, live research medium, bulk low.
|
|
61
|
-
- **Three inputs, not one:** role picks the agent, complexity moves the effort,
|
|
61
|
+
- **Three inputs, not one:** role picks the agent, complexity moves the effort, stakes move the tier and who reads it. A one-line auth change is simple and high-stakes at once, and the stakes decide. See `TIERS.md`.
|
|
62
62
|
- **Pin the route when it matters:** a lane with no `--model`/`--effort` and no `defaults` entry in `bin/lanes.json` runs on its own config, which may be nothing like what this file describes. `cli-run --doctor` prints what each lane is pinned to, and every run logs the value requested and where it came from.
|
|
63
63
|
|
|
64
64
|
## Example routings
|
|
@@ -47,21 +47,25 @@ reasoning than the reviewer judging its output.** When the plan is airtight the
|
|
|
47
47
|
spec is carrying the thinking, so builder drops to medium. When the plan is
|
|
48
48
|
vague, fix the plan; do not buy reasoning to paper over it.
|
|
49
49
|
|
|
50
|
-
**
|
|
50
|
+
**Stakes move the tier and the reader, never just the effort.** These four are
|
|
51
51
|
the ones worth naming, because their failures are not recoverable by editing the
|
|
52
52
|
code afterwards.
|
|
53
53
|
|
|
54
|
-
|
|
54
|
+
Stakes means what a mistake would cost: a security hole, leaked personal data,
|
|
55
|
+
lost data, or something you can't undo. Most tasks are low-stakes and route
|
|
56
|
+
normally.
|
|
57
|
+
|
|
58
|
+
| Stakes | Present when the change touches | What it buys |
|
|
55
59
|
|---|---|---|
|
|
56
|
-
| security | auth, tokens, sessions, routes, untrusted input | the
|
|
60
|
+
| security | auth, tokens, sessions, routes, untrusted input | the challenge pass, ideally a different model family |
|
|
57
61
|
| privacy | personal data, anything leaving the machine | the local lane, and a named check on what is sent |
|
|
58
62
|
| data loss | deletion, bulk mutation, migrations, overwrites | a reviewed rollback path before the change is written |
|
|
59
63
|
| irreversible | publishing, sending, rotating, anything with an audience | a human yes at Stage 5b, never an agent's |
|
|
60
64
|
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
synonym for difficulty: a one-line change to an auth check is simple and
|
|
64
|
-
high-
|
|
65
|
+
High stakes raise code-reviewer to xhigh, and a security-shaped diff goes to
|
|
66
|
+
the challenge lane rather than to a second read by the same family. Stakes are
|
|
67
|
+
not a synonym for difficulty: a one-line change to an auth check is simple and
|
|
68
|
+
high-stakes at the same time, and it is the stakes that decide the route.
|
|
65
69
|
|
|
66
70
|
**Reserve the top of the ladder for evidence.** xhigh and the escalation tier are
|
|
67
71
|
bought with a named reason: a reproduced failure, a checkpoint that came back
|
|
@@ -70,7 +74,7 @@ task, not an escalation.
|
|
|
70
74
|
|
|
71
75
|
## Why split tiers: robustness first, cost second
|
|
72
76
|
|
|
73
|
-
The split produces better work. The deep tier steers every build twice, and what it steers is **judgment, never retrieval**: the orchestrator sweeps
|
|
77
|
+
The split produces better work. The deep tier steers every build twice, and what it steers is **judgment, never retrieval**: the orchestrator sweeps everything it touches itself and hands the deep tier a finished map. Paying deep-tier rates for a file list is the most expensive routing mistake available.
|
|
74
78
|
|
|
75
79
|
Against a baseline of "standard tier with no consults", default checkpoints are a spend increase. That is the accepted trade, not a saving to claim.
|
|
76
80
|
|
|
@@ -36,7 +36,7 @@ The tools cannot help a model that never reaches for them. codecalc ships `SKILL
|
|
|
36
36
|
|
|
37
37
|
## What it is not
|
|
38
38
|
|
|
39
|
-
Not a cloud sandbox for multi-tenant loads, not a replacement for a vendor's built-in interpreter when zero setup matters more than measurement.
|
|
39
|
+
Not a cloud sandbox for multi-tenant loads, not a replacement for a vendor's built-in interpreter when zero setup matters more than measurement. It assumes a single-operator, local, stdio setup. It earns its keep when the correctness of a claim, not "it ran", is the point.
|
|
40
40
|
|
|
41
41
|
## On a box (level 3)
|
|
42
42
|
|
|
@@ -11,7 +11,7 @@ A durable, searchable, governed store that the protocols can call by name:
|
|
|
11
11
|
| Need in the protocols | obsidian-tc tool |
|
|
12
12
|
|---|---|
|
|
13
13
|
| find what exists before writing (deep research dedupe, gap analysis) | `semantic_search`, `search_text`, `search_regex` |
|
|
14
|
-
| map a rename
|
|
14
|
+
| map everything a rename touches (propagate) | `get_backlinks`, `find_unresolved_links`, `rewrite_link` |
|
|
15
15
|
| record the end-to-end doc (build Stage 7) | `write_note` (compare-and-swap, confirmation on overwrite), `patch_note`, `append_note` |
|
|
16
16
|
| keep inferred content honest | `write_note` with `provenance: "agent_synthesis"` runs a poison scan before the write lands |
|
|
17
17
|
| keep a shared vault safe for several agents | JWT scopes, per-vault folder ACLs, a read-only kill switch, human-in-the-loop tokens |
|
|
@@ -56,9 +56,9 @@ Merge the block; do not replace the file.
|
|
|
56
56
|
|
|
57
57
|
## Security posture, read before a second agent touches it
|
|
58
58
|
|
|
59
|
-
Zero-config mode boots with **auth off and no folder ACL**: anything that can reach the server has the same authority as raw filesystem access to the vault. That is acceptable only because the surface is local-only (the config
|
|
59
|
+
Zero-config mode boots with **auth off and no folder ACL**: anything that can reach the server has the same authority as raw filesystem access to the vault. That is acceptable only because the surface is local-only (the config refuses by default if you enable HTTP on a non-loopback host with auth off, and a DNS-rebinding guard protects loopback). Before exposing it to partially-trusted, remote or multi-agent callers, turn on `auth.mode: "jwt"` and set `acl.readPaths` / `writePaths` / `deletePaths` in the config file. Upstream `SECURITY.md` has the security notes and a private disclosure path.
|
|
60
60
|
|
|
61
|
-
Track record worth knowing: an independent code audit of v1.8.1 (July 2026) found three security-relevant gaps (an ACL
|
|
61
|
+
Track record worth knowing: an independent code audit of v1.8.1 (July 2026) found three security-relevant gaps (an ACL bypass that let enumeration tools skip its refuse-by-default rule, a compare-and-swap bypass through `upsert`, a poison-eligibility gap in preference extraction). All three were fixed upstream before they were filed; verified against the v1.25.0 source on 2026-09-03.
|
|
62
62
|
|
|
63
63
|
## Level 3
|
|
64
64
|
|