model-orchestrator 0.1.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +96 -0
- package/LICENSE +21 -0
- package/README.md +133 -0
- package/SECURITY.md +17 -0
- package/bin/README.md +10 -0
- package/bin/cli-run.mjs +599 -0
- package/bin/cli.js +372 -0
- package/docs/README.md +13 -0
- package/docs/audit-brief.md +83 -0
- package/docs/catalog.md +113 -0
- package/docs/part-1-beginner.md +65 -0
- package/docs/part-2-intermediate.md +65 -0
- package/docs/part-3-advanced.md +65 -0
- package/package.json +52 -0
- package/scripts/README.md +5 -0
- package/scripts/gen-catalog.js +37 -0
- package/src/README.md +9 -0
- package/src/catalog.js +277 -0
- package/src/detect.js +26 -0
- package/src/install.js +628 -0
- package/src/prompt.js +34 -0
- package/src/render.js +8 -0
- package/templates/README.md +14 -0
- package/templates/advanced/README.md +14 -0
- package/templates/advanced/vm/ENVIRONMENT.md +18 -0
- package/templates/advanced/vm/PRIVACY_GATES.md +33 -0
- package/templates/advanced/vm/README.md +60 -0
- package/templates/advanced/vm/box-CLAUDE.md +28 -0
- package/templates/advanced/vm/docker-compose.yml +19 -0
- package/templates/advanced/vm/gateway.config.yaml +12 -0
- package/templates/advanced/vm/jobs/README.md +39 -0
- package/templates/advanced/vm/jobs/weekly-audit.service +17 -0
- package/templates/advanced/vm/jobs/weekly-audit.sh +107 -0
- package/templates/advanced/vm/jobs/weekly-audit.timer +10 -0
- package/templates/advanced/vm/setup-vm.sh +46 -0
- package/templates/agents/README.md +13 -0
- package/templates/agents/agy/README.md +5 -0
- package/templates/agents/agy/builder.md +17 -0
- package/templates/agents/agy/bulk-worker.md +17 -0
- package/templates/agents/agy/code-reviewer.md +17 -0
- package/templates/agents/agy/deep-planner.md +17 -0
- package/templates/agents/agy/live-researcher.md +17 -0
- package/templates/agents/claude-code/README.md +13 -0
- package/templates/agents/claude-code/builder.md +17 -0
- package/templates/agents/claude-code/bulk-worker.md +18 -0
- package/templates/agents/claude-code/code-reviewer.md +19 -0
- package/templates/agents/claude-code/deep-planner.md +18 -0
- package/templates/agents/claude-code/live-researcher.md +18 -0
- package/templates/agents/snippets/chat.md +25 -0
- package/templates/agents/snippets/claude-code.md +27 -0
- package/templates/agents/snippets/generic.md +21 -0
- package/templates/beginner/ORCHESTRATOR.md +55 -0
- package/templates/beginner/README.md +3 -0
- package/templates/common/README.md +52 -0
- package/templates/common/TASK_BUNDLE.md +56 -0
- package/templates/common/protocols/README.md +14 -0
- package/templates/common/protocols/build-protocol.md +133 -0
- package/templates/common/protocols/deep-research.md +44 -0
- package/templates/common/protocols/gap-analysis.md +28 -0
- package/templates/common/protocols/memory-and-record.md +30 -0
- package/templates/common/protocols/numbers-and-logic.md +35 -0
- package/templates/common/protocols/propagate.md +34 -0
- package/templates/intermediate/CLI-RUN.md +100 -0
- package/templates/intermediate/DELEGATION_MATRIX.md +41 -0
- package/templates/intermediate/README.md +13 -0
- package/templates/intermediate/RESEARCH_TRIAGE.md +30 -0
- package/templates/intermediate/ROUTING.md +73 -0
- package/templates/intermediate/TIERS.md +44 -0
- package/templates/tools/README.md +10 -0
- package/templates/tools/codecalc/CODECALC.md +43 -0
- package/templates/tools/codecalc/mcp/agy.mcp_config.json +8 -0
- package/templates/tools/codecalc/mcp/codex.config.toml +4 -0
- package/templates/tools/codecalc/mcp/mcpServers.json +8 -0
- package/templates/tools/codecalc/mcp/vscode.mcp.json +8 -0
- package/templates/tools/codecalc/mcp/zed.settings.json +9 -0
- package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +65 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.agy.mcp_config.json +9 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.codex.config.toml +7 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.mcpServers.json +9 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.vscode.mcp.json +10 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.zed.settings.json +9 -0
|
@@ -0,0 +1,100 @@
|
|
|
1
|
+
# CLI-RUN.md: exit 0 means a structurally accepted, non-empty response
|
|
2
|
+
|
|
3
|
+
`bin/cli-run.mjs` is one entrypoint for the agent CLI lanes. It builds the right invocation per lane, reads that lane's native terminal event, and exits non-zero unless a structurally accepted, non-empty response came back.
|
|
4
|
+
|
|
5
|
+
## The guarantee, exactly
|
|
6
|
+
|
|
7
|
+
Exit 0 means: the lane's native terminal event says it finished, the response is non-empty, and the lane-specific error checks passed. **It does not mean the task was done.** A refusal that parses cleanly is exit 0. When your task has a real contract, state it:
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
node bin/cli-run.mjs codex "Write the report to out/report.md" --expect-file out/report.md # must exist, be non-empty, and be written during this run
|
|
11
|
+
node bin/cli-run.mjs grok "Return the table as JSON" --expect-json # the response must parse as JSON
|
|
12
|
+
```
|
|
13
|
+
|
|
14
|
+
An unmet contract is exit 10 with reason `contract_unmet`. `--expect-file` snapshots the target before the lane starts (existence, size, mtime, content hash) and afterwards requires a non-empty regular file that is new or changed: a different hash, or a later mtime. A file that existed before and was not touched fails, however recent it is; a rewrite with identical bytes and an unchanged mtime also fails, because nothing distinguishes it from no write. Timestamps and hashes are evidence of change, not proof of authorship: if another process could write the same path during the run, use a per-attempt path. Text-only callers need nothing new: without a contract flag the behaviour is the structural check above.
|
|
15
|
+
|
|
16
|
+
Enabled lanes (edit `bin/lanes.json`): {{CLI_RUN_LANES}}
|
|
17
|
+
|
|
18
|
+
```bash
|
|
19
|
+
node bin/cli-run.mjs <grok|codex|agy|hermes|qwen> "<prompt>" [--brief FILE] [--timeout SECS] [--quiet]
|
|
20
|
+
node bin/cli-run.mjs codex --audit "<prompt>" # read-only sandbox, the audit shape
|
|
21
|
+
node bin/cli-run.mjs qwen [--model ID] [--safe-mode] "<prompt>"
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
Put it on your PATH if you like: `ln -s "$PWD/bin/cli-run.mjs" ~/.local/bin/cli-run`.
|
|
25
|
+
|
|
26
|
+
## First run: `--doctor`
|
|
27
|
+
|
|
28
|
+
```bash
|
|
29
|
+
node bin/cli-run.mjs --doctor # which lanes are enabled and which binaries are on PATH; runs nothing
|
|
30
|
+
node bin/cli-run.mjs --doctor --run # also sends each enabled lane one tiny prompt and judges the reply (uses a little quota)
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
A lane that is enabled but not on PATH, or that answers with no deliverable, shows up here before it shows up mid-task.
|
|
34
|
+
|
|
35
|
+
## Why it exists
|
|
36
|
+
|
|
37
|
+
Every agent CLI can exit 0 having produced nothing. The symptom (confident preamble, exit 0, no deliverable) is indistinguishable from a model failure, so it gets blamed on the model. Four wrong diagnoses in one week came from exactly that.
|
|
38
|
+
|
|
39
|
+
Byte count is not a deliverable check either: a run can emit hundreds of kilobytes and contain no conclusion.
|
|
40
|
+
|
|
41
|
+
## The success signal per lane
|
|
42
|
+
|
|
43
|
+
| Lane | Invocation built | Success = |
|
|
44
|
+
|---|---|---|
|
|
45
|
+
| grok | `--output-format json -p` | `stopReason == "end_turn"` and non-empty `text` |
|
|
46
|
+
| codex | `exec --json --color never --skip-git-repo-check -o FILE` | terminal `{"type":"turn.completed"}` and non-empty FILE |
|
|
47
|
+
| agy | `--print-timeout Nm --output-format stream-json -p` | terminal `{"event":"result"}`, `status == "SUCCESS"`, non-empty `response` |
|
|
48
|
+
| hermes | `-z … --usage-file FILE` | its exit code is already honest: 0 response · 1 none · 2 bad args |
|
|
49
|
+
| qwen | `-o json [-m ID] [--safe-mode] -p` | terminal `{"type":"result"}`, `subtype == "success"`, `is_error` false, non-empty `result` not starting with `[API Error:`, and every `stats.models.*.api.totalErrors == 0` |
|
|
50
|
+
|
|
51
|
+
qwen is the lane whose own success flags lie: an upstream 400 comes back as exit 0, `subtype: success`, `is_error: false`, with the error text inside `result`. The two extra checks are the honest ones. Absent telemetry is refused, not read as zero.
|
|
52
|
+
|
|
53
|
+
## Exit codes
|
|
54
|
+
|
|
55
|
+
| Code | Meaning |
|
|
56
|
+
|---|---|
|
|
57
|
+
| 0 | structurally accepted non-empty response, every `--expect-*` contract met |
|
|
58
|
+
| 10 | ran and produced no deliverable, or a contract was unmet, or the lane was killed by a signal, or output overran the 16 MiB buffer |
|
|
59
|
+
| 11 | no output at all |
|
|
60
|
+
| 12 | timed out; the lane and every descendant in its process group were killed |
|
|
61
|
+
| 13 | lane unavailable: binary missing, disabled in `lanes.json`, or `lanes.json` malformed |
|
|
62
|
+
| 130 / 143 | cli-run itself received SIGINT / SIGTERM; the lane's process group was killed first, then the temp dir removed |
|
|
63
|
+
| 2 | usage error in cli-run itself |
|
|
64
|
+
| N | the lane exited N != 0: passed through unchanged, verdict `exit_nonzero`, even when parseable text came back. The bounded head of the lane's stderr is shown on your terminal so an auth failure reads as one |
|
|
65
|
+
|
|
66
|
+
## Permissions are a separate layer
|
|
67
|
+
|
|
68
|
+
`cli-run` never injects permission flags. Each CLI carries its own config, so every caller gets the same behaviour. Use each vendor's deny-list as the base layer; allow-lists only hold if every binary is enumerable in advance.
|
|
69
|
+
|
|
70
|
+
## Log
|
|
71
|
+
|
|
72
|
+
`~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, rc, the lane's own exit code, signal, seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument.
|
|
73
|
+
|
|
74
|
+
## The prompt travels in argv
|
|
75
|
+
|
|
76
|
+
That is each vendor's documented headless shape (`-p`, `exec`). Two consequences: argv is visible to other processes on the machine, so a prompt is never the place for a key; and argv is bounded by the OS (`ARG_MAX`), so a very large brief should be referenced by path inside the prompt rather than pasted whole.
|
|
77
|
+
|
|
78
|
+
## lanes.json fails closed
|
|
79
|
+
|
|
80
|
+
Absent: every lane enabled. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled.
|
|
81
|
+
|
|
82
|
+
## A killed lane is not a deliverable
|
|
83
|
+
|
|
84
|
+
A lane that dies by signal has no honest exit status. Whatever it printed first is discarded; the run reports `killed` with exit 10.
|
|
85
|
+
|
|
86
|
+
## Interrupting cli-run kills the lane too
|
|
87
|
+
|
|
88
|
+
Ctrl-C or a `kill` on the wrapper kills the lane's whole process group before the wrapper exits (130 for SIGINT, 143 for SIGTERM). A second signal during cleanup kills again and exits at once. Handlers are installed per run and removed when it finishes, so `--doctor --run` does not accumulate them. Uncatchable SIGKILL to the wrapper leaves the lane running; that is the operating system, not a promise this tool can make. Under systemd, `KillMode=control-group` covers that case.
|
|
89
|
+
|
|
90
|
+
## Output is decoded as a UTF-8 stream
|
|
91
|
+
|
|
92
|
+
Vendor output is decoded with a streaming decoder, so a multibyte character split across two chunks is preserved byte for byte. The 16 MiB cap and the `raw_bytes` field count bytes, not characters.
|
|
93
|
+
|
|
94
|
+
## Timeouts kill the whole process group
|
|
95
|
+
|
|
96
|
+
The lane is started detached, as the leader of its own process group. On timeout, or when output overruns the buffer, the group is killed, so a tool the agent shelled out to cannot keep writing after the wrapper reported 12. A child that calls `setsid()` itself escapes this boundary; nothing user-space can promise more without a cgroup, which is what the level 3 systemd unit adds.
|
|
97
|
+
|
|
98
|
+
## Lane choice is not automated
|
|
99
|
+
|
|
100
|
+
`cli-run` runs the lane it is given. Which lane fits the job is `ROUTING.md` and `DELEGATION_MATRIX.md`, or a question to the human.
|
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
# DELEGATION_MATRIX.md: task → lane → pick
|
|
2
|
+
|
|
3
|
+
Generated {{DATE}} from the AIs you said you have: `{{AI_IDS}}`.
|
|
4
|
+
|
|
5
|
+
## Your lanes
|
|
6
|
+
|
|
7
|
+
{{LANES_TABLE}}
|
|
8
|
+
|
|
9
|
+
## Task → lane
|
|
10
|
+
|
|
11
|
+
| Task type | Pick | Why |
|
|
12
|
+
|---|---|---|
|
|
13
|
+
| Bulk classify / extract / summarize, data may leave the machine | the cheapest metered lane, then the fast tier | cost gap is an order of magnitude; batch APIs add more |
|
|
14
|
+
| Bulk work on data that must stay local | the local lane | a privacy lane, never a cost lane; route here for confinement, not to save money |
|
|
15
|
+
| Many independent items each needing its own agent turn | a concurrent fan-out lane | one call, N children, on a subscription |
|
|
16
|
+
| Live web or social reads | the live-data CLI | subscription-covered; the same search on the API bills per call |
|
|
17
|
+
| Code review, no changes | standard tier, or the second-coder CLI | a different model family catches what one misses |
|
|
18
|
+
| Adversarial audit of a security-shaped diff | the second-coder CLI in read-only audit mode | Claude writes, a second family attacks, the orchestrator reproduces |
|
|
19
|
+
| Deep architecture / planning | deep tier | expensive to get wrong |
|
|
20
|
+
| Well-specified execution | the orchestrator | execution does not need the top tier |
|
|
21
|
+
| Long-document analysis | the largest-context lane, or caching on the primary | window size vs re-query cost |
|
|
22
|
+
| Routing decisions themselves | the cheapest lane you have, or none | never spend deep tokens deciding not to use deep |
|
|
23
|
+
| Rough drafts, divergent reads, first-pass summaries | the free tier | $0, and disagreement with the primary is information |
|
|
24
|
+
| Anything citing a line, a number, or a source | never the cheapest metered lane without a full verification pass | measured: conclusions right, every supporting number invented |
|
|
25
|
+
|
|
26
|
+
## Install and sign-in
|
|
27
|
+
|
|
28
|
+
{{INSTALL_TABLE}}
|
|
29
|
+
|
|
30
|
+
## Cost playbook
|
|
31
|
+
|
|
32
|
+
1. Prompt caching everywhere it fits: frozen prefix first, volatile text last.
|
|
33
|
+
2. Cascade: cheapest capable tier first, escalate on signal.
|
|
34
|
+
3. Batch APIs for anything not latency-sensitive.
|
|
35
|
+
4. A free or local model for routing decisions.
|
|
36
|
+
5. Effort and reasoning knobs before model swaps; often the bigger lever.
|
|
37
|
+
6. Alias-based config so a vendor rename is a one-line repoint.
|
|
38
|
+
|
|
39
|
+
## Privacy gate
|
|
40
|
+
|
|
41
|
+
No private notes, client data, or personal records go to a metered third-party bulk lane or a fan-out lane. Name the barred lanes explicitly in your own rules; an unnamed bar is not enforced.
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
# templates/intermediate/
|
|
2
|
+
|
|
3
|
+
Written at level 2 and above, on top of `common/` and `beginner/`.
|
|
4
|
+
|
|
5
|
+
| File | What it adds |
|
|
6
|
+
|---|---|
|
|
7
|
+
| `ROUTING.md` | the multi-lane decision tree; supersedes `ORCHESTRATOR.md` when present |
|
|
8
|
+
| `TIERS.md` | capability tiers, the three cost levers, the escalation rule, slotting rules |
|
|
9
|
+
| `DELEGATION_MATRIX.md` | task → lane → pick, generated from the user's selection |
|
|
10
|
+
| `RESEARCH_TRIAGE.md` | three engines in parallel, one triager |
|
|
11
|
+
| `CLI-RUN.md` | how `bin/cli-run.mjs` judges each lane |
|
|
12
|
+
|
|
13
|
+
`bin/cli-run.mjs` and `bin/lanes.json` are written by the installer from `bin/cli-run.mjs` in this repo and the user's selection; they are not templates.
|
|
@@ -0,0 +1,30 @@
|
|
|
1
|
+
# RESEARCH_TRIAGE.md: engines in parallel, one triager
|
|
2
|
+
|
|
3
|
+
The deep-research lane at level 2: fan the same plan out to different model families through their CLIs, then triage against primary sources you open yourself.
|
|
4
|
+
|
|
5
|
+
Your `cli-run` lanes: {{CLI_RUN_LANES}} ({{RESEARCH_ENGINES}} research engine(s) below). Everything in this file was rendered from that selection; a lane that is not listed is not one you have.
|
|
6
|
+
|
|
7
|
+
## Roles
|
|
8
|
+
|
|
9
|
+
| Role | Typical lane | Job |
|
|
10
|
+
|---|---|---|
|
|
11
|
+
{{RESEARCH_ROLES}} opens primary sources, marks every claim, writes the artifact |
|
|
12
|
+
|
|
13
|
+
Run each engine as one `cli-run` call with a task bundle in `--brief`. A run that produced nothing exits 10 and is a missing engine, not an empty finding.
|
|
14
|
+
|
|
15
|
+
## One run
|
|
16
|
+
|
|
17
|
+
```bash
|
|
18
|
+
BRIEF=research/brief.md # purpose, sub-questions, source standard, report contract, exit parameters
|
|
19
|
+
{{RESEARCH_RUN}}
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
Then the orchestrator reads the three outputs, opens every primary source that carries a decision, and writes one dated brief with marks: **CONFIRMED** (two engines + primary source) · **DISAGREEMENT** (both readings kept) · **REPORTED** (someone's own post, quoted not trusted) · **UNVERIFIED**.
|
|
23
|
+
|
|
24
|
+
## Triage discipline
|
|
25
|
+
|
|
26
|
+
- Plant one deliberately wrong figure in one brief. An engine that does not correct it has confirmations worth less than they look.
|
|
27
|
+
- Expect one engine to return confident unsourced numerics and claim full coverage. Downgrade to hypothesis. Weight the engines that report their own gaps.
|
|
28
|
+
- Agreement is weak evidence. Disagreement is the signal.
|
|
29
|
+
- Only the orchestrator writes the durable record. Every other engine proposes.
|
|
30
|
+
- Count dispositions, not briefs.
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
# ROUTING.md: the multi-lane decision tree
|
|
2
|
+
|
|
3
|
+
Primary agent (the orchestrator): **{{PRIMARY_NAME}}**. It routes, maps, builds, verifies and records. Every other AI is a lane it calls.
|
|
4
|
+
|
|
5
|
+
Your lanes:
|
|
6
|
+
|
|
7
|
+
{{LANES_TABLE}}
|
|
8
|
+
|
|
9
|
+
Two kinds of lane. **Lane A** = subscription CLIs: $0 marginal, already paid for, used for interactive and agentic work. **Lane B** = metered APIs: per token, used for programmatic bulk where a subscription CLI cannot serve. **Local** = stays on the machine; a privacy lane, never a cost lane.
|
|
10
|
+
|
|
11
|
+
Rule of thumb: never spend a frontier token on a task a cheap tier finishes correctly. Escalate on signal (low confidence, explicit complexity, a failed verification), not by default. And an external lane must earn the hop with a real strength; when in doubt, stay in-house.
|
|
12
|
+
|
|
13
|
+
## Decision tree (first match wins)
|
|
14
|
+
|
|
15
|
+
0. **Is there a cheaper or better external lane for this?** Check `DELEGATION_MATRIX.md`. Your enabled lanes, every one called through `bin/cli-run.mjs`:
|
|
16
|
+
{{LANE_STEP0}}
|
|
17
|
+
1. **Bulk and mechanical?** → fast tier{{BULK_LANE}}. Many independent items each needing its own agent turn → a concurrent fan-out lane if you have one.
|
|
18
|
+
2. **Needs live data?** → {{LIVE_LANE}} standard tier with web tools.
|
|
19
|
+
3. **Reviewing without changing?** → standard tier read-only. Security-critical → {{ATTACK_LANE}}.
|
|
20
|
+
4. **Ambiguous, strategic, expensive to get wrong?** → deep tier (deep-planner). Then hand the plan down.
|
|
21
|
+
5. **Everything else that changes files** → the orchestrator builds it directly. Bounded sub-parts go to cheaper tiers; the main build is never handed off whole.
|
|
22
|
+
|
|
23
|
+
## Who builds
|
|
24
|
+
|
|
25
|
+
**The orchestrator owns the main build.** It is the only surface that holds these rules: a subagent or a second CLI starts with none of them and cannot route. Handing the main build to one hands it to something the router cannot reach.
|
|
26
|
+
|
|
27
|
+
Delegate: background and long-running tasks, small tasks, scoping, verification, research, bounded sub-parts. Never delegate: the main build, or any step that must carry a house rule (secrets handling, the loud-negative verification, the durable record).
|
|
28
|
+
|
|
29
|
+
Every delegation carries `TASK_BUNDLE.md`. Its brief must restate every convention the delegate needs.
|
|
30
|
+
|
|
31
|
+
## The Build Protocol, with lanes bound
|
|
32
|
+
|
|
33
|
+
| Stage | Binding |
|
|
34
|
+
|---|---|
|
|
35
|
+
| 0 Route | live probe for access; `cli-run` lanes are $0 and uncapped |
|
|
36
|
+
| 1 Map | the orchestrator sweeps{{STAGE1_LANES}} |
|
|
37
|
+
| 2 Judge | deep tier, on the finished map: a named risk and a named flaw |
|
|
38
|
+
| 3 Build | the orchestrator, against the installed dependency's source |
|
|
39
|
+
| 4 Scan | secret + static + dependency scanners, diff-scoped, fail closed |
|
|
40
|
+
| 5 Attack | security-shaped diff → {{ATTACK_LANE}}. Architecture-shaped → deep tier, build against plan. Never both |
|
|
41
|
+
| 5b Ship | rollback id recorded, explicit human yes |
|
|
42
|
+
| 6 Verify | real test, negative test seen red, old identifier re-grepped to zero |
|
|
43
|
+
| 7 Record | one end-to-end doc, tracker Done with evidence, plan doc deleted |
|
|
44
|
+
|
|
45
|
+
Caps: two deep-tier checkpoints per build. CLI lanes are $0 and do not count.
|
|
46
|
+
|
|
47
|
+
## Numbers and logic
|
|
48
|
+
|
|
49
|
+
Every number, comparison, complexity or equivalence claim goes through a tool that computes (`protocols/numbers-and-logic.md`; companion: codecalc, {{CODECALC_STATUS}}). A lane's figure is re-derived before it is repeated: the cheapest metered lane measured 0 of 11 line citations correct while its conclusions were right.
|
|
50
|
+
|
|
51
|
+
## Memory and record
|
|
52
|
+
|
|
53
|
+
One writer per run; every other lane proposes. Search before writing, index in the same pass (`protocols/memory-and-record.md`; companion, optional: obsidian-tc, {{OBSIDIAN_TC_STATUS}}).
|
|
54
|
+
|
|
55
|
+
## Modifier rules
|
|
56
|
+
|
|
57
|
+
- **Plan big, execute small**, within a build: deep tier plans at Checkpoint 1, the orchestrator executes, bulk and wide searches go down.
|
|
58
|
+
- **Escalation:** never silently retry at the same tier. Escalate one tier or consult deep once, and say which. Two consults that do not unstick it → stop and tell the human.
|
|
59
|
+
- **De-escalation:** a request that sounds deep but is a lookup routes down.
|
|
60
|
+
- **Long context:** mechanical digestion → fast tier in chunks; judgment over a long input → standard tier.
|
|
61
|
+
- **Token discipline on every delegation:** pass only the context the delegate needs, never the conversation.
|
|
62
|
+
- **Effort per agent:** deep xhigh, review and build high, live research medium, bulk low.
|
|
63
|
+
|
|
64
|
+
## Example routings
|
|
65
|
+
|
|
66
|
+
| Task | Route |
|
|
67
|
+
|---|---|
|
|
68
|
+
| "Design the architecture for X" | deep-planner |
|
|
69
|
+
| "Review this service for bugs" | code-reviewer |
|
|
70
|
+
| "Add an endpoint" | the orchestrator builds it |
|
|
71
|
+
| "Why does this silently drop rows sometimes" | deep-planner (unknown cause), then build the fix directly |
|
|
72
|
+
| "Summarize these 30 notes into one index" | bulk-worker |
|
|
73
|
+
{{LANE_EXAMPLES}}
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# TIERS.md: capability tiers and the three cost levers
|
|
2
|
+
|
|
3
|
+
The router routes by capability tier, not model name. Each tier maps to a family alias where the vendor offers one, so a routine version bump needs no file change.
|
|
4
|
+
|
|
5
|
+
| Tier | Purpose | On {{PRIMARY_NAME}} |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| deep | ambiguous planning, architecture, strategy, hard debugging | {{PRIMARY_DEEP}} |
|
|
8
|
+
| standard | code writing, code review, execution, live research synthesis | {{PRIMARY_STANDARD}} |
|
|
9
|
+
| fast | classification, extraction, formatting, bulk summarization | {{PRIMARY_FAST}} |
|
|
10
|
+
| escalation | above deep: only when the human asks, or when deep has already run, the call is still unresolved, and the change is irreversible | your vendor's strongest model, if you have one |
|
|
11
|
+
|
|
12
|
+
Non-primary lanes are owned by `DELEGATION_MATRIX.md`.
|
|
13
|
+
|
|
14
|
+
## Effort per agent (the third lever)
|
|
15
|
+
|
|
16
|
+
Tier sets the price per token. Token discipline sets how many tokens. **Effort sets how hard each call thinks.**
|
|
17
|
+
|
|
18
|
+
| Agent | Tier | Effort | Why |
|
|
19
|
+
|---|---|---|---|
|
|
20
|
+
| deep-planner | deep | xhigh | judges every build twice; expensive to get wrong |
|
|
21
|
+
| code-reviewer | standard | high | every endpoint is internet-facing |
|
|
22
|
+
| builder | standard | high | a botched deploy is the costly failure |
|
|
23
|
+
| live-researcher | standard | medium | tools do the retrieval |
|
|
24
|
+
| bulk-worker | fast | low | the biggest cost win |
|
|
25
|
+
|
|
26
|
+
Dials: drop builder to medium when the plan is airtight; raise code-reviewer to xhigh for a security-critical audit.
|
|
27
|
+
|
|
28
|
+
## Why split tiers: robustness first, cost second
|
|
29
|
+
|
|
30
|
+
The split produces better work. The deep tier steers every build twice, and what it steers is **judgment, never retrieval**: the orchestrator sweeps the blast radius itself and hands the deep tier a finished map. Paying deep-tier rates for a file list is the most expensive routing mistake available.
|
|
31
|
+
|
|
32
|
+
Against a baseline of "standard tier with no consults", default checkpoints are a spend increase. That is the accepted trade, not a saving to claim.
|
|
33
|
+
|
|
34
|
+
## Escalation above deep
|
|
35
|
+
|
|
36
|
+
Fires on exactly two conditions: (a) the human asks for it directly; or (b) all three of: deep has already run on this task, the decision is still unresolved, and the change is irreversible or rewrites a standing rule. (b) is a conjunction, not a mood. A failed attempt is an escalation-ladder event; a hard problem is a deep-tier event; neither reaches the top model alone. It replaces the second deep consult, never adds a third. Say so whenever it fires.
|
|
37
|
+
|
|
38
|
+
## Slotting a new model
|
|
39
|
+
|
|
40
|
+
1. Newer version of an existing family: same tier; aliases pick it up.
|
|
41
|
+
2. New family above your deep model: candidate for deep. Confirm with the human before touching agent files.
|
|
42
|
+
3. New family between tiers: slot by the vendor's own positioning; confirm if it would change who handles a task type.
|
|
43
|
+
4. New cheap family below fast: candidate for fast if quality holds.
|
|
44
|
+
5. Deprecation notice on a slotted model: move the tier immediately and note it here.
|
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
# templates/tools/
|
|
2
|
+
|
|
3
|
+
Companion tools: not AIs, but things the AIs call. Each subfolder is written only when the user selects that tool (`--tools codecalc`, or the interactive question).
|
|
4
|
+
|
|
5
|
+
| Folder | Written | Contents |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| `codecalc/` | when codecalc is selected (recommended, default yes) | `CODECALC.md` (install, per-client registration, the skill) and `mcp/` snippets for the agents its own `setup --write` does not cover |
|
|
8
|
+
| `obsidian-tc/` | when obsidian-tc is selected (optional, default no; needs an Obsidian vault, Node 24+, Ollama or a cloud embeddings key) | `OBSIDIAN-TC.md` (what you need first, install, per-agent registration, security posture) and `mcp/` snippets |
|
|
9
|
+
|
|
10
|
+
The rules the tools serve, `protocols/numbers-and-logic.md` and `protocols/memory-and-record.md`, are in `common/` and are written at every level whether or not a tool was selected: the rule binds, the tool makes it cheap to follow. `src/install.js` writes `templates/tools/<id>/` for every selected tool that has a folder here.
|
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
# codecalc: the calculator, code runner and logic checker your agent calls
|
|
2
|
+
|
|
3
|
+
Repo: https://github.com/The-40-Thieves/codecalc · offline, self-hosted, no key, no telemetry · 52 MCP tools
|
|
4
|
+
|
|
5
|
+
What it gives every agent in this folder: exact arithmetic (`evaluate_expression`, `solve_linear`), code execution in 31 languages (`execute_code`, sandboxed), logic (`z3_check`, `truth_table`), complexity (`analyze_complexity`, `benchmark`), and two proofs nobody else offers cleanly: `verify_translation` (a port behaves identically) and `verify_optimization` (an optimization preserved behaviour). The rule that makes it worth installing is `protocols/numbers-and-logic.md`.
|
|
6
|
+
|
|
7
|
+
## Install (one command)
|
|
8
|
+
|
|
9
|
+
Needs `uv` (https://docs.astral.sh/uv/) and Python 3.10+.
|
|
10
|
+
|
|
11
|
+
```bash
|
|
12
|
+
uvx 'codecalc[full]' setup # prints what it would do, changes nothing
|
|
13
|
+
uvx 'codecalc[full]' setup --write # merges the codecalc entry into your client's config, copies the skill
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
`setup` detects Claude Desktop, Claude Code, Cursor, VS Code and Zed (`--client=NAME` if several), runs two real canaries and ends in one verdict: `ready` / `degraded` / `not-ready`. It backs up the client config it touches. `uvx 'codecalc[full]' doctor` prints a config block with your absolute paths.
|
|
17
|
+
|
|
18
|
+
Pinned form, if you want the version this installer was released with: `uvx 'codecalc[full]=={{CODECALC_PIN}}' setup --write`. `[full]` is the edition that actually runs everything documented (about 120 MB). Base `codecalc` is execution only; symbolic tools then return a `dependency_missing` error naming the extra, never a silent failure.
|
|
19
|
+
|
|
20
|
+
## Agents `setup` does not register (snippets in `mcp/`)
|
|
21
|
+
|
|
22
|
+
| Agent | File to edit | Snippet |
|
|
23
|
+
|---|---|---|
|
|
24
|
+
| Codex CLI | `~/.codex/config.toml` | `mcp/codex.config.toml` |
|
|
25
|
+
| Antigravity `agy` | `~/.gemini/config/mcp_config.json` (remote servers use `serverUrl`; this one is local stdio) | `mcp/agy.mcp_config.json` |
|
|
26
|
+
| Qwen Code | `~/.qwen/settings.json` under `mcpServers` | `mcp/mcpServers.json` |
|
|
27
|
+
| Any client with a `mcpServers` map | its MCP config | `mcp/mcpServers.json` |
|
|
28
|
+
| VS Code | `.vscode/mcp.json` (key is `servers`) | `mcp/vscode.mcp.json` |
|
|
29
|
+
| Zed | `~/.config/zed/settings.json` (key is `context_servers`) | `mcp/zed.settings.json` |
|
|
30
|
+
|
|
31
|
+
Merge the block; do not replace the file. Every other server you have stays as it was.
|
|
32
|
+
|
|
33
|
+
## Install the skill too
|
|
34
|
+
|
|
35
|
+
The tools cannot help a model that never reaches for them. codecalc ships `SKILL.md` inside the package and `setup --write` copies it for Claude Code. For other agents, copy it into that agent's skills folder (Antigravity and Qwen Code read the same `SKILL.md` format). `protocols/numbers-and-logic.md` in this folder is the house rule that points at it.
|
|
36
|
+
|
|
37
|
+
## What it is not
|
|
38
|
+
|
|
39
|
+
Not a cloud sandbox for multi-tenant loads, not a replacement for a vendor's built-in interpreter when zero setup matters more than measurement. Its threat model is single-operator, local, stdio. It earns its keep when the correctness of a claim, not "it ran", is the point.
|
|
40
|
+
|
|
41
|
+
## On a box (level 3)
|
|
42
|
+
|
|
43
|
+
Runs as a stdio server next to the orchestrator CLI; nothing to expose, nothing to bind. Its executor is the Rust sandbox when the platform wheel carries it, and `doctor` tells you which backend you are on before a tool call surprises you.
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
# obsidian-tc: governed memory for your agents (optional)
|
|
2
|
+
|
|
3
|
+
Repo: https://github.com/The-40-Thieves/obsidian-tc · "Obsidian Turbocharged" · 163 tools · TypeScript + Rust · AGPL-3.0 · local by default, no cloud account
|
|
4
|
+
|
|
5
|
+
**Optional.** Skip this if you do not keep notes in Obsidian. The rule it serves, `protocols/memory-and-record.md`, binds either way.
|
|
6
|
+
|
|
7
|
+
## What it gives every agent in this folder
|
|
8
|
+
|
|
9
|
+
A durable, searchable, governed store that the protocols can call by name:
|
|
10
|
+
|
|
11
|
+
| Need in the protocols | obsidian-tc tool |
|
|
12
|
+
|---|---|
|
|
13
|
+
| find what exists before writing (deep research dedupe, gap analysis) | `semantic_search`, `search_text`, `search_regex` |
|
|
14
|
+
| map a rename's blast radius (propagate) | `get_backlinks`, `find_unresolved_links`, `rewrite_link` |
|
|
15
|
+
| record the end-to-end doc (build Stage 7) | `write_note` (compare-and-swap, confirmation on overwrite), `patch_note`, `append_note` |
|
|
16
|
+
| keep inferred content honest | `write_note` with `provenance: "agent_synthesis"` runs a poison scan before the write lands |
|
|
17
|
+
| keep a shared vault safe for several agents | JWT scopes, per-vault folder ACLs, a read-only kill switch, human-in-the-loop tokens |
|
|
18
|
+
|
|
19
|
+
## What you need first (contingent tools)
|
|
20
|
+
|
|
21
|
+
| Requirement | Why | Notes |
|
|
22
|
+
|---|---|---|
|
|
23
|
+
| **An Obsidian vault folder** | it is the store | a folder of markdown files; the Obsidian app itself is only needed for the live plugin bridges |
|
|
24
|
+
| **Node 24+ or Bun 1.1+** | its runtime | stricter than this installer's Node 18; check `node -v` |
|
|
25
|
+
| **Ollama with `nomic-embed-text`** | local embeddings for semantic search | `ollama pull nomic-embed-text`; or configure a cloud embeddings provider (OpenAI, Voyage, Cohere, any OpenAI-shaped endpoint) with a key in your environment |
|
|
26
|
+
| The Obsidian app + its **Local REST API** plugin | only for bridge tools (Dataview, Templater, Excalidraw, OCR, Obsidian Git, the command palette) | optional; without them every filesystem tool still works and bridge tools return a typed `requires_live_obsidian` |
|
|
27
|
+
| Docker (alternative) | run it from `docker-compose.yml` against a bind-mounted vault with no npm install | optional |
|
|
28
|
+
|
|
29
|
+
## Install
|
|
30
|
+
|
|
31
|
+
```bash
|
|
32
|
+
npm install -g obsidian-tc@{{OBSIDIAN_TC_PIN}} # the version this installer was released with; drop the pin for latest
|
|
33
|
+
ollama pull nomic-embed-text
|
|
34
|
+
obsidian-tc /path/to/your/vault # zero-config: one vault named "main", local only
|
|
35
|
+
obsidian-tc plugin install --vault /path/to/your/vault # optional companion plugin, then enable it in Obsidian
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
For more than one vault, auth, ACLs or custom embeddings, write `obsidian-tc.config.json` and pass its path (or set `OBSIDIAN_TC_CONFIG` to it). `obsidian-tc config show <file>` prints the effective config with secrets redacted.
|
|
39
|
+
|
|
40
|
+
## Register it with your agent (snippets in `mcp/`)
|
|
41
|
+
|
|
42
|
+
Every snippet points `OBSIDIAN_TC_CONFIG` at your config file. Replace `/ABSOLUTE/PATH/TO/obsidian-tc.config.json`; the value is a path, not a secret.
|
|
43
|
+
|
|
44
|
+
| Agent | File to edit | Snippet |
|
|
45
|
+
|---|---|---|
|
|
46
|
+
| Claude Code, Claude Desktop, any client with a `mcpServers` map | its MCP config (`.mcp.json` for Claude Code) | `mcp/obsidian-tc.mcpServers.json` |
|
|
47
|
+
| Cursor | one-click badge in the upstream README, or `~/.cursor/mcp.json` | `mcp/obsidian-tc.mcpServers.json` |
|
|
48
|
+
| VS Code | one-click badge, or `.vscode/mcp.json` (key is `servers`) | `mcp/obsidian-tc.vscode.mcp.json` |
|
|
49
|
+
| Zed | `~/.config/zed/settings.json` (key is `context_servers`) | `mcp/obsidian-tc.zed.settings.json` |
|
|
50
|
+
| Codex CLI | `~/.codex/config.toml` | `mcp/obsidian-tc.codex.config.toml` |
|
|
51
|
+
| Antigravity `agy` | `~/.gemini/config/mcp_config.json` | `mcp/obsidian-tc.agy.mcp_config.json` |
|
|
52
|
+
| Qwen Code | `~/.qwen/settings.json` under `mcpServers` | `mcp/obsidian-tc.mcpServers.json` |
|
|
53
|
+
| Claude Desktop, other MCPB hosts | a prebuilt `.mcpb` bundle from the upstream build | see upstream README |
|
|
54
|
+
|
|
55
|
+
Merge the block; do not replace the file.
|
|
56
|
+
|
|
57
|
+
## Security posture, read before a second agent touches it
|
|
58
|
+
|
|
59
|
+
Zero-config mode boots with **auth off and no folder ACL**: anything that can reach the server has the same authority as raw filesystem access to the vault. That is acceptable only because the surface is local-only (the config fail-closes if you enable HTTP on a non-loopback host with auth off, and a DNS-rebinding guard protects loopback). Before exposing it to partially-trusted, remote or multi-agent callers, turn on `auth.mode: "jwt"` and set `acl.readPaths` / `writePaths` / `deletePaths` in the config file. Upstream `SECURITY.md` has the threat model and a private disclosure path.
|
|
60
|
+
|
|
61
|
+
Track record worth knowing: an independent code audit of v1.8.1 (July 2026) found three security-relevant gaps (an ACL fail-closed bypass in enumeration tools, a compare-and-swap bypass through `upsert`, a poison-eligibility gap in preference extraction). All three were fixed upstream before they were filed; verified against the v1.25.0 source on 2026-09-03.
|
|
62
|
+
|
|
63
|
+
## Level 3
|
|
64
|
+
|
|
65
|
+
On a box it runs as a stdio server next to the orchestrator, or as the Docker service, against the vault the box holds. Keep the HTTP transport off unless every caller is on your private mesh and auth is on. Its embeddings run on the box's Ollama, so nothing leaves the machine.
|