model-orchestrator 0.1.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +96 -0
- package/LICENSE +21 -0
- package/README.md +133 -0
- package/SECURITY.md +17 -0
- package/bin/README.md +10 -0
- package/bin/cli-run.mjs +599 -0
- package/bin/cli.js +372 -0
- package/docs/README.md +13 -0
- package/docs/audit-brief.md +83 -0
- package/docs/catalog.md +113 -0
- package/docs/part-1-beginner.md +65 -0
- package/docs/part-2-intermediate.md +65 -0
- package/docs/part-3-advanced.md +65 -0
- package/package.json +52 -0
- package/scripts/README.md +5 -0
- package/scripts/gen-catalog.js +37 -0
- package/src/README.md +9 -0
- package/src/catalog.js +277 -0
- package/src/detect.js +26 -0
- package/src/install.js +628 -0
- package/src/prompt.js +34 -0
- package/src/render.js +8 -0
- package/templates/README.md +14 -0
- package/templates/advanced/README.md +14 -0
- package/templates/advanced/vm/ENVIRONMENT.md +18 -0
- package/templates/advanced/vm/PRIVACY_GATES.md +33 -0
- package/templates/advanced/vm/README.md +60 -0
- package/templates/advanced/vm/box-CLAUDE.md +28 -0
- package/templates/advanced/vm/docker-compose.yml +19 -0
- package/templates/advanced/vm/gateway.config.yaml +12 -0
- package/templates/advanced/vm/jobs/README.md +39 -0
- package/templates/advanced/vm/jobs/weekly-audit.service +17 -0
- package/templates/advanced/vm/jobs/weekly-audit.sh +107 -0
- package/templates/advanced/vm/jobs/weekly-audit.timer +10 -0
- package/templates/advanced/vm/setup-vm.sh +46 -0
- package/templates/agents/README.md +13 -0
- package/templates/agents/agy/README.md +5 -0
- package/templates/agents/agy/builder.md +17 -0
- package/templates/agents/agy/bulk-worker.md +17 -0
- package/templates/agents/agy/code-reviewer.md +17 -0
- package/templates/agents/agy/deep-planner.md +17 -0
- package/templates/agents/agy/live-researcher.md +17 -0
- package/templates/agents/claude-code/README.md +13 -0
- package/templates/agents/claude-code/builder.md +17 -0
- package/templates/agents/claude-code/bulk-worker.md +18 -0
- package/templates/agents/claude-code/code-reviewer.md +19 -0
- package/templates/agents/claude-code/deep-planner.md +18 -0
- package/templates/agents/claude-code/live-researcher.md +18 -0
- package/templates/agents/snippets/chat.md +25 -0
- package/templates/agents/snippets/claude-code.md +27 -0
- package/templates/agents/snippets/generic.md +21 -0
- package/templates/beginner/ORCHESTRATOR.md +55 -0
- package/templates/beginner/README.md +3 -0
- package/templates/common/README.md +52 -0
- package/templates/common/TASK_BUNDLE.md +56 -0
- package/templates/common/protocols/README.md +14 -0
- package/templates/common/protocols/build-protocol.md +133 -0
- package/templates/common/protocols/deep-research.md +44 -0
- package/templates/common/protocols/gap-analysis.md +28 -0
- package/templates/common/protocols/memory-and-record.md +30 -0
- package/templates/common/protocols/numbers-and-logic.md +35 -0
- package/templates/common/protocols/propagate.md +34 -0
- package/templates/intermediate/CLI-RUN.md +100 -0
- package/templates/intermediate/DELEGATION_MATRIX.md +41 -0
- package/templates/intermediate/README.md +13 -0
- package/templates/intermediate/RESEARCH_TRIAGE.md +30 -0
- package/templates/intermediate/ROUTING.md +73 -0
- package/templates/intermediate/TIERS.md +44 -0
- package/templates/tools/README.md +10 -0
- package/templates/tools/codecalc/CODECALC.md +43 -0
- package/templates/tools/codecalc/mcp/agy.mcp_config.json +8 -0
- package/templates/tools/codecalc/mcp/codex.config.toml +4 -0
- package/templates/tools/codecalc/mcp/mcpServers.json +8 -0
- package/templates/tools/codecalc/mcp/vscode.mcp.json +8 -0
- package/templates/tools/codecalc/mcp/zed.settings.json +9 -0
- package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +65 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.agy.mcp_config.json +9 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.codex.config.toml +7 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.mcpServers.json +9 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.vscode.mcp.json +10 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.zed.settings.json +9 -0
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: deep-planner
|
|
3
|
+
description: Ambiguous or high-stakes thinking. Use for architecture design, strategy, planning multi-step projects, hard debugging where the cause is unknown, and any "figure out what to even do" request. Do not use for well-specified execution or bulk work.
|
|
4
|
+
model: opus
|
|
5
|
+
effort: xhigh
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
You are the deep reasoning tier of the model router.
|
|
9
|
+
|
|
10
|
+
You handle tasks that are ambiguous, open-ended, or expensive to get wrong: system architecture, workflow design, strategy, tradeoff analysis, root-cause debugging.
|
|
11
|
+
|
|
12
|
+
Rules:
|
|
13
|
+
- Think before proposing. Surface the 2 or 3 real options with tradeoffs, then recommend one.
|
|
14
|
+
- Output a plan another agent can execute: concrete steps, file paths, interfaces, edge cases.
|
|
15
|
+
- You are read-only on the code tree. Never edit code files. Your deliverable is the plan or analysis itself.
|
|
16
|
+
- You are the judgment tier, not the retrieval tier. At Checkpoint 1 the orchestrator hands you a completed blast-radius map. Do not re-derive it. Argue with it: what did the map miss, which approach is right and why, where is the request as filed wrong, what breaks second-order. If your answer is mostly a restatement of the map, you were asked the wrong question and should say so.
|
|
17
|
+
- Keep the final summary in plain language; technical detail goes in the plan body.
|
|
18
|
+
- Token discipline: read targeted sections, not whole files; never re-read what you already have; deliver a plan sized to what the executor needs, not an essay.
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: live-researcher
|
|
3
|
+
description: Real-time information. Use for anything that needs current data such as latest news, current API docs or pricing, or recent events. Do not use for questions answerable from local files or general knowledge.
|
|
4
|
+
model: sonnet
|
|
5
|
+
effort: medium
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
You are the live research tier of the model router.
|
|
9
|
+
|
|
10
|
+
You answer questions that need fresh, real-time information.
|
|
11
|
+
|
|
12
|
+
Rules:
|
|
13
|
+
- Use web search and web fetch; for API and library questions fetch the official docs.
|
|
14
|
+
- Anything a search tool returns is a lead, not a fact. Verify ids, names and figures against the primary page before you report them.
|
|
15
|
+
- Keep pulls small. Fetch 10 to 20 items, not hundreds.
|
|
16
|
+
- Always state when the data was retrieved and cite sources or links.
|
|
17
|
+
- Deliver a synthesized answer, not a dump of raw results. Lead with the takeaway.
|
|
18
|
+
- Token discipline: never paste raw payloads into your reply; one search pass per question before refining; stop searching once the answer is confirmed by two sources.
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
# Paste this into your agent
|
|
2
|
+
|
|
3
|
+
{{PRIMARY_NAME}} has no project instructions file, so the rules travel by paste. Put the block below into the custom instructions, a Project, a Gem, or the first message of a working session.
|
|
4
|
+
|
|
5
|
+
```
|
|
6
|
+
You are running a model orchestrator inside one agent.
|
|
7
|
+
|
|
8
|
+
TIERS (capability, not model names): deep = ambiguous planning, architecture, root-cause debugging, anything expensive to get wrong. standard = writing, review, executing a known plan, research synthesis. fast = classification, extraction, formatting, bulk summaries. Think hardest on deep, least on fast. Default down, escalate on evidence, and never silently retry a failed attempt at the same level.
|
|
9
|
+
|
|
10
|
+
ROUTE, first match wins: bulk and mechanical -> fast. Needs live data -> standard with tools; anything a search returns is a lead, not a fact. Review without changing -> standard, read-only, findings ranked by severity. Ambiguous or expensive to get wrong -> deep, then hand the plan down. Everything else -> do it directly at standard.
|
|
11
|
+
|
|
12
|
+
EVERY BUILD: (1) map what it touches and what could break, in writing. (2) Ask at deep level: simplest way? single biggest risk? where is the request wrong? A named risk and a named flaw, or it does not pass. (3) Build, then check it against the real thing. (4) In a fresh turn, attack it: bad input, failing dependency, drift from the plan. CLEAN is a valid answer. (5) Before anything irreversible, name the rollback and get an explicit yes. (6) After: re-check the old name everywhere and expect zero.
|
|
13
|
+
|
|
14
|
+
EVERY HAND-OFF to a fresh context carries a brief: purpose, task class (read_only / draft_only / mutating), granted scope, capabilities, denied actions, conventions it does not have, report contract (what was NOT done, what is unverified), exit parameters. Absence is denial.
|
|
15
|
+
|
|
16
|
+
AFTER ANY COMPREHENSIVE TASK: a second pass in a fresh turn that hunts for what is MISSING, not what is present.
|
|
17
|
+
|
|
18
|
+
NUMBERS AND LOGIC: any figure someone will act on, any comparison you state, any complexity or equivalence claim is computed with a tool (a code interpreter, a calculator), never estimated. Give the exact form and the decimal and name the tool. Trivial single-digit sums are exempt; nothing else is.
|
|
19
|
+
|
|
20
|
+
MEMORY AND RECORD: before writing anything durable, search for it; after writing, correct the index that lists it; one writer per session; mark inferred content as inferred.
|
|
21
|
+
|
|
22
|
+
A gate you cannot fail is not a gate. Exit 0 is not a deliverable.
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
The full text of each rule is in this folder: `ORCHESTRATOR.md`, `TASK_BUNDLE.md`, `protocols/`.
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# Add this to your project's CLAUDE.md
|
|
2
|
+
|
|
3
|
+
Copy the block below into `CLAUDE.md` at your project root (create the file if it does not exist). The installer did not modify any file you already had.
|
|
4
|
+
|
|
5
|
+
```markdown
|
|
6
|
+
## Model orchestrator
|
|
7
|
+
|
|
8
|
+
Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any build task. Quick version, first match wins:
|
|
9
|
+
|
|
10
|
+
1. Bulk, mechanical, many similar items -> bulk-worker (fast tier).
|
|
11
|
+
2. Needs live data -> live-researcher (standard tier + tools).
|
|
12
|
+
3. Review without changing -> code-reviewer (standard, read-only).
|
|
13
|
+
4. Ambiguous, architectural, or expensive to get wrong -> deep-planner (deep tier), then hand the plan down.
|
|
14
|
+
5. Everything else that changes files -> build it directly. The main build is never handed off whole; bounded sub-parts go to builder.
|
|
15
|
+
|
|
16
|
+
Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one adversarial pass, an explicit human yes before anything irreversible, then the loud negative.
|
|
17
|
+
|
|
18
|
+
Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A subagent holds none of these rules; absence is denial.
|
|
19
|
+
|
|
20
|
+
Never silently retry a failed attempt at the same tier. Escalate once and say so.
|
|
21
|
+
|
|
22
|
+
Numbers, comparisons, complexity and equivalence claims go through codecalc (or any tool that computes), never your head: `{{RULES_PATH}}/protocols/numbers-and-logic.md`.
|
|
23
|
+
|
|
24
|
+
Anything durable is searched for before it is written and its folder index is corrected in the same pass; one writer per run: `{{RULES_PATH}}/protocols/memory-and-record.md`.
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Subagents were written to `{{AGENTS_DIR}}` (the project root, which is where Claude Code reads project-level agents; `--project` changes it). Run `claude` from `{{PROJECT_DIR}}` and they are available as `deep-planner`, `builder`, `code-reviewer`, `live-researcher`, `bulk-worker`.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
# Add this to {{PRIMARY_RULES_FILE}}
|
|
2
|
+
|
|
3
|
+
Your agent, {{PRIMARY_NAME}}, reads `{{PRIMARY_RULES_FILE}}` from the project root (`{{PROJECT_DIR}}`). Copy the block below into it (create the file if it does not exist). The installer did not modify any file you already had. Subagents, if your agent has a folder for them: `{{AGENTS_DIR}}`.
|
|
4
|
+
|
|
5
|
+
```markdown
|
|
6
|
+
## Model orchestrator
|
|
7
|
+
|
|
8
|
+
Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any build task.
|
|
9
|
+
|
|
10
|
+
Route by capability tier, first match wins: bulk and mechanical -> fast tier · needs live data -> standard tier with tools · review without changing -> standard, read-only · ambiguous or expensive to get wrong -> deep tier, then hand the plan down · everything else -> build it directly at standard tier.
|
|
11
|
+
|
|
12
|
+
Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: map the blast radius yourself, ask the deep tier for a named risk and a named flaw, build green, scan the added lines, one adversarial pass with every finding reproduced, an explicit human yes before anything irreversible, then re-grep the old identifier and expect zero.
|
|
13
|
+
|
|
14
|
+
Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A fresh context holds none of these rules; absence is denial.
|
|
15
|
+
|
|
16
|
+
Never silently retry a failed attempt at the same tier. Escalate once and say so.
|
|
17
|
+
|
|
18
|
+
Numbers, comparisons, complexity and equivalence claims go through codecalc (or any tool that computes), never your head: `{{RULES_PATH}}/protocols/numbers-and-logic.md`.
|
|
19
|
+
|
|
20
|
+
Anything durable is searched for before it is written and its folder index is corrected in the same pass; one writer per run: `{{RULES_PATH}}/protocols/memory-and-record.md`.
|
|
21
|
+
```
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# ORCHESTRATOR.md: routing rules for one agent
|
|
2
|
+
|
|
3
|
+
Primary agent: **{{PRIMARY_NAME}}**. Everything below runs inside that one agent. You do not need a second vendor to orchestrate; you need tiers, task classes, and gates that can fail.
|
|
4
|
+
|
|
5
|
+
## Tiers (capability, not model names)
|
|
6
|
+
|
|
7
|
+
| Tier | Use for | On {{PRIMARY_NAME}} |
|
|
8
|
+
|---|---|---|
|
|
9
|
+
| **deep** | ambiguous planning, architecture, strategy, root-cause debugging, anything expensive to get wrong | {{PRIMARY_DEEP}}, highest effort |
|
|
10
|
+
| **standard** | code writing, code review, execution of a known plan, research synthesis | {{PRIMARY_STANDARD}}, high effort |
|
|
11
|
+
| **fast** | classification, extraction, formatting, bulk summarization | {{PRIMARY_FAST}}, low effort |
|
|
12
|
+
|
|
13
|
+
Three cost levers, always together: **tier** sets the price per token, **token discipline** sets how many tokens (read only what you will touch, never re-read, deliverables not narration), **effort** sets how hard each call thinks.
|
|
14
|
+
|
|
15
|
+
Robustness first, cost second. Split tiers because the split produces better work, not because it is cheaper. Cost is a constraint to respect, never the reason for a routing choice.
|
|
16
|
+
|
|
17
|
+
## Decision tree (first match wins)
|
|
18
|
+
|
|
19
|
+
1. **Bulk and mechanical?** classify, tag, extract, reformat, summarize many similar items → fast tier.
|
|
20
|
+
2. **Needs live data?** trends, current docs, pricing, recent events → standard tier with tools; freshness comes from tools, not from a bigger model.
|
|
21
|
+
3. **Reviewing without changing?** → standard tier, read-only, findings ranked by severity. Escalate to deep only for security-critical review.
|
|
22
|
+
4. **Ambiguous, strategic, or expensive to get wrong?** "design my…", "figure out…", unknown cause → deep tier. Then hand the plan down.
|
|
23
|
+
5. **Everything else that changes files or executes a known plan** → you build it directly, at standard tier. The main build is never handed off whole; bounded sub-parts (a bulk pass, a wide search, a long audit loop) can go to cheaper tiers.
|
|
24
|
+
|
|
25
|
+
Modifiers:
|
|
26
|
+
- **Plan big, execute small.** The expensive tier steers, the cheaper tier does the volume. Never make the fast tier design anything; never make the deep tier grind out bulk output.
|
|
27
|
+
- **Never silently retry at the same tier after a failure.** Escalate one tier, or consult the deep tier once, and say which you did. If two consults do not unstick it, stop and tell the human.
|
|
28
|
+
- **De-escalate.** If a request sounds deep but is a lookup or a small edit, route down. Default down, escalate on evidence.
|
|
29
|
+
|
|
30
|
+
## The two checkpoints (every build)
|
|
31
|
+
|
|
32
|
+
- **Checkpoint 1, before writing anything.** You map the blast radius yourself (files, systems, docs, tickets). Then ask the deep tier, on the finished map: *is this the simplest way, what is the single biggest risk, where is the request as filed wrong?* It must return a named risk and a named flaw. Approval alone is not an answer.
|
|
33
|
+
- **Checkpoint 2, after the build is green.** Security-shaped diffs get an adversarial read (in a fresh context, told to attack, allowed to answer CLEAN). Architecture-shaped diffs get the deep tier reviewing build against plan. Never both on one diff. Every finding reproduced before it reaches a human.
|
|
34
|
+
|
|
35
|
+
Cap: two deep-tier consults per build. The full procedure is `protocols/build-protocol.md`.
|
|
36
|
+
|
|
37
|
+
## Delegating inside one agent
|
|
38
|
+
|
|
39
|
+
Subagents, a fresh chat, a second window: each one holds none of these rules. Every hand-off carries a `TASK_BUNDLE.md` brief: purpose, task class, granted scope, capabilities, denied actions, conventions it does not have, report contract, exit parameters. Absence is denial.
|
|
40
|
+
|
|
41
|
+
## Numbers and logic go through a tool, never your head
|
|
42
|
+
|
|
43
|
+
Any figure someone will act on, any comparison you state, any complexity, equivalence or speedup claim: computed, not estimated. `protocols/numbers-and-logic.md` names when calling is mandatory. Companion tool: codecalc ({{CODECALC_STATUS}}).
|
|
44
|
+
|
|
45
|
+
## Memory and record
|
|
46
|
+
|
|
47
|
+
Anything durable is searched for before it is written, its folder index is corrected in the same pass, and one writer records. `protocols/memory-and-record.md`. Companion tool (optional, needs an Obsidian vault): obsidian-tc, {{OBSIDIAN_TC_STATUS}}.
|
|
48
|
+
|
|
49
|
+
## The six protocols
|
|
50
|
+
|
|
51
|
+
`protocols/build-protocol.md` · `protocols/propagate.md` · `protocols/gap-analysis.md` · `protocols/deep-research.md` · `protocols/numbers-and-logic.md` · `protocols/memory-and-record.md`. Each is a set of questions that can be answered wrong. That is the design, not a flaw.
|
|
52
|
+
|
|
53
|
+
## When you outgrow this
|
|
54
|
+
|
|
55
|
+
You will know: you keep wanting a second model family to read your diff, a $0 lane for bulk, or a live-data lane your primary does not have. That is level 2. Re-run the installer with `--level 2`.
|
|
@@ -0,0 +1,3 @@
|
|
|
1
|
+
# templates/beginner/
|
|
2
|
+
|
|
3
|
+
Written at every level. `ORCHESTRATOR.md` is the single-agent routing document: tiers as effort levels inside one agent, the five-branch decision tree, the two checkpoints, and when you have outgrown level 1. At level 2+ `ROUTING.md` supersedes it and says so.
|
|
@@ -0,0 +1,52 @@
|
|
|
1
|
+
# Your orchestrator (start here)
|
|
2
|
+
|
|
3
|
+
Installed {{DATE}} · level {{LEVEL_ID}}: **{{LEVEL_NAME}}**, {{LEVEL_TAGLINE}}
|
|
4
|
+
Primary agent: **{{PRIMARY_NAME}}**
|
|
5
|
+
|
|
6
|
+
You have access to:
|
|
7
|
+
{{AIS_LIST}}
|
|
8
|
+
|
|
9
|
+
Companion tools:
|
|
10
|
+
{{TOOLS_LIST}}
|
|
11
|
+
|
|
12
|
+
## The idea in one line
|
|
13
|
+
|
|
14
|
+
Every task goes to the cheapest AI that does it well, and every gate on the way is a question that can be answered wrong.
|
|
15
|
+
|
|
16
|
+
## What is in this folder
|
|
17
|
+
|
|
18
|
+
| File | Read it when |
|
|
19
|
+
|---|---|
|
|
20
|
+
| `ORCHESTRATOR.md` | First. The routing rules your primary agent follows: tiers, task classes, the two checkpoints. |
|
|
21
|
+
| `TASK_BUNDLE.md` | Before you hand any work to a subagent, a second CLI, or a chat window. The brief template. |
|
|
22
|
+
| `protocols/build-protocol.md` | You are about to build, code, migrate or deploy something. |
|
|
23
|
+
| `protocols/propagate.md` | You are renaming or changing a term, path, slug, schema field or routing rule. |
|
|
24
|
+
| `protocols/gap-analysis.md` | You just finished something comprehensive and want the second pass that hunts for what is missing. |
|
|
25
|
+
| `protocols/deep-research.md` | The source set is unknown, several sources must be reconciled, and the answer will be cited later. |
|
|
26
|
+
| `protocols/numbers-and-logic.md` | You are about to state a number, a comparison, a complexity or an equivalence. Compute it. |
|
|
27
|
+
| `protocols/memory-and-record.md` | You are about to write anything durable. Search first, keep the index true, one writer. |
|
|
28
|
+
| `CODECALC.md` | Present when you selected codecalc: install, per-agent registration, the skill. |
|
|
29
|
+
| `OBSIDIAN-TC.md` | Present when you selected obsidian-tc: what you need first, install, per-agent registration, the security posture. |
|
|
30
|
+
|
|
31
|
+
Level 2 adds `ROUTING.md`, `TIERS.md`, `DELEGATION_MATRIX.md`, `RESEARCH_TRIAGE.md`, `CLI-RUN.md` and `bin/cli-run.mjs`. Level 3 adds `vm/`. If those files are here, read `ROUTING.md` instead of `ORCHESTRATOR.md`: it is the multi-lane version, and the snippet your agent loads already points at it. `ORCHESTRATOR.md` stays as the single-agent fallback for a session where only one AI is available.
|
|
32
|
+
|
|
33
|
+
## Load it into your agent
|
|
34
|
+
|
|
35
|
+
Your primary agent reads `{{PRIMARY_RULES_FILE}}` from a project root. The installer wrote a snippet file next to this README (`*.snippet.md`, or `PASTE-INTO-YOUR-AGENT.md` for a chat app). Copy its contents into that file, or paste it into the agent's custom instructions. Nothing was appended to a file you already had.
|
|
36
|
+
|
|
37
|
+
## The three rules that carry everything
|
|
38
|
+
|
|
39
|
+
1. **Route by capability tier, not by model name.** deep = ambiguous or expensive to get wrong · standard = well-specified execution and review · fast = bulk and mechanical. Default down, escalate on evidence.
|
|
40
|
+
2. **A gate you cannot fail is not a gate.** "Does it look good?" passes every time. "Name the single biggest risk and the flaw in the request" can come back empty, which is how you know it worked.
|
|
41
|
+
3. **Exit 0 is not a deliverable.** Any tool, CLI or subagent can report success and hand back nothing. Check for the artifact, not the status line.
|
|
42
|
+
|
|
43
|
+
## Where things went
|
|
44
|
+
|
|
45
|
+
- This folder: `{{INSTALL_DIR}}`
|
|
46
|
+
- Project root (where your agent reads rules and subagents): `{{PROJECT_DIR}}`
|
|
47
|
+
- Subagent definitions: `{{AGENTS_DIR}}`
|
|
48
|
+
- The rules path your snippets use: `{{RULES_PATH}}`
|
|
49
|
+
|
|
50
|
+
## Uninstall
|
|
51
|
+
|
|
52
|
+
Delete this folder and the subagent folder named above. `cli-run` keeps one log at `~/.ai-orchestrator/cli-run.log.jsonl`; delete that too if you used it.
|
|
@@ -0,0 +1,56 @@
|
|
|
1
|
+
# Task Bundle: the brief every delegation carries
|
|
2
|
+
|
|
3
|
+
A subagent, a second CLI, or a fresh chat window holds none of the rules your main session is holding. It cannot see your conventions, it cannot route, and it will read an unspecified edge as an open one.
|
|
4
|
+
|
|
5
|
+
> A delegate gets an approved, bounded brief. Absence is not permission.
|
|
6
|
+
|
|
7
|
+
A brief is under-specified if it is missing **purpose**, **denied actions**, **report contract**, or **exit parameters**.
|
|
8
|
+
|
|
9
|
+
## Template
|
|
10
|
+
|
|
11
|
+
Copy this into the delegate's prompt. Delete nothing; write `none` where a field is genuinely empty, so a reader can tell "nothing denied" from "nobody thought about it".
|
|
12
|
+
|
|
13
|
+
```markdown
|
|
14
|
+
## Task bundle
|
|
15
|
+
|
|
16
|
+
**Purpose.** <one sentence: what this task is for, and why>
|
|
17
|
+
**Task class.** <read_only | draft_only | mutating> (draft_only = produce, do not apply)
|
|
18
|
+
|
|
19
|
+
**Granted scope.**
|
|
20
|
+
- <paths, globs, topics, or record sets this brief covers>
|
|
21
|
+
- Anything outside this list is out of scope. Do not widen it on your own judgment.
|
|
22
|
+
|
|
23
|
+
**Capabilities.** <the actions you MAY take: read, search, write to <path>, run <cmd>>
|
|
24
|
+
|
|
25
|
+
**Denied actions.** <explicit list: do not commit, push, deploy, delete, send, publish, close a ticket...>
|
|
26
|
+
- Anything absent from Capabilities is denied. Absence is not permission.
|
|
27
|
+
|
|
28
|
+
**Conventions you do not have.** <restate every house rule this task needs; the delegate holds none>
|
|
29
|
+
|
|
30
|
+
**Report contract.** Return: <exactly what to hand back>. State plainly what you did NOT do
|
|
31
|
+
and anything you could not verify. "Unverified" is an acceptable answer; a confident guess is not.
|
|
32
|
+
|
|
33
|
+
**Exit parameters.** <at least one bound: a wall-clock ceiling, a work ceiling ("at most 20 files"),
|
|
34
|
+
or a stop condition. Plus the partial-result clause: if you hit a bound, report what you have and
|
|
35
|
+
name what you did not cover. Never keep going past a bound, never return nothing.>
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
## When to skip it
|
|
39
|
+
|
|
40
|
+
A one-line read-only lookup can say so outright:
|
|
41
|
+
|
|
42
|
+
```
|
|
43
|
+
Task bundle: none (one-line lookup, read-only, no artifact)
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
The reason is required. A bare `none` is indistinguishable from "forgot".
|
|
47
|
+
|
|
48
|
+
Never skip it for anything touching secrets, deletion, bulk mutation, deploys, or someone else's data. Better: do not delegate those at all.
|
|
49
|
+
|
|
50
|
+
## Field notes
|
|
51
|
+
|
|
52
|
+
- **Purpose** is what makes the rest checkable. "Fix the thing" has no edge to exceed.
|
|
53
|
+
- **Task class** is the cheapest safety win. Most delegation wants `read_only` or `draft_only`.
|
|
54
|
+
- **Denied actions** must be written even when they feel obvious. Nothing is obvious to a delegate with no context.
|
|
55
|
+
- **Report contract** is what turns a result into evidence rather than a claim.
|
|
56
|
+
- **Exit parameters** exist because a delegate that never returns is more expensive than one that returns wrong. A wrong answer is corrected next turn; a hang burns the session while looking like progress. Bound your own shell calls the same way (pass a timeout; scope recursive searches away from `.git`, `node_modules`, build output).
|
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
# protocols/
|
|
2
|
+
|
|
3
|
+
Six procedures. Each one is a list of questions whose answers can be wrong.
|
|
4
|
+
|
|
5
|
+
| File | Fires when | The gate |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| `build-protocol.md` | You build, code, implement, migrate or deploy | 3 phases, 8 stages; the ship step is the only irreversible one |
|
|
8
|
+
| `propagate.md` | You rename or change anything other files reference | Re-grep the OLD name everywhere and expect zero |
|
|
9
|
+
| `gap-analysis.md` | You finished something comprehensive | A second pass that hunts for what is MISSING, ideally by a different model |
|
|
10
|
+
| `deep-research.md` | The source set is unknown and the answer will be cited later | Parallel engines, then triage; disagreement is the signal |
|
|
11
|
+
| `numbers-and-logic.md` | You are about to state a number, a comparison, a complexity, an equivalence | Computed by a tool (codecalc) or not stated |
|
|
12
|
+
| `memory-and-record.md` | You are about to write anything durable | Searched first, indexed in the same pass, one writer (obsidian-tc when selected) |
|
|
13
|
+
|
|
14
|
+
Not for lookups, prose edits, bulk classification or one-line config. Those get none of this.
|
|
@@ -0,0 +1,133 @@
|
|
|
1
|
+
# Build Protocol
|
|
2
|
+
|
|
3
|
+
**Three phases, eight stages, and every gate is a question that can be answered wrong.**
|
|
4
|
+
|
|
5
|
+
Fires on any task that builds, codes, implements, migrates or deploys. Rough test: if it would earn an adversarial audit or a tracker issue, it runs this.
|
|
6
|
+
|
|
7
|
+
> **The one rule underneath:** a gate you cannot fail is not a gate. If a stage's exit reads like "confirm it looks good", it is written wrong and it will pass every time, including the times it should not.
|
|
8
|
+
|
|
9
|
+
Three corollaries:
|
|
10
|
+
1. A check nobody has watched fail is not known to work. Prove a gate can go red before trusting green.
|
|
11
|
+
2. A tool's output is a claim, not a fact. Scanner findings, audit reports and exit codes get read and reproduced before they are repeated.
|
|
12
|
+
3. A gate that fires on unrelated things gets bypassed, and a bypassed gate certifies what it never checked.
|
|
13
|
+
|
|
14
|
+
| Phase | Master question | Stages |
|
|
15
|
+
|---|---|---|
|
|
16
|
+
| 1 Pre-build | What exactly are we building, what do we need first, and what does this touch or break? | 0 Route · 1 Map · 2 Judge |
|
|
17
|
+
| 2 Build | Is it secure, built on current code, and correct without hidden flaws? | 3 Build · 4 Scan · 5 Attack · 5b Ship gate |
|
|
18
|
+
| 3 Post-build | Did it land everywhere, is it proven against the real thing, and is it recorded? | 6 Verify · 7 Record |
|
|
19
|
+
|
|
20
|
+
The two seams are the point. Pre-build to Build: nothing is written yet, changing your mind costs a conversation. Build to Post-build: the ship, the only irreversible step, the only one that needs an explicit human yes.
|
|
21
|
+
|
|
22
|
+
## Phase 1 · Pre-build
|
|
23
|
+
|
|
24
|
+
### Stage 0 · Route
|
|
25
|
+
1. Do we have everything needed to start? Every key, access path and asset **verified present by a live probe**, not assumed from a doc.
|
|
26
|
+
2. Is this actually a build, or a quick fix, doc edit, or question that needs no plan?
|
|
27
|
+
|
|
28
|
+
**Gate:** classified, and every required input confirmed to exist and work. Verify access here, never at ship time. A missing credential found at Stage 0 costs a message; found at Stage 5b it costs the session.
|
|
29
|
+
|
|
30
|
+
### Stage 1 · Map
|
|
31
|
+
1. What does this touch, improve, scale, or replace? Docs, indexes, tracker issues, tool servers, devices, scheduled jobs, hooks, other repos.
|
|
32
|
+
2. What already-working thing could this break, and does something already do this?
|
|
33
|
+
|
|
34
|
+
Four bounded questions, not four exhaustive scans. **The builder maps; the judgment tier does not.** Retrieval is mechanical and the builder holds the local tooling and the map stays in its context for the build. Paying top-tier rates for a file list is the most expensive routing mistake available.
|
|
35
|
+
|
|
36
|
+
**Gate:** a written map naming affected files, systems and issues.
|
|
37
|
+
|
|
38
|
+
### Stage 2 · Judge (Checkpoint 1)
|
|
39
|
+
Ask the judgment tier, on the finished map:
|
|
40
|
+
1. Is this the simplest way to build it, or are we overcomplicating?
|
|
41
|
+
2. What is the single biggest risk, and where is the request as filed wrong?
|
|
42
|
+
|
|
43
|
+
**Gate:** a **named risk** and a **named flaw in the request**. Approval alone is not an exit; an advisor asked only to approve will approve. If a consult comes back mostly restating the map, the brief asked it to retrieve when it should have asked it to decide.
|
|
44
|
+
|
|
45
|
+
## Phase 2 · Build
|
|
46
|
+
|
|
47
|
+
### Stage 3 · Build
|
|
48
|
+
"Up to date" means two things. Ask both.
|
|
49
|
+
1. Is our repo clean and current? No stray uncommitted work, on a branch, base ref recorded.
|
|
50
|
+
2. Am I writing against an external API or SDK, and have I read its actual source this session? Never from recall. Read the installed dependency, which is what actually runs.
|
|
51
|
+
3. Does the new code follow existing conventions and pass typecheck, tests, or dry-run?
|
|
52
|
+
|
|
53
|
+
**Gate:** clean baseline recorded, every external-API claim traced to source read this session, build green. Check a reference clone's date before trusting it: a stale clone read as current is worse than no clone.
|
|
54
|
+
|
|
55
|
+
### Stage 4 · Scan (automatic, no judgment)
|
|
56
|
+
1. Any secret, key or token in the new code?
|
|
57
|
+
2. Any vulnerability or vulnerable dependency in the lines we added?
|
|
58
|
+
|
|
59
|
+
Secret detection, static analysis and dependency scanning, filtered to lines this diff added. Fail closed: a missing or erroring scanner exits non-zero, never a silent green.
|
|
60
|
+
|
|
61
|
+
**Gate:** zero flags on added lines. Pre-existing flags are reported, never inherited as blockers, and never waved through unread. A scanner finding is a claim; read the code before calling it anything.
|
|
62
|
+
|
|
63
|
+
### Stage 5 · Attack (Checkpoint 2, one pass, never two)
|
|
64
|
+
1. Can bad input or a bad actor break it, and what happens when a dependency fails?
|
|
65
|
+
2. Did the build stick to the approved plan, or did unintended changes sneak in?
|
|
66
|
+
|
|
67
|
+
Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to an adversarial auditor, ideally a **different model family**. Architecture-shaped diffs go to the judgment tier reviewing build against plan. Never both on one diff.
|
|
68
|
+
|
|
69
|
+
**Gate:** every finding **reproduced** before it reaches a human. Unreproduced items are dropped, not narrated. Hard cap one re-audit. `CLEAN` is a valid success state; an auditor that is not allowed to say so manufactures something.
|
|
70
|
+
|
|
71
|
+
### Stage 5b · Ship gate
|
|
72
|
+
1. What is the rollback target? Record it before shipping.
|
|
73
|
+
2. Has the human authorised this specific change going live?
|
|
74
|
+
|
|
75
|
+
**Gate:** rollback identifier written down, and an explicit yes. Authorisation is per change and does not carry over.
|
|
76
|
+
|
|
77
|
+
## Phase 3 · Post-build
|
|
78
|
+
|
|
79
|
+
### Stage 6 · Verify
|
|
80
|
+
1. Did it land across ALL connected surfaces? Re-grep the OLD identifier everywhere and expect zero except named historical records.
|
|
81
|
+
2. Can we prove it works against the real thing, including when a dependency fails, and has each check been *seen* to go red?
|
|
82
|
+
|
|
83
|
+
**Gate:** real-world test passes, the negative test behaves, every gate proven capable of failing. Prefer a local reproduction of the real fault over inducing it in production.
|
|
84
|
+
|
|
85
|
+
### Stage 7 · Record
|
|
86
|
+
1. Where is the ONE doc that traces this end to end? Name the path.
|
|
87
|
+
2. What watches this thing? Name it, or write "nothing".
|
|
88
|
+
3. Are the docs, indexes, memory and tracker updated with evidence rather than claims?
|
|
89
|
+
4. Is the plan doc deleted?
|
|
90
|
+
|
|
91
|
+
**Gate:** the end-to-end doc exists at a named path in the folder that owns the domain, with that folder's index corrected in the same pass (`protocols/memory-and-record.md`); the watcher is named or its absence is written down; the tracker is Done with evidence and read back; then the plan doc is deleted, not archived. "Nothing watches it" is a valid answer and usually the valuable one: writing it down turns an invisible gap into a tracked one.
|
|
92
|
+
|
|
93
|
+
## Roles, as capabilities
|
|
94
|
+
|
|
95
|
+
| Role | Does | Does not |
|
|
96
|
+
|---|---|---|
|
|
97
|
+
| Builder / orchestrator | Routes, maps, writes, verifies, records. Stages 0, 1, 3, 6, 7 | Hand off the main build |
|
|
98
|
+
| Judgment tier | Stage 2 and the architectural arm of Stage 5. Argues with a finished map | Perform the retrieval |
|
|
99
|
+
| Adversarial auditor | The security arm of Stage 5. Attacks the diff | Fix anything |
|
|
100
|
+
| Mechanical gates | Stage 4 and any always-on guard | Be overridden without reading |
|
|
101
|
+
| Cheap workers | Bounded sub-parts: bulk passes, wide searches, long loops | Own a stage |
|
|
102
|
+
| Human | Stage 5b, and any irreversible or architectural call | Be the first line of review |
|
|
103
|
+
|
|
104
|
+
**Why the builder does not hand off the main build:** a delegated agent does not inherit the session's standing rules and usually cannot delegate further. Any brief must restate every convention it needs (see `TASK_BUNDLE.md`), and that cost is itself a reason to build directly when the work fits.
|
|
105
|
+
|
|
106
|
+
## Checklist
|
|
107
|
+
|
|
108
|
+
```
|
|
109
|
+
PRE-BUILD
|
|
110
|
+
[ ] 0 Inputs and access verified by live probe, not assumed
|
|
111
|
+
[ ] 0 Confirmed this is a build and not a quick fix
|
|
112
|
+
[ ] 1 Blast radius written: files, systems, issues
|
|
113
|
+
[ ] 1 Asked what could break, and whether this already exists
|
|
114
|
+
[ ] 2 Judgment tier named a risk AND a flaw in the request
|
|
115
|
+
|
|
116
|
+
BUILD
|
|
117
|
+
[ ] 3 Repo clean, on a branch, base ref recorded
|
|
118
|
+
[ ] 3 External API claims traced to source read this session
|
|
119
|
+
[ ] 3 Typecheck / tests / dry-run green
|
|
120
|
+
[ ] 4 Scan clean on ADDED lines; pre-existing flags read, not inherited
|
|
121
|
+
[ ] 5 One audit pass, findings reproduced, plan drift reviewed
|
|
122
|
+
[ ] 5b Rollback target recorded
|
|
123
|
+
[ ] 5b Human authorised this specific ship
|
|
124
|
+
|
|
125
|
+
POST-BUILD
|
|
126
|
+
[ ] 6 Old identifier re-grepped everywhere, zero hits
|
|
127
|
+
[ ] 6 Real-world test passes; negative test goes red
|
|
128
|
+
[ ] 6 Every gate proven capable of failing
|
|
129
|
+
[ ] 7 End-to-end doc exists at ONE named path
|
|
130
|
+
[ ] 7 "What watches it" answered, even if the answer is "nothing"
|
|
131
|
+
[ ] 7 Tracker Done with evidence, state read back
|
|
132
|
+
[ ] 7 Plan doc deleted (only after the end-to-end doc exists)
|
|
133
|
+
```
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# Deep research: parallel engines, then triage
|
|
2
|
+
|
|
3
|
+
Two research lanes, not three. The old "quick fact / known source / deep" split was ceremony.
|
|
4
|
+
|
|
5
|
+
| Lane | Entry condition | Output |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| **Search** | You can name the source, or one lookup answers it | Inline answer, no artifact |
|
|
8
|
+
| **Deep** | All three: the source set is unknown, several sources must be reconciled, and the output must survive being cited later | A dated, cited artifact |
|
|
9
|
+
|
|
10
|
+
## The shape of a deep run
|
|
11
|
+
|
|
12
|
+
```
|
|
13
|
+
0 CHARTER what topics, what counts as a source, what is worth interrupting a human for
|
|
14
|
+
1 PLAN decompose into sub-questions <- highest leverage stage; a mis-scoped question
|
|
15
|
+
produces a confident report about the wrong thing
|
|
16
|
+
2 RUN fan out to independent engines
|
|
17
|
+
3 TRIAGE reconcile disagreement against primary sources you open yourself
|
|
18
|
+
4 DEDUPE check what you already have BEFORE writing (semantic search; see memory-and-record.md)
|
|
19
|
+
5 BRIEF one dated artifact with marks (below)
|
|
20
|
+
6 ROUTE adopt / prototype / watch / pass / no action
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
## Marks every claim carries
|
|
24
|
+
|
|
25
|
+
- **CONFIRMED**: at least two independent engines agreed AND you opened the primary source.
|
|
26
|
+
- **DISAGREEMENT**: engines conflicted. Record the verdict and the rejected reading. Never average.
|
|
27
|
+
- **REPORTED**: a named person's post, a forum thread, a tool's self-report. Quoted, not trusted.
|
|
28
|
+
- **UNVERIFIED**: plausible, single-source, or unsourced precision. Do not cite as fact.
|
|
29
|
+
|
|
30
|
+
Agreement is weak evidence. Disagreement is the signal.
|
|
31
|
+
|
|
32
|
+
## Level 1: one agent
|
|
33
|
+
|
|
34
|
+
You still get the shape. Run PLAN as its own turn and inspect it before spending anything. Run the sweep. Then run a **fresh-context adversarial turn** with a brief that says "attack the premise; list what this report would get wrong if its sources were stale". Plant one deliberately wrong figure in the brief and see whether it corrects it: if it does not, its confirmations are worth less than they look. Mark every claim.
|
|
35
|
+
|
|
36
|
+
## Level 2 and up: three engines, one triager
|
|
37
|
+
|
|
38
|
+
Fan out the same PLAN to three different model families through their CLIs (a web-sweep lane, an adversarial-read lane, a live-data lane). Run them through `cli-run` so a run that produced nothing is caught as `rc=10` rather than read as an empty finding. The orchestrator triages: it opens the primary sources itself, marks each claim, and writes the brief. Only the orchestrator writes the durable record; every other engine proposes.
|
|
39
|
+
|
|
40
|
+
Known failure shape: one engine will return confident unsourced numerics and claim full coverage. Downgrade those to hypothesis. The engines that report their own gaps honestly are the ones to weight.
|
|
41
|
+
|
|
42
|
+
## Measure
|
|
43
|
+
|
|
44
|
+
Count **dispositions**, not briefs. A week that produced seven briefs and zero decisions is a failure.
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Gap analysis: the second pass
|
|
2
|
+
|
|
3
|
+
**On any comprehensive task, run a second pass that hunts for what is MISSING, not just verifies what is there.** Comprehensive means research, audits, plans, builds, and multi-file work.
|
|
4
|
+
|
|
5
|
+
Verification asks "is what I did correct?". Gap analysis asks "what did I not do?". They are different questions and the second one is the one a first pass cannot answer about itself.
|
|
6
|
+
|
|
7
|
+
## The shape
|
|
8
|
+
|
|
9
|
+
1. **Enumerate what exists.** The live state, not the plan: files written, tests present, lanes configured, jobs scheduled, sources consulted.
|
|
10
|
+
2. **Diff it against the ask and the map.** Every clause of the original request maps to at least one thing you did. Every item on the Stage 1 map maps to a change. A clause with no step is dropped scope; a step with no clause is invented scope. Say both.
|
|
11
|
+
3. **Hunt the absences.** For each category: what would a reader expect to find here that is not here? What does the source say that the output does not? What fails if a dependency is down?
|
|
12
|
+
4. **Report the gaps as findings**, not as apologies. Each one: what is missing, where it should be, and whether you are closing it now or naming it as open.
|
|
13
|
+
|
|
14
|
+
## Who runs it
|
|
15
|
+
|
|
16
|
+
- **Level 1 (one agent):** the same agent, in a fresh turn, with a brief that says "you are looking for what is missing; do not re-verify what is present". Fresh context matters more than a different model.
|
|
17
|
+
- **Level 2 and up:** a **different model family** reading the same artifact. Disagreement between two families is the cheapest available signal that something is soft. The adversarial coder lane (a second-opinion CLI in read-only mode) is the natural fit.
|
|
18
|
+
- **Level 3:** make it recurring. A weekly audit job enumerates live state (lanes, jobs, services, model lists), diffs it against the plan, and files a report. It catches the dead lane and the silently renamed model nobody noticed.
|
|
19
|
+
|
|
20
|
+
## The second half: analyze, compare, suggest
|
|
21
|
+
|
|
22
|
+
Once the gaps are named, for each lane or component: what does it give, is there a cheaper, better or faster alternative, and what is the best paid option beside the free default. Parked is fine; unshown is not.
|
|
23
|
+
|
|
24
|
+
## Anti-patterns
|
|
25
|
+
|
|
26
|
+
- Treating a green test suite as a gap analysis. Tests verify presence; they cannot see absence.
|
|
27
|
+
- Asking the model that wrote the thing whether it is complete, in the same context. It will say yes.
|
|
28
|
+
- Reporting gaps you did not verify. A gap is a finding; it needs the same evidence a bug does.
|
|
@@ -0,0 +1,30 @@
|
|
|
1
|
+
# Memory and record: the store is part of the change
|
|
2
|
+
|
|
3
|
+
Every protocol here ends in a write: the end-to-end doc, the rename that lands everywhere, the research brief, the gap report. A write that nothing indexes is a note in a drawer. This protocol says how the store is kept honest, whatever the store is.
|
|
4
|
+
|
|
5
|
+
Companion tool for this rule: **obsidian-tc**, {{OBSIDIAN_TC_STATUS}}.
|
|
6
|
+
|
|
7
|
+
## Rules
|
|
8
|
+
|
|
9
|
+
1. **Search before you write.** A research brief, a decision, a rule: check whether it already exists (`semantic_search` for the concept, `search_text` for the exact phrase). Duplicates are how a store starts lying: two notes, two answers, and a reader picks one.
|
|
10
|
+
2. **The folder index is part of the change, not a follow-up.** Every folder has one index file that says what is in it and what state it is in. Any write, edit or delete reopens that index in the same pass and corrects whatever the change made untrue. A stale index is worse than a missing one because agents believe it.
|
|
11
|
+
3. **One writer per run.** Several agents may propose; one records. If you are not the writer, produce the file and name it in your report.
|
|
12
|
+
4. **Machine output stays out of the index.** Scan dumps, logs, traces embed well and outrank the thing they describe. Keep them outside the searchable store, or in a folder the index excludes.
|
|
13
|
+
5. **A record is not present state.** A note, a ticket, a checkbox is a dated observation. Re-read the live thing before you act on it.
|
|
14
|
+
6. **Inferred content is marked as inferred.** A conclusion an agent reached, rather than copied from a source, carries `source: agent-synthesis` (and, with obsidian-tc, goes through its poison scan before it lands). A reader must be able to tell a quote from a guess.
|
|
15
|
+
7. **Compare-and-swap on overwrite.** Read, then write with the hash you read. A blind overwrite of a note someone else changed is a lost update nobody notices.
|
|
16
|
+
|
|
17
|
+
## Where the other protocols touch the store
|
|
18
|
+
|
|
19
|
+
| Protocol | Store call |
|
|
20
|
+
|---|---|
|
|
21
|
+
| Propagate, step 1 | `get_backlinks` on the thing being renamed; `search_text` for the literal old term; after the change, `find_unresolved_links` |
|
|
22
|
+
| Build, Stage 7 | the end-to-end doc goes in the folder that owns the domain; its index is updated in the same pass |
|
|
23
|
+
| Deep research, step 4 | `semantic_search` before the brief is written; a hit means append to the existing note, not a second note |
|
|
24
|
+
| Gap analysis | the "what exists" enumeration starts from the store, then diffs against the live state |
|
|
25
|
+
|
|
26
|
+
## Without obsidian-tc
|
|
27
|
+
|
|
28
|
+
The rules still bind. `grep -rn` is your literal search, a folder README is your index, `git` is your compare-and-swap. What is not allowed is a write nobody can find again.
|
|
29
|
+
|
|
30
|
+
Source: https://github.com/The-40-Thieves/obsidian-tc
|
|
@@ -0,0 +1,35 @@
|
|
|
1
|
+
# Numbers and logic: compute, never guess
|
|
2
|
+
|
|
3
|
+
**Never do arithmetic in your head when being wrong would matter.** Money, rates, percentages, margins, budgets, token and cost counts, any comparison you are about to state, any complexity or equivalence or speedup claim, anything past 2^53. A wrong number that looks right is worse than no number, because it gets acted on.
|
|
4
|
+
|
|
5
|
+
A model that is confident about `0.1 + 0.2` does not feel uncertain. It feels finished. The tool exists so the feeling is not the check.
|
|
6
|
+
|
|
7
|
+
Companion tool for this rule: **codecalc**, {{CODECALC_STATUS}}.
|
|
8
|
+
|
|
9
|
+
## When calling is mandatory
|
|
10
|
+
|
|
11
|
+
| You are about to | Use |
|
|
12
|
+
|---|---|
|
|
13
|
+
| state a number someone will act on | `evaluate_expression` (exact rationals, no float drift) |
|
|
14
|
+
| say A is bigger, cheaper, faster than B | compute both, then compare; never eyeball |
|
|
15
|
+
| claim two programs behave the same (a port, a rewrite) | `verify_translation` |
|
|
16
|
+
| claim an optimization preserved behaviour | `verify_optimization` |
|
|
17
|
+
| state a Big-O, or that something scales | `analyze_complexity` (static) or `benchmark` (measured) |
|
|
18
|
+
| assert a logical property holds, or that a set of constraints is satisfiable | `z3_check` (SMT) or `truth_table` |
|
|
19
|
+
| run a snippet to see what it actually does | `execute_code` (31 languages, sandboxed) |
|
|
20
|
+
|
|
21
|
+
Trivial single-digit sums are exempt. Everything else is not.
|
|
22
|
+
|
|
23
|
+
## How to report a computed figure
|
|
24
|
+
|
|
25
|
+
Give the exact form and the decimal, and say which tool produced it. `37/210 = 17.62%` reads differently from `about 18%`; the first can be checked, the second cannot. If a tool returned `unenforced` or a grade below its top, say so beside the number.
|
|
26
|
+
|
|
27
|
+
## Logic flow
|
|
28
|
+
|
|
29
|
+
Reasoning scaffolds help models that have no native reasoning mode and add nothing to models that already think before answering. So: do not narrate a chain of thought as evidence. **The authorising evidence is the computed outcome, not the thought log.** When a decision hangs on a logical claim, encode the claim and check it (`z3_check`); when it hangs on behaviour, run it (`execute_code`); when it hangs on a figure, compute it. A trace that was never checked is a story.
|
|
30
|
+
|
|
31
|
+
## Without codecalc
|
|
32
|
+
|
|
33
|
+
The rule still binds. Use whatever computes: a shell `python3 -c`, a spreadsheet, the vendor's built-in interpreter. What is not allowed is the number that came from nowhere.
|
|
34
|
+
|
|
35
|
+
Source: https://github.com/The-40-Thieves/codecalc
|
|
@@ -0,0 +1,34 @@
|
|
|
1
|
+
# Propagate: change completeness
|
|
2
|
+
|
|
3
|
+
**A rename is a refactor, not a single-file edit.** Any change to a name, term, path, slug, schema field, routing rule or shared convention has a blast radius, and the goal is zero silent strays.
|
|
4
|
+
|
|
5
|
+
This is retrieval work. It stays with the orchestrator (or a cheap worker for the grep sweep). It never goes to the deep tier: a judgment model re-deriving a file list is the most expensive routing mistake there is.
|
|
6
|
+
|
|
7
|
+
## 1. Map the blast radius (before editing anything)
|
|
8
|
+
|
|
9
|
+
- **Docs and notes:** backlinks to the thing being renamed; literal search for the old term and its link forms. With obsidian-tc: `get_backlinks`, `search_text`, then `find_unresolved_links` after the change (`protocols/memory-and-record.md`).
|
|
10
|
+
- **Memory / instructions:** grep every instructions file your agents read (`CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, `QWEN.md`, custom instructions) and any memory store.
|
|
11
|
+
- **Code / config:** grep the repos, settings files, hooks, scheduled jobs, CI, and the files this installer wrote.
|
|
12
|
+
- **Other people's surfaces:** anything that consumes the old name from outside (webhooks, dashboards, bookmarks).
|
|
13
|
+
|
|
14
|
+
```bash
|
|
15
|
+
grep -rniE "<old-term>" --exclude-dir=.git --exclude-dir=node_modules --exclude-dir=dist .
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
## 2. Build the checklist
|
|
19
|
+
|
|
20
|
+
Every hit becomes a line, grouped by surface. Nothing closes until the list is empty. Mark which items are approval-gated (production deploys, database migrations, anything public).
|
|
21
|
+
|
|
22
|
+
## 3. Execute on all fronts
|
|
23
|
+
|
|
24
|
+
Change every item. Use the tool's own governed rename where one exists (a link rewriter, an IDE refactor) over hand edits. If an index or README lists the renamed thing, that index is part of the change, not a follow-up.
|
|
25
|
+
|
|
26
|
+
## 4. Verify: the loud negative (non-negotiable)
|
|
27
|
+
|
|
28
|
+
Re-grep the OLD identifier across every surface. **Expect zero**, except historical records you name explicitly. Then check for dangling references the rename created (unresolved links, 404s, failing imports).
|
|
29
|
+
|
|
30
|
+
Paste the final grep output as proof. If any stray survives, it is not done.
|
|
31
|
+
|
|
32
|
+
## Why the last step is the one that matters
|
|
33
|
+
|
|
34
|
+
Steps 1 to 3 find what you thought of. Step 4 finds what you did not. A rename that "looks complete" and leaves one stray is worse than an unstarted one, because the stray is now believed.
|