model-orchestrator 0.1.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +96 -0
- package/LICENSE +21 -0
- package/README.md +133 -0
- package/SECURITY.md +17 -0
- package/bin/README.md +10 -0
- package/bin/cli-run.mjs +599 -0
- package/bin/cli.js +372 -0
- package/docs/README.md +13 -0
- package/docs/audit-brief.md +83 -0
- package/docs/catalog.md +113 -0
- package/docs/part-1-beginner.md +65 -0
- package/docs/part-2-intermediate.md +65 -0
- package/docs/part-3-advanced.md +65 -0
- package/package.json +52 -0
- package/scripts/README.md +5 -0
- package/scripts/gen-catalog.js +37 -0
- package/src/README.md +9 -0
- package/src/catalog.js +277 -0
- package/src/detect.js +26 -0
- package/src/install.js +628 -0
- package/src/prompt.js +34 -0
- package/src/render.js +8 -0
- package/templates/README.md +14 -0
- package/templates/advanced/README.md +14 -0
- package/templates/advanced/vm/ENVIRONMENT.md +18 -0
- package/templates/advanced/vm/PRIVACY_GATES.md +33 -0
- package/templates/advanced/vm/README.md +60 -0
- package/templates/advanced/vm/box-CLAUDE.md +28 -0
- package/templates/advanced/vm/docker-compose.yml +19 -0
- package/templates/advanced/vm/gateway.config.yaml +12 -0
- package/templates/advanced/vm/jobs/README.md +39 -0
- package/templates/advanced/vm/jobs/weekly-audit.service +17 -0
- package/templates/advanced/vm/jobs/weekly-audit.sh +107 -0
- package/templates/advanced/vm/jobs/weekly-audit.timer +10 -0
- package/templates/advanced/vm/setup-vm.sh +46 -0
- package/templates/agents/README.md +13 -0
- package/templates/agents/agy/README.md +5 -0
- package/templates/agents/agy/builder.md +17 -0
- package/templates/agents/agy/bulk-worker.md +17 -0
- package/templates/agents/agy/code-reviewer.md +17 -0
- package/templates/agents/agy/deep-planner.md +17 -0
- package/templates/agents/agy/live-researcher.md +17 -0
- package/templates/agents/claude-code/README.md +13 -0
- package/templates/agents/claude-code/builder.md +17 -0
- package/templates/agents/claude-code/bulk-worker.md +18 -0
- package/templates/agents/claude-code/code-reviewer.md +19 -0
- package/templates/agents/claude-code/deep-planner.md +18 -0
- package/templates/agents/claude-code/live-researcher.md +18 -0
- package/templates/agents/snippets/chat.md +25 -0
- package/templates/agents/snippets/claude-code.md +27 -0
- package/templates/agents/snippets/generic.md +21 -0
- package/templates/beginner/ORCHESTRATOR.md +55 -0
- package/templates/beginner/README.md +3 -0
- package/templates/common/README.md +52 -0
- package/templates/common/TASK_BUNDLE.md +56 -0
- package/templates/common/protocols/README.md +14 -0
- package/templates/common/protocols/build-protocol.md +133 -0
- package/templates/common/protocols/deep-research.md +44 -0
- package/templates/common/protocols/gap-analysis.md +28 -0
- package/templates/common/protocols/memory-and-record.md +30 -0
- package/templates/common/protocols/numbers-and-logic.md +35 -0
- package/templates/common/protocols/propagate.md +34 -0
- package/templates/intermediate/CLI-RUN.md +100 -0
- package/templates/intermediate/DELEGATION_MATRIX.md +41 -0
- package/templates/intermediate/README.md +13 -0
- package/templates/intermediate/RESEARCH_TRIAGE.md +30 -0
- package/templates/intermediate/ROUTING.md +73 -0
- package/templates/intermediate/TIERS.md +44 -0
- package/templates/tools/README.md +10 -0
- package/templates/tools/codecalc/CODECALC.md +43 -0
- package/templates/tools/codecalc/mcp/agy.mcp_config.json +8 -0
- package/templates/tools/codecalc/mcp/codex.config.toml +4 -0
- package/templates/tools/codecalc/mcp/mcpServers.json +8 -0
- package/templates/tools/codecalc/mcp/vscode.mcp.json +8 -0
- package/templates/tools/codecalc/mcp/zed.settings.json +9 -0
- package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +65 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.agy.mcp_config.json +9 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.codex.config.toml +7 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.mcpServers.json +9 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.vscode.mcp.json +10 -0
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.zed.settings.json +9 -0
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
# templates/
|
|
2
|
+
|
|
3
|
+
Everything the installer can write, organized by the level that adds it. Files are rendered with `{{PLACEHOLDERS}}` filled from `src/catalog.js` and the user's answers (`src/install.js` computes every value; templates contain no logic).
|
|
4
|
+
|
|
5
|
+
| Folder | Written at | Contents |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| `common/` | every level | the start-here README, `TASK_BUNDLE.md`, `protocols/` (build, propagate, gap analysis, deep research, numbers and logic, memory and record) |
|
|
8
|
+
| `beginner/` | every level | `ORCHESTRATOR.md`, the single-agent routing rules |
|
|
9
|
+
| `agents/` | every level, one variant | the primary agent's loading surface: Claude Code subagents, Antigravity custom agents, or a paste snippet |
|
|
10
|
+
| `intermediate/` | level 2+ | `ROUTING.md`, `TIERS.md`, `DELEGATION_MATRIX.md`, `RESEARCH_TRIAGE.md`, `CLI-RUN.md` |
|
|
11
|
+
| `advanced/` | level 3 | `vm/`: gateway config, compose file, box rules, privacy gates, scheduled jobs |
|
|
12
|
+
| `tools/` | when selected | companion tools the AIs call: `codecalc/` and `obsidian-tc/` (install doc + MCP snippets each). See `tools/README.md` |
|
|
13
|
+
|
|
14
|
+
Agent definitions under `agents/claude-code/` and `agents/agy/` are written to the PROJECT root (`--project`), not `--dir`, because that is where those CLIs read them. A `README.md` at the root of a tier folder (like this one) documents the repo and is not installed. `common/README.md` is the exception: it is the user's start-here file. READMEs deeper in (`protocols/`, `vm/`, `vm/jobs/`) are installed as folder indexes.
|
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
# templates/advanced/
|
|
2
|
+
|
|
3
|
+
Written at level 3 only, on top of everything below it. Everything lands under `vm/` in the user's install. This tier is documentation and config templates for a Linux box; the installer never touches a remote machine.
|
|
4
|
+
|
|
5
|
+
| File | What it is |
|
|
6
|
+
|---|---|
|
|
7
|
+
| `vm/README.md` | start-here for the box: what runs where, the gateway-holds-the-keys rule, the closed loop |
|
|
8
|
+
| `vm/setup-vm.sh` | idempotent setup for a fresh Ubuntu box: deps, the selected CLIs; prints every vendor script instead of running it |
|
|
9
|
+
| `vm/docker-compose.yml` | the gateway (and a local model runtime if selected), bound to loopback |
|
|
10
|
+
| `vm/gateway.config.yaml` | one lane per selected provider, keys referenced by environment variable NAME only |
|
|
11
|
+
| `vm/ENVIRONMENT.md` | which variable names the gateway expects, and where to keep the values (a secrets manager, never a file in the repo) |
|
|
12
|
+
| `vm/box-CLAUDE.md` | the rules a Claude Code session on the box inherits: cheapest tier that does the job, what never goes to the cheap tier, no public bind |
|
|
13
|
+
| `vm/PRIVACY_GATES.md` | what data never leaves the machine, which lanes are barred by name |
|
|
14
|
+
| `vm/jobs/` | a systemd timer + service pair for the weekly gap-analysis audit, plus an index |
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
# ENVIRONMENT.md: names the gateway expects
|
|
2
|
+
|
|
3
|
+
Generated {{DATE}} from the metered API keys you said you hold (asked separately from your subscription CLIs, because a Claude Code plan is not an Anthropic API key):
|
|
4
|
+
|
|
5
|
+
{{APIS_LIST}}
|
|
6
|
+
|
|
7
|
+
These are variable **names**. The values live in a secrets manager and are injected at start time (for example `<manager> run -- docker compose up -d`, or a systemd `EnvironmentFile=` that lives outside this folder with mode 600).
|
|
8
|
+
|
|
9
|
+
{{ENV_NAMES}}
|
|
10
|
+
|
|
11
|
+
`GATEWAY_MASTER_KEY` is the bearer every client presents to the gateway. Generate it once (`openssl rand -hex 32`), store it in the manager, never paste it into a file here. It must be a single token matching `^[A-Za-z0-9._-]+$`: the audit job interpolates it into curl's config grammar and refuses anything else.
|
|
12
|
+
|
|
13
|
+
## Rules
|
|
14
|
+
|
|
15
|
+
- Never print a value in a terminal or a log. Verify by length or by a hash prefix.
|
|
16
|
+
- Never pass a value on a command line; argv is world-readable on Linux. Use `curl --config -` fed from `printf`, or a tool's own env-var option.
|
|
17
|
+
- Never commit a file that contains a value. Add a secret scanner as a pre-push hook.
|
|
18
|
+
- Subscription CLIs (`claude`, `codex`, `agy`, `grok`, `hermes`) keep their own sign-in state; they need none of these names.
|
|
@@ -0,0 +1,33 @@
|
|
|
1
|
+
# PRIVACY_GATES.md: what never leaves the machine
|
|
2
|
+
|
|
3
|
+
A bar that is not named is not enforced. Fill in the names.
|
|
4
|
+
|
|
5
|
+
## Data classes
|
|
6
|
+
|
|
7
|
+
| Class | Examples | May go to |
|
|
8
|
+
|---|---|---|
|
|
9
|
+
| Public | published posts, open docs, public repos | any lane |
|
|
10
|
+
| Working | your own notes, drafts, code you will publish | your primary vendor's lanes, subscription CLIs you trust with it |
|
|
11
|
+
| Confidential | client data, other people's records, contracts | your primary vendor only, or the local lane |
|
|
12
|
+
| Personal | health, identity, private journals | the local lane only, or nowhere |
|
|
13
|
+
|
|
14
|
+
## Lanes barred by name for Confidential and Personal
|
|
15
|
+
|
|
16
|
+
- metered third-party bulk lanes (the cheapest-tier API you use for volume)
|
|
17
|
+
- concurrent fan-out lanes on a consumer subscription
|
|
18
|
+
- any shared compute lane (free GPU tiers, notebook services)
|
|
19
|
+
- any tool that stores conversation history on its own servers without a retention control you have read
|
|
20
|
+
|
|
21
|
+
Write your own lane names here, in this file, so the bar is checkable:
|
|
22
|
+
|
|
23
|
+
```
|
|
24
|
+
BARRED: <lane>, <lane>, <lane>
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
## The local lane is a privacy lane, not a cost lane
|
|
28
|
+
|
|
29
|
+
Route to a local model when confinement is the requirement. Never to save money: the accuracy gap is real, and pennies saved are not worth a wrong answer that looks right.
|
|
30
|
+
|
|
31
|
+
## The check
|
|
32
|
+
|
|
33
|
+
Before any bulk call: which class is this data, and is the lane in the barred list? If you cannot answer both, it is Confidential.
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# vm/: the box that runs it unattended
|
|
2
|
+
|
|
3
|
+
Level 3 = levels 1 and 2 plus a machine that is always on. A small Linux VM (any cloud's free ARM tier is enough) that holds the CLIs, a model gateway, and the scheduled jobs. Your laptop stays the interactive driver; the box owns the schedule.
|
|
4
|
+
|
|
5
|
+
Generated {{DATE}} for: `{{AI_IDS}}`. Installed at `{{INSTALL_DIR}}`; the systemd unit and the audit script carry that path.
|
|
6
|
+
|
|
7
|
+
## The one architectural property
|
|
8
|
+
|
|
9
|
+
**Only the gateway holds provider credentials.** Nothing else on the box does: not the orchestrator, not the scheduler, not a job. Every surface reaches models through the gateway, so rotating a key is a change in exactly one place. The gateway is bound to loopback (or a private mesh network), never to `0.0.0.0`.
|
|
10
|
+
|
|
11
|
+
## What runs where
|
|
12
|
+
|
|
13
|
+
| Surface | Role | Reaches models via |
|
|
14
|
+
|---|---|---|
|
|
15
|
+
| The orchestrator CLI ({{PRIMARY_NAME}}) | interactive driver when you SSH in; dispatch brain for jobs | its own subscription, off the gateway |
|
|
16
|
+
| `cli-run` lanes ({{CLI_RUN_LANES}}) | the other agent CLIs, headless | their own subscriptions (Lane A) |
|
|
17
|
+
| The gateway (`docker-compose.yml`) | one OpenAI-compatible endpoint fronting every metered provider | provider keys from the environment |
|
|
18
|
+
| A local model runtime (if selected) | the privacy lane | nothing leaves the box |
|
|
19
|
+
| Scheduled jobs (`jobs/`) | the weekly gap-analysis audit (lane: `{{AUDIT_LANE}}`), and anything else recurring | the gateway, or `cli-run` |
|
|
20
|
+
| codecalc (if selected) | the calculator, code runner and logic checker every agent here calls; stdio, offline, no key | nothing; it computes locally |
|
|
21
|
+
|
|
22
|
+
## Setup, in order
|
|
23
|
+
|
|
24
|
+
1. Provision a box. Ubuntu, 2+ vCPU, 8 GB is comfortable. Put it on a private mesh network if you can; do not open ports to the internet.
|
|
25
|
+
2. `bash setup-vm.sh`. It installs system deps and the npm-installable CLIs, then **prints** the vendor shell installers for the rest. Read those scripts before running them.
|
|
26
|
+
3. Sign each CLI in with its device-code flow (`codex login --device-auth`, `grok login --device-auth`, `agy` on first run). Run these inside `tmux` so a dropped SSH session does not kill the prompt. Headless Linux has no keyring by default; `setup-vm.sh` installs one so the CLIs stop re-prompting.
|
|
27
|
+
4. Put provider keys in your secrets manager and export the names listed in `ENVIRONMENT.md` into the gateway's environment at start time. Never write a value into a file in this folder. The gateway config was rendered from the API keys you said you hold, not from your CLI subscriptions: those are different entitlements.
|
|
28
|
+
5. `docker compose up -d`, then list the lanes without putting the key in argv (the key must be a single token, `^[A-Za-z0-9._-]+$`, because it is interpolated into curl's config grammar):
|
|
29
|
+
```bash
|
|
30
|
+
printf 'header = "Authorization: Bearer %s"\n' "$GATEWAY_MASTER_KEY" | curl -s --config - http://127.0.0.1:4000/v1/models
|
|
31
|
+
```
|
|
32
|
+
6. Install the weekly audit: `jobs/README.md`.
|
|
33
|
+
7. Copy `box-CLAUDE.md` to `~/CLAUDE.md` on the box (or your agent's equivalent rules file) so a session there inherits the house rules without you present.
|
|
34
|
+
|
|
35
|
+
## The dispatch shape
|
|
36
|
+
|
|
37
|
+
1. **Deterministic pre-triage, zero tokens:** a keyword table sends the obvious cases (bulk patterns → the cheap lane, URLs and current events → the live lane, "review this" → the reviewer, "write this up" → the orchestrator).
|
|
38
|
+
2. **Judgment dispatch:** everything else is routed by the orchestrator against `ROUTING.md` and `DELEGATION_MATRIX.md`, with a logged reason.
|
|
39
|
+
3. **Free lane first, escalate on signal.** Every job starts on its $0 lane and climbs only on failure, low confidence, or an explicit "expensive to get wrong".
|
|
40
|
+
4. **Unattended means no escalation to a human-gated tier.** An unresolved irreversible call is surfaced (a message, a ticket comment), never executed.
|
|
41
|
+
5. **Writes stay locked to one writer.** Every other engine proposes; one writer records.
|
|
42
|
+
|
|
43
|
+
## The closed loop (name what watches it)
|
|
44
|
+
|
|
45
|
+
| Thing | Ran? | Watched by | On failure |
|
|
46
|
+
|---|---|---|---|
|
|
47
|
+
| gateway | `docker compose ps`, `/v1/models` | the weekly audit job | audit report names the dead lane |
|
|
48
|
+
| weekly audit | `systemctl --user list-timers` | **nothing** unless you wire a notifier | write "nothing" here until you do; that line is the useful one |
|
|
49
|
+
| each `cli-run` job | exit code + `~/.ai-orchestrator/cli-run.log.jsonl` | the job's own caller | rc 10/12/13 in the log |
|
|
50
|
+
|
|
51
|
+
"Nothing watches it" is a valid answer and usually the valuable one. Writing it down turns an invisible gap into a tracked one.
|
|
52
|
+
|
|
53
|
+
## Never on this box
|
|
54
|
+
|
|
55
|
+
- Vendor scripts run blind. Read first.
|
|
56
|
+
- An unpinned image or package. `docker-compose.yml` and `setup-vm.sh` pin versions; bump them on purpose, never by restarting.
|
|
57
|
+
- A provider key in a file in this folder, in shell history, in argv, or in a container image.
|
|
58
|
+
- A service bound to `0.0.0.0`.
|
|
59
|
+
- Private notes, client data, or personal records sent to a metered bulk lane. See `PRIVACY_GATES.md`.
|
|
60
|
+
- A payment card attached to a compute lane "to unlock a tier". Free credit only unless a human says otherwise.
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# CLAUDE.md for the box
|
|
2
|
+
|
|
3
|
+
Copy to `~/CLAUDE.md` on the machine (or your agent's equivalent rules file). A session here inherits these without a human present.
|
|
4
|
+
|
|
5
|
+
## Cost rule
|
|
6
|
+
Prefer the cheapest tier that does the job well. Delegate grunt work through the dispatch layer; keep the reasoning in-session.
|
|
7
|
+
|
|
8
|
+
Send to the cheap tier via the gateway or `cli-run`: bulk classification and tagging, reformatting, extraction, first-pass summaries, mechanical transforms, explicit rough drafts.
|
|
9
|
+
|
|
10
|
+
Never send to the cheap tier: anything that will be published in a person's own voice, code that gets committed, anything needing current vendor-specific knowledge, anything where being wrong is expensive, anything time-sensitive.
|
|
11
|
+
|
|
12
|
+
## On dispatch failure
|
|
13
|
+
Report it and fall back to doing the work in-session. Never silently retry the same lane.
|
|
14
|
+
|
|
15
|
+
## Zero ingress
|
|
16
|
+
Nothing binds to `0.0.0.0`. Nothing publishes a container port to the public interface. New services go on loopback or the private mesh.
|
|
17
|
+
|
|
18
|
+
## Unattended means no human-gated escalation
|
|
19
|
+
An unresolved call that is irreversible or rewrites a standing rule gets surfaced (a message, a ticket comment) and stops. It is never executed on the strength of a model's confidence.
|
|
20
|
+
|
|
21
|
+
## One writer
|
|
22
|
+
Scheduled jobs and other engines propose. One writer records. If you are not that writer, produce a file and name it in your report.
|
|
23
|
+
|
|
24
|
+
## Secrets
|
|
25
|
+
Names in the environment, values in the secrets manager. Never print one, never pass one in argv, never write one to disk here.
|
|
26
|
+
|
|
27
|
+
## Routing
|
|
28
|
+
`{{RULES_PATH}}/ROUTING.md` and `{{RULES_PATH}}/DELEGATION_MATRIX.md` are the rules. `{{RULES_PATH}}/protocols/` are the procedures.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# The gateway is the ONLY holder of provider credentials on this box.
|
|
2
|
+
# It is bound to loopback. Put it behind a private mesh network if other
|
|
3
|
+
# machines need it; never publish it on 0.0.0.0.
|
|
4
|
+
services:
|
|
5
|
+
gateway:
|
|
6
|
+
image: {{LITELLM_IMAGE}} # pinned on purpose; bump deliberately after reading the release notes
|
|
7
|
+
container_name: gateway
|
|
8
|
+
restart: unless-stopped
|
|
9
|
+
ports:
|
|
10
|
+
- "127.0.0.1:4000:4000"
|
|
11
|
+
volumes:
|
|
12
|
+
- ./gateway.config.yaml:/app/config.yaml:ro
|
|
13
|
+
command: ["--config", "/app/config.yaml", "--port", "4000"]
|
|
14
|
+
environment:
|
|
15
|
+
# Names only. Values come from the environment you start compose in
|
|
16
|
+
# (a secrets manager's `run` wrapper, or a systemd EnvironmentFile outside this repo).
|
|
17
|
+
- LITELLM_MASTER_KEY=${GATEWAY_MASTER_KEY}
|
|
18
|
+
{{COMPOSE_ENV}}
|
|
19
|
+
{{COMPOSE_OLLAMA}}
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
# LiteLLM gateway config. One lane per provider you said you have.
|
|
2
|
+
# Every key is `os.environ/<NAME>`: the value lives in the environment, never here.
|
|
3
|
+
# Bind to ALIASES (model_name), not raw provider ids, so a vendor rename is a one-line repoint.
|
|
4
|
+
model_list:
|
|
5
|
+
{{GATEWAY_MODELS}}
|
|
6
|
+
|
|
7
|
+
litellm_settings:
|
|
8
|
+
drop_params: true
|
|
9
|
+
set_verbose: false
|
|
10
|
+
|
|
11
|
+
general_settings:
|
|
12
|
+
master_key: os.environ/LITELLM_MASTER_KEY
|
|
@@ -0,0 +1,39 @@
|
|
|
1
|
+
# jobs/
|
|
2
|
+
|
|
3
|
+
Scheduled work on the box, as user-level systemd timers. Each job is a timer + service pair and gets one line in the table below. **Each line names what watches it.** "nothing" is an honest answer and the one that tells you where to look.
|
|
4
|
+
|
|
5
|
+
| Job | Schedule | Does | Lane | Watched by |
|
|
6
|
+
|---|---|---|---|---|
|
|
7
|
+
| `weekly-audit` | Monday 09:00 | collects live state (gateway lanes, timers, CLI versions), composes a brief with the protocol and `DELEGATION_MATRIX.md`, and asks a cli-run lane for the gap report | `{{AUDIT_LANE}}` (first enabled lane at install time; edit `AUDIT_LANE` in the script to change it) | nothing yet: wire a notifier and update this line |
|
|
8
|
+
|
|
9
|
+
Paths in the service and the script were rendered for this install: `{{INSTALL_DIR}}`. If you move the folder, re-run the installer or edit both files.
|
|
10
|
+
|
|
11
|
+
## Install a job
|
|
12
|
+
|
|
13
|
+
```bash
|
|
14
|
+
mkdir -p ~/.config/systemd/user
|
|
15
|
+
cp weekly-audit.service weekly-audit.timer ~/.config/systemd/user/
|
|
16
|
+
systemctl --user daemon-reload
|
|
17
|
+
systemctl --user enable --now weekly-audit.timer
|
|
18
|
+
systemctl --user list-timers # it should be listed with a next-run time
|
|
19
|
+
loginctl enable-linger "$USER" # so user timers run without a login session
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
The service reads `GATEWAY_MASTER_KEY` from an `EnvironmentFile` that lives outside this folder (mode 600). Point the `EnvironmentFile=` line at yours before installing. The key must be a single token matching `^[A-Za-z0-9._-]+$`; the script refuses anything else.
|
|
23
|
+
|
|
24
|
+
## Verify a job ran
|
|
25
|
+
|
|
26
|
+
```bash
|
|
27
|
+
systemctl --user status weekly-audit.service
|
|
28
|
+
journalctl --user -u weekly-audit.service -n 50
|
|
29
|
+
ls -la {{INSTALL_DIR}}/reports/
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
## What the job guarantees
|
|
33
|
+
|
|
34
|
+
- **Bounded:** every probe (gateway, `systemctl`, each CLI `--version`) runs under a 10 s watchdog; the model call under 600 s; the unit under `TimeoutStartSec=900`, which kills the whole cgroup.
|
|
35
|
+
- **Previous report preserved:** output goes to a temp file and is renamed over `audit-<date>.md` only on a clean, non-empty run. A failed run leaves `failed-audit-<stamp>-rc<N>.md` beside it and the last good report untouched.
|
|
36
|
+
- **Boundary:** the lane runs with the strongest restriction it offers ({{AUDIT_LANE_BOUNDARY_NOTE}}). The brief's denied-actions list is an instruction, not an enforcement, for lanes without a sandbox flag.
|
|
37
|
+
- **Honest unknowns:** a probe that times out writes an `UNVERIFIED` line, which the brief tells the lane to treat as unknown, never clean.
|
|
38
|
+
|
|
39
|
+
A timer that has never been seen to fire is not known to work. Run `systemctl --user start weekly-audit.service` once by hand and read the journal before trusting the schedule.
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
[Unit]
|
|
2
|
+
Description=ai-orchestrator weekly gap-analysis audit
|
|
3
|
+
After=network-online.target
|
|
4
|
+
|
|
5
|
+
[Service]
|
|
6
|
+
Type=oneshot
|
|
7
|
+
WorkingDirectory={{INSTALL_DIR_SYSTEMD}}
|
|
8
|
+
# Names come from a file OUTSIDE this repo, mode 600. Edit the path.
|
|
9
|
+
EnvironmentFile=%h/.config/ai-orchestrator/gateway.env
|
|
10
|
+
ExecStart=/bin/bash "{{INSTALL_DIR_SYSTEMD}}/vm/jobs/weekly-audit.sh"
|
|
11
|
+
# Whole-job deadline: collection probes (6 x 10 s) + the 600 s model call + cleanup.
|
|
12
|
+
# On expiry systemd kills the whole cgroup, so nothing the job spawned survives.
|
|
13
|
+
TimeoutStartSec=900
|
|
14
|
+
KillMode=control-group
|
|
15
|
+
|
|
16
|
+
[Install]
|
|
17
|
+
WantedBy=default.target
|
|
@@ -0,0 +1,107 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# weekly-audit.sh: collect live state, then let a cli-run lane draft the gap report.
|
|
3
|
+
# Rendered at install time: the install directory and the audit lane came from your selection.
|
|
4
|
+
#
|
|
5
|
+
# Guarantees this script makes, and how:
|
|
6
|
+
# - every collection step is BOUNDED (a watchdog around curl and each --version probe)
|
|
7
|
+
# - the previous successful report is NEVER truncated: output goes to a temp file and is
|
|
8
|
+
# renamed into place only on a clean exit; failed output is kept beside it for diagnosis
|
|
9
|
+
# - the lane runs with the strongest boundary it offers ({{AUDIT_LANE_BOUNDARY_NOTE}})
|
|
10
|
+
# - rc 10/12/13 from cli-run means no report was produced; the timer's journal shows it
|
|
11
|
+
set -uo pipefail
|
|
12
|
+
INSTALL_DIR={{INSTALL_DIR_SH}}
|
|
13
|
+
AUDIT_LANE="{{AUDIT_LANE}}"
|
|
14
|
+
AUDIT_LANE_FLAGS="{{AUDIT_LANE_FLAGS}}"
|
|
15
|
+
PROBE_SECS="${PROBE_SECS:-10}" # per collection probe
|
|
16
|
+
RUNNER_SECS="${RUNNER_SECS:-600}" # the model call; TimeoutStartSec in the unit covers the whole job
|
|
17
|
+
{{AUDIT_LANE_GUARD}}
|
|
18
|
+
cd "$INSTALL_DIR" || { echo "weekly-audit: $INSTALL_DIR missing" >&2; exit 2; }
|
|
19
|
+
mkdir -p reports
|
|
20
|
+
|
|
21
|
+
# bounded SECS cmd... : run cmd, and after SECS kill it AND every descendant
|
|
22
|
+
# (a probe that forks, or a stub that ignores its own flags, must not hold a
|
|
23
|
+
# pipe open). Process groups do not help here: bash disables job control inside
|
|
24
|
+
# pipeline subshells, so `kill -- -pid` would kill nothing. A recursive tree
|
|
25
|
+
# kill via pgrep works on macOS and Linux alike; `timeout(1)` is not on macOS.
|
|
26
|
+
killtree() {
|
|
27
|
+
local p="$1" c
|
|
28
|
+
for c in $(pgrep -P "$p" 2>/dev/null); do killtree "$c"; done
|
|
29
|
+
kill -KILL "$p" 2>/dev/null
|
|
30
|
+
}
|
|
31
|
+
bounded() {
|
|
32
|
+
local secs="$1"; shift
|
|
33
|
+
( "$@" ) & local pid=$!
|
|
34
|
+
( sleep "$secs"; killtree "$pid" ) >/dev/null 2>&1 & local wd=$!
|
|
35
|
+
wait "$pid" 2>/dev/null; local rc=$?
|
|
36
|
+
killtree "$wd" >/dev/null 2>&1; wait "$wd" 2>/dev/null
|
|
37
|
+
return $rc
|
|
38
|
+
}
|
|
39
|
+
|
|
40
|
+
# The gateway key must be a single token: it is interpolated into curl's config
|
|
41
|
+
# grammar, and a quote or newline in it would become a second directive.
|
|
42
|
+
KEY="${GATEWAY_MASTER_KEY:-}"
|
|
43
|
+
if [ -n "$KEY" ] && ! printf '%s' "$KEY" | grep -Eq '^[A-Za-z0-9._-]+$'; then
|
|
44
|
+
echo "weekly-audit: GATEWAY_MASTER_KEY must match ^[A-Za-z0-9._-]+$ (generate it with: openssl rand -hex 32)" >&2
|
|
45
|
+
exit 2
|
|
46
|
+
fi
|
|
47
|
+
|
|
48
|
+
DATE="$(date -u +%F)"
|
|
49
|
+
STAMP="$(date -u +%Y%m%dT%H%M%SZ)"
|
|
50
|
+
{
|
|
51
|
+
echo "# live state $DATE"; echo
|
|
52
|
+
echo "## gateway lanes"
|
|
53
|
+
if [ -n "$KEY" ]; then
|
|
54
|
+
# The key never enters argv: curl reads the header from a 0600 config file that
|
|
55
|
+
# exists only for this probe. curl's own timeouts AND the watchdog bound it.
|
|
56
|
+
CFG="$(umask 077 && mktemp "${TMPDIR:-/tmp}/audit-curl-XXXXXX")"
|
|
57
|
+
printf 'header = "Authorization: Bearer %s"\n' "$KEY" > "$CFG"
|
|
58
|
+
if ! bounded "$PROBE_SECS" curl -s --connect-timeout 5 --max-time "$PROBE_SECS" --max-filesize 1048576 --config "$CFG" http://127.0.0.1:4000/v1/models \
|
|
59
|
+
| jq -r '.data[].id' 2>/dev/null; then
|
|
60
|
+
echo "UNVERIFIED: gateway unreachable or timed out within ${PROBE_SECS}s"
|
|
61
|
+
fi
|
|
62
|
+
rm -f "$CFG"
|
|
63
|
+
else
|
|
64
|
+
echo "UNVERIFIED: GATEWAY_MASTER_KEY not set; gateway not queried"
|
|
65
|
+
fi
|
|
66
|
+
echo; echo "## timers"
|
|
67
|
+
bounded "$PROBE_SECS" systemctl --user list-timers --no-pager 2>/dev/null || echo "UNVERIFIED: systemd user session unavailable or timed out"
|
|
68
|
+
echo; echo "## cli versions"
|
|
69
|
+
for b in claude codex agy grok hermes qwen; do
|
|
70
|
+
if command -v "$b" >/dev/null 2>&1; then
|
|
71
|
+
printf '%s ' "$b"
|
|
72
|
+
bounded "$PROBE_SECS" "$b" --version 2>/dev/null | head -1 || echo "UNVERIFIED: --version timed out after ${PROBE_SECS}s"
|
|
73
|
+
fi
|
|
74
|
+
done
|
|
75
|
+
} > "reports/live-state-$STAMP.md"
|
|
76
|
+
ln -sfn "live-state-$STAMP.md" reports/live-state.md
|
|
77
|
+
|
|
78
|
+
# The brief the lane actually reads: the protocol, the intended configuration,
|
|
79
|
+
# and the live state it is meant to diff against.
|
|
80
|
+
BRIEF="reports/audit-brief-$STAMP.md"
|
|
81
|
+
{
|
|
82
|
+
echo "## Task bundle"
|
|
83
|
+
echo "**Purpose.** Weekly gap analysis: compare the live state below with the intended configuration and name what is missing, dead, or drifted."
|
|
84
|
+
echo "**Task class.** read_only"
|
|
85
|
+
echo "**Denied actions.** Do not run commands, do not modify files, do not call any network service. Enforcement: {{AUDIT_LANE_BOUNDARY_NOTE}}."
|
|
86
|
+
echo "**Report contract.** A ranked list of gaps (what is missing, where it should be, evidence line), then a CLEAN line per area with no gap, then what you could not assess. Treat every UNVERIFIED line below as unknown, never as clean."
|
|
87
|
+
echo "**Exit parameters.** Stop after one pass over the three sections below."
|
|
88
|
+
echo; echo "# Protocol"; cat protocols/gap-analysis.md
|
|
89
|
+
echo; echo "# Intended configuration"; cat DELEGATION_MATRIX.md 2>/dev/null || cat ORCHESTRATOR.md
|
|
90
|
+
echo; echo "# Live state"; cat "reports/live-state-$STAMP.md"
|
|
91
|
+
} > "$BRIEF"
|
|
92
|
+
|
|
93
|
+
# Write to a temp file; the dated report is replaced only by a clean, non-empty run.
|
|
94
|
+
FINAL="reports/audit-$DATE.md"
|
|
95
|
+
TMP="$(mktemp "reports/.audit-$STAMP-XXXXXX")"
|
|
96
|
+
# shellcheck disable=SC2086
|
|
97
|
+
node bin/cli-run.mjs "$AUDIT_LANE" $AUDIT_LANE_FLAGS --brief "$BRIEF" --timeout "$RUNNER_SECS" --quiet < /dev/null > "$TMP"
|
|
98
|
+
rc=$?
|
|
99
|
+
if [ "$rc" -eq 0 ] && [ -s "$TMP" ]; then
|
|
100
|
+
mv -f "$TMP" "$FINAL"
|
|
101
|
+
echo "audit rc=0 lane=$AUDIT_LANE report=$FINAL"
|
|
102
|
+
else
|
|
103
|
+
FAILED="reports/failed-audit-$STAMP-rc$rc.md"
|
|
104
|
+
mv -f "$TMP" "$FAILED" 2>/dev/null || rm -f "$TMP"
|
|
105
|
+
echo "audit rc=$rc lane=$AUDIT_LANE; previous report kept; partial output (if any) at $FAILED" >&2
|
|
106
|
+
fi
|
|
107
|
+
exit $rc
|
|
@@ -0,0 +1,46 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# setup-vm.sh: idempotent setup for a fresh Ubuntu box.
|
|
3
|
+
# Installs system deps and the npm-installable agent CLIs.
|
|
4
|
+
# PRINTS the vendor shell installers for the rest; never pipes a remote script into bash for you.
|
|
5
|
+
# Never writes a secret. Sign-ins are the vendors' own device-code flows.
|
|
6
|
+
set -euo pipefail
|
|
7
|
+
|
|
8
|
+
say() { printf '\n[setup-vm] %s\n' "$*"; }
|
|
9
|
+
have() { command -v "$1" >/dev/null 2>&1; }
|
|
10
|
+
|
|
11
|
+
say "system packages"
|
|
12
|
+
sudo apt-get update -y
|
|
13
|
+
sudo apt-get install -y curl git tmux jq ca-certificates gnupg build-essential \
|
|
14
|
+
dbus-x11 libsecret-1-0 gnome-keyring # a keyring, or headless CLIs re-prompt for auth on every launch
|
|
15
|
+
|
|
16
|
+
if ! have node; then
|
|
17
|
+
say "node 22: download the NodeSource setup script, read it, then run it yourself:"
|
|
18
|
+
say " curl -fsSL https://deb.nodesource.com/setup_22.x -o /tmp/nodesource.sh && less /tmp/nodesource.sh"
|
|
19
|
+
say " sudo -E bash /tmp/nodesource.sh && sudo apt-get install -y nodejs"
|
|
20
|
+
exit 1
|
|
21
|
+
fi
|
|
22
|
+
say "node present: $(node -v)"
|
|
23
|
+
|
|
24
|
+
if ! have docker; then
|
|
25
|
+
say "docker: follow https://docs.docker.com/engine/install/ubuntu/ (read the script before running it), then re-run this script"
|
|
26
|
+
else
|
|
27
|
+
say "docker present: $(docker --version)"
|
|
28
|
+
fi
|
|
29
|
+
|
|
30
|
+
say "npm-installable CLIs (installed only if missing)"
|
|
31
|
+
# Versions are pinned to what this installer was released with. Check for newer before trusting a pin forever.
|
|
32
|
+
for pkg in {{NPM_PACKAGES}}; do
|
|
33
|
+
[ -z "$pkg" ] && continue # nothing npm-installable was selected
|
|
34
|
+
case "$pkg" in
|
|
35
|
+
@anthropic-ai/claude-code*) bin=claude ;;
|
|
36
|
+
@openai/codex*) bin=codex ;;
|
|
37
|
+
@qwen-code/qwen-code*) bin=qwen ;;
|
|
38
|
+
*) bin="${pkg##*/}"; bin="${bin%%@*}" ;;
|
|
39
|
+
esac
|
|
40
|
+
if have "$bin"; then say "$bin present"; else say "npm install -g $pkg"; npm install -g "$pkg"; fi
|
|
41
|
+
done
|
|
42
|
+
|
|
43
|
+
say "vendor shell installers (read, then run yourself):"
|
|
44
|
+
{{SCRIPT_INSTALLERS}}
|
|
45
|
+
|
|
46
|
+
say "next: sign in to each CLI inside tmux (device-code flows), export the names in ENVIRONMENT.md, then: docker compose up -d"
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
# templates/agents/
|
|
2
|
+
|
|
3
|
+
Loading surfaces for the primary agent. The installer writes exactly one of these, based on `--primary`:
|
|
4
|
+
|
|
5
|
+
| Primary | Written | Why |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| `claude-code` | `.claude/agents/*.md` + `CLAUDE.snippet.md` | Claude Code loads project-level subagents from that folder |
|
|
8
|
+
| `agy` | `.agents/agents/*.md` + `GEMINI.snippet.md` | Antigravity custom agents live there |
|
|
9
|
+
| `codex`, `qwen` | `AGENTS.snippet.md` / `QWEN.snippet.md` | those CLIs read a rules file but have no subagent folder |
|
|
10
|
+
| `grok`, `hermes` | nothing agent-specific | rules travel with the prompt or the task bundle |
|
|
11
|
+
| a chat app | `PASTE-INTO-YOUR-AGENT.md` | no files to load; paste into custom instructions |
|
|
12
|
+
|
|
13
|
+
`snippets/` are rendered with the chosen agent's name and rules file. Nothing here is appended to a file the user already has.
|
|
@@ -0,0 +1,5 @@
|
|
|
1
|
+
# .agents/agents/
|
|
2
|
+
|
|
3
|
+
Antigravity CLI custom agents, one per tier, in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
|
|
4
|
+
|
|
5
|
+
`commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; `auto` keeps deletes and other high-risk commands gated) and `off` for the read-only agents. `model` is a tier: `pro` for deep-planner, `flash` for the rest.
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: builder
|
|
3
|
+
description: Well-specified execution of a bounded sub-part of a build.
|
|
4
|
+
model: flash
|
|
5
|
+
subagent: true
|
|
6
|
+
mainAgent: true
|
|
7
|
+
commandExecutionPolicy: auto # standard build/test commands run unattended; high-risk commands stay gated
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# builder
|
|
11
|
+
|
|
12
|
+
Well-specified execution of a bounded sub-part of a build.
|
|
13
|
+
|
|
14
|
+
Rules:
|
|
15
|
+
- Stay inside the task bundle you were given. Anything not granted is denied.
|
|
16
|
+
- Report what you did, what you did not do, and what you could not verify. "Unverified" is acceptable; a confident guess is not.
|
|
17
|
+
- Token discipline: read only what the task needs, never re-read, hand back deliverables not narration.
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: bulk-worker
|
|
3
|
+
description: High-volume mechanical work: classify, tag, extract, reformat, summarize many items.
|
|
4
|
+
model: flash
|
|
5
|
+
subagent: true
|
|
6
|
+
mainAgent: true
|
|
7
|
+
commandExecutionPolicy: off
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# bulk-worker
|
|
11
|
+
|
|
12
|
+
High-volume mechanical work: classify, tag, extract, reformat, summarize many items.
|
|
13
|
+
|
|
14
|
+
Rules:
|
|
15
|
+
- Stay inside the task bundle you were given. Anything not granted is denied.
|
|
16
|
+
- Report what you did, what you did not do, and what you could not verify. "Unverified" is acceptable; a confident guess is not.
|
|
17
|
+
- Token discipline: read only what the task needs, never re-read, hand back deliverables not narration.
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: code-reviewer
|
|
3
|
+
description: Read-only code review; findings ranked by severity with a concrete failure scenario each.
|
|
4
|
+
model: flash
|
|
5
|
+
subagent: true
|
|
6
|
+
mainAgent: true
|
|
7
|
+
commandExecutionPolicy: off
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# code-reviewer
|
|
11
|
+
|
|
12
|
+
Read-only code review; findings ranked by severity with a concrete failure scenario each.
|
|
13
|
+
|
|
14
|
+
Rules:
|
|
15
|
+
- Stay inside the task bundle you were given. Anything not granted is denied.
|
|
16
|
+
- Report what you did, what you did not do, and what you could not verify. "Unverified" is acceptable; a confident guess is not.
|
|
17
|
+
- Token discipline: read only what the task needs, never re-read, hand back deliverables not narration.
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: deep-planner
|
|
3
|
+
description: Ambiguous or high-stakes thinking: architecture, strategy, hard debugging. Returns a plan; never edits code.
|
|
4
|
+
model: pro
|
|
5
|
+
subagent: true
|
|
6
|
+
mainAgent: true
|
|
7
|
+
commandExecutionPolicy: off
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# deep-planner
|
|
11
|
+
|
|
12
|
+
Ambiguous or high-stakes thinking: architecture, strategy, hard debugging. Returns a plan; never edits code.
|
|
13
|
+
|
|
14
|
+
Rules:
|
|
15
|
+
- Stay inside the task bundle you were given. Anything not granted is denied.
|
|
16
|
+
- Report what you did, what you did not do, and what you could not verify. "Unverified" is acceptable; a confident guess is not.
|
|
17
|
+
- Token discipline: read only what the task needs, never re-read, hand back deliverables not narration.
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: live-researcher
|
|
3
|
+
description: Fresh information through search_web and read_url_content; cites sources and retrieval time.
|
|
4
|
+
model: flash
|
|
5
|
+
subagent: true
|
|
6
|
+
mainAgent: true
|
|
7
|
+
commandExecutionPolicy: off
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# live-researcher
|
|
11
|
+
|
|
12
|
+
Fresh information through search_web and read_url_content; cites sources and retrieval time.
|
|
13
|
+
|
|
14
|
+
Rules:
|
|
15
|
+
- Stay inside the task bundle you were given. Anything not granted is denied.
|
|
16
|
+
- Report what you did, what you did not do, and what you could not verify. "Unverified" is acceptable; a confident guess is not.
|
|
17
|
+
- Token discipline: read only what the task needs, never re-read, hand back deliverables not narration.
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
# .claude/agents/
|
|
2
|
+
|
|
3
|
+
Five subagents, one per tier. Claude Code loads project-level agents from this folder automatically.
|
|
4
|
+
|
|
5
|
+
| Agent | Tier | Model alias | Effort | Job |
|
|
6
|
+
|---|---|---|---|---|
|
|
7
|
+
| deep-planner | deep | opus | xhigh | judges every build twice; never retrieves |
|
|
8
|
+
| builder | standard | sonnet | high | bounded sub-parts of a build |
|
|
9
|
+
| code-reviewer | standard | sonnet | high | read-only findings |
|
|
10
|
+
| live-researcher | standard | sonnet | medium | fresh data through tools |
|
|
11
|
+
| bulk-worker | fast | haiku | low | mechanical volume |
|
|
12
|
+
|
|
13
|
+
Aliases resolve to the newest model in each family, so a version bump needs no edit here. Each agent carries its own token-discipline rule; the `effort` field is the third cost lever.
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: builder
|
|
3
|
+
description: Well-specified execution of a bounded sub-part. Use for writing code, editing files, wiring configs, and implementing a plan that already exists. Do not use for open-ended architecture questions, bulk classification, or the main build itself.
|
|
4
|
+
model: sonnet
|
|
5
|
+
effort: high
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
You are the execution tier of the model router.
|
|
9
|
+
|
|
10
|
+
You implement specs and plans: write code, edit files, run commands.
|
|
11
|
+
|
|
12
|
+
Rules:
|
|
13
|
+
- Follow the spec you were given. If the spec has a real gap, state the assumption you chose and proceed; do not redesign the architecture.
|
|
14
|
+
- Lightweight, concise code. No heavy dependencies.
|
|
15
|
+
- Verify your work runs (typecheck, test, or dry-run) before reporting done.
|
|
16
|
+
- Report plainly: what you changed, file paths, and proof it works.
|
|
17
|
+
- Token discipline: read only the files you will touch; never dump full file contents into replies, reference paths and the changed lines instead; do not re-read files you just wrote.
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: bulk-worker
|
|
3
|
+
description: Cheap high-volume work. Use for classifying, tagging, extracting, reformatting, or summarizing many items such as posts, rows, files, or notes. Fast and low cost. Do not use for tasks needing deep judgment or code changes.
|
|
4
|
+
tools: Read, Glob, Grep, Write
|
|
5
|
+
model: haiku
|
|
6
|
+
effort: low
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
You are the fast tier of the model router.
|
|
10
|
+
|
|
11
|
+
You do high-volume mechanical work: classify, tag, extract, reformat, summarize lists.
|
|
12
|
+
|
|
13
|
+
Rules:
|
|
14
|
+
- Be consistent. Define your categories or format once, then apply uniformly to every item.
|
|
15
|
+
- Output structured results: a markdown table or list, one row per item.
|
|
16
|
+
- Do not editorialize per item. One short summary line at the end is enough.
|
|
17
|
+
- If more than roughly 20 percent of items do not fit the given categories, stop and report that instead of forcing them.
|
|
18
|
+
- Token discipline: identify items by index or a short stub, never echo full item text back; output the table and the one summary line, nothing else.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: code-reviewer
|
|
3
|
+
description: Code review. Use when asked to review code, a diff, or a repo for bugs, security issues, or quality. Read-only, returns findings. Do not use for writing or fixing code.
|
|
4
|
+
tools: Read, Glob, Grep, Bash
|
|
5
|
+
model: sonnet
|
|
6
|
+
effort: high
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
You are the review tier of the model router.
|
|
10
|
+
|
|
11
|
+
You review code for real bugs, security problems, and correctness issues.
|
|
12
|
+
|
|
13
|
+
Rules:
|
|
14
|
+
- Report only findings you can defend with a concrete failure scenario. No style nitpicks unless asked.
|
|
15
|
+
- Rank by severity. For each: file, line, what breaks, and the fix in one or two sentences.
|
|
16
|
+
- Security findings (auth, secrets, injection, exposed endpoints) always rank first. Treat every endpoint as internet-facing.
|
|
17
|
+
- You are read-only. Suggest fixes; do not apply them.
|
|
18
|
+
- If the code is clean, say so plainly. Do not invent findings.
|
|
19
|
+
- Token discipline: read only the files under review, targeted sections where possible; report findings without restating the code; quote at most the few lines a finding needs.
|