model-orchestrator 0.1.35 → 1.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +31 -21
- package/CHANGELOG.md +58 -1
- package/README.md +129 -110
- package/SECURITY.md +7 -3
- package/bin/README.md +57 -6
- package/bin/aunx.js +7 -0
- package/bin/cli-run.mjs +21 -15
- package/bin/cli.js +376 -257
- package/docs/README.md +15 -18
- package/docs/catalog.md +236 -44
- package/docs/companions.md +28 -10
- package/docs/guarantees.md +21 -12
- package/docs/how-it-routes.md +49 -42
- package/docs/install.md +141 -33
- package/docs/part-1-beginner.md +37 -45
- package/docs/part-2-intermediate.md +34 -52
- package/docs/part-3-advanced.md +36 -26
- package/docs/security-review-history.md +39 -0
- package/llms.txt +24 -25
- package/package.json +15 -8
- package/proof/README.md +100 -0
- package/proof/gate-demo.cast +9 -0
- package/proof/gate-demo.gif +0 -0
- package/proof/results.json +198 -0
- package/proof/scripts/check-gate.js +26 -0
- package/proof/scripts/install-time.js +16 -0
- package/proof/scripts/lib.js +73 -0
- package/proof/scripts/measure.js +15 -0
- package/proof/scripts/missing-results.js +30 -0
- package/proof/scripts/record-gate.js +38 -0
- package/proof/scripts/render.js +18 -0
- package/proof/scripts/runner-overhead.js +21 -0
- package/src/README.md +10 -3
- package/src/activation-ownership.js +19 -0
- package/src/apply-companions.js +104 -0
- package/src/apply-snippets.js +60 -28
- package/src/aunx.js +272 -0
- package/src/bounded-file.js +31 -0
- package/src/catalog.js +257 -121
- package/src/install.js +483 -212
- package/src/plugin.js +13 -4
- package/src/postinstall.js +57 -0
- package/src/roles.js +184 -0
- package/src/uninstall.js +128 -10
- package/templates/README.md +19 -2
- package/templates/advanced/README.md +2 -2
- package/templates/advanced/vm/ENVIRONMENT.md +8 -0
- package/templates/advanced/vm/PRIVACY_GATES.md +17 -19
- package/templates/advanced/vm/README.md +25 -20
- package/templates/advanced/vm/box-CLAUDE.md +19 -18
- package/templates/advanced/vm/docker-compose.yml +2 -1
- package/templates/advanced/vm/jobs/README.md +31 -2
- package/templates/advanced/vm/jobs/weekly-audit.service +7 -2
- package/templates/advanced/vm/jobs/weekly-audit.sh +24 -17
- package/templates/advanced/vm/setup-vm.sh +49 -2
- package/templates/agents/README.md +2 -2
- package/templates/agents/agy/README.md +20 -3
- package/templates/agents/agy/builder.md +11 -7
- package/templates/agents/agy/bulk-worker.md +9 -7
- package/templates/agents/agy/code-reviewer.md +13 -7
- package/templates/agents/agy/deep-planner.md +10 -7
- package/templates/agents/agy/done-verifier.md +13 -22
- package/templates/agents/agy/finding-verifier.md +14 -22
- package/templates/agents/agy/live-researcher.md +10 -7
- package/templates/agents/agy/reader.md +10 -12
- package/templates/agents/claude-code/README.md +18 -14
- package/templates/agents/claude-code/builder.md +10 -15
- package/templates/agents/claude-code/bulk-worker.md +8 -10
- package/templates/agents/claude-code/code-reviewer.md +11 -17
- package/templates/agents/claude-code/deep-planner.md +9 -11
- package/templates/agents/claude-code/done-verifier.md +12 -33
- package/templates/agents/claude-code/finding-verifier.md +13 -39
- package/templates/agents/claude-code/live-researcher.md +9 -11
- package/templates/agents/claude-code/reader.md +9 -18
- package/templates/agents/snippets/chat.md +9 -10
- package/templates/agents/snippets/claude-code.md +17 -18
- package/templates/agents/snippets/generic.md +9 -11
- package/templates/agents/snippets/route-gate.mjs +2 -2
- package/templates/agents/snippets/route-metrics.mjs +1 -1
- package/templates/agents/snippets/subagent-context.mjs +4 -4
- package/templates/beginner/ORCHESTRATOR.md +31 -36
- package/templates/beginner/README.md +1 -1
- package/templates/common/ACCEPTANCE_CHECKS.json +12 -0
- package/templates/common/CONTEXT.md +37 -0
- package/templates/common/DECISIONS.md +11 -0
- package/templates/common/README.md +24 -11
- package/templates/common/TASK_BRIEF.md +84 -0
- package/templates/common/protocols/README.md +14 -11
- package/templates/common/protocols/acceptance-checks.md +15 -0
- package/templates/common/protocols/build-protocol.md +91 -106
- package/templates/common/protocols/context-file.md +10 -0
- package/templates/common/protocols/decision-log.md +9 -0
- package/templates/common/protocols/deep-research.md +20 -34
- package/templates/common/protocols/docs-then-prove.md +13 -18
- package/templates/common/protocols/gap-analysis.md +15 -21
- package/templates/common/protocols/memory-and-record.md +21 -20
- package/templates/common/protocols/numbers-and-logic.md +20 -26
- package/templates/common/protocols/propagate.md +18 -27
- package/templates/intermediate/CLI-RUN.md +83 -113
- package/templates/intermediate/DELEGATION_MATRIX.md +9 -3
- package/templates/intermediate/README.md +3 -3
- package/templates/intermediate/RESEARCH_TRIAGE.md +23 -15
- package/templates/intermediate/ROUTING.md +54 -51
- package/templates/intermediate/TIERS.md +37 -76
- package/templates/tools/README.md +1 -1
- package/templates/tools/codecalc/CODECALC.md +4 -4
- package/templates/tools/codecalc/mcp/agy.mcp_config.json +1 -1
- package/templates/tools/codecalc/mcp/codex.config.toml +1 -1
- package/templates/tools/codecalc/mcp/mcpServers.json +1 -1
- package/templates/tools/codecalc/mcp/vscode.mcp.json +1 -1
- package/templates/tools/codecalc/mcp/zed.settings.json +1 -1
- package/templates/tools/context7/CONTEXT7.md +6 -10
- package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +3 -3
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.agy.mcp_config.json +1 -1
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.codex.config.toml +1 -1
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.mcpServers.json +1 -1
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.vscode.mcp.json +1 -1
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.zed.settings.json +1 -1
- package/docs/audit-brief.md +0 -148
- package/scripts/README.md +0 -7
- package/scripts/gen-catalog.js +0 -81
- package/scripts/gen-plugin.js +0 -16
- package/scripts/record-demo.sh +0 -45
- package/templates/common/TASK_BUNDLE.md +0 -56
|
@@ -1,19 +1,19 @@
|
|
|
1
|
-
# vm/:
|
|
1
|
+
# vm/: templates for your always-on Linux machine
|
|
2
2
|
|
|
3
|
-
Level 3
|
|
3
|
+
Level 3 adds deployment templates to levels 1 and 2. Provide a Linux machine sized for your workload to hold the CLIs, a model gateway, and the scheduled jobs. Your laptop stays the interactive driver; the machine owns the schedule after you configure it.
|
|
4
4
|
|
|
5
5
|
Generated {{DATE}} for: `{{AI_IDS}}`. Installed at `{{INSTALL_DIR}}`; the systemd unit and the audit script carry that path.
|
|
6
6
|
|
|
7
|
-
##
|
|
7
|
+
## Keep provider credentials in the gateway
|
|
8
8
|
|
|
9
|
-
|
|
9
|
+
For metered API calls, inject provider credentials into the gateway and have jobs use its authenticated endpoint. Subscription CLIs keep their own vendor login state. Bind the gateway to loopback or the configured private network; never expose it publicly without explicit authorization and access controls.
|
|
10
10
|
|
|
11
11
|
## What runs where
|
|
12
12
|
|
|
13
13
|
| Surface | Role | Reaches models via |
|
|
14
14
|
|---|---|---|
|
|
15
15
|
| The orchestrator CLI ({{PRIMARY_NAME}}) | interactive driver when you SSH in; dispatch brain for jobs | its own subscription, off the gateway |
|
|
16
|
-
| `cli-run` lanes ({{CLI_RUN_LANES}}) | the other agent CLIs, headless | their own subscriptions (
|
|
16
|
+
| `cli-run` lanes ({{CLI_RUN_LANES}}) | the other agent CLIs, headless | their own subscriptions (subscription lanes) |
|
|
17
17
|
| The gateway (`docker-compose.yml`) | one OpenAI-compatible endpoint fronting every metered provider | provider keys from the environment |
|
|
18
18
|
| A local model runtime (if selected) | the privacy lane | nothing leaves the box |
|
|
19
19
|
| Scheduled jobs (`jobs/`) | the weekly gap-analysis audit (lane: `{{AUDIT_LANE}}`), and anything else recurring | the gateway, or `cli-run` |
|
|
@@ -22,25 +22,29 @@ Generated {{DATE}} for: `{{AI_IDS}}`. Installed at `{{INSTALL_DIR}}`; the system
|
|
|
22
22
|
|
|
23
23
|
## Setup, in order
|
|
24
24
|
|
|
25
|
-
|
|
25
|
+
The model-orchestrator installer writes these files only. When you separately invoke the generated `setup-vm.sh`, that manual deployment script installs the configured system dependencies and npm vendor CLIs. Review it before running it.
|
|
26
|
+
|
|
27
|
+
1. Provision a machine sized for your workload and put it on the configured private network.
|
|
26
28
|
2. `bash setup-vm.sh`. It installs system deps and the npm-installable CLIs, then **prints** the vendor shell installers for the rest. Read those scripts before running them.
|
|
27
29
|
3. Sign each CLI in, using the flow its vendor gives you. Run these inside `tmux` so a dropped SSH session does not kill the prompt. Headless Linux has no keyring by default; `setup-vm.sh` installs one so the CLIs stop re-prompting.
|
|
28
30
|
{{VM_SIGNIN}}
|
|
29
31
|
4. Put provider keys in your secrets manager and export the names listed in `ENVIRONMENT.md` into the gateway's environment at start time. Never write a value into a file in this folder. The gateway config was rendered from the API keys you said you hold, not from your CLI subscriptions: those are different entitlements.
|
|
30
|
-
5. `
|
|
32
|
+
5. From this `vm/` folder, run `bash setup-vm.sh --start-services` with those environment variables injected. It starts Compose and validates `GATEWAY_MASTER_KEY` as a single token matching `^[A-Za-z0-9._-]+$`. {{VM_LOCAL_SETUP}}
|
|
33
|
+
|
|
34
|
+
To inspect configured aliases afterward, keep the key out of argv:
|
|
31
35
|
```bash
|
|
32
36
|
printf 'header = "Authorization: Bearer %s"\n' "$GATEWAY_MASTER_KEY" | curl -s --config - http://127.0.0.1:4000/v1/models
|
|
33
37
|
```
|
|
34
|
-
6. Install the weekly audit: `jobs/README.md`.
|
|
38
|
+
6. Install the weekly audit: `jobs/README.md`. Set the service's literal `PATH` to include the directories that hold your Node and selected CLI executables before enabling the timer.
|
|
35
39
|
7. Copy `box-CLAUDE.md` to `~/CLAUDE.md` on the box (or your agent's equivalent rules file) so a session there inherits the house rules without you present.
|
|
36
40
|
|
|
37
41
|
## The dispatch shape
|
|
38
42
|
|
|
39
|
-
1.
|
|
40
|
-
2.
|
|
41
|
-
3.
|
|
42
|
-
4.
|
|
43
|
-
5.
|
|
43
|
+
1. When the task is obvious, use the routing table or `aunx route` suggestion to identify the candidate tier, then verify tools and scope.
|
|
44
|
+
2. When judgment is needed, apply `ROUTING.md` and `DELEGATION_MATRIX.md` and record the selected route with a reason.
|
|
45
|
+
3. When a cheaper eligible route can satisfy the checks, select it; when checks fail or required capability is absent, diagnose and choose an authorized fallback.
|
|
46
|
+
4. When an unattended action exceeds the existing mandate, preserve the result and return the needed approval through the configured channel.
|
|
47
|
+
5. When recording shared state, keep one writer and have other workers return proposed updates.
|
|
44
48
|
|
|
45
49
|
## The closed loop (name what watches it)
|
|
46
50
|
|
|
@@ -52,11 +56,12 @@ Generated {{DATE}} for: `{{AI_IDS}}`. Installed at `{{INSTALL_DIR}}`; the system
|
|
|
52
56
|
|
|
53
57
|
"Nothing watches it" is a valid answer and usually the valuable one. Writing it down turns an invisible gap into a tracked one.
|
|
54
58
|
|
|
55
|
-
##
|
|
59
|
+
## Keep the server within its scope
|
|
56
60
|
|
|
57
|
-
-
|
|
58
|
-
-
|
|
59
|
-
-
|
|
60
|
-
-
|
|
61
|
-
-
|
|
62
|
-
-
|
|
61
|
+
- Read vendor scripts before executing them.
|
|
62
|
+
- Update pinned packages and images deliberately, with a verification plan.
|
|
63
|
+
- Keep provider secrets in the manager and out of argv, logs and generated files.
|
|
64
|
+
- Bind services to loopback or the approved private network.
|
|
65
|
+
- Apply the named data-processing permissions in `PRIVACY_GATES.md` before dispatch.
|
|
66
|
+
- Obtain authorization before adding a paid resource or a payment method.
|
|
67
|
+
- When optional companion tools are absent, use the local runtime, official documentation and project files specified by the protocols.
|
|
@@ -1,28 +1,29 @@
|
|
|
1
|
-
#
|
|
1
|
+
# Project instructions for the server
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
When activating an unattended machine, copy these rules to the agent's documented instructions path and verify that a fresh session loads them.
|
|
4
4
|
|
|
5
|
-
##
|
|
6
|
-
Prefer the cheapest tier that does the job well. Delegate grunt work through the dispatch layer; keep the reasoning in-session.
|
|
5
|
+
## Choose the route
|
|
7
6
|
|
|
8
|
-
|
|
7
|
+
When work is classification, extraction, formatting or other bounded volume, select an eligible cheap model. When it requires architecture, writing code, review or current sources, use the tier and tools named in `{{RULES_PATH}}/ROUTING.md` and `{{RULES_PATH}}/DELEGATION_MATRIX.md`.
|
|
9
8
|
|
|
10
|
-
|
|
9
|
+
When dispatch fails, inspect its class and report the cause. Continue locally only when the session has the required scope, tools and capacity.
|
|
11
10
|
|
|
12
|
-
##
|
|
13
|
-
Report it and fall back to doing the work in-session. Never silently retry the same lane.
|
|
11
|
+
## Keep service ingress private
|
|
14
12
|
|
|
15
|
-
|
|
16
|
-
Nothing binds to `0.0.0.0`. Nothing publishes a container port to the public interface. New services go on loopback or the private mesh.
|
|
13
|
+
Bind services to loopback or the configured private network. Never publish a container port on the public interface without explicit authorization and the required access controls.
|
|
17
14
|
|
|
18
|
-
##
|
|
19
|
-
An unresolved call that is irreversible or rewrites a standing rule gets surfaced (a message, a ticket comment) and stops. It is never executed on the strength of a model's confidence.
|
|
15
|
+
## Handle unattended decisions
|
|
20
16
|
|
|
21
|
-
|
|
22
|
-
Scheduled jobs and other engines propose. One writer records. If you are not that writer, produce a file and name it in your report.
|
|
17
|
+
When an action exceeds the existing mandate, preserve the checked result and return the needed approval through the configured channel. Continue independent authorized work. Never execute an irreversible action solely on a model's confidence.
|
|
23
18
|
|
|
24
|
-
##
|
|
25
|
-
Names in the environment, values in the secrets manager. Never print one, never pass one in argv, never write one to disk here.
|
|
19
|
+
## Record with one writer
|
|
26
20
|
|
|
27
|
-
|
|
28
|
-
|
|
21
|
+
When scheduled jobs or other engines produce evidence, return proposed updates to the designated record writer. Keep raw machine logs separate from curated records.
|
|
22
|
+
|
|
23
|
+
## Protect secrets
|
|
24
|
+
|
|
25
|
+
Load secret values from the configured manager at runtime. Never print them, put them in argv, or write them into this generated folder.
|
|
26
|
+
|
|
27
|
+
## Run and verify
|
|
28
|
+
|
|
29
|
+
When building, follow `{{RULES_PATH}}/protocols/build-protocol.md`. For background work, arm the five-minute heartbeat at launch and diagnose two checks without progress. When optional companions are absent, use the local runtime, official docs and project records described by each protocol.
|
|
@@ -1,4 +1,5 @@
|
|
|
1
|
-
#
|
|
1
|
+
# This gateway receives provider credentials from its own launch environment.
|
|
2
|
+
# The weekly audit uses a separate environment containing only the gateway bearer.
|
|
2
3
|
# It is bound to loopback. Put it behind a private mesh network if other
|
|
3
4
|
# machines need it; never publish it on 0.0.0.0.
|
|
4
5
|
services:
|
|
@@ -4,12 +4,14 @@ Scheduled work on the box, as user-level systemd timers. Each job is a timer + s
|
|
|
4
4
|
|
|
5
5
|
| Job | Schedule | Does | Lane | Watched by |
|
|
6
6
|
|---|---|---|---|---|
|
|
7
|
-
| `weekly-audit` | Monday 09:00 | collects live state (gateway lanes, timers, CLI versions), composes a brief with the protocol and `DELEGATION_MATRIX.md`, and asks a cli-run lane for the gap report | `{{AUDIT_LANE}}` (first enabled
|
|
7
|
+
| `weekly-audit` | Monday 09:00 | collects live state (gateway lanes, timers, CLI versions), composes a brief with the protocol and `DELEGATION_MATRIX.md`, and asks a cli-run lane for the gap report | `{{AUDIT_LANE}}` (first enabled route at install time; edit `AUDIT_LANE` in the script to change it) | nothing yet: wire a notifier and update this line |
|
|
8
8
|
|
|
9
9
|
Paths in the service and the script were rendered for this install: `{{INSTALL_DIR}}`. If you move the folder, re-run the installer or edit both files.
|
|
10
10
|
|
|
11
11
|
## Install a job
|
|
12
12
|
|
|
13
|
+
Before copying the unit, run `command -v node` and `command -v <selected-cli>` in the account that will run the timer. The service sets `PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin`; add the absolute parent directories of your actual Node and vendor CLI executables to its `Environment="PATH=..."` line when needed. The script resolves `node` and vendor commands through that PATH. systemd does not expand `$PATH`, `$HOME` or `~` in this setting, so write complete directories and retain the system entries. Shell profile files and interactive version-manager initialization are not loaded.
|
|
14
|
+
|
|
13
15
|
```bash
|
|
14
16
|
mkdir -p ~/.config/systemd/user
|
|
15
17
|
cp weekly-audit.service weekly-audit.timer ~/.config/systemd/user/
|
|
@@ -19,7 +21,27 @@ systemctl --user list-timers # it should be listed with a next-run time
|
|
|
19
21
|
loginctl enable-linger "$USER" # so user timers run without a login session
|
|
20
22
|
```
|
|
21
23
|
|
|
22
|
-
The service reads `GATEWAY_MASTER_KEY` from
|
|
24
|
+
The service reads only `GATEWAY_MASTER_KEY` from `~/.config/ai-orchestrator/weekly-audit.env`, outside this folder with mode 600. Provision that audit-only file from your secrets manager and point `EnvironmentFile=` at it before installing. Keep the gateway's provider-key environment separate. Existing installs must replace or edit their copied service, then run `systemctl --user daemon-reload`; updating the source template alone does not update the installed unit. The key must be a single token matching `^[A-Za-z0-9._-]+$`; any embedded or trailing newline is refused.
|
|
25
|
+
|
|
26
|
+
Sign the selected vendor CLI in under the same user before enabling the timer. The job preserves `HOME`, `PATH`, stored sign-in state and unrelated authentication variables, but removes `GATEWAY_MASTER_KEY`, `LITELLM_MASTER_KEY`, `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GEMINI_API_KEY`, `XAI_API_KEY`, and `OPENROUTER_API_KEY` from all child environments. An API-only lane dependent on a removed variable needs vendor-supported stored authentication, such as `hermes auth add <provider>`, before scheduling it. Those provider keys belong in the gateway launch environment.
|
|
27
|
+
|
|
28
|
+
## Invocation, dependencies, reads and writes
|
|
29
|
+
|
|
30
|
+
The Monday timer starts the user service, which runs Bash on `vm/jobs/weekly-audit.sh`. The script queries the local gateway using curl, collects user timers and installed CLI versions, assembles a brief, and invokes `node bin/cli-run.mjs` with the selected audit lane. Dependencies are Bash, Node, curl, jq, pgrep, standard shell utilities, the systemd user session, and the selected signed-in CLI. The gateway probe may fail without preventing the report: that section becomes `UNVERIFIED`.
|
|
31
|
+
|
|
32
|
+
The script reads `protocols/gap-analysis.md`, `DELEGATION_MATRIX.md` (falling back to `ORCHESTRATOR.md`), and collected live state. It writes `reports/live-state-<stamp>.md`, the `live-state.md` symlink, `audit-brief-<stamp>.md`, and a dated report or retained failed output. Credential headers travel only through a pipe into curl's stdin. No gateway credential temp file is created.
|
|
33
|
+
|
|
34
|
+
## The closed loop
|
|
35
|
+
|
|
36
|
+
The timer schedules the job; the script records unknown probes in its brief; the worker returns a report; a clean, non-empty result replaces the dated report. The journal records its exit code and report path. **Watched by: nothing.** The operator must inspect the journal and report after a failed service or add a notifier, then update the table above. Automatic notification and remediation are not configured.
|
|
37
|
+
|
|
38
|
+
## Failure modes
|
|
39
|
+
|
|
40
|
+
- **Invalid key:** exit 2 before collection; correct the audit-only environment through the secrets manager.
|
|
41
|
+
- **Missing environment file:** systemd cannot start the service; provision the file and check its path and permissions.
|
|
42
|
+
- **Failed probe:** the live-state section says `UNVERIFIED`; check the named dependency and retry.
|
|
43
|
+
- **Missing lane or sign-in:** cli-run returns its failure code, retaining failed output and preserving the last good report. Configure that user's vendor sign-in and PATH.
|
|
44
|
+
- **Timeout:** the watchdog stops probe descendants, and the service's whole-job deadline kills its cgroup. A killed run can leave a non-secret report temp file; a later run sweeps report temps older than a day.
|
|
23
45
|
|
|
24
46
|
## Verify a job ran
|
|
25
47
|
|
|
@@ -29,11 +51,18 @@ journalctl --user -u weekly-audit.service -n 50
|
|
|
29
51
|
ls -la {{INSTALL_DIR}}/reports/
|
|
30
52
|
```
|
|
31
53
|
|
|
54
|
+
For a manual check, start `systemctl --user start weekly-audit.service`, then run the checks above. Confirm a new dated report contains the expected gap analysis and read any `UNVERIFIED` sections before treating the run as healthy. Verify the service uses the audit-only environment by inspecting its `EnvironmentFile=` path, never by printing environment values. Test coverage in the package's `test/security-vm.test.js` uses synthetic runtime credentials and stub CLIs to check stdin delivery, child environment isolation, invalid keys, and report preservation; it calls no vendor service.
|
|
55
|
+
|
|
32
56
|
## What the job guarantees
|
|
33
57
|
|
|
34
58
|
- **Bounded:** every probe (gateway, `systemctl`, each CLI `--version`) runs under a 10 s watchdog; the model call under 600 s; the unit under `TimeoutStartSec=900`, which kills the whole cgroup.
|
|
35
59
|
- **Previous report preserved:** output goes to a temp file and is renamed over `audit-<date>.md` only on a clean, non-empty run. A failed run leaves `failed-audit-<stamp>-rc<N>.md` beside it and the last good report untouched.
|
|
36
60
|
- **Boundary:** the lane runs with the strongest restriction it offers ({{AUDIT_LANE_BOUNDARY_NOTE}}). The brief's denied-actions list is an instruction, not an enforcement, for lanes without a sandbox flag.
|
|
37
61
|
- **Honest unknowns:** a probe that times out writes an `UNVERIFIED` line, which the brief tells the lane to treat as unknown, never clean.
|
|
62
|
+
- **Credential separation:** the gateway bearer is never written to a temp file or passed on argv, and gateway/provider keys are absent from version probes and the report worker's environment. Stored vendor sign-ins remain available.
|
|
38
63
|
|
|
39
64
|
A timer that has never been seen to fire is not known to work. Run `systemctl --user start weekly-audit.service` once by hand and read the journal before trusting the schedule.
|
|
65
|
+
|
|
66
|
+
## Source of truth
|
|
67
|
+
|
|
68
|
+
The installed `vm/jobs/weekly-audit.sh`, `weekly-audit.service`, and `weekly-audit.timer` define behavior; the copied unit in `~/.config/systemd/user/` is what systemd runs. This README describes that job end to end. Live scheduling and vendor authentication remain UNVERIFIED until the manual run succeeds on your VM.
|
|
@@ -5,8 +5,13 @@ After=network-online.target
|
|
|
5
5
|
[Service]
|
|
6
6
|
Type=oneshot
|
|
7
7
|
WorkingDirectory={{INSTALL_DIR_SYSTEMD}}
|
|
8
|
-
#
|
|
9
|
-
|
|
8
|
+
# Add absolute Node and vendor CLI directories if they live outside these defaults.
|
|
9
|
+
# systemd does not expand shell variables in Environment=.
|
|
10
|
+
Environment="PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
|
|
11
|
+
# Audit-only file OUTSIDE this repo, mode 600: GATEWAY_MASTER_KEY only.
|
|
12
|
+
# Provider credentials stay in the separate gateway/Compose environment.
|
|
13
|
+
# Vendor CLIs use this user's stored sign-in state. Edit the path if needed.
|
|
14
|
+
EnvironmentFile=%h/.config/ai-orchestrator/weekly-audit.env
|
|
10
15
|
ExecStart=/bin/bash "{{INSTALL_DIR_SYSTEMD}}/vm/jobs/weekly-audit.sh"
|
|
11
16
|
# Whole-job deadline: collection probes (6 x 10 s) + the 600 s model call + cleanup.
|
|
12
17
|
# On expiry systemd kills the whole cgroup, so nothing the job spawned survives.
|
|
@@ -14,6 +14,20 @@ AUDIT_LANE="{{AUDIT_LANE}}"
|
|
|
14
14
|
AUDIT_LANE_FLAGS="{{AUDIT_LANE_FLAGS}}"
|
|
15
15
|
PROBE_SECS="${PROBE_SECS:-10}" # per collection probe
|
|
16
16
|
RUNNER_SECS="${RUNNER_SECS:-600}" # the model call; TimeoutStartSec in the unit covers the whole job
|
|
17
|
+
# Keep the probe credential in this shell only. Gateway provider keys belong to
|
|
18
|
+
# Compose, not to collection tools or the scheduled vendor CLI. Stored vendor
|
|
19
|
+
# sign-ins, HOME, PATH, and unrelated authentication variables remain available.
|
|
20
|
+
KEY="${GATEWAY_MASTER_KEY:-}"
|
|
21
|
+
export -n KEY
|
|
22
|
+
unset GATEWAY_MASTER_KEY LITELLM_MASTER_KEY ANTHROPIC_API_KEY OPENAI_API_KEY GEMINI_API_KEY XAI_API_KEY OPENROUTER_API_KEY
|
|
23
|
+
|
|
24
|
+
# A shell pattern checks the whole value, including embedded/trailing newlines.
|
|
25
|
+
# Line-oriented grep accepts a valid line even when another line is malformed.
|
|
26
|
+
case "$KEY" in
|
|
27
|
+
*[!A-Za-z0-9._-]*)
|
|
28
|
+
echo "weekly-audit: GATEWAY_MASTER_KEY must match ^[A-Za-z0-9._-]+$ (generate it with: openssl rand -hex 32)" >&2
|
|
29
|
+
exit 2 ;;
|
|
30
|
+
esac
|
|
17
31
|
{{AUDIT_LANE_GUARD}}
|
|
18
32
|
cd "$INSTALL_DIR" || { echo "weekly-audit: $INSTALL_DIR missing" >&2; exit 2; }
|
|
19
33
|
mkdir -p reports
|
|
@@ -47,7 +61,9 @@ killtree() {
|
|
|
47
61
|
bounded() {
|
|
48
62
|
local secs="$1"; shift
|
|
49
63
|
local fired; fired="$(mktemp "${TMPDIR:-/tmp}/wa-fired.XXXXXX" 2>/dev/null)" && rm -f "$fired"
|
|
50
|
-
|
|
64
|
+
# Bash otherwise replaces stdin with /dev/null for asynchronous commands.
|
|
65
|
+
# Preserve the caller's pipe explicitly so curl can read its config on stdin.
|
|
66
|
+
( "$@" ) <&0 & local pid=$!
|
|
51
67
|
( sleep "$secs"; [ -n "$fired" ] && : > "$fired"; killtree "$pid" ) >/dev/null 2>&1 & local wd=$!
|
|
52
68
|
wait "$pid" 2>/dev/null; local rc=$?
|
|
53
69
|
killtree "$wd" >/dev/null 2>&1; wait "$wd" 2>/dev/null
|
|
@@ -55,29 +71,19 @@ bounded() {
|
|
|
55
71
|
return $rc
|
|
56
72
|
}
|
|
57
73
|
|
|
58
|
-
# The gateway key must be a single token: it is interpolated into curl's config
|
|
59
|
-
# grammar, and a quote or newline in it would become a second directive.
|
|
60
|
-
KEY="${GATEWAY_MASTER_KEY:-}"
|
|
61
|
-
if [ -n "$KEY" ] && ! printf '%s' "$KEY" | grep -Eq '^[A-Za-z0-9._-]+$'; then
|
|
62
|
-
echo "weekly-audit: GATEWAY_MASTER_KEY must match ^[A-Za-z0-9._-]+$ (generate it with: openssl rand -hex 32)" >&2
|
|
63
|
-
exit 2
|
|
64
|
-
fi
|
|
65
|
-
|
|
66
74
|
DATE="$(date -u +%F)"
|
|
67
75
|
STAMP="$(date -u +%Y%m%dT%H%M%SZ)"
|
|
68
76
|
{
|
|
69
77
|
echo "# live state $DATE"; echo
|
|
70
78
|
echo "## gateway lanes"
|
|
71
79
|
if [ -n "$KEY" ]; then
|
|
72
|
-
#
|
|
73
|
-
#
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
if ! bounded "$PROBE_SECS" curl -s --connect-timeout 5 --max-time "$PROBE_SECS" --max-filesize 1048576 --config "$CFG" http://127.0.0.1:4000/v1/models \
|
|
80
|
+
# No secret file or argv value: the builtin printf feeds curl through a pipe.
|
|
81
|
+
# curl's own timeouts AND the watchdog bound it, including a hung reader.
|
|
82
|
+
if ! printf 'header = "Authorization: Bearer %s"\n' "$KEY" \
|
|
83
|
+
| bounded "$PROBE_SECS" curl -s --connect-timeout 5 --max-time "$PROBE_SECS" --max-filesize 1048576 --config - http://127.0.0.1:4000/v1/models \
|
|
77
84
|
| jq -r '.data[].id' 2>/dev/null; then
|
|
78
85
|
echo "UNVERIFIED: gateway unreachable or timed out within ${PROBE_SECS}s"
|
|
79
86
|
fi
|
|
80
|
-
rm -f "$CFG"
|
|
81
87
|
else
|
|
82
88
|
echo "UNVERIFIED: GATEWAY_MASTER_KEY not set; gateway not queried"
|
|
83
89
|
fi
|
|
@@ -91,13 +97,14 @@ STAMP="$(date -u +%Y%m%dT%H%M%SZ)"
|
|
|
91
97
|
fi
|
|
92
98
|
done
|
|
93
99
|
} > "reports/live-state-$STAMP.md"
|
|
100
|
+
unset KEY
|
|
94
101
|
ln -sfn "live-state-$STAMP.md" reports/live-state.md
|
|
95
102
|
|
|
96
|
-
# The brief the
|
|
103
|
+
# The brief the worker actually reads: the protocol, the intended configuration,
|
|
97
104
|
# and the live state it is meant to diff against.
|
|
98
105
|
BRIEF="reports/audit-brief-$STAMP.md"
|
|
99
106
|
{
|
|
100
|
-
echo "## Task
|
|
107
|
+
echo "## Task brief"
|
|
101
108
|
echo "**Purpose.** Weekly gap analysis: compare the live state below with the intended configuration and name what is missing, dead, or drifted."
|
|
102
109
|
echo "**Task class.** read_only"
|
|
103
110
|
echo "**Denied actions.** Do not run commands, do not modify files, do not call any network service. Enforcement: {{AUDIT_LANE_BOUNDARY_NOTE}}."
|
|
@@ -8,6 +8,53 @@ set -euo pipefail
|
|
|
8
8
|
say() { printf '\n[setup-vm] %s\n' "$*"; }
|
|
9
9
|
have() { command -v "$1" >/dev/null 2>&1; }
|
|
10
10
|
|
|
11
|
+
# Run this phase explicitly after the dependency setup, sign-ins and secret injection.
|
|
12
|
+
# The package installer only writes this script; it never invokes either phase.
|
|
13
|
+
if [ "${1:-}" = "--start-services" ]; then
|
|
14
|
+
LOCAL_MODEL={{VM_LOCAL_MODEL_SH}}
|
|
15
|
+
KEY="${GATEWAY_MASTER_KEY:-}"
|
|
16
|
+
if [[ ! "$KEY" =~ ^[A-Za-z0-9._-]+$ ]]; then
|
|
17
|
+
echo 'setup-vm: GATEWAY_MASTER_KEY must match ^[A-Za-z0-9._-]+$' >&2
|
|
18
|
+
exit 2
|
|
19
|
+
fi
|
|
20
|
+
for cmd in docker curl jq; do
|
|
21
|
+
have "$cmd" || { echo "setup-vm: install $cmd before --start-services" >&2; exit 1; }
|
|
22
|
+
done
|
|
23
|
+
cd -- "$(dirname -- "${BASH_SOURCE[0]}")"
|
|
24
|
+
docker compose up -d
|
|
25
|
+
if [ -n "$LOCAL_MODEL" ]; then
|
|
26
|
+
ready=0
|
|
27
|
+
for attempt in {1..12}; do
|
|
28
|
+
if docker compose exec -T ollama ollama list >/dev/null 2>&1; then ready=1; break; fi
|
|
29
|
+
sleep 5
|
|
30
|
+
done
|
|
31
|
+
[ "$ready" -eq 1 ] || { echo 'setup-vm: Ollama service did not become ready' >&2; exit 1; }
|
|
32
|
+
docker compose exec -T ollama ollama pull "$LOCAL_MODEL"
|
|
33
|
+
ready=0
|
|
34
|
+
for attempt in {1..12}; do
|
|
35
|
+
# The key goes to curl on stdin, never argv, a file or printed output.
|
|
36
|
+
# A model list alone proves only that an alias is configured, not that it can answer.
|
|
37
|
+
if printf 'header = "Authorization: Bearer %s"\n' "$KEY" | \
|
|
38
|
+
curl --silent --fail --connect-timeout 5 --max-time 60 --config - \
|
|
39
|
+
--header 'Content-Type: application/json' \
|
|
40
|
+
--data '{"model":"local-small","messages":[{"role":"user","content":"Say ready."}],"max_tokens":16,"stream":false}' \
|
|
41
|
+
http://127.0.0.1:4000/v1/chat/completions | \
|
|
42
|
+
jq -e '.choices[0].message.content | type == "string" and length > 0' >/dev/null 2>&1; then
|
|
43
|
+
ready=1; break
|
|
44
|
+
fi
|
|
45
|
+
sleep 5
|
|
46
|
+
done
|
|
47
|
+
[ "$ready" -eq 1 ] || { echo 'setup-vm: local-small inference did not become ready; inspect docker compose logs' >&2; exit 1; }
|
|
48
|
+
say 'local-small inference verified'
|
|
49
|
+
fi
|
|
50
|
+
say 'services started; see jobs/README.md for the scheduled job'
|
|
51
|
+
exit 0
|
|
52
|
+
fi
|
|
53
|
+
if [ "$#" -ne 0 ]; then
|
|
54
|
+
echo 'usage: bash setup-vm.sh [--start-services]' >&2
|
|
55
|
+
exit 2
|
|
56
|
+
fi
|
|
57
|
+
|
|
11
58
|
say "system packages"
|
|
12
59
|
sudo apt-get update -y
|
|
13
60
|
sudo apt-get install -y curl git tmux jq ca-certificates gnupg build-essential \
|
|
@@ -41,6 +88,6 @@ for pkg in {{NPM_PACKAGES}}; do
|
|
|
41
88
|
done
|
|
42
89
|
|
|
43
90
|
say "vendor shell installers (read, then run yourself):"
|
|
44
|
-
{{
|
|
91
|
+
{{VM_SCRIPT_INSTALLERS}}
|
|
45
92
|
|
|
46
|
-
say "next: sign in to each CLI inside tmux (device-code flows),
|
|
93
|
+
say "next: sign in to each CLI inside tmux (device-code flows), inject the names in ENVIRONMENT.md, then: bash setup-vm.sh --start-services"
|
|
@@ -1,13 +1,13 @@
|
|
|
1
1
|
# templates/agents/
|
|
2
2
|
|
|
3
|
-
Loading surfaces for the
|
|
3
|
+
Loading surfaces for the main agent. The installer writes exactly one of these, based on `--primary`:
|
|
4
4
|
|
|
5
5
|
| Primary | Written | Why |
|
|
6
6
|
|---|---|---|
|
|
7
7
|
| `claude-code` | `.claude/agents/*.md` + `CLAUDE.snippet.md` | Claude Code loads project-level subagents from that folder |
|
|
8
8
|
| `agy` | `.agents/agents/*.md` + `GEMINI.snippet.md` | Antigravity custom agents live there |
|
|
9
9
|
| `codex`, `qwen` | `AGENTS.snippet.md` / `QWEN.snippet.md` | those CLIs read a rules file but have no subagent folder |
|
|
10
|
-
| `grok`, `hermes` | nothing agent-specific | rules travel with the prompt or the task
|
|
10
|
+
| `grok`, `hermes` | nothing agent-specific | rules travel with the prompt or the task brief |
|
|
11
11
|
| a chat app | `PASTE-INTO-YOUR-AGENT.md` | no files to load; paste into custom instructions |
|
|
12
12
|
|
|
13
13
|
`snippets/` are rendered with the chosen agent's name and rules file. Nothing here is appended to a file the user already has. `snippets/route-gate.mjs`, `snippets/subagent-context.mjs`, `snippets/route-metrics.mjs`, and `snippets/settings.hooks.snippet.json` are claude-code only: three hooks and the settings block that wires them, installed to `.claude/hooks/` and next to `CLAUDE.snippet.md`.
|
|
@@ -1,5 +1,22 @@
|
|
|
1
|
-
#
|
|
1
|
+
# Antigravity project agents
|
|
2
2
|
|
|
3
|
-
Antigravity
|
|
3
|
+
When using Antigravity custom agents, load these definitions from `.agents/agents/<name>.md`. A coordinator can call them through `invoke_subagent`; `mainAgent: true` also supports `agy --agent <name>`.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
The definitions omit the optional model field and inherit your configuration. To pin an available alias, add `model:` to a definition: Antigravity exposes `pro` for planning and `flash` for working and cheap tiers. Check your access, the live roster and tool reach before assigning a build.
|
|
6
|
+
|
|
7
|
+
| Agent | Tier | Effort guidance |
|
|
8
|
+
|---|---|---|
|
|
9
|
+
| deep-planner | planning model | xhigh where supported |
|
|
10
|
+
| builder | working model | high |
|
|
11
|
+
| code-reviewer | working model | high |
|
|
12
|
+
| finding-verifier | working model | high |
|
|
13
|
+
| live-researcher | working model | medium |
|
|
14
|
+
| bulk-worker | cheap model | low |
|
|
15
|
+
| done-verifier | cheap model | low |
|
|
16
|
+
| reader | cheap model | low |
|
|
17
|
+
|
|
18
|
+
`builder` has `commandExecutionPolicy: auto` so standard builds and checks can run, while destructive operations remain subject to the vendor's permission policy. Review and reading agents use `commandExecutionPolicy: off`.
|
|
19
|
+
|
|
20
|
+
`code-reviewer`, `finding-verifier`, `done-verifier` and `reader` use read-only tools with command execution disabled. Their Claude Code counterparts carrying Bash have a prompt-enforced read-only boundary instead. When an Antigravity agent needs a shell probe, hand the exact probe to an authorized worker and report the unverified check until its evidence returns.
|
|
21
|
+
|
|
22
|
+
When optional companion software is absent, use available read, fetch and local-runtime capabilities through the authorized owner. `bulk-worker` owns classification and transformation; `reader` returns source facts and digests.
|
|
@@ -1,17 +1,21 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: builder
|
|
3
|
-
description:
|
|
4
|
-
model: flash
|
|
3
|
+
description: Implements the section assigned by the task brief; writes code, edits files and runs the required checks.
|
|
5
4
|
subagent: true
|
|
6
5
|
mainAgent: true
|
|
7
6
|
commandExecutionPolicy: auto # standard build/test commands run unattended; destructive commands, like deletes, still ask before running
|
|
8
7
|
---
|
|
9
8
|
|
|
9
|
+
Tier: working model. This agent inherits the model your Antigravity configuration selects. Antigravity exposes `pro` and `flash`; the working and cheap tiers both map to `flash` when you choose an explicit alias. To pin one, add a `model:` line here after checking your access.
|
|
10
|
+
|
|
10
11
|
# builder
|
|
11
12
|
|
|
12
|
-
|
|
13
|
+
When a task brief assigns implementation, read its context file and acceptance checks first. Confirm the assigned paths, interfaces, capabilities and current runtime access.
|
|
13
14
|
|
|
14
|
-
|
|
15
|
-
-
|
|
16
|
-
-
|
|
17
|
-
-
|
|
15
|
+
- When a plan has an implementation gap within scope, state the assumption and verify it. When the gap changes architecture or authority, return the needed decision.
|
|
16
|
+
- Write the assigned section using the project's conventions and existing dependencies.
|
|
17
|
+
- When the build depends on a changing interface, consult current official docs or installed source and run a check.
|
|
18
|
+
- When the sandbox refuses a write, hand the required patch to an authorized writer and continue independent work.
|
|
19
|
+
- When authorized to split work, give each child the whole scope and its own section. Merge the result and name conflicts.
|
|
20
|
+
- When checks pass, report changed paths, coverage against the brief and evidence. Leave independent audit to the assigned reviewer.
|
|
21
|
+
- Keep context targeted and return concise results with source paths.
|
|
@@ -1,17 +1,19 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: bulk-worker
|
|
3
|
-
description:
|
|
4
|
-
model: flash
|
|
3
|
+
description: Classifies, tags, extracts, reformats or summarizes many similar items with a cheap model and bounded scope.
|
|
5
4
|
subagent: true
|
|
6
5
|
mainAgent: true
|
|
7
6
|
commandExecutionPolicy: off
|
|
8
7
|
---
|
|
9
8
|
|
|
9
|
+
Tier: cheap model. This agent inherits the model your Antigravity configuration selects. Antigravity exposes `pro` and `flash`; the working and cheap tiers both map to `flash` when you choose an explicit alias. To pin one, add a `model:` line here after checking your access.
|
|
10
|
+
|
|
10
11
|
# bulk-worker
|
|
11
12
|
|
|
12
|
-
|
|
13
|
+
When a brief assigns many similar items, use its categories or output schema consistently across the full authorized set.
|
|
13
14
|
|
|
14
|
-
|
|
15
|
-
-
|
|
16
|
-
-
|
|
17
|
-
-
|
|
15
|
+
- Read the context and scope before processing.
|
|
16
|
+
- When the categories are unclear or items stop fitting, report the mismatch and the affected items before continuing dependent work.
|
|
17
|
+
- Return structured output with one row or item per input, using short identifiers instead of repeating full input text.
|
|
18
|
+
- Write only to destinations the brief authorizes.
|
|
19
|
+
- Check input coverage and output shape, then report omissions and unverified items.
|
|
@@ -1,17 +1,23 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: code-reviewer
|
|
3
|
-
description:
|
|
4
|
-
model: flash
|
|
3
|
+
description: Reviews code for concrete security and correctness failures; command execution disabled; read-only tools.
|
|
5
4
|
subagent: true
|
|
6
5
|
mainAgent: true
|
|
7
6
|
commandExecutionPolicy: off
|
|
8
7
|
---
|
|
9
8
|
|
|
9
|
+
Tier: working model. This agent inherits the model your Antigravity configuration selects. Antigravity exposes `pro` and `flash`; the working and cheap tiers both map to `flash` when you choose an explicit alias. To pin one, add a `model:` line here after checking your access.
|
|
10
|
+
|
|
10
11
|
# code-reviewer
|
|
11
12
|
|
|
12
|
-
|
|
13
|
+
When assigned a review, read the task brief, context file, final diff and acceptance checks. Review the merged artifact against scope in the single audit step.
|
|
14
|
+
|
|
15
|
+
Command execution is disabled by `commandExecutionPolicy: off`. Use available read and fetch tools. When a check needs a command, return the needed authorized probe as UNVERIFIABLE or INCONCLUSIVE rather than running it.
|
|
13
16
|
|
|
14
|
-
|
|
15
|
-
-
|
|
16
|
-
-
|
|
17
|
-
-
|
|
17
|
+
- Trace each suspected failure to concrete input, state, caller and affected behavior.
|
|
18
|
+
- Check guards, tests and framework behavior that could disprove the claim.
|
|
19
|
+
- Rank reproducible security and correctness findings by severity; cite the file and line, trigger, consequence and proposed fix.
|
|
20
|
+
- When a scanner flags a line, inspect the actual object before repeating the finding.
|
|
21
|
+
- When reviewing code you authored, hand the review to an independent author and model family.
|
|
22
|
+
- When the code is clean, return CLEAN with the checked scope and limits.
|
|
23
|
+
- Suggest fixes and return evidence; fixes are assigned separately.
|
|
@@ -1,17 +1,20 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: deep-planner
|
|
3
|
-
description:
|
|
4
|
-
model: pro
|
|
3
|
+
description: Resolves architecture, strategy and unknown causes from a prepared context file; returns an executable plan.
|
|
5
4
|
subagent: true
|
|
6
5
|
mainAgent: true
|
|
7
6
|
commandExecutionPolicy: off
|
|
8
7
|
---
|
|
9
8
|
|
|
9
|
+
Tier: planning model. This agent inherits the model your Antigravity configuration selects. Antigravity exposes `pro` and `flash`; the working and cheap tiers both map to `flash` when you choose an explicit alias. To pin one, add a `model:` line here after checking your access.
|
|
10
|
+
|
|
10
11
|
# deep-planner
|
|
11
12
|
|
|
12
|
-
|
|
13
|
+
When the task needs architecture, strategy or an unknown cause resolved, read the prepared context file and acceptance checks, then test the key assumptions.
|
|
13
14
|
|
|
14
|
-
|
|
15
|
-
-
|
|
16
|
-
-
|
|
17
|
-
-
|
|
15
|
+
- Compare the mechanism-distinct options that fit the request and recommend one with concrete tradeoffs.
|
|
16
|
+
- Use the prepared map for retrieval evidence; when a claim is uncertain, request a targeted probe.
|
|
17
|
+
- At Assign, compare available lanes by reasoning, tool reach, context window and capacity, then record the choice and reason.
|
|
18
|
+
- Return a plan with file boundaries, interfaces, risky assumptions, verification and order of work.
|
|
19
|
+
- Keep this session read-only. Your result is a plan or analysis; code changes belong to the assigned builder.
|
|
20
|
+
- Cite the evidence supporting decisions and keep the report sized to the executor's needs.
|