model-orchestrator 1.0.2 → 1.0.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,21 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [1.0.3] - 2026-09-28
8
+
9
+ ### Security
10
+
11
+ - Weekly audit jobs use an explicit child environment allowlist, keep selected worker configuration overrides out of probes, and preserve stored sign-in paths. See the generated jobs README for migration (#45).
12
+ - Catalog package and container pins get advisory checks with explicit coverage, dated machine-readable results and expiring advisory-specific exceptions. See [catalog advisory checks](docs/catalog-advisories.md) for operation and limitations (#46).
13
+
14
+ ### Added
15
+
16
+ - `aunx route-metrics --summary` compares the route your agent named in each reply with the delegation that actually followed, from data the log already records: the share of sessions that match, sessions that named a route and dispatched nothing, and sessions that dispatched without naming one.
17
+
18
+ ### Fixed
19
+
20
+ - The product website navigation follows the section being read, keeps the active sidebar link visible, and labels its portfolio breadcrumb Home.
21
+
7
22
  ## [1.0.2] - 2026-09-28
8
23
 
9
24
  ### Fixed
@@ -525,7 +540,8 @@ First release.
525
540
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
526
541
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
527
542
 
528
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.2...HEAD
543
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.3...HEAD
544
+ [1.0.3]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.2...v1.0.3
529
545
  [1.0.2]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.1...v1.0.2
530
546
  [1.0.1]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.0...v1.0.1
531
547
  [1.0.0]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.35...v1.0.0
package/docs/README.md CHANGED
@@ -12,6 +12,7 @@ Use these pages for setup, operation and evidence. The [front page](../README.md
12
12
  | [Intermediate](part-2-intermediate.md) | Delegation across several AI CLIs |
13
13
  | [Advanced](part-3-advanced.md) | Gateway templates and scheduled work on a Linux host |
14
14
  | [Catalog](catalog.md) | Supported tools, capability facts, unverified values, installation and sign-in notes |
15
+ | [Catalog advisory checks](catalog-advisories.md) | Package and container coverage, report statuses, exceptions and manual verification |
15
16
  | [Security review history](security-review-history.md) | Review rounds, reproduced findings, fixes and regression tests |
16
17
  | [Proof](../proof/README.md) | Dated measurements, methods, sample sizes and reproduction scripts |
17
18
  | [Commands](../bin/README.md) | `aunx` subcommands and lane-runner exit codes |
@@ -0,0 +1,114 @@
1
+ # Catalog advisory checks
2
+
3
+ ## What and why
4
+
5
+ The repository checks exact executable package and container pins in `src/catalog.js`. OSV API v1 answers direct npm and PyPI package/version queries. Trivy examines OS and language packages in the selected container image. This check gives dated advisory evidence for its stated coverage. A clean result does not prove safety.
6
+
7
+ The owning files are `scripts/catalog-advisory-inventory.mjs`, `scripts/check-catalog-advisories.mjs`, `scripts/catalog-advisory-http.mjs`, `scripts/catalog-advisory-containers.mjs`, `docs/catalog-advisory-exceptions.json` and `.github/workflows/catalog-advisories.yml`. This document describes the complete operating procedure.
8
+
9
+ ## Trigger
10
+
11
+ The `catalog advisories` workflow runs on every push to main, on pull requests changing the catalog, checker, exceptions, tests or workflow, and by manual dispatch. It has read-only repository permissions. There is no periodic schedule; a maintainer can dispatch a fresh check when advisory data changes. Existing Dependabot automation continues to update GitHub Actions separately.
12
+
13
+ ## Invocation chain
14
+
15
+ 1. The workflow checks out the source and selects Node 22.
16
+ 2. The inventory imports `AIS`, `TOOLS`, `IMAGES` and `PROVIDERS` directly from the catalog. npm names come from `install.npm`; companion ecosystem, package and extras come from their `advisory` metadata. Every version comes from the existing pin field.
17
+ 3. The package adapter posts exact package/version queries to `https://api.osv.dev/v1/querybatch`, then retrieves each returned advisory from `/v1/vulns/<id>` to retain ranges and fixed versions.
18
+ 4. The container adapter resolves target tags to a platform-specific SHA-256 manifest digest, checks the manifest and image configuration, and records the platform. Docker Hub and GHCR public images are supported; another registry is unknown until its adapter is implemented.
19
+ 5. The adapter resolves the Trivy scanner image to a digest too, asserts its reported version, and runs that scanner against the target digest. It starts no target image or vendor service. Scanning uses registry access, with no host mounts or Docker socket passed into the scanner.
20
+ 6. Exact, unexpired exceptions are applied. The checker writes `report.json` and `summary.md`, then exits with its result code. The workflow appends the summary to the job and uploads both files even if the scan step fails.
21
+
22
+ ## Dependencies
23
+
24
+ A repository checkout, Node 18+ with built-in fetch, Docker with a running Linux-container daemon, public HTTPS registry access, OSV, and the Trivy vulnerability databases are needed for a complete local check. CI supplies Node 22 and Docker. Package checks need only Node and OSV. A missing scanner, daemon or database produces unknown container results while retaining available package results.
25
+
26
+ Trivy is version-pinned in `scripts/catalog-advisory-containers.mjs`. Each run resolves that scanner tag to an immutable digest and records it in each container result. Runtime version verification rejects a binary that reports a different version. The version probe has a 60-second deadline; image scans have Trivy's 10-minute deadline and a 12-minute client deadline, followed by bounded container cleanup. The scanner tag itself can move between runs; the recorded digest is the reproducible scanner identity for that run. For exact reproduction, fetch the scanner and target digests from the report before investigating.
27
+
28
+ Pin provenance, checked 2026-09-28 against upstream: `actions/checkout` `3d3c42e` is tag v7.0.1, `actions/setup-node` `8207627` is tag v7.0.0, and `actions/upload-artifact` `043fb46` is tag v7.0.1 in each action owner's repository. Trivy 0.74.0 is the latest release in Aqua Security's official repository (2026-08-14) and its public image tag resolves to `sha256:62b1e65e8869bc4b4c6aa4fa2b21595256c7c2f6018a9d9ad61caf87187c1969`. The first successful main-branch artifact is the check that the JSON parser matches this Trivy version.
29
+
30
+ To update the scanner, verify the desired release in Aqua Security's official Trivy repository, resolve its public registry tag to a digest, and inspect the pinned version's `image --help`, JSON report schema and database requirements. Change `TRIVY` in `scripts/catalog-advisory-containers.mjs`, run the offline fixtures, then run the real checker and inspect its recorded scanner digest and version. This builder could not compare the JSON parser with current official Trivy source; that compatibility is **UNVERIFIED** until a real scan passes. To update an action, compare its full commit SHA against the release tag in the action owner's repository and update the SHA and version comment together. Run the workflow and download the artifact before accepting either upgrade.
31
+
32
+ ## Reads
33
+
34
+ - `src/catalog.js`: package names, ecosystems, extras, exact versions, image references and excluded inputs. A new companion without advisory metadata becomes unknown.
35
+ - `docs/catalog-advisory-exceptions.json`: the reviewed exception array, initially empty.
36
+ - OSV: current direct-package advisory matches and their details. No project files are sent.
37
+ - Public container registries: manifests, configurations, layers and anonymous pull authorization. Anonymous registry tokens stay in memory. User sign-ins are not read by the resolver; the Docker client receives an empty temporary configuration directory.
38
+ - Trivy databases: current OS and language advisories. All severities and unfixed findings are retained.
39
+
40
+ Direct package checks do not resolve dependency trees. Each such input has an explicit transitive-dependency exclusion. `codecalc[full]` also has an optional-extras exclusion: querying the base distribution does not cover dependencies selected by `full`. Vendor installer scripts, unpinned downloads/package-manager inputs, chat apps and hosted model services are named exclusions. Compatibility snapshots such as `builtAgainst` are not installation pins. Image coverage is the OS and language packages Trivy recognizes; unsupported OS images and empty package results are unknown.
41
+
42
+ ## Writes
43
+
44
+ The default output directory is `.release-work/catalog-advisories/`. The checker writes:
45
+
46
+ - `report.json`: source commit, timestamp, scanner versions, catalog input, ecosystem, exact version, image/scanner digest, platform, coverage, status, advisory IDs, ranges, fixed versions and matched exception expiry.
47
+ - `summary.md`: the same result in a concise human-readable form, with explicit exclusions.
48
+
49
+ CI retains these as `catalog-advisories-<commit>-<attempt>` for 30 days and shows the summary in the job. An early workflow failure creates an unknown fallback report instead of presenting a missing report as success. The checker creates and removes an empty temporary Docker configuration directory. Scanner cache data lives under `/tmp/trivy` in its disposable writable container layer; `--rm` removes that layer. A disk-backed layer accommodates large image scans without a small RAM-backed cache cap. The scanner runs as a nonroot user with capabilities dropped, no new privileges, and no host filesystem or socket mounts. Live disk usage is **UNVERIFIED** here. Catalog pins and installed dependencies are never changed.
50
+
51
+ ## The closed loop
52
+
53
+ GitHub Actions is the watcher. A maintainer reviews every affected or unknown result before merging or releasing. No automatic issue, upgrade, exception or release is created by this workflow. Branch-protection configuration is repository administration: **UNVERIFIED** here. To make a green check mandatory, require the `catalog advisories / scan` check in the repository's branch rules.
54
+
55
+ | Status | Meaning | Exit behavior |
56
+ |---|---|---|
57
+ | clean | Complete adapter response with no advisory match in stated coverage | 0 if every result is clean or excepted |
58
+ | affected | At least one advisory match lacks a valid exception | 1 unless another result is unknown |
59
+ | unknown | Missing, unavailable, unsupported or malformed evidence, or an invalid exception policy | 2 |
60
+ | excepted | Every advisory match has an exact, reviewed, unexpired exception | 0 if every result is clean or excepted |
61
+
62
+ Any unknown takes precedence over affected for the exit code. Known advisory IDs are still retained when a detail lookup fails. Empty results never pass. Read the exclusions alongside the status; exclusions are never counted as clean pins.
63
+
64
+ ## Failure modes
65
+
66
+ - **OSV timeout, HTTP error or malformed response:** package results become unknown. A valid empty query result (`{}`) means no returned match; a missing result array or missing position is unknown.
67
+ - **Registry failure, bad image reference, missing platform or digest mismatch:** the image is unknown. The checker retains the other results.
68
+ - **Docker, scanner or database failure:** the image is unknown. No exception suppresses unknown status. Check network access and the daemon, then rerun the same source.
69
+ - **Incomplete scanner JSON, mismatched digest/platform or unsupported image OS:** the image is unknown even if the process exited successfully.
70
+ - **Advisory outage:** keep the failing result, retry after recovery, and inspect the provider's status through its normal support channel. An outage is not an exception for a package.
71
+ - **Expired or malformed exception:** expired entries stop matching at 00:00 UTC on their expiry date; malformed policy blocks success. Expired entries are retained in `expiredExceptions` for review.
72
+ - **Interrupted job:** `always()` upload and summary steps preserve available evidence. A runner outage can still prevent artifact upload; verify artifact presence in the run itself.
73
+
74
+ ## Exception review
75
+
76
+ The exception file is a JSON array. Every entry requires `ecosystem`, `package`, `version`, `advisory`, `rationale` and `expires` (`YYYY-MM-DD`). Wildcards and unknown fields are rejected. Match the exact advisory ID emitted by the adapter, not an alias. A package exception optionally names its exact catalog `source` in `target`. A container exception requires `target` equal to the exact catalog image reference and identifies the vulnerable component package and installed version, not every component in the image.
77
+
78
+ For example, a metadata-only example is:
79
+
80
+ ```json
81
+ [
82
+ {
83
+ "ecosystem": "npm",
84
+ "package": "fixture-package",
85
+ "version": "1.0.0",
86
+ "advisory": "TEST-2026-0001",
87
+ "rationale": "Example only: replace with a reviewed reason and tracking reference.",
88
+ "expires": "2026-10-01"
89
+ }
90
+ ]
91
+ ```
92
+
93
+ A maintainer reviews the affected range, fixed versions, applicability, expiry and rationale in a normal pull request. Use a short explicit expiry and a tracking reference in the rationale. Never add an exception merely because the database or scanner is unavailable. Review expired entries and remove those whose pin or advisory is no longer relevant. A matched exception remains visible as excepted with its expiry; it is never relabeled clean. Fixing a package requires a separate deliberate catalog-pin change and the usual installer validation.
94
+
95
+ ## Run and verify by hand
96
+
97
+ From the repository checkout:
98
+
99
+ ```bash
100
+ node --test test/catalog-advisories.test.js
101
+ node scripts/check-catalog-advisories.mjs --report-dir .release-work/catalog-advisories
102
+ ```
103
+
104
+ The fixture tests use metadata only and make no network request. The second command makes real advisory and registry requests and may start the pinned scanner container. Use `--platform linux/arm64` to inspect another declared image platform; the default is `linux/amd64`. Use `--exceptions <path>` for a reviewed alternate policy and retain that file with the evidence.
105
+
106
+ 1. Read both output files and check `sourceCommit`, `generatedAt`, every catalog source, scanner identity, platform and exclusions.
107
+ 2. For a package finding, retrieve its recorded OSV detail URL and compare the affected package, ranges and fixed versions. The fixture's affected result must exit 1; an unavailable-data result must exit 2.
108
+ 3. For an image finding, pull the recorded scanner digest and run its `image --scanners vuln --list-all-pkgs --format json --platform <platform> --image-src remote <image>@<digest>` command. Compare the component name, installed version, advisory ID and database evidence. A later database can change matches while the image stays identical.
109
+ 4. Inspect the GitHub run on main and download its report artifact. A local report is not proof of main-branch retention. Re-run the workflow after an outage and confirm a new timestamp.
110
+ 5. Run `npm test` before releasing. Existing OIDC publication and installer flags are unchanged by this check.
111
+
112
+ ## Source of truth
113
+
114
+ Catalog pins and companion metadata: `src/catalog.js`. Exception decisions: `docs/catalog-advisory-exceptions.json` and its reviewed change. Operational behavior: this document and the checker modules. Dated evidence: the uploaded GitHub run artifact. Upstream evidence: each report's OSV URL and Trivy data-source URL. A report describes the source commit, selected platform and database responses at its recorded time; it is not a permanent certificate.
@@ -0,0 +1 @@
1
+ []
@@ -50,7 +50,7 @@ Optional companions can help: codecalc for execution and calculations, obsidian-
50
50
 
51
51
  ## Measure your own routing
52
52
 
53
- `aunx route-metrics --summary` reads your local Claude Code routing log. It reports where work went, route-marker coverage and subagent durations. Your own measurements are the basis for changing assignments and checking whether the rules are being followed.
53
+ `aunx route-metrics --summary` reads your local Claude Code routing log. It reports where work went, route-marker coverage, subagent durations, and whether the route each reply named matches the delegation that followed. Your own measurements are the basis for changing assignments and checking whether the rules are being followed.
54
54
 
55
55
  ## What the installer gives you at this level
56
56
 
package/llms.txt CHANGED
@@ -29,6 +29,7 @@ Pick a model proxy (LiteLLM, Portkey, OpenRouter, claude-code-router) for per-re
29
29
  - [Claude Code plugin](https://github.com/aunysillyme/model-orchestrator/blob/main/plugin/README.md): installing read-only routing hooks and subagents through the plugin marketplace
30
30
  - [Proof](https://github.com/aunysillyme/model-orchestrator/blob/main/proof/README.md): dated measurements, methods, sample sizes, expiry and reproduction scripts
31
31
  - [Security review history](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/security-review-history.md): completed review rounds, incident summaries and regression evidence
32
+ - [Catalog advisory checks](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/catalog-advisories.md): package and container inventory, coverage, report statuses, exceptions and manual verification
32
33
  - [Changelog](https://github.com/aunysillyme/model-orchestrator/blob/main/CHANGELOG.md): release changes and upgrade notes
33
34
  - [Agent instructions](https://github.com/aunysillyme/model-orchestrator/blob/main/AGENTS.md): headless setup and contributor checks
34
35
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "1.0.2",
3
+ "version": "1.0.3",
4
4
  "description": "Model router for AI coding agents: installs routing rules, 8 subagents, hooks and a CLI runner so your AI picks model and effort per task and saves tokens",
5
5
  "type": "module",
6
6
  "bin": {
package/src/README.md CHANGED
@@ -3,7 +3,7 @@
3
3
  | File | Job |
4
4
  |---|---|
5
5
  | `bounded-file.js` | shared regular-file reader for manifests and check configuration: no-follow/nonblocking open, identity checks and a fixed byte cap even if a file grows. Unsafe files are refused; callers decide how to handle missing or malformed data. |
6
- | `catalog.js` | the single list of levels and AIs. Add an AI here and the prompts, docs tables, delegation matrix, gateway config and installer all pick it up. Nothing else lists AIs. |
6
+ | `catalog.js` | the single list of levels and AIs. Add an AI here and the prompts, docs tables, delegation matrix, gateway config and installer all pick it up. Nothing else lists AIs. Package pins, companion advisory metadata and image references also supply the repository's advisory inventory. |
7
7
  | `roles.js` | pure role assignment from selected catalog capability facts, billing and selection order. Renders the stack table and manifest roles, and infers the main agent from its supported surfaces. Unknown facts remain unverified; review requires a known different model family and private work requires local execution. |
8
8
  | `aunx.js` | command dispatch for briefs, context, checks, routing and runner calls. Route suggestions read manifest roles through a capped regular-file JSON reader; symlinks and malformed files are ignored. Route lookup executes no project code. A project's runner requires explicit `--dir`. |
9
9
  | `detect.js` | PATH lookup for a binary, plus the few places vendor installers drop binaries without touching PATH. No shell-outs. |
package/src/catalog.js CHANGED
@@ -402,6 +402,7 @@ export const TOOLS = [
402
402
  role: 'exact arithmetic, code execution in 31 languages, SMT logic checks, complexity and equivalence proofs; offline, no key, no telemetry',
403
403
  get install() { return `uvx 'codecalc[full]==${this.pin}' setup --write`; },
404
404
  pin: '0.5.0',
405
+ advisory: { ecosystem: 'PyPI', package: 'codecalc', extras: ['full'] },
405
406
  mcpSnippets: { 'claude-code': 'mcp/mcpServers.json', codex: 'mcp/codex.config.toml', agy: 'mcp/agy.mcp_config.json', qwen: 'mcp/mcpServers.json' },
406
407
  requires: 'uv (https://docs.astral.sh/uv/) and Python 3.10+',
407
408
  autoClients: ['Claude Code', 'Claude Desktop', 'Cursor', 'VS Code', 'Zed'],
@@ -415,6 +416,7 @@ export const TOOLS = [
415
416
  role: 'durable memory and record for your agents: hybrid retrieval (BM25 + dense + link graph), backlinks, compare-and-swap writes with a confirmation gate, folder ACLs, a poison scan on inferred writes; 163 tools, local by default',
416
417
  get install() { return `npm install -g obsidian-tc@${this.pin} && obsidian-tc /path/to/your/vault`; },
417
418
  pin: '1.26.0',
419
+ advisory: { ecosystem: 'npm', package: 'obsidian-tc' },
418
420
  mcpSnippets: { 'claude-code': 'mcp/obsidian-tc.mcpServers.json', codex: 'mcp/obsidian-tc.codex.config.toml', agy: 'mcp/obsidian-tc.agy.mcp_config.json', qwen: 'mcp/obsidian-tc.mcpServers.json' },
419
421
  requires: 'an Obsidian vault folder (the Obsidian app itself is only needed for live plugin bridges); Node 24+ or Bun 1.1+ (stricter than this installer); Ollama with `nomic-embed-text` for local embeddings, or a cloud embeddings key; the Local REST API plugin only for bridge tools',
420
422
  autoClients: ['Cursor', 'VS Code'],
@@ -428,6 +430,7 @@ export const TOOLS = [
428
430
  role: 'up-to-date, version-specific documentation and code examples for libraries, SDKs, APIs and CLIs, pulled into the prompt; tells the agent what the code is SUPPOSED to do. Paired with codecalc, which runs the code and proves what it actually does: docs never stand as proof, and where they disagree the run wins',
429
431
  get install() { return `npx -y @upstash/context7-mcp@${this.pin}`; },
430
432
  pin: '4.1.1',
433
+ advisory: { ecosystem: 'npm', package: '@upstash/context7-mcp' },
431
434
  mcpSnippets: { 'claude-code': 'mcp/context7.claude-code.mcp.json', codex: 'mcp/context7.codex.config.toml', agy: 'mcp/context7.agy.mcp_config.json', qwen: 'mcp/context7.qwen.settings.json' },
432
435
  requires: 'Node.js 18+ for the local server or the ctx7 CLI; a free CONTEXT7_API_KEY is optional, for higher rate limits (it works anonymously at the base rate)',
433
436
  autoClients: [], // The pinned MCP server does not register itself; merge its snippets.
@@ -12,11 +12,17 @@ These are variable **names**. The values live in a secrets manager and are injec
12
12
 
13
13
  ## Gateway and scheduled audit environments
14
14
 
15
- Inject the provider names above only into the environment used to launch the gateway with Compose. The weekly audit service instead reads `~/.config/ai-orchestrator/weekly-audit.env`, outside the installation and mode 600, containing only `GATEWAY_MASTER_KEY`. Do not point that service at the gateway's provider-key file. Existing installations must update and reload their copied systemd unit when adopting this template.
15
+ Inject the provider names above only into the environment used to launch the gateway with Compose. The weekly audit service instead reads `~/.config/ai-orchestrator/weekly-audit.env`, outside the installation and mode 600, containing `GATEWAY_MASTER_KEY` and optional documented runtime/location settings. Do not point that service at the gateway's provider-key file. Existing installations must update and reload their copied systemd unit when adopting this template.
16
16
 
17
- The audit script removes `GATEWAY_MASTER_KEY`, `LITELLM_MASTER_KEY`, and the gateway provider names `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GEMINI_API_KEY`, `XAI_API_KEY`, and `OPENROUTER_API_KEY` from child environments. It supplies the gateway header to curl through stdin, without a credential temp file or an argv value. Newlines anywhere in the key are rejected before collection.
17
+ The audit script builds an explicit allowed environment with shell builtins before its first external command. It supplies the gateway header to curl through stdin, without a credential temp file or an argv value. Newlines anywhere in the key are rejected before collection. All provider keys, unrelated exported names, token variables and exported shell functions are absent from child environments.
18
18
 
19
- Vendor version probes and the report worker retain `HOME`, `PATH`, stored vendor sign-ins, and unrelated authentication variables. Configure the selected lane's own sign-in in the systemd user's account, such as `hermes auth add <provider>` for Hermes. An API-only lane relying solely on one of the removed provider variables needs vendor-supported stored authentication before this scheduled job can run; gateway keys are not a substitute for that setup.
19
+ The allowed runtime names are `HOME`, `PATH`, `USER`, `LOGNAME`, `SHELL`; `LANG`, `LANGUAGE`, `TZ`; `LC_ALL`, `LC_CTYPE`, `LC_COLLATE`, `LC_MESSAGES`, `LC_MONETARY`, `LC_NUMERIC`, `LC_TIME`, `LC_ADDRESS`, `LC_IDENTIFICATION`, `LC_MEASUREMENT`, `LC_NAME`, `LC_PAPER`, `LC_TELEPHONE`; `TMPDIR`, `TMP`, `TEMP`; `XDG_CONFIG_HOME`, `XDG_DATA_HOME`, `XDG_STATE_HOME`, `XDG_CACHE_HOME`, `XDG_RUNTIME_DIR`, `DBUS_SESSION_BUS_ADDRESS`; and the Windows shell runtime names `SystemRoot`, `SYSTEMROOT`, `WINDIR`. User-service access and stored sign-ins use those same locations.
20
+
21
+ Only the selected report worker additionally receives its location override: `CODEX_HOME` for `codex`, `HERMES_HOME` for `hermes`. These names are removed from collection/version probes. `agy`, `grok` and `qwen` use stored sign-ins in the common user/configuration locations, with no additional exported authentication name. Configure the selected vendor's sign-in in the systemd user's account, such as `hermes auth add <provider>` for Hermes. API-only setups that depend on exported provider keys need vendor-supported stored authentication before this job can run. Claude Code is not a cli-run worker in this catalog, so its session-token export is removed as well.
22
+
23
+ `PROBE_SECS` and `RUNNER_SECS` remain shell-only deadline overrides. Set documented runtime and location overrides in the service's `Environment=` settings or audit-only environment file. The key and any selected worker settings stay out of command arguments and logs. There is no arbitrary variable pass-through option.
24
+
25
+ When migrating, use `--update-docs` for unchanged managed documents, review preserved edited files, apply the new script and docs, update the copied service, then reload systemd. Reconfigure any sign-in that depended on an unrelated export and manually start the service. [The jobs README](jobs/README.md) owns the invocation chain, exact environment contract, migration and verification steps. The environment boundary preserves access to files under HOME/XDG; it does not isolate those files or undo Bash startup files. Real vendor sign-in and systemd behavior remain unverified until the manual run.
20
26
 
21
27
  ## Rules
22
28
 
@@ -21,9 +21,28 @@ systemctl --user list-timers # it should be listed with a next-run time
21
21
  loginctl enable-linger "$USER" # so user timers run without a login session
22
22
  ```
23
23
 
24
- The service reads only `GATEWAY_MASTER_KEY` from `~/.config/ai-orchestrator/weekly-audit.env`, outside this folder with mode 600. Provision that audit-only file from your secrets manager and point `EnvironmentFile=` at it before installing. Keep the gateway's provider-key environment separate. Existing installs must replace or edit their copied service, then run `systemctl --user daemon-reload`; updating the source template alone does not update the installed unit. The key must be a single token matching `^[A-Za-z0-9._-]+$`; any embedded or trailing newline is refused.
24
+ The service reads `GATEWAY_MASTER_KEY` and optional documented runtime/location settings from `~/.config/ai-orchestrator/weekly-audit.env`, outside this folder with mode 600. Provision that audit-only file from your secrets manager and point `EnvironmentFile=` at it before installing. Keep the gateway's provider-key environment separate. Existing installs must replace or edit their copied service, then run `systemctl --user daemon-reload`; updating the source template alone does not update the installed unit. The key must be a single token matching `^[A-Za-z0-9._-]+$`; any embedded or trailing newline is refused.
25
25
 
26
- Sign the selected vendor CLI in under the same user before enabling the timer. The job preserves `HOME`, `PATH`, stored sign-in state and unrelated authentication variables, but removes `GATEWAY_MASTER_KEY`, `LITELLM_MASTER_KEY`, `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GEMINI_API_KEY`, `XAI_API_KEY`, and `OPENROUTER_API_KEY` from all child environments. An API-only lane dependent on a removed variable needs vendor-supported stored authentication, such as `hermes auth add <provider>`, before scheduling it. Those provider keys belong in the gateway launch environment.
26
+ Sign the selected vendor CLI in under the same user before enabling the timer. Before its first external command, the script uses shell builtins to remove every export except the runtime settings below. It also removes exported shell functions and disables inherited tracing and automatic export. Collection commands and version probes receive this same allowed environment.
27
+
28
+ | Purpose | Allowed names |
29
+ |---|---|
30
+ | User and executable lookup | `HOME`, `PATH`, `USER`, `LOGNAME`, `SHELL` |
31
+ | Locale and time | `LANG`, `LANGUAGE`, `TZ`, `LC_ALL`, `LC_CTYPE`, `LC_COLLATE`, `LC_MESSAGES`, `LC_MONETARY`, `LC_NUMERIC`, `LC_TIME`, `LC_ADDRESS`, `LC_IDENTIFICATION`, `LC_MEASUREMENT`, `LC_NAME`, `LC_PAPER`, `LC_TELEPHONE` |
32
+ | Temporary directories | `TMPDIR`, `TMP`, `TEMP` |
33
+ | Stored configuration and sign-ins | `XDG_CONFIG_HOME`, `XDG_DATA_HOME`, `XDG_STATE_HOME`, `XDG_CACHE_HOME` |
34
+ | User service and keyring session | `XDG_RUNTIME_DIR`, `DBUS_SESSION_BUS_ADDRESS` |
35
+ | Windows shell runtime compatibility | `SystemRoot`, `SYSTEMROOT`, `WINDIR` |
36
+
37
+ The report worker additionally receives `CODEX_HOME` when the selected worker is `codex`, or `HERMES_HOME` when it is `hermes`, if that name was set. These location overrides are absent from collection and version probes. The `agy`, `grok` and `qwen` workers use their stored sign-ins under the common user/configuration directories. No token or provider-key variable is allowed for any current worker; `CLAUDE_CODE_OAUTH_TOKEN` is also removed because Claude Code is not a supported cli-run worker in this catalog. Configure authentication with the selected vendor's sign-in flow, such as `codex login`, `grok login`, `hermes auth add <provider>`, `agy`, or Qwen's `/auth`.
38
+
39
+ `GATEWAY_MASTER_KEY` stays in a private shell variable until the stdin probe finishes and is then removed. `PROBE_SECS` and `RUNNER_SECS` override the shell's deadlines without becoming child exports. Configure the allowed names and worker location overrides through the service's `Environment=` settings or its audit-only environment file. Keep that file limited to the gateway key and the documented runtime/location settings; keep provider credentials in the gateway launch environment. The script provides no arbitrary extra-variable override.
40
+
41
+ ### Migrate an existing job
42
+
43
+ Preview your existing selection and paths with `--update-docs --dry-run`, then apply the update. Unchanged managed runtime files upgrade automatically; `--update-docs` refreshes unchanged documents. Review any edited files the installer preserves and merge the environment change into those copies, or use `--upgrade-runtime` when you intend to replace edited runtime files. Apply the new script and both environment documents to your installation, then replace or edit the copied service and run `systemctl --user daemon-reload`. Move sign-in setups that depend on other exported names to the vendor's stored sign-in flow under the service user. Reapply only the documented runtime/location overrides, confirm the selected CLI is on the service's explicit PATH, then perform the manual service check below. Changing `AUDIT_LANE` also requires matching that worker's runner configuration and flags.
44
+
45
+ This boundary limits child environment inheritance. Files under `HOME` and XDG directories remain readable, including stored sign-ins, and Bash startup files run before the script can filter exports. Live vendor authentication must be verified in the service user's account.
27
46
 
28
47
  ## Invocation, dependencies, reads and writes
29
48
 
@@ -59,7 +78,7 @@ For a manual check, start `systemctl --user start weekly-audit.service`, then ru
59
78
  - **Previous report preserved:** output goes to a temp file and is renamed over `audit-<date>.md` only on a clean, non-empty run. A failed run leaves `failed-audit-<stamp>-rc<N>.md` beside it and the last good report untouched.
60
79
  - **Boundary:** the lane runs with the strongest restriction it offers ({{AUDIT_LANE_BOUNDARY_NOTE}}). The brief's denied-actions list is an instruction, not an enforcement, for lanes without a sandbox flag.
61
80
  - **Honest unknowns:** a probe that times out writes an `UNVERIFIED` line, which the brief tells the lane to treat as unknown, never clean.
62
- - **Credential separation:** the gateway bearer is never written to a temp file or passed on argv, and gateway/provider keys are absent from version probes and the report worker's environment. Stored vendor sign-ins remain available.
81
+ - **Credential separation:** the gateway bearer is never written to a temp file or passed on argv. Only the named runtime environment reaches probes, and only the selected worker receives its documented location override. Stored vendor sign-ins remain available.
63
82
 
64
83
  A timer that has never been seen to fire is not known to work. Run `systemctl --user start weekly-audit.service` once by hand and read the journal before trusting the schedule.
65
84
 
@@ -8,7 +8,8 @@ WorkingDirectory={{INSTALL_DIR_SYSTEMD}}
8
8
  # Add absolute Node and vendor CLI directories if they live outside these defaults.
9
9
  # systemd does not expand shell variables in Environment=.
10
10
  Environment="PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
11
- # Audit-only file OUTSIDE this repo, mode 600: GATEWAY_MASTER_KEY only.
11
+ # Audit-only file OUTSIDE this repo, mode 600: GATEWAY_MASTER_KEY and optional
12
+ # runtime/location settings from jobs/README.md. Other exports are removed.
12
13
  # Provider credentials stay in the separate gateway/Compose environment.
13
14
  # Vendor CLIs use this user's stored sign-in state. Edit the path if needed.
14
15
  EnvironmentFile=%h/.config/ai-orchestrator/weekly-audit.env
@@ -8,18 +8,32 @@
8
8
  # renamed into place only on a clean exit; failed output is kept beside it for diagnosis
9
9
  # - the lane runs with the strongest boundary it offers ({{AUDIT_LANE_BOUNDARY_NOTE}})
10
10
  # - any nonzero rc from cli-run (10 to 18) means no report was produced; the timer's journal shows it
11
+ # Disable inherited tracing and automatic export before handling credentials.
12
+ set +a +x +v
11
13
  set -uo pipefail
12
14
  INSTALL_DIR={{INSTALL_DIR_SH}}
13
15
  AUDIT_LANE="{{AUDIT_LANE}}"
14
16
  AUDIT_LANE_FLAGS="{{AUDIT_LANE_FLAGS}}"
15
17
  PROBE_SECS="${PROBE_SECS:-10}" # per collection probe
16
18
  RUNNER_SECS="${RUNNER_SECS:-600}" # the model call; TimeoutStartSec in the unit covers the whole job
17
- # Keep the probe credential in this shell only. Gateway provider keys belong to
18
- # Compose, not to collection tools or the scheduled vendor CLI. Stored vendor
19
- # sign-ins, HOME, PATH, and unrelated authentication variables remain available.
20
- KEY="${GATEWAY_MASTER_KEY:-}"
19
+ # Builtins only until the exported environment is reduced to these named runtime
20
+ # settings. De-exporting also handles shell-owned readonly variables. Keep home
21
+ # overrides private until the selected worker runs; stored sign-ins stay on disk.
21
22
  export -n KEY
22
- unset GATEWAY_MASTER_KEY LITELLM_MASTER_KEY ANTHROPIC_API_KEY OPENAI_API_KEY GEMINI_API_KEY XAI_API_KEY OPENROUTER_API_KEY
23
+ KEY="${GATEWAY_MASTER_KEY:-}"
24
+ while IFS= read -r name; do
25
+ case "$name" in
26
+ HOME|PATH|USER|LOGNAME|SHELL|LANG|LANGUAGE|TZ|TMPDIR|TMP|TEMP|\
27
+ LC_ALL|LC_CTYPE|LC_COLLATE|LC_MESSAGES|LC_MONETARY|LC_NUMERIC|LC_TIME|\
28
+ LC_ADDRESS|LC_IDENTIFICATION|LC_MEASUREMENT|LC_NAME|LC_PAPER|LC_TELEPHONE|\
29
+ XDG_CONFIG_HOME|XDG_DATA_HOME|XDG_STATE_HOME|XDG_CACHE_HOME|\
30
+ XDG_RUNTIME_DIR|DBUS_SESSION_BUS_ADDRESS|SystemRoot|SYSTEMROOT|WINDIR) ;;
31
+ *) export -n "$name" ;;
32
+ esac
33
+ done < <(compgen -e)
34
+ # Exported shell functions are also outside the child environment contract.
35
+ while IFS= read -r name; do export -nf "$name"; done < <(compgen -A function)
36
+ unset name GATEWAY_MASTER_KEY LITELLM_MASTER_KEY ANTHROPIC_API_KEY OPENAI_API_KEY GEMINI_API_KEY XAI_API_KEY OPENROUTER_API_KEY
23
37
 
24
38
  # A shell pattern checks the whole value, including embedded/trailing newlines.
25
39
  # Line-oriented grep accepts a valid line even when another line is malformed.
@@ -118,8 +132,16 @@ BRIEF="reports/audit-brief-$STAMP.md"
118
132
  # Write to a temp file; the dated report is replaced only by a clean, non-empty run.
119
133
  FINAL="reports/audit-$DATE.md"
120
134
  TMP="$(mktemp "reports/.audit-$STAMP-XXXXXX")"
121
- # shellcheck disable=SC2086
122
- node bin/cli-run.mjs "$AUDIT_LANE" $AUDIT_LANE_FLAGS --brief "$BRIEF" --timeout "$RUNNER_SECS" --quiet < /dev/null > "$TMP"
135
+ (
136
+ # These are configuration locations, not provider keys. All current scheduled
137
+ # lanes use stored sign-ins; no token variable is part of this contract.
138
+ case "$AUDIT_LANE" in
139
+ codex) if [ "${CODEX_HOME+x}" ]; then export CODEX_HOME; fi ;;
140
+ hermes) if [ "${HERMES_HOME+x}" ]; then export HERMES_HOME; fi ;;
141
+ esac
142
+ # shellcheck disable=SC2086
143
+ exec node bin/cli-run.mjs "$AUDIT_LANE" $AUDIT_LANE_FLAGS --brief "$BRIEF" --timeout "$RUNNER_SECS" --quiet
144
+ ) < /dev/null > "$TMP"
123
145
  rc=$?
124
146
  if [ "$rc" -eq 0 ] && [ -s "$TMP" ]; then
125
147
  mv -f "$TMP" "$FINAL"
@@ -342,6 +342,40 @@ function runSummary(args) {
342
342
  const starts = records.filter((r) => r.event === 'start').length;
343
343
  const noMatchingStart = Math.max(0, dispatches.length - starts);
344
344
 
345
+ // Reconciliation: does the marker's claim match what the session actually did?
346
+ // Both halves are already in this log -- the Stop marker says which lane the turn
347
+ // DECLARED, and the PreToolUse event says what was actually DISPATCHED -- and until
348
+ // now nothing compared them. A declared lane is a claim; a dispatch is the act.
349
+ // Grouped by session because a marker is written once per turn while a dispatch can
350
+ // land on any turn of the same session, so turn-level pairing would report drift
351
+ // that is only ordering.
352
+ // PRESENCE, not identity, and the name says so: this asks whether a session that
353
+ // NAMED a lane went on to dispatch at all, not whether it dispatched the lane it named. A
354
+ // session that named builder and dispatched reader matches here. Lane identity is
355
+ // not recoverable from this log, because the marker names a lane from the user's own
356
+ // ROUTING.md while a dispatch names a subagent_type, and the two vocabularies do not
357
+ // have to line up.
358
+ const sessions = new Map();
359
+ for (const r of records) {
360
+ if (!r.session_id) continue;
361
+ if (!sessions.has(r.session_id)) sessions.set(r.session_id, { declaredOff: false, dispatched: false, sawLane: false });
362
+ const acc = sessions.get(r.session_id);
363
+ if (r.event === 'dispatch') acc.dispatched = true;
364
+ if (r.event === 'route' && Array.isArray(r.lane)) {
365
+ // 'missing' is what the hook writes when a turn carried NO marker, so it is the
366
+ // absence of a claim, not a claim of inline. Counting it here would double-report
367
+ // the gap the coverage line above already reports.
368
+ const named = r.lane.filter((l) => l !== 'missing');
369
+ if (named.length > 0) acc.sawLane = true;
370
+ if (named.some((l) => !inlineNames.has(String(l).trim().toLowerCase()))) acc.declaredOff = true;
371
+ }
372
+ }
373
+ const reconcilable = [...sessions.values()].filter((a) => a.sawLane);
374
+ const claimedNotDone = reconcilable.filter((a) => a.declaredOff && !a.dispatched).length;
375
+ const doneNotClaimed = reconcilable.filter((a) => !a.declaredOff && a.dispatched).length;
376
+ const agreed = reconcilable.length - claimedNotDone - doneNotClaimed;
377
+ const agreedPct = reconcilable.length > 0 ? (agreed / reconcilable.length) * 100 : null;
378
+
345
379
  const ends = records.filter((r) => r.event === 'end' && r.agent_type && typeof r.duration_s === 'number');
346
380
  const durationsByType = new Map();
347
381
  for (const r of ends) {
@@ -366,6 +400,15 @@ function runSummary(args) {
366
400
  if (dispatchCounts.size === 0) lines.push(' (none)');
367
401
  for (const [type, count] of [...dispatchCounts.entries()].sort((a, b) => b[1] - a[1])) lines.push(' ' + type + ': ' + count);
368
402
  lines.push('dispatches with no matching start: ' + noMatchingStart + ' (a hook or guard blocked them before launch)');
403
+ lines.push(
404
+ 'delegation claimed vs observed: ' +
405
+ (agreedPct === null ? 'no session named a lane yet' : formatNumber(agreedPct) + '% of sessions match') +
406
+ ' (' + agreed + '/' + reconcilable.length + ' sessions that named a lane)'
407
+ );
408
+ lines.push(' named a lane, no dispatch in this window: ' + claimedNotDone);
409
+ lines.push(' dispatched, but every named lane was inline: ' + doneNotClaimed);
410
+ lines.push(' presence only: a session that named one lane, then dispatched a different one, still counts as matching,');
411
+ lines.push(' and --since or a rotated log can split a session so one half lands in the counts above.');
369
412
  lines.push('duration by agent_type (mean / max, seconds):');
370
413
  if (durationsByType.size === 0) lines.push(' (none)');
371
414
  for (const [type, durs] of durationsByType) {