model-orchestrator 1.0.1 → 1.0.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +30 -1
- package/README.md +1 -1
- package/docs/README.md +1 -0
- package/docs/catalog-advisories.md +114 -0
- package/docs/catalog-advisory-exceptions.json +1 -0
- package/docs/part-2-intermediate.md +1 -1
- package/docs/security-review-history.md +1 -0
- package/llms.txt +1 -0
- package/package.json +1 -1
- package/src/README.md +1 -1
- package/src/bounded-file.js +5 -2
- package/src/catalog.js +3 -0
- package/templates/advanced/vm/ENVIRONMENT.md +9 -3
- package/templates/advanced/vm/jobs/README.md +22 -3
- package/templates/advanced/vm/jobs/weekly-audit.service +2 -1
- package/templates/advanced/vm/jobs/weekly-audit.sh +29 -7
- package/templates/agents/snippets/route-metrics.mjs +43 -0
- package/templates/common/protocols/acceptance-checks.md +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,33 @@ All notable changes to this project are documented here. The format follows [Kee
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [1.0.3] - 2026-09-28
|
|
8
|
+
|
|
9
|
+
### Security
|
|
10
|
+
|
|
11
|
+
- Weekly audit jobs use an explicit child environment allowlist, keep selected worker configuration overrides out of probes, and preserve stored sign-in paths. See the generated jobs README for migration (#45).
|
|
12
|
+
- Catalog package and container pins get advisory checks with explicit coverage, dated machine-readable results and expiring advisory-specific exceptions. See [catalog advisory checks](docs/catalog-advisories.md) for operation and limitations (#46).
|
|
13
|
+
|
|
14
|
+
### Added
|
|
15
|
+
|
|
16
|
+
- `aunx route-metrics --summary` compares the route your agent named in each reply with the delegation that actually followed, from data the log already records: the share of sessions that match, sessions that named a route and dispatched nothing, and sessions that dispatched without naming one.
|
|
17
|
+
|
|
18
|
+
### Fixed
|
|
19
|
+
|
|
20
|
+
- The product website navigation follows the section being read, keeps the active sidebar link visible, and labels its portfolio breadcrumb Home.
|
|
21
|
+
|
|
22
|
+
## [1.0.2] - 2026-09-28
|
|
23
|
+
|
|
24
|
+
### Fixed
|
|
25
|
+
|
|
26
|
+
- A refused symlinked or changing file (such as `MANIFEST.json`) now names its path and the fix: replace the link with the file it points to. Since 1.0.1, rerunning the installer over a symlinked manifest stops instead of treating it as a fresh install.
|
|
27
|
+
- A path shown in a refusal escapes control characters, so a crafted file name cannot write terminal escape sequences; a directory or other non-file gets "expected a regular file" without symlink advice.
|
|
28
|
+
- The acceptance-check protocol states the timeout's reach: it stops the command's process group, and a descendant that starts its own session (for example with `setsid`) keeps running.
|
|
29
|
+
|
|
30
|
+
### Added
|
|
31
|
+
|
|
32
|
+
- Product website at `model-orchestrator.aunysillyme.dev`, with repository-generated documentation, changelog, and the model-orchestrator trailer. Main-branch pushes rebuild the site through Vercel; the published version is checked against npm.
|
|
33
|
+
|
|
7
34
|
## [1.0.1] - 2026-09-27
|
|
8
35
|
|
|
9
36
|
### Security
|
|
@@ -513,7 +540,9 @@ First release.
|
|
|
513
540
|
- Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
|
|
514
541
|
- Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
|
|
515
542
|
|
|
516
|
-
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.
|
|
543
|
+
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.3...HEAD
|
|
544
|
+
[1.0.3]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.2...v1.0.3
|
|
545
|
+
[1.0.2]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.1...v1.0.2
|
|
517
546
|
[1.0.1]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.0...v1.0.1
|
|
518
547
|
[1.0.0]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.35...v1.0.0
|
|
519
548
|
[0.1.35]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.34...v0.1.35
|
package/README.md
CHANGED
|
@@ -209,7 +209,7 @@ Run `npx model-orchestrator --uninstall --dir ./ai-orchestrator --project .` (ad
|
|
|
209
209
|
|
|
210
210
|
Node 18 or newer, with zero runtime dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22), including lane execution end to end through `cli-run` against a fake CLI installed the same way npm installs a real one (a `.cmd` shim). `cli-run` never runs a lane through `cmd.exe`: it resolves the shim to the Node script underneath and spawns Node directly, so a prompt reaching a real lane never passes through a Windows shell. A `.cmd` or `.bat` lane that cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, rather than run through `cmd.exe`: a batch file re-reads its arguments after `cmd.exe` has parsed them once, and no escaping fully contains a prompt through both passes. Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested on Windows too, including SIGTERM/SIGINT to the wrapper (Windows has no OS-level signals: both terminate it unconditionally, verified there rather than treated the same as POSIX). The Windows skip list covers POSIX behavior, with each skip pinned by `test/prose.test.js`: `statSync().mode`'s executable bit (NTFS has none, so that one assertion is conditional inside a test that otherwise runs everywhere); a lane dying mid-run from a real POSIX signal (a real Windows lane cannot die "by signal"); running `weekly-audit.sh`'s watchdog functions for real under Git Bash's job control, both the end-to-end run and the `bounded()` timeout check (the script itself only ever runs on the Ubuntu box it targets); and a `mkfifo` FIFO at the rules path, the one case that proves `route-gate.mjs` cannot HANG on a non-regular file, since Windows has no `mkfifo` to build one (the guard behind it is covered on every OS by a directory at the same path); and an untracked `mkfifo` FIFO in the repository `cli-run --audit` sizes, the case that proves `--effort auto` never opens a non-regular file (the symlink half of that test runs on every OS). `test/prose.test.js` counts every `skip:` in the suite and requires this list to document each one.
|
|
211
211
|
|
|
212
|
-
Additional security regressions skip Windows for the project-hook symlink and manifest FIFO fixtures (symlink privileges and POSIX special files), the POSIX shell descendant timeout fixture (the argv equivalent still runs on Windows), and four weekly credential and report lifecycle runtime checks (the Ubuntu watchdog requires POSIX process-tree semantics). Their configuration and generated syntax remain covered on every platform.
|
|
212
|
+
Additional security regressions skip Windows for the project-hook symlink and manifest FIFO fixtures and the symlinked-manifest refusal message (symlink privileges and POSIX special files), the POSIX shell descendant timeout fixture (the argv equivalent still runs on Windows), and four weekly credential and report lifecycle runtime checks (the Ubuntu watchdog requires POSIX process-tree semantics). Their configuration and generated syntax remain covered on every platform.
|
|
213
213
|
|
|
214
214
|
**Privacy.** The installer sends no telemetry and makes no network call of its own once it is running. Two things around that are worth being exact about:
|
|
215
215
|
|
package/docs/README.md
CHANGED
|
@@ -12,6 +12,7 @@ Use these pages for setup, operation and evidence. The [front page](../README.md
|
|
|
12
12
|
| [Intermediate](part-2-intermediate.md) | Delegation across several AI CLIs |
|
|
13
13
|
| [Advanced](part-3-advanced.md) | Gateway templates and scheduled work on a Linux host |
|
|
14
14
|
| [Catalog](catalog.md) | Supported tools, capability facts, unverified values, installation and sign-in notes |
|
|
15
|
+
| [Catalog advisory checks](catalog-advisories.md) | Package and container coverage, report statuses, exceptions and manual verification |
|
|
15
16
|
| [Security review history](security-review-history.md) | Review rounds, reproduced findings, fixes and regression tests |
|
|
16
17
|
| [Proof](../proof/README.md) | Dated measurements, methods, sample sizes and reproduction scripts |
|
|
17
18
|
| [Commands](../bin/README.md) | `aunx` subcommands and lane-runner exit codes |
|
|
@@ -0,0 +1,114 @@
|
|
|
1
|
+
# Catalog advisory checks
|
|
2
|
+
|
|
3
|
+
## What and why
|
|
4
|
+
|
|
5
|
+
The repository checks exact executable package and container pins in `src/catalog.js`. OSV API v1 answers direct npm and PyPI package/version queries. Trivy examines OS and language packages in the selected container image. This check gives dated advisory evidence for its stated coverage. A clean result does not prove safety.
|
|
6
|
+
|
|
7
|
+
The owning files are `scripts/catalog-advisory-inventory.mjs`, `scripts/check-catalog-advisories.mjs`, `scripts/catalog-advisory-http.mjs`, `scripts/catalog-advisory-containers.mjs`, `docs/catalog-advisory-exceptions.json` and `.github/workflows/catalog-advisories.yml`. This document describes the complete operating procedure.
|
|
8
|
+
|
|
9
|
+
## Trigger
|
|
10
|
+
|
|
11
|
+
The `catalog advisories` workflow runs on every push to main, on pull requests changing the catalog, checker, exceptions, tests or workflow, and by manual dispatch. It has read-only repository permissions. There is no periodic schedule; a maintainer can dispatch a fresh check when advisory data changes. Existing Dependabot automation continues to update GitHub Actions separately.
|
|
12
|
+
|
|
13
|
+
## Invocation chain
|
|
14
|
+
|
|
15
|
+
1. The workflow checks out the source and selects Node 22.
|
|
16
|
+
2. The inventory imports `AIS`, `TOOLS`, `IMAGES` and `PROVIDERS` directly from the catalog. npm names come from `install.npm`; companion ecosystem, package and extras come from their `advisory` metadata. Every version comes from the existing pin field.
|
|
17
|
+
3. The package adapter posts exact package/version queries to `https://api.osv.dev/v1/querybatch`, then retrieves each returned advisory from `/v1/vulns/<id>` to retain ranges and fixed versions.
|
|
18
|
+
4. The container adapter resolves target tags to a platform-specific SHA-256 manifest digest, checks the manifest and image configuration, and records the platform. Docker Hub and GHCR public images are supported; another registry is unknown until its adapter is implemented.
|
|
19
|
+
5. The adapter resolves the Trivy scanner image to a digest too, asserts its reported version, and runs that scanner against the target digest. It starts no target image or vendor service. Scanning uses registry access, with no host mounts or Docker socket passed into the scanner.
|
|
20
|
+
6. Exact, unexpired exceptions are applied. The checker writes `report.json` and `summary.md`, then exits with its result code. The workflow appends the summary to the job and uploads both files even if the scan step fails.
|
|
21
|
+
|
|
22
|
+
## Dependencies
|
|
23
|
+
|
|
24
|
+
A repository checkout, Node 18+ with built-in fetch, Docker with a running Linux-container daemon, public HTTPS registry access, OSV, and the Trivy vulnerability databases are needed for a complete local check. CI supplies Node 22 and Docker. Package checks need only Node and OSV. A missing scanner, daemon or database produces unknown container results while retaining available package results.
|
|
25
|
+
|
|
26
|
+
Trivy is version-pinned in `scripts/catalog-advisory-containers.mjs`. Each run resolves that scanner tag to an immutable digest and records it in each container result. Runtime version verification rejects a binary that reports a different version. The version probe has a 60-second deadline; image scans have Trivy's 10-minute deadline and a 12-minute client deadline, followed by bounded container cleanup. The scanner tag itself can move between runs; the recorded digest is the reproducible scanner identity for that run. For exact reproduction, fetch the scanner and target digests from the report before investigating.
|
|
27
|
+
|
|
28
|
+
Pin provenance, checked 2026-09-28 against upstream: `actions/checkout` `3d3c42e` is tag v7.0.1, `actions/setup-node` `8207627` is tag v7.0.0, and `actions/upload-artifact` `043fb46` is tag v7.0.1 in each action owner's repository. Trivy 0.74.0 is the latest release in Aqua Security's official repository (2026-08-14) and its public image tag resolves to `sha256:62b1e65e8869bc4b4c6aa4fa2b21595256c7c2f6018a9d9ad61caf87187c1969`. The first successful main-branch artifact is the check that the JSON parser matches this Trivy version.
|
|
29
|
+
|
|
30
|
+
To update the scanner, verify the desired release in Aqua Security's official Trivy repository, resolve its public registry tag to a digest, and inspect the pinned version's `image --help`, JSON report schema and database requirements. Change `TRIVY` in `scripts/catalog-advisory-containers.mjs`, run the offline fixtures, then run the real checker and inspect its recorded scanner digest and version. This builder could not compare the JSON parser with current official Trivy source; that compatibility is **UNVERIFIED** until a real scan passes. To update an action, compare its full commit SHA against the release tag in the action owner's repository and update the SHA and version comment together. Run the workflow and download the artifact before accepting either upgrade.
|
|
31
|
+
|
|
32
|
+
## Reads
|
|
33
|
+
|
|
34
|
+
- `src/catalog.js`: package names, ecosystems, extras, exact versions, image references and excluded inputs. A new companion without advisory metadata becomes unknown.
|
|
35
|
+
- `docs/catalog-advisory-exceptions.json`: the reviewed exception array, initially empty.
|
|
36
|
+
- OSV: current direct-package advisory matches and their details. No project files are sent.
|
|
37
|
+
- Public container registries: manifests, configurations, layers and anonymous pull authorization. Anonymous registry tokens stay in memory. User sign-ins are not read by the resolver; the Docker client receives an empty temporary configuration directory.
|
|
38
|
+
- Trivy databases: current OS and language advisories. All severities and unfixed findings are retained.
|
|
39
|
+
|
|
40
|
+
Direct package checks do not resolve dependency trees. Each such input has an explicit transitive-dependency exclusion. `codecalc[full]` also has an optional-extras exclusion: querying the base distribution does not cover dependencies selected by `full`. Vendor installer scripts, unpinned downloads/package-manager inputs, chat apps and hosted model services are named exclusions. Compatibility snapshots such as `builtAgainst` are not installation pins. Image coverage is the OS and language packages Trivy recognizes; unsupported OS images and empty package results are unknown.
|
|
41
|
+
|
|
42
|
+
## Writes
|
|
43
|
+
|
|
44
|
+
The default output directory is `.release-work/catalog-advisories/`. The checker writes:
|
|
45
|
+
|
|
46
|
+
- `report.json`: source commit, timestamp, scanner versions, catalog input, ecosystem, exact version, image/scanner digest, platform, coverage, status, advisory IDs, ranges, fixed versions and matched exception expiry.
|
|
47
|
+
- `summary.md`: the same result in a concise human-readable form, with explicit exclusions.
|
|
48
|
+
|
|
49
|
+
CI retains these as `catalog-advisories-<commit>-<attempt>` for 30 days and shows the summary in the job. An early workflow failure creates an unknown fallback report instead of presenting a missing report as success. The checker creates and removes an empty temporary Docker configuration directory. Scanner cache data lives under `/tmp/trivy` in its disposable writable container layer; `--rm` removes that layer. A disk-backed layer accommodates large image scans without a small RAM-backed cache cap. The scanner runs as a nonroot user with capabilities dropped, no new privileges, and no host filesystem or socket mounts. Live disk usage is **UNVERIFIED** here. Catalog pins and installed dependencies are never changed.
|
|
50
|
+
|
|
51
|
+
## The closed loop
|
|
52
|
+
|
|
53
|
+
GitHub Actions is the watcher. A maintainer reviews every affected or unknown result before merging or releasing. No automatic issue, upgrade, exception or release is created by this workflow. Branch-protection configuration is repository administration: **UNVERIFIED** here. To make a green check mandatory, require the `catalog advisories / scan` check in the repository's branch rules.
|
|
54
|
+
|
|
55
|
+
| Status | Meaning | Exit behavior |
|
|
56
|
+
|---|---|---|
|
|
57
|
+
| clean | Complete adapter response with no advisory match in stated coverage | 0 if every result is clean or excepted |
|
|
58
|
+
| affected | At least one advisory match lacks a valid exception | 1 unless another result is unknown |
|
|
59
|
+
| unknown | Missing, unavailable, unsupported or malformed evidence, or an invalid exception policy | 2 |
|
|
60
|
+
| excepted | Every advisory match has an exact, reviewed, unexpired exception | 0 if every result is clean or excepted |
|
|
61
|
+
|
|
62
|
+
Any unknown takes precedence over affected for the exit code. Known advisory IDs are still retained when a detail lookup fails. Empty results never pass. Read the exclusions alongside the status; exclusions are never counted as clean pins.
|
|
63
|
+
|
|
64
|
+
## Failure modes
|
|
65
|
+
|
|
66
|
+
- **OSV timeout, HTTP error or malformed response:** package results become unknown. A valid empty query result (`{}`) means no returned match; a missing result array or missing position is unknown.
|
|
67
|
+
- **Registry failure, bad image reference, missing platform or digest mismatch:** the image is unknown. The checker retains the other results.
|
|
68
|
+
- **Docker, scanner or database failure:** the image is unknown. No exception suppresses unknown status. Check network access and the daemon, then rerun the same source.
|
|
69
|
+
- **Incomplete scanner JSON, mismatched digest/platform or unsupported image OS:** the image is unknown even if the process exited successfully.
|
|
70
|
+
- **Advisory outage:** keep the failing result, retry after recovery, and inspect the provider's status through its normal support channel. An outage is not an exception for a package.
|
|
71
|
+
- **Expired or malformed exception:** expired entries stop matching at 00:00 UTC on their expiry date; malformed policy blocks success. Expired entries are retained in `expiredExceptions` for review.
|
|
72
|
+
- **Interrupted job:** `always()` upload and summary steps preserve available evidence. A runner outage can still prevent artifact upload; verify artifact presence in the run itself.
|
|
73
|
+
|
|
74
|
+
## Exception review
|
|
75
|
+
|
|
76
|
+
The exception file is a JSON array. Every entry requires `ecosystem`, `package`, `version`, `advisory`, `rationale` and `expires` (`YYYY-MM-DD`). Wildcards and unknown fields are rejected. Match the exact advisory ID emitted by the adapter, not an alias. A package exception optionally names its exact catalog `source` in `target`. A container exception requires `target` equal to the exact catalog image reference and identifies the vulnerable component package and installed version, not every component in the image.
|
|
77
|
+
|
|
78
|
+
For example, a metadata-only example is:
|
|
79
|
+
|
|
80
|
+
```json
|
|
81
|
+
[
|
|
82
|
+
{
|
|
83
|
+
"ecosystem": "npm",
|
|
84
|
+
"package": "fixture-package",
|
|
85
|
+
"version": "1.0.0",
|
|
86
|
+
"advisory": "TEST-2026-0001",
|
|
87
|
+
"rationale": "Example only: replace with a reviewed reason and tracking reference.",
|
|
88
|
+
"expires": "2026-10-01"
|
|
89
|
+
}
|
|
90
|
+
]
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
A maintainer reviews the affected range, fixed versions, applicability, expiry and rationale in a normal pull request. Use a short explicit expiry and a tracking reference in the rationale. Never add an exception merely because the database or scanner is unavailable. Review expired entries and remove those whose pin or advisory is no longer relevant. A matched exception remains visible as excepted with its expiry; it is never relabeled clean. Fixing a package requires a separate deliberate catalog-pin change and the usual installer validation.
|
|
94
|
+
|
|
95
|
+
## Run and verify by hand
|
|
96
|
+
|
|
97
|
+
From the repository checkout:
|
|
98
|
+
|
|
99
|
+
```bash
|
|
100
|
+
node --test test/catalog-advisories.test.js
|
|
101
|
+
node scripts/check-catalog-advisories.mjs --report-dir .release-work/catalog-advisories
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
The fixture tests use metadata only and make no network request. The second command makes real advisory and registry requests and may start the pinned scanner container. Use `--platform linux/arm64` to inspect another declared image platform; the default is `linux/amd64`. Use `--exceptions <path>` for a reviewed alternate policy and retain that file with the evidence.
|
|
105
|
+
|
|
106
|
+
1. Read both output files and check `sourceCommit`, `generatedAt`, every catalog source, scanner identity, platform and exclusions.
|
|
107
|
+
2. For a package finding, retrieve its recorded OSV detail URL and compare the affected package, ranges and fixed versions. The fixture's affected result must exit 1; an unavailable-data result must exit 2.
|
|
108
|
+
3. For an image finding, pull the recorded scanner digest and run its `image --scanners vuln --list-all-pkgs --format json --platform <platform> --image-src remote <image>@<digest>` command. Compare the component name, installed version, advisory ID and database evidence. A later database can change matches while the image stays identical.
|
|
109
|
+
4. Inspect the GitHub run on main and download its report artifact. A local report is not proof of main-branch retention. Re-run the workflow after an outage and confirm a new timestamp.
|
|
110
|
+
5. Run `npm test` before releasing. Existing OIDC publication and installer flags are unchanged by this check.
|
|
111
|
+
|
|
112
|
+
## Source of truth
|
|
113
|
+
|
|
114
|
+
Catalog pins and companion metadata: `src/catalog.js`. Exception decisions: `docs/catalog-advisory-exceptions.json` and its reviewed change. Operational behavior: this document and the checker modules. Dated evidence: the uploaded GitHub run artifact. Upstream evidence: each report's OSV URL and Trivy data-source URL. A report describes the source commit, selected platform and database responses at its recorded time; it is not a permanent certificate.
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
[]
|
|
@@ -50,7 +50,7 @@ Optional companions can help: codecalc for execution and calculations, obsidian-
|
|
|
50
50
|
|
|
51
51
|
## Measure your own routing
|
|
52
52
|
|
|
53
|
-
`aunx route-metrics --summary` reads your local Claude Code routing log. It reports where work went, route-marker coverage and
|
|
53
|
+
`aunx route-metrics --summary` reads your local Claude Code routing log. It reports where work went, route-marker coverage, subagent durations, and whether the route each reply named matches the delegation that followed. Your own measurements are the basis for changing assignments and checking whether the rules are being followed.
|
|
54
54
|
|
|
55
55
|
## What the installer gives you at this level
|
|
56
56
|
|
|
@@ -6,6 +6,7 @@ This page summarizes completed reviews recorded in the [changelog](../CHANGELOG.
|
|
|
6
6
|
|
|
7
7
|
| Release | Review recorded | What the review found | What changed and where it is checked |
|
|
8
8
|
|---|---|---|---|
|
|
9
|
+
| 1.0.2 | Independent review of 1.0.1 by a second model family, then an audit of its fix | A detached check descendant outlived the timeout while the notes said descendants stop; a symlinked-manifest refusal named no file or fix; refusal paths could carry terminal escapes; a directory got symlink advice | The protocol states the process-group limit; refusals name the path, escape control characters and give advice by file type; `test/security-messages.test.js` |
|
|
9
10
|
| 1.0.1 | Codex Security scan and regression-backed remediation | A metrics summary executed project code; manifest reads and check descendants needed bounds; weekly jobs inherited gateway credentials | Packaged metrics dispatch, bounded file readers, check process-tree cleanup, stdin-only gateway probe and separate audit environment; `test/security-cli.test.js`, `test/security-vm.test.js`, `test/security-pins.test.js` |
|
|
10
11
|
| 0.1.0 | Initial review and follow-up round, with a second model-family review | Paths could escape the write roots, partial writes could remain, malformed arguments could proceed, and empty results could appear successful | Containment preflight, exclusive writes with rollback, strict flag parsing and vendor-specific result checks; `test/install.test.js`, `test/cli.test.js`, `test/judges.test.js` |
|
|
11
12
|
| 0.1.1 | Follow-up on issues #1 through #10 | Child processes could survive timeouts; a failed scheduled run could replace a good report; logs could contain provider text | Process-group cleanup, temporary report plus rename, bounded probes and fixed log codes; `test/cli.test.js`, `test/install.test.js` |
|
package/llms.txt
CHANGED
|
@@ -29,6 +29,7 @@ Pick a model proxy (LiteLLM, Portkey, OpenRouter, claude-code-router) for per-re
|
|
|
29
29
|
- [Claude Code plugin](https://github.com/aunysillyme/model-orchestrator/blob/main/plugin/README.md): installing read-only routing hooks and subagents through the plugin marketplace
|
|
30
30
|
- [Proof](https://github.com/aunysillyme/model-orchestrator/blob/main/proof/README.md): dated measurements, methods, sample sizes, expiry and reproduction scripts
|
|
31
31
|
- [Security review history](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/security-review-history.md): completed review rounds, incident summaries and regression evidence
|
|
32
|
+
- [Catalog advisory checks](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/catalog-advisories.md): package and container inventory, coverage, report statuses, exceptions and manual verification
|
|
32
33
|
- [Changelog](https://github.com/aunysillyme/model-orchestrator/blob/main/CHANGELOG.md): release changes and upgrade notes
|
|
33
34
|
- [Agent instructions](https://github.com/aunysillyme/model-orchestrator/blob/main/AGENTS.md): headless setup and contributor checks
|
|
34
35
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "model-orchestrator",
|
|
3
|
-
"version": "1.0.
|
|
3
|
+
"version": "1.0.3",
|
|
4
4
|
"description": "Model router for AI coding agents: installs routing rules, 8 subagents, hooks and a CLI runner so your AI picks model and effort per task and saves tokens",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
package/src/README.md
CHANGED
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
| File | Job |
|
|
4
4
|
|---|---|
|
|
5
5
|
| `bounded-file.js` | shared regular-file reader for manifests and check configuration: no-follow/nonblocking open, identity checks and a fixed byte cap even if a file grows. Unsafe files are refused; callers decide how to handle missing or malformed data. |
|
|
6
|
-
| `catalog.js` | the single list of levels and AIs. Add an AI here and the prompts, docs tables, delegation matrix, gateway config and installer all pick it up. Nothing else lists AIs. |
|
|
6
|
+
| `catalog.js` | the single list of levels and AIs. Add an AI here and the prompts, docs tables, delegation matrix, gateway config and installer all pick it up. Nothing else lists AIs. Package pins, companion advisory metadata and image references also supply the repository's advisory inventory. |
|
|
7
7
|
| `roles.js` | pure role assignment from selected catalog capability facts, billing and selection order. Renders the stack table and manifest roles, and infers the main agent from its supported surfaces. Unknown facts remain unverified; review requires a known different model family and private work requires local execution. |
|
|
8
8
|
| `aunx.js` | command dispatch for briefs, context, checks, routing and runner calls. Route suggestions read manifest roles through a capped regular-file JSON reader; symlinks and malformed files are ignored. Route lookup executes no project code. A project's runner requires explicit `--dir`. |
|
|
9
9
|
| `detect.js` | PATH lookup for a binary, plus the few places vendor installers drop binaries without touching PATH. No shell-outs. |
|
package/src/bounded-file.js
CHANGED
|
@@ -2,6 +2,8 @@ import { closeSync, constants, fstatSync, lstatSync, openSync, readSync } from '
|
|
|
2
2
|
|
|
3
3
|
export const MANIFEST_BYTE_CAP = 1024 * 1024;
|
|
4
4
|
const refused = message => Object.assign(new Error(message), { code: 'UNSAFE_FILE' });
|
|
5
|
+
// A path is printed to a terminal; control characters (escape sequences) are shown escaped.
|
|
6
|
+
const shown = path => String(path).replace(/[\u0000-\u001f\u007f-\u009f]/g, c => `\\x${c.charCodeAt(0).toString(16).padStart(2, '0')}`);
|
|
5
7
|
|
|
6
8
|
// A fixed buffer also bounds files that grow after the initial size check.
|
|
7
9
|
export function readBounded(fd, maxBytes) {
|
|
@@ -21,11 +23,12 @@ export function readBounded(fd, maxBytes) {
|
|
|
21
23
|
|
|
22
24
|
export function readRegularFile(path, maxBytes) {
|
|
23
25
|
const before = lstatSync(path);
|
|
24
|
-
if (
|
|
26
|
+
if (before.isSymbolicLink()) throw refused(`${shown(path)}: expected a regular file, not a symlink; replace the link with the file it points to, then run the command again`);
|
|
27
|
+
if (!before.isFile()) throw refused(`${shown(path)}: expected a regular file; point the command at a file`);
|
|
25
28
|
const fd = openSync(path, constants.O_RDONLY | (constants.O_NOFOLLOW || 0) | (constants.O_NONBLOCK || 0));
|
|
26
29
|
try {
|
|
27
30
|
const after = fstatSync(fd);
|
|
28
|
-
if (!after.isFile() || after.dev !== before.dev || after.ino !== before.ino) throw refused(
|
|
31
|
+
if (!after.isFile() || after.dev !== before.dev || after.ino !== before.ino) throw refused(`${shown(path)}: file changed during inspection; run the command again`);
|
|
29
32
|
return readBounded(fd, maxBytes);
|
|
30
33
|
} finally { closeSync(fd); }
|
|
31
34
|
}
|
package/src/catalog.js
CHANGED
|
@@ -402,6 +402,7 @@ export const TOOLS = [
|
|
|
402
402
|
role: 'exact arithmetic, code execution in 31 languages, SMT logic checks, complexity and equivalence proofs; offline, no key, no telemetry',
|
|
403
403
|
get install() { return `uvx 'codecalc[full]==${this.pin}' setup --write`; },
|
|
404
404
|
pin: '0.5.0',
|
|
405
|
+
advisory: { ecosystem: 'PyPI', package: 'codecalc', extras: ['full'] },
|
|
405
406
|
mcpSnippets: { 'claude-code': 'mcp/mcpServers.json', codex: 'mcp/codex.config.toml', agy: 'mcp/agy.mcp_config.json', qwen: 'mcp/mcpServers.json' },
|
|
406
407
|
requires: 'uv (https://docs.astral.sh/uv/) and Python 3.10+',
|
|
407
408
|
autoClients: ['Claude Code', 'Claude Desktop', 'Cursor', 'VS Code', 'Zed'],
|
|
@@ -415,6 +416,7 @@ export const TOOLS = [
|
|
|
415
416
|
role: 'durable memory and record for your agents: hybrid retrieval (BM25 + dense + link graph), backlinks, compare-and-swap writes with a confirmation gate, folder ACLs, a poison scan on inferred writes; 163 tools, local by default',
|
|
416
417
|
get install() { return `npm install -g obsidian-tc@${this.pin} && obsidian-tc /path/to/your/vault`; },
|
|
417
418
|
pin: '1.26.0',
|
|
419
|
+
advisory: { ecosystem: 'npm', package: 'obsidian-tc' },
|
|
418
420
|
mcpSnippets: { 'claude-code': 'mcp/obsidian-tc.mcpServers.json', codex: 'mcp/obsidian-tc.codex.config.toml', agy: 'mcp/obsidian-tc.agy.mcp_config.json', qwen: 'mcp/obsidian-tc.mcpServers.json' },
|
|
419
421
|
requires: 'an Obsidian vault folder (the Obsidian app itself is only needed for live plugin bridges); Node 24+ or Bun 1.1+ (stricter than this installer); Ollama with `nomic-embed-text` for local embeddings, or a cloud embeddings key; the Local REST API plugin only for bridge tools',
|
|
420
422
|
autoClients: ['Cursor', 'VS Code'],
|
|
@@ -428,6 +430,7 @@ export const TOOLS = [
|
|
|
428
430
|
role: 'up-to-date, version-specific documentation and code examples for libraries, SDKs, APIs and CLIs, pulled into the prompt; tells the agent what the code is SUPPOSED to do. Paired with codecalc, which runs the code and proves what it actually does: docs never stand as proof, and where they disagree the run wins',
|
|
429
431
|
get install() { return `npx -y @upstash/context7-mcp@${this.pin}`; },
|
|
430
432
|
pin: '4.1.1',
|
|
433
|
+
advisory: { ecosystem: 'npm', package: '@upstash/context7-mcp' },
|
|
431
434
|
mcpSnippets: { 'claude-code': 'mcp/context7.claude-code.mcp.json', codex: 'mcp/context7.codex.config.toml', agy: 'mcp/context7.agy.mcp_config.json', qwen: 'mcp/context7.qwen.settings.json' },
|
|
432
435
|
requires: 'Node.js 18+ for the local server or the ctx7 CLI; a free CONTEXT7_API_KEY is optional, for higher rate limits (it works anonymously at the base rate)',
|
|
433
436
|
autoClients: [], // The pinned MCP server does not register itself; merge its snippets.
|
|
@@ -12,11 +12,17 @@ These are variable **names**. The values live in a secrets manager and are injec
|
|
|
12
12
|
|
|
13
13
|
## Gateway and scheduled audit environments
|
|
14
14
|
|
|
15
|
-
Inject the provider names above only into the environment used to launch the gateway with Compose. The weekly audit service instead reads `~/.config/ai-orchestrator/weekly-audit.env`, outside the installation and mode 600, containing
|
|
15
|
+
Inject the provider names above only into the environment used to launch the gateway with Compose. The weekly audit service instead reads `~/.config/ai-orchestrator/weekly-audit.env`, outside the installation and mode 600, containing `GATEWAY_MASTER_KEY` and optional documented runtime/location settings. Do not point that service at the gateway's provider-key file. Existing installations must update and reload their copied systemd unit when adopting this template.
|
|
16
16
|
|
|
17
|
-
The audit script
|
|
17
|
+
The audit script builds an explicit allowed environment with shell builtins before its first external command. It supplies the gateway header to curl through stdin, without a credential temp file or an argv value. Newlines anywhere in the key are rejected before collection. All provider keys, unrelated exported names, token variables and exported shell functions are absent from child environments.
|
|
18
18
|
|
|
19
|
-
|
|
19
|
+
The allowed runtime names are `HOME`, `PATH`, `USER`, `LOGNAME`, `SHELL`; `LANG`, `LANGUAGE`, `TZ`; `LC_ALL`, `LC_CTYPE`, `LC_COLLATE`, `LC_MESSAGES`, `LC_MONETARY`, `LC_NUMERIC`, `LC_TIME`, `LC_ADDRESS`, `LC_IDENTIFICATION`, `LC_MEASUREMENT`, `LC_NAME`, `LC_PAPER`, `LC_TELEPHONE`; `TMPDIR`, `TMP`, `TEMP`; `XDG_CONFIG_HOME`, `XDG_DATA_HOME`, `XDG_STATE_HOME`, `XDG_CACHE_HOME`, `XDG_RUNTIME_DIR`, `DBUS_SESSION_BUS_ADDRESS`; and the Windows shell runtime names `SystemRoot`, `SYSTEMROOT`, `WINDIR`. User-service access and stored sign-ins use those same locations.
|
|
20
|
+
|
|
21
|
+
Only the selected report worker additionally receives its location override: `CODEX_HOME` for `codex`, `HERMES_HOME` for `hermes`. These names are removed from collection/version probes. `agy`, `grok` and `qwen` use stored sign-ins in the common user/configuration locations, with no additional exported authentication name. Configure the selected vendor's sign-in in the systemd user's account, such as `hermes auth add <provider>` for Hermes. API-only setups that depend on exported provider keys need vendor-supported stored authentication before this job can run. Claude Code is not a cli-run worker in this catalog, so its session-token export is removed as well.
|
|
22
|
+
|
|
23
|
+
`PROBE_SECS` and `RUNNER_SECS` remain shell-only deadline overrides. Set documented runtime and location overrides in the service's `Environment=` settings or audit-only environment file. The key and any selected worker settings stay out of command arguments and logs. There is no arbitrary variable pass-through option.
|
|
24
|
+
|
|
25
|
+
When migrating, use `--update-docs` for unchanged managed documents, review preserved edited files, apply the new script and docs, update the copied service, then reload systemd. Reconfigure any sign-in that depended on an unrelated export and manually start the service. [The jobs README](jobs/README.md) owns the invocation chain, exact environment contract, migration and verification steps. The environment boundary preserves access to files under HOME/XDG; it does not isolate those files or undo Bash startup files. Real vendor sign-in and systemd behavior remain unverified until the manual run.
|
|
20
26
|
|
|
21
27
|
## Rules
|
|
22
28
|
|
|
@@ -21,9 +21,28 @@ systemctl --user list-timers # it should be listed with a next-run time
|
|
|
21
21
|
loginctl enable-linger "$USER" # so user timers run without a login session
|
|
22
22
|
```
|
|
23
23
|
|
|
24
|
-
The service reads
|
|
24
|
+
The service reads `GATEWAY_MASTER_KEY` and optional documented runtime/location settings from `~/.config/ai-orchestrator/weekly-audit.env`, outside this folder with mode 600. Provision that audit-only file from your secrets manager and point `EnvironmentFile=` at it before installing. Keep the gateway's provider-key environment separate. Existing installs must replace or edit their copied service, then run `systemctl --user daemon-reload`; updating the source template alone does not update the installed unit. The key must be a single token matching `^[A-Za-z0-9._-]+$`; any embedded or trailing newline is refused.
|
|
25
25
|
|
|
26
|
-
Sign the selected vendor CLI in under the same user before enabling the timer.
|
|
26
|
+
Sign the selected vendor CLI in under the same user before enabling the timer. Before its first external command, the script uses shell builtins to remove every export except the runtime settings below. It also removes exported shell functions and disables inherited tracing and automatic export. Collection commands and version probes receive this same allowed environment.
|
|
27
|
+
|
|
28
|
+
| Purpose | Allowed names |
|
|
29
|
+
|---|---|
|
|
30
|
+
| User and executable lookup | `HOME`, `PATH`, `USER`, `LOGNAME`, `SHELL` |
|
|
31
|
+
| Locale and time | `LANG`, `LANGUAGE`, `TZ`, `LC_ALL`, `LC_CTYPE`, `LC_COLLATE`, `LC_MESSAGES`, `LC_MONETARY`, `LC_NUMERIC`, `LC_TIME`, `LC_ADDRESS`, `LC_IDENTIFICATION`, `LC_MEASUREMENT`, `LC_NAME`, `LC_PAPER`, `LC_TELEPHONE` |
|
|
32
|
+
| Temporary directories | `TMPDIR`, `TMP`, `TEMP` |
|
|
33
|
+
| Stored configuration and sign-ins | `XDG_CONFIG_HOME`, `XDG_DATA_HOME`, `XDG_STATE_HOME`, `XDG_CACHE_HOME` |
|
|
34
|
+
| User service and keyring session | `XDG_RUNTIME_DIR`, `DBUS_SESSION_BUS_ADDRESS` |
|
|
35
|
+
| Windows shell runtime compatibility | `SystemRoot`, `SYSTEMROOT`, `WINDIR` |
|
|
36
|
+
|
|
37
|
+
The report worker additionally receives `CODEX_HOME` when the selected worker is `codex`, or `HERMES_HOME` when it is `hermes`, if that name was set. These location overrides are absent from collection and version probes. The `agy`, `grok` and `qwen` workers use their stored sign-ins under the common user/configuration directories. No token or provider-key variable is allowed for any current worker; `CLAUDE_CODE_OAUTH_TOKEN` is also removed because Claude Code is not a supported cli-run worker in this catalog. Configure authentication with the selected vendor's sign-in flow, such as `codex login`, `grok login`, `hermes auth add <provider>`, `agy`, or Qwen's `/auth`.
|
|
38
|
+
|
|
39
|
+
`GATEWAY_MASTER_KEY` stays in a private shell variable until the stdin probe finishes and is then removed. `PROBE_SECS` and `RUNNER_SECS` override the shell's deadlines without becoming child exports. Configure the allowed names and worker location overrides through the service's `Environment=` settings or its audit-only environment file. Keep that file limited to the gateway key and the documented runtime/location settings; keep provider credentials in the gateway launch environment. The script provides no arbitrary extra-variable override.
|
|
40
|
+
|
|
41
|
+
### Migrate an existing job
|
|
42
|
+
|
|
43
|
+
Preview your existing selection and paths with `--update-docs --dry-run`, then apply the update. Unchanged managed runtime files upgrade automatically; `--update-docs` refreshes unchanged documents. Review any edited files the installer preserves and merge the environment change into those copies, or use `--upgrade-runtime` when you intend to replace edited runtime files. Apply the new script and both environment documents to your installation, then replace or edit the copied service and run `systemctl --user daemon-reload`. Move sign-in setups that depend on other exported names to the vendor's stored sign-in flow under the service user. Reapply only the documented runtime/location overrides, confirm the selected CLI is on the service's explicit PATH, then perform the manual service check below. Changing `AUDIT_LANE` also requires matching that worker's runner configuration and flags.
|
|
44
|
+
|
|
45
|
+
This boundary limits child environment inheritance. Files under `HOME` and XDG directories remain readable, including stored sign-ins, and Bash startup files run before the script can filter exports. Live vendor authentication must be verified in the service user's account.
|
|
27
46
|
|
|
28
47
|
## Invocation, dependencies, reads and writes
|
|
29
48
|
|
|
@@ -59,7 +78,7 @@ For a manual check, start `systemctl --user start weekly-audit.service`, then ru
|
|
|
59
78
|
- **Previous report preserved:** output goes to a temp file and is renamed over `audit-<date>.md` only on a clean, non-empty run. A failed run leaves `failed-audit-<stamp>-rc<N>.md` beside it and the last good report untouched.
|
|
60
79
|
- **Boundary:** the lane runs with the strongest restriction it offers ({{AUDIT_LANE_BOUNDARY_NOTE}}). The brief's denied-actions list is an instruction, not an enforcement, for lanes without a sandbox flag.
|
|
61
80
|
- **Honest unknowns:** a probe that times out writes an `UNVERIFIED` line, which the brief tells the lane to treat as unknown, never clean.
|
|
62
|
-
- **Credential separation:** the gateway bearer is never written to a temp file or passed on argv
|
|
81
|
+
- **Credential separation:** the gateway bearer is never written to a temp file or passed on argv. Only the named runtime environment reaches probes, and only the selected worker receives its documented location override. Stored vendor sign-ins remain available.
|
|
63
82
|
|
|
64
83
|
A timer that has never been seen to fire is not known to work. Run `systemctl --user start weekly-audit.service` once by hand and read the journal before trusting the schedule.
|
|
65
84
|
|
|
@@ -8,7 +8,8 @@ WorkingDirectory={{INSTALL_DIR_SYSTEMD}}
|
|
|
8
8
|
# Add absolute Node and vendor CLI directories if they live outside these defaults.
|
|
9
9
|
# systemd does not expand shell variables in Environment=.
|
|
10
10
|
Environment="PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
|
|
11
|
-
# Audit-only file OUTSIDE this repo, mode 600: GATEWAY_MASTER_KEY
|
|
11
|
+
# Audit-only file OUTSIDE this repo, mode 600: GATEWAY_MASTER_KEY and optional
|
|
12
|
+
# runtime/location settings from jobs/README.md. Other exports are removed.
|
|
12
13
|
# Provider credentials stay in the separate gateway/Compose environment.
|
|
13
14
|
# Vendor CLIs use this user's stored sign-in state. Edit the path if needed.
|
|
14
15
|
EnvironmentFile=%h/.config/ai-orchestrator/weekly-audit.env
|
|
@@ -8,18 +8,32 @@
|
|
|
8
8
|
# renamed into place only on a clean exit; failed output is kept beside it for diagnosis
|
|
9
9
|
# - the lane runs with the strongest boundary it offers ({{AUDIT_LANE_BOUNDARY_NOTE}})
|
|
10
10
|
# - any nonzero rc from cli-run (10 to 18) means no report was produced; the timer's journal shows it
|
|
11
|
+
# Disable inherited tracing and automatic export before handling credentials.
|
|
12
|
+
set +a +x +v
|
|
11
13
|
set -uo pipefail
|
|
12
14
|
INSTALL_DIR={{INSTALL_DIR_SH}}
|
|
13
15
|
AUDIT_LANE="{{AUDIT_LANE}}"
|
|
14
16
|
AUDIT_LANE_FLAGS="{{AUDIT_LANE_FLAGS}}"
|
|
15
17
|
PROBE_SECS="${PROBE_SECS:-10}" # per collection probe
|
|
16
18
|
RUNNER_SECS="${RUNNER_SECS:-600}" # the model call; TimeoutStartSec in the unit covers the whole job
|
|
17
|
-
#
|
|
18
|
-
#
|
|
19
|
-
#
|
|
20
|
-
KEY="${GATEWAY_MASTER_KEY:-}"
|
|
19
|
+
# Builtins only until the exported environment is reduced to these named runtime
|
|
20
|
+
# settings. De-exporting also handles shell-owned readonly variables. Keep home
|
|
21
|
+
# overrides private until the selected worker runs; stored sign-ins stay on disk.
|
|
21
22
|
export -n KEY
|
|
22
|
-
|
|
23
|
+
KEY="${GATEWAY_MASTER_KEY:-}"
|
|
24
|
+
while IFS= read -r name; do
|
|
25
|
+
case "$name" in
|
|
26
|
+
HOME|PATH|USER|LOGNAME|SHELL|LANG|LANGUAGE|TZ|TMPDIR|TMP|TEMP|\
|
|
27
|
+
LC_ALL|LC_CTYPE|LC_COLLATE|LC_MESSAGES|LC_MONETARY|LC_NUMERIC|LC_TIME|\
|
|
28
|
+
LC_ADDRESS|LC_IDENTIFICATION|LC_MEASUREMENT|LC_NAME|LC_PAPER|LC_TELEPHONE|\
|
|
29
|
+
XDG_CONFIG_HOME|XDG_DATA_HOME|XDG_STATE_HOME|XDG_CACHE_HOME|\
|
|
30
|
+
XDG_RUNTIME_DIR|DBUS_SESSION_BUS_ADDRESS|SystemRoot|SYSTEMROOT|WINDIR) ;;
|
|
31
|
+
*) export -n "$name" ;;
|
|
32
|
+
esac
|
|
33
|
+
done < <(compgen -e)
|
|
34
|
+
# Exported shell functions are also outside the child environment contract.
|
|
35
|
+
while IFS= read -r name; do export -nf "$name"; done < <(compgen -A function)
|
|
36
|
+
unset name GATEWAY_MASTER_KEY LITELLM_MASTER_KEY ANTHROPIC_API_KEY OPENAI_API_KEY GEMINI_API_KEY XAI_API_KEY OPENROUTER_API_KEY
|
|
23
37
|
|
|
24
38
|
# A shell pattern checks the whole value, including embedded/trailing newlines.
|
|
25
39
|
# Line-oriented grep accepts a valid line even when another line is malformed.
|
|
@@ -118,8 +132,16 @@ BRIEF="reports/audit-brief-$STAMP.md"
|
|
|
118
132
|
# Write to a temp file; the dated report is replaced only by a clean, non-empty run.
|
|
119
133
|
FINAL="reports/audit-$DATE.md"
|
|
120
134
|
TMP="$(mktemp "reports/.audit-$STAMP-XXXXXX")"
|
|
121
|
-
|
|
122
|
-
|
|
135
|
+
(
|
|
136
|
+
# These are configuration locations, not provider keys. All current scheduled
|
|
137
|
+
# lanes use stored sign-ins; no token variable is part of this contract.
|
|
138
|
+
case "$AUDIT_LANE" in
|
|
139
|
+
codex) if [ "${CODEX_HOME+x}" ]; then export CODEX_HOME; fi ;;
|
|
140
|
+
hermes) if [ "${HERMES_HOME+x}" ]; then export HERMES_HOME; fi ;;
|
|
141
|
+
esac
|
|
142
|
+
# shellcheck disable=SC2086
|
|
143
|
+
exec node bin/cli-run.mjs "$AUDIT_LANE" $AUDIT_LANE_FLAGS --brief "$BRIEF" --timeout "$RUNNER_SECS" --quiet
|
|
144
|
+
) < /dev/null > "$TMP"
|
|
123
145
|
rc=$?
|
|
124
146
|
if [ "$rc" -eq 0 ] && [ -s "$TMP" ]; then
|
|
125
147
|
mv -f "$TMP" "$FINAL"
|
|
@@ -342,6 +342,40 @@ function runSummary(args) {
|
|
|
342
342
|
const starts = records.filter((r) => r.event === 'start').length;
|
|
343
343
|
const noMatchingStart = Math.max(0, dispatches.length - starts);
|
|
344
344
|
|
|
345
|
+
// Reconciliation: does the marker's claim match what the session actually did?
|
|
346
|
+
// Both halves are already in this log -- the Stop marker says which lane the turn
|
|
347
|
+
// DECLARED, and the PreToolUse event says what was actually DISPATCHED -- and until
|
|
348
|
+
// now nothing compared them. A declared lane is a claim; a dispatch is the act.
|
|
349
|
+
// Grouped by session because a marker is written once per turn while a dispatch can
|
|
350
|
+
// land on any turn of the same session, so turn-level pairing would report drift
|
|
351
|
+
// that is only ordering.
|
|
352
|
+
// PRESENCE, not identity, and the name says so: this asks whether a session that
|
|
353
|
+
// NAMED a lane went on to dispatch at all, not whether it dispatched the lane it named. A
|
|
354
|
+
// session that named builder and dispatched reader matches here. Lane identity is
|
|
355
|
+
// not recoverable from this log, because the marker names a lane from the user's own
|
|
356
|
+
// ROUTING.md while a dispatch names a subagent_type, and the two vocabularies do not
|
|
357
|
+
// have to line up.
|
|
358
|
+
const sessions = new Map();
|
|
359
|
+
for (const r of records) {
|
|
360
|
+
if (!r.session_id) continue;
|
|
361
|
+
if (!sessions.has(r.session_id)) sessions.set(r.session_id, { declaredOff: false, dispatched: false, sawLane: false });
|
|
362
|
+
const acc = sessions.get(r.session_id);
|
|
363
|
+
if (r.event === 'dispatch') acc.dispatched = true;
|
|
364
|
+
if (r.event === 'route' && Array.isArray(r.lane)) {
|
|
365
|
+
// 'missing' is what the hook writes when a turn carried NO marker, so it is the
|
|
366
|
+
// absence of a claim, not a claim of inline. Counting it here would double-report
|
|
367
|
+
// the gap the coverage line above already reports.
|
|
368
|
+
const named = r.lane.filter((l) => l !== 'missing');
|
|
369
|
+
if (named.length > 0) acc.sawLane = true;
|
|
370
|
+
if (named.some((l) => !inlineNames.has(String(l).trim().toLowerCase()))) acc.declaredOff = true;
|
|
371
|
+
}
|
|
372
|
+
}
|
|
373
|
+
const reconcilable = [...sessions.values()].filter((a) => a.sawLane);
|
|
374
|
+
const claimedNotDone = reconcilable.filter((a) => a.declaredOff && !a.dispatched).length;
|
|
375
|
+
const doneNotClaimed = reconcilable.filter((a) => !a.declaredOff && a.dispatched).length;
|
|
376
|
+
const agreed = reconcilable.length - claimedNotDone - doneNotClaimed;
|
|
377
|
+
const agreedPct = reconcilable.length > 0 ? (agreed / reconcilable.length) * 100 : null;
|
|
378
|
+
|
|
345
379
|
const ends = records.filter((r) => r.event === 'end' && r.agent_type && typeof r.duration_s === 'number');
|
|
346
380
|
const durationsByType = new Map();
|
|
347
381
|
for (const r of ends) {
|
|
@@ -366,6 +400,15 @@ function runSummary(args) {
|
|
|
366
400
|
if (dispatchCounts.size === 0) lines.push(' (none)');
|
|
367
401
|
for (const [type, count] of [...dispatchCounts.entries()].sort((a, b) => b[1] - a[1])) lines.push(' ' + type + ': ' + count);
|
|
368
402
|
lines.push('dispatches with no matching start: ' + noMatchingStart + ' (a hook or guard blocked them before launch)');
|
|
403
|
+
lines.push(
|
|
404
|
+
'delegation claimed vs observed: ' +
|
|
405
|
+
(agreedPct === null ? 'no session named a lane yet' : formatNumber(agreedPct) + '% of sessions match') +
|
|
406
|
+
' (' + agreed + '/' + reconcilable.length + ' sessions that named a lane)'
|
|
407
|
+
);
|
|
408
|
+
lines.push(' named a lane, no dispatch in this window: ' + claimedNotDone);
|
|
409
|
+
lines.push(' dispatched, but every named lane was inline: ' + doneNotClaimed);
|
|
410
|
+
lines.push(' presence only: a session that named one lane, then dispatched a different one, still counts as matching,');
|
|
411
|
+
lines.push(' and --since or a rotated log can split a session so one half lands in the counts above.');
|
|
369
412
|
lines.push('duration by agent_type (mean / max, seconds):');
|
|
370
413
|
if (durationsByType.size === 0) lines.push(' (none)');
|
|
371
414
|
for (const [type, durs] of durationsByType) {
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
- Scaffold a checks file with `aunx checks`. Replace the intentionally failing example with a real verifier before using the file.
|
|
5
5
|
- Prefer an argv array for `command`, for example `["node", "--test", "test/example.test.js"]`. String commands run through the local shell, so review them as executable code before running a checks file.
|
|
6
6
|
- Set `cwd` relative to the checks file. Keep checks inside the task's authorized scope.
|
|
7
|
-
- Checks run sequentially with inherited terminal output. A timeout stops the command and its ordinary descendants; interrupts stop the run. A finished check must not leave a background service running. This lifecycle cleanup is not a sandbox for
|
|
7
|
+
- Checks run sequentially with inherited terminal output. A timeout stops the command and its ordinary descendants; interrupts stop the run. A finished check must not leave a background service running. This lifecycle cleanup is not a sandbox: a descendant that starts its own session (for example with `setsid`) leaves the process group and keeps running after the timeout.
|
|
8
8
|
- Demonstrate a failing case for each new gate before trusting a passing result.
|
|
9
9
|
- Include availability facts when later work depends on a tool, permission or service remaining accessible.
|
|
10
10
|
- Run `aunx checks run ACCEPTANCE_CHECKS.json` against the final artifact. A failed command gives the gate exit code 1.
|