model-orchestrator 0.1.9 → 0.1.11

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,29 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.11] - 2026-09-07
8
+
9
+ ### Fixed
10
+
11
+ - **The cost claim the audit rejected was still live in three places outside the README.** #11 asked for the opening promise to be narrowed, and 0.1.2 narrowed it in `README.md` only. "Route every task to the cheapest AI that does it well" survived in `package.json`'s `description`, which is what **npm search results show**, in the repository's GitHub description, which is what **GitHub search shows**, and in the installer's own banner, printed to **every user on every run**. All three now say what the package generates instead of what it guarantees: "Routing instructions and a CLI runner for your AI tools." The cheapest-capable-lane *guidance* in the docs and templates is untouched; that is the product's advice, not a promise about what the code enforces.
12
+
13
+ ## [0.1.10] - 2026-09-07
14
+
15
+ Closes the last two verification items on #11. Both had been described as needing vendor sign-ins or infrastructure that was not available. Both turned out to be doable with what was already here, and doing them found a real defect.
16
+
17
+ ### Fixed
18
+
19
+ - **The generated weekly audit orphaned a temp file on every timeout.** A run killed by the unit's `TimeoutStartSec` dies on SIGKILL, so no trap and no cleanup line of ours can run, and its `reports/.audit-<stamp>-XXXXXX` file was left behind forever. The job now sweeps `.audit-*` older than a day at start. A day is far outside the unit's own 900s deadline, so a temp belonging to a run still in flight can never be swept. Found by actually starting the unit on Ubuntu; the previous text-only assertion could not see it.
20
+
21
+ ### Added
22
+
23
+ - `test/fixtures/`: raw output captured from **real vendor CLI runs**, with `manifest.json` recording the vendor version, the exact flags, the exit code, and what each fixture proves. Every other judge test in this repository uses shapes written by hand. The capture earned itself immediately: real codex 0.153.4 emits an `item.completed` whose item is `type:"error"` (a skills-budget warning) *before* `turn.completed`, which no synthetic fixture contained, and real agy returns `"OK\n"` with a trailing newline. `test/fixtures.test.js` runs every judge against them.
24
+ - `test/systemd/run-on-ubuntu.sh` and its README: starts the generated job as a **real systemd user unit** and proves that the timeout kills the whole cgroup (a detached grandchild does not survive it), that a failed rerun preserves the previous report, that a malformed gateway key exits 2 before anything is written, and that the new sweep removes an aged orphan while leaving a fresh one alone. 11/11 on Ubuntu 24.04.4 LTS, systemd 255. Not part of `npm test`, which has no systemd to run against.
25
+
26
+ ### Notes
27
+
28
+ - `test/fixtures/README.md` states its own gaps rather than hiding them: qwen's success shape is still synthetic because its key was not present in the capture environment, and `claude` and `ollama` have no judge, so no fixture.
29
+
7
30
  ## [0.1.9] - 2026-09-06
8
31
 
9
32
  ### Added
@@ -144,6 +167,8 @@ First release.
144
167
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
145
168
 
146
169
  [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.7...HEAD
170
+ [0.1.11]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.10...v0.1.11
171
+ [0.1.10]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.9...v0.1.10
147
172
  [0.1.9]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.8...v0.1.9
148
173
  [0.1.8]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.7...v0.1.8
149
174
  [0.1.7]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.6...v0.1.7
package/bin/cli.js CHANGED
@@ -139,7 +139,10 @@ function bad(msg) {
139
139
  }
140
140
 
141
141
  async function main() {
142
- console.log('\nmodel-orchestrator\nRoute every task to the cheapest AI that does it well.\n');
142
+ // Says what this generates, not what it guarantees. The old line promised
143
+ // routing this package does not perform: lane choice is an instruction an
144
+ // agent follows, never something enforced here (#11).
145
+ console.log('\nmodel-orchestrator\nRouting instructions and a CLI runner for the AIs you actually have.\n');
143
146
 
144
147
  // 1. Level
145
148
  let level = Number(opt('level'));
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.9",
4
- "description": "A model orchestrator you can npm run: route every task to the cheapest AI that does it well, across one agent, many CLIs, or a whole virtual machine. Three levels, one installer that asks what you have access to.",
3
+ "version": "0.1.11",
4
+ "description": "Routing instructions and a CLI runner for your AI tools. One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine.",
5
5
  "type": "module",
6
6
  "bin": {
7
7
  "model-orchestrator": "bin/cli.js"
@@ -18,6 +18,14 @@ RUNNER_SECS="${RUNNER_SECS:-600}" # the model call; TimeoutStartSec in the uni
18
18
  cd "$INSTALL_DIR" || { echo "weekly-audit: $INSTALL_DIR missing" >&2; exit 2; }
19
19
  mkdir -p reports
20
20
 
21
+ # Sweep temp files a previous run could not clean up. A run killed by the unit's
22
+ # TimeoutStartSec dies on SIGKILL, so no trap and no cleanup line of ours can
23
+ # run, and its `.audit-<stamp>-XXXXXX` file is orphaned in reports/ forever.
24
+ # Measured on Ubuntu 24.04 / systemd 255, 2026-09-07: one orphan per timeout.
25
+ # Anything older than a day cannot belong to a live run (the unit's own deadline
26
+ # is 900s), so removing it is safe even if another run is in flight.
27
+ find reports -maxdepth 1 -name '.audit-*' -type f -mtime +0 -delete 2>/dev/null || true
28
+
21
29
  # bounded SECS cmd... : run cmd, and after SECS kill it AND every descendant
22
30
  # (a probe that forks, or a stub that ignores its own flags, must not hold a
23
31
  # pipe open). Process groups do not help here: bash disables job control inside