model-orchestrator 0.1.9 → 0.1.11
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +25 -0
- package/bin/cli.js +4 -1
- package/package.json +2 -2
- package/templates/advanced/vm/jobs/weekly-audit.sh +8 -0
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,29 @@ All notable changes to this project are documented here. The format follows [Kee
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [0.1.11] - 2026-09-07
|
|
8
|
+
|
|
9
|
+
### Fixed
|
|
10
|
+
|
|
11
|
+
- **The cost claim the audit rejected was still live in three places outside the README.** #11 asked for the opening promise to be narrowed, and 0.1.2 narrowed it in `README.md` only. "Route every task to the cheapest AI that does it well" survived in `package.json`'s `description`, which is what **npm search results show**, in the repository's GitHub description, which is what **GitHub search shows**, and in the installer's own banner, printed to **every user on every run**. All three now say what the package generates instead of what it guarantees: "Routing instructions and a CLI runner for your AI tools." The cheapest-capable-lane *guidance* in the docs and templates is untouched; that is the product's advice, not a promise about what the code enforces.
|
|
12
|
+
|
|
13
|
+
## [0.1.10] - 2026-09-07
|
|
14
|
+
|
|
15
|
+
Closes the last two verification items on #11. Both had been described as needing vendor sign-ins or infrastructure that was not available. Both turned out to be doable with what was already here, and doing them found a real defect.
|
|
16
|
+
|
|
17
|
+
### Fixed
|
|
18
|
+
|
|
19
|
+
- **The generated weekly audit orphaned a temp file on every timeout.** A run killed by the unit's `TimeoutStartSec` dies on SIGKILL, so no trap and no cleanup line of ours can run, and its `reports/.audit-<stamp>-XXXXXX` file was left behind forever. The job now sweeps `.audit-*` older than a day at start. A day is far outside the unit's own 900s deadline, so a temp belonging to a run still in flight can never be swept. Found by actually starting the unit on Ubuntu; the previous text-only assertion could not see it.
|
|
20
|
+
|
|
21
|
+
### Added
|
|
22
|
+
|
|
23
|
+
- `test/fixtures/`: raw output captured from **real vendor CLI runs**, with `manifest.json` recording the vendor version, the exact flags, the exit code, and what each fixture proves. Every other judge test in this repository uses shapes written by hand. The capture earned itself immediately: real codex 0.153.4 emits an `item.completed` whose item is `type:"error"` (a skills-budget warning) *before* `turn.completed`, which no synthetic fixture contained, and real agy returns `"OK\n"` with a trailing newline. `test/fixtures.test.js` runs every judge against them.
|
|
24
|
+
- `test/systemd/run-on-ubuntu.sh` and its README: starts the generated job as a **real systemd user unit** and proves that the timeout kills the whole cgroup (a detached grandchild does not survive it), that a failed rerun preserves the previous report, that a malformed gateway key exits 2 before anything is written, and that the new sweep removes an aged orphan while leaving a fresh one alone. 11/11 on Ubuntu 24.04.4 LTS, systemd 255. Not part of `npm test`, which has no systemd to run against.
|
|
25
|
+
|
|
26
|
+
### Notes
|
|
27
|
+
|
|
28
|
+
- `test/fixtures/README.md` states its own gaps rather than hiding them: qwen's success shape is still synthetic because its key was not present in the capture environment, and `claude` and `ollama` have no judge, so no fixture.
|
|
29
|
+
|
|
7
30
|
## [0.1.9] - 2026-09-06
|
|
8
31
|
|
|
9
32
|
### Added
|
|
@@ -144,6 +167,8 @@ First release.
|
|
|
144
167
|
- Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
|
|
145
168
|
|
|
146
169
|
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.7...HEAD
|
|
170
|
+
[0.1.11]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.10...v0.1.11
|
|
171
|
+
[0.1.10]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.9...v0.1.10
|
|
147
172
|
[0.1.9]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.8...v0.1.9
|
|
148
173
|
[0.1.8]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.7...v0.1.8
|
|
149
174
|
[0.1.7]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.6...v0.1.7
|
package/bin/cli.js
CHANGED
|
@@ -139,7 +139,10 @@ function bad(msg) {
|
|
|
139
139
|
}
|
|
140
140
|
|
|
141
141
|
async function main() {
|
|
142
|
-
|
|
142
|
+
// Says what this generates, not what it guarantees. The old line promised
|
|
143
|
+
// routing this package does not perform: lane choice is an instruction an
|
|
144
|
+
// agent follows, never something enforced here (#11).
|
|
145
|
+
console.log('\nmodel-orchestrator\nRouting instructions and a CLI runner for the AIs you actually have.\n');
|
|
143
146
|
|
|
144
147
|
// 1. Level
|
|
145
148
|
let level = Number(opt('level'));
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "model-orchestrator",
|
|
3
|
-
"version": "0.1.
|
|
4
|
-
"description": "
|
|
3
|
+
"version": "0.1.11",
|
|
4
|
+
"description": "Routing instructions and a CLI runner for your AI tools. One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
7
7
|
"model-orchestrator": "bin/cli.js"
|
|
@@ -18,6 +18,14 @@ RUNNER_SECS="${RUNNER_SECS:-600}" # the model call; TimeoutStartSec in the uni
|
|
|
18
18
|
cd "$INSTALL_DIR" || { echo "weekly-audit: $INSTALL_DIR missing" >&2; exit 2; }
|
|
19
19
|
mkdir -p reports
|
|
20
20
|
|
|
21
|
+
# Sweep temp files a previous run could not clean up. A run killed by the unit's
|
|
22
|
+
# TimeoutStartSec dies on SIGKILL, so no trap and no cleanup line of ours can
|
|
23
|
+
# run, and its `.audit-<stamp>-XXXXXX` file is orphaned in reports/ forever.
|
|
24
|
+
# Measured on Ubuntu 24.04 / systemd 255, 2026-09-07: one orphan per timeout.
|
|
25
|
+
# Anything older than a day cannot belong to a live run (the unit's own deadline
|
|
26
|
+
# is 900s), so removing it is safe even if another run is in flight.
|
|
27
|
+
find reports -maxdepth 1 -name '.audit-*' -type f -mtime +0 -delete 2>/dev/null || true
|
|
28
|
+
|
|
21
29
|
# bounded SECS cmd... : run cmd, and after SECS kill it AND every descendant
|
|
22
30
|
# (a probe that forks, or a stub that ignores its own flags, must not hold a
|
|
23
31
|
# pipe open). Process groups do not help here: bash disables job control inside
|