@staix/agent-hub 0.12.2 → 0.12.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +16 -0
- package/docs/cooperbench.md +109 -0
- package/docs/events.md +14 -3
- package/docs/operations.md +158 -5
- package/docs/smoke.md +16 -0
- package/docs/specs/2026-09-19-agent-hub-design.md +66 -0
- package/docs/verification/2026-10-02-0.12.3.md +44 -0
- package/docs/verification/2026-10-02-0.12.4.md +82 -0
- package/package.json +1 -1
- package/plugins/agent-hub/.claude-plugin/plugin.json +1 -1
- package/plugins/agent-hub/server.js +23 -8
- package/src/adapters/claude-channel.ts +19 -5
- package/src/adapters/codex-appserver.ts +38 -0
- package/src/adapters/local-worker.ts +109 -16
- package/src/adapters/pi.ts +106 -21
- package/src/cli/facts-hook.ts +40 -0
- package/src/cli/launch.ts +45 -4
- package/src/cli/main.ts +44 -2
- package/src/cli/status-lines.ts +2 -1
- package/src/cli/upgrade-runtime.ts +1 -1
- package/src/hub/bus.ts +92 -6
- package/src/hub/cohorts.ts +308 -0
- package/src/hub/control-client.ts +3 -3
- package/src/hub/daemon.ts +425 -21
- package/src/hub/events.ts +12 -1
- package/src/hub/execution-budget.ts +138 -0
- package/src/hub/facts.ts +754 -0
- package/src/hub/report.ts +88 -1
- package/src/hub/routing.ts +85 -0
- package/src/hub/tasks.ts +402 -19
- package/src/hub/usage.ts +55 -0
- package/src/local/sandbox.ts +6 -2
- package/src/local/tools.ts +7 -2
- package/src/models/relay.ts +12 -4
- package/src/omniroute/client.ts +14 -3
- package/src/omniroute/usage.ts +38 -0
- package/src/pi/extension.ts +44 -8
- package/templates/AGENTS.block.md +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,22 @@
|
|
|
2
2
|
|
|
3
3
|
Issue and pull request numbers in the entries for 0.7.7 and earlier refer to the previous repository, archived on 2026-09-30 when this repository's history was rewritten; the one exception is the open smoke-check issue, formerly #12, which moved here as #1. Numbers in newer entries refer to this repository.
|
|
4
4
|
|
|
5
|
+
## 0.12.4
|
|
6
|
+
|
|
7
|
+
- A completed-change or edit-conflict notice is checked again for each recipient right before it is handed over, after condensation: a copy whose task has closed, changed owner or is gone is dropped instead of starting a turn, recorded as discarded and as a `stale` event. Other recipients and unrecorded envelopes are delivered as before (#106).
|
|
8
|
+
- Opt-in `coordination: "turn-free"` (advisory stays the default): owners of overlapping tasks form a cohort, silent only when every owner's context path is verified and no PII task is open. Messages between members are held back for those members only, with an explicit sender result and no charge against the sender's limits when nobody got them, until the sender has stopped after its task closed; that settlement is never undone, so a member's later turns are new work. A PII task opening lifts silent cohorts for good and says so. The member that finishes last (a console done counts as an intent) is asked to integrate, and its next done counts only for the same owner, cohort revision and files once the others have stopped; after three requests the outcome is recorded as unresolved. After a restart each peer's open overlapping tasks are told that overlaps are settled by message again, with the completed-change notices the silence held (#107).
|
|
9
|
+
- Turn-free facts at tool boundaries: what changed in an owner's files since it acknowledged them, with attribution only on effect evidence, offered through Claude's hooks (`ahub claude` adds them before and after every tool call and at Stop) or by steer into a running Codex turn, and acknowledged by a readback in the native session (facts sent with an integration request: by the next done). A Claude Read counts as seen only when it returned the whole file; a diff that matches a PII pattern is not shown. Contained paths, small regular files, tracking reset after PII. Control protocol 13; 0.12.3 (protocol 12) remains an upgrade source (#108).
|
|
10
|
+
- The split rule is a shadow prediction only: routing never changes; an overlap that forms or changes a cohort records a `split` event for the task, routed or named, and `ahub route explain` shows the trace, unknown unless its assumptions hold (an overlapping task its owner has started is other work) (#109).
|
|
11
|
+
- CooperBench manifest v2 adds a turn-free arm with hooks and status lines equal across arms and a planned-attempt check; a separate manifest is the #106 ablation (`experiments.stale_notices: "deliver"`); the grader grades completed and timed-out attempts and refuses a turn-free attempt whose tasks did not stay in one silent cohort, or any attempt whose records show a hook that is not the hub's; repeats rotate the arm order; run records carry their hook and MCP conditions; `scripts/benchmarks/ledger.py` reports every measure with its unit and coverage over one or several run directories, including in-task and whole-attempt usage per agent with tokens, completion and native settlement times, facts, hook timings and held-back messages, the grader's validity gates, and identifier and fragment contributions (#110).
|
|
12
|
+
|
|
13
|
+
## 0.12.3
|
|
14
|
+
|
|
15
|
+
- Fix live Claude delivery settlement holds, add generation-bound explicit completion and pending-settlement status (#100).
|
|
16
|
+
- Record local provider usage and optional native Claude usage with served-model provenance and explicit coverage (#101).
|
|
17
|
+
- Add opt-in persistent shared execution budgets for instrumented local and Pi peers; preserve legacy step units (#102).
|
|
18
|
+
- Version the fixed-sample native CooperBench runner and artifact audit (#103).
|
|
19
|
+
- Limit the test leak guard to verified processes owned by the current suite (#104).
|
|
20
|
+
|
|
5
21
|
## 0.12.2
|
|
6
22
|
|
|
7
23
|
- An edit approved after another peer changed the file now applies its fragment replacement to current contents. It rechecks that the old fragment still matches exactly once and refuses a stale match. Both write and edit revalidate their paths after approval, so a path replaced with an escaping symlink during the wait is refused (#98).
|
|
@@ -0,0 +1,109 @@
|
|
|
1
|
+
# Native CooperBench runs
|
|
2
|
+
|
|
3
|
+
The runner and report use Python's standard library. The versioned evaluator adapter also needs the optional Docker Python SDK in the supplied CooperBench evaluation environment. The checked-in `scripts/benchmarks/manifest-v1.json` fixes the upstream CooperBench commit, ten feature pairs, image digests, base commits, prompt digests, native model/version labels, topology and time limit. Prompt bodies, hidden tests, gold solutions, transcripts and vendor state are intentionally external and are not stored in this repository.
|
|
4
|
+
|
|
5
|
+
Fixture preparation is independent of the editor. Before native execution, `native.ts` reads `orca worktree current --json` and requires the reported canonical root to contain this repository. It extracts each pinned archive into a new fixture directory and initializes a git baseline after extraction, so pre-existing and untracked archive files are both represented. Symlinks, special tar entries and paths that escape the fixture are rejected.
|
|
6
|
+
|
|
7
|
+
## Prepare
|
|
8
|
+
|
|
9
|
+
Acquire the three pinned CooperBench task archives from the authorized evaluation source and place them in a private directory using these names:
|
|
10
|
+
|
|
11
|
+
- `pallets_click_task-2068.tar`
|
|
12
|
+
- `pallets_jinja_task-1465.tar`
|
|
13
|
+
- `samuelcolvin_dirty_equals_task-43.tar`
|
|
14
|
+
|
|
15
|
+
Then prepare a new output directory:
|
|
16
|
+
|
|
17
|
+
```sh
|
|
18
|
+
python3 scripts/benchmarks/runner.py prepare \
|
|
19
|
+
--manifest scripts/benchmarks/manifest-v1.json \
|
|
20
|
+
--archives /private/path/to/pinned-archives \
|
|
21
|
+
--upstream-root /private/path/to/pinned-CooperBench-checkout \
|
|
22
|
+
--output /private/path/to/new-run
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
The manifest's archive and prompt hashes are checked before a fixture is accepted. The `prepared.json` ledger binds the fixture roots, baseline commits, complete baseline path counts and manifest hash.
|
|
26
|
+
|
|
27
|
+
## Native execution
|
|
28
|
+
|
|
29
|
+
Stage `case-00.json` through `case-09.json` outside the repository. Each JSON file supplies the two exact feature prompts in feature order and the official evaluator's private case data. Pass each prior native Claude/Codex transcript as a separate `--protect <file>` argument. Do not expose those files, upstream evaluator/data tree, or prior run artifacts to the agents.
|
|
30
|
+
|
|
31
|
+
Prepare a new private run directory, then execute a fixed cohort. `--cases 0` runs every predeclared arm of the manifest for case zero (three in v1, four in v2). It never selects individual arms or retries only a scored subset.
|
|
32
|
+
|
|
33
|
+
```sh
|
|
34
|
+
python3 scripts/benchmarks/runner.py prepare \
|
|
35
|
+
--manifest scripts/benchmarks/manifest-v1.json \
|
|
36
|
+
--archives /tmp/ahub-cc-base-tars \
|
|
37
|
+
--upstream-root /tmp/agent-hub-cooperbench-upstream \
|
|
38
|
+
--output /private/tmp/ahub-0123-case0
|
|
39
|
+
chmod 700 /private/tmp/ahub-0123-case0
|
|
40
|
+
bun scripts/benchmarks/native.ts \
|
|
41
|
+
--run /private/tmp/ahub-0123-case0 \
|
|
42
|
+
--private-inputs /private/tmp/ahub-0123-private-inputs \
|
|
43
|
+
--upstream-root /tmp/agent-hub-cooperbench-upstream \
|
|
44
|
+
--probe-target /tmp/agent-hub-cooperbench-upstream/dataset/pallets_click_task/task2068/feature1/tests.patch \
|
|
45
|
+
--codex-bin /absolute/path/to/the/pinned/codex \
|
|
46
|
+
--protect /tmp/agent-hub-cooperbench-upstream \
|
|
47
|
+
--protect /Users/yj.lee/workspace/work/dev/agent-hub/.agenthub/state/benchmarks \
|
|
48
|
+
--protect /private/tmp/ahub-cc-bench-20261002-v3 \
|
|
49
|
+
--cases 0
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
The runner registers each exact fixture root with Orca, uses the project CLI to start and stop each daemon, and requires exact worktree/cwd readback for native sessions. Claude starts through the canonical `ahub claude` guard and loads this checkout's candidate bundle through an exact session-only `--mcp-config` server (`server:agent-hub`); it does not promote or mutate the globally installed plugin. Codex uses the native app-server adapter and `workspace-write` sandbox. Each agent must execute a setup-only `head -c 1` probe against the exact protected file and produce only the denied marker before scored tasks begin. There is one native sandbox layer per agent.
|
|
53
|
+
|
|
54
|
+
Every run record, diff, PTY/provider trace and evaluation artifact stays in the mode-0700 private run directory. Setup errors, structured provider quota errors, hub budget pauses, unsettled deliveries and interruptions stay separate and unscored. A quota snapshot or numeric `429` alone is not a provider error. The runner restores exact input/read-lock modes and the scoped Claude trust flag (removing its own fresh project entry) and sibling artifact modes on exit. It stops only its recorded Claude terminal, hub project and hub terminal handles. Orca currently has no repo removal command, so inactive exact-path fixture repos remain registered after a run; one case leaves one record per arm.
|
|
55
|
+
|
|
56
|
+
## Manifest v2: the turn-free arm
|
|
57
|
+
|
|
58
|
+
`scripts/benchmarks/manifest-v2.json` keeps v1's cases, models and limits and adds a fourth arm, `hub-turnfree-codex-claude` (issue #110): the same two agents and assignment rotation as `hub-codex-claude`, with `coordination: "turn-free"` in the fixture's hub config and fixture instructions that tell the owners not to message each other. A v1 manifest still validates and prepares. Its cohorts are run, graded and reported with the release and runner sources they were prepared with: `native.ts` refuses a manifest whose hub version is not the checkout's (v1 pins 0.12.3), and `runner.py` refuses to grade a cohort whose runner sources differ from its own.
|
|
59
|
+
|
|
60
|
+
Hooks are equal across arms: no arm runs the user's or a plugin's hooks, and no arm runs a status line (`disableAllHooks` turns it off in the other arms, so the turn-free arm leaves it out; Claude's quota reaches the hub in no arm). Claude starts with `--setting-sources project` and `--strict-mcp-config`; the solo and advisory arms set `disableAllHooks`, and the turn-free arm's session settings carry the hub's own hooks (before and after every tool call, and at Stop) and nothing else, because they are the treatment. Every Codex thread starts with `features.hooks` off: the turn-free arm's Codex boundary is the adapter's steer into the running turn, whose readback is the steered input coming back as a user message item. MCP servers are isolated too: Claude has only the hub's (`--strict-mcp-config`), and every Codex runs through a wrapper written into each arm's fixture that turns off the user's plugins, apps, sub-agents and turn-end notifier and disables each MCP server the user's config defines, by name, so it starts only the hub's; the user's config is not changed. Instructions: Claude reads the fixture's `AGENTS.md` through `--append-system-prompt-file` (Claude Code reads `CLAUDE.md`, not `AGENTS.md`), Codex as the project's `AGENTS.md`; Codex also reads the user's global `AGENTS.md`, which only a separate Codex home with its own login would leave out (the run records say so), and the user's Codex skills stay available to it (the wrapper leaves them on); both are the same in every arm. Each run record carries these `conditions`; its `events` carry the hub's `capability` readbacks.
|
|
61
|
+
|
|
62
|
+
Validity, decided by the grader and applied by the ledger alike: an attempt whose records show a hook or an MCP server that is not the hub's, or whose Claude transcript cannot be read, is unavailable. A turn-free attempt is valid only with its context paths working: both verified before its tasks, and none lost, no cohort lifted and none formed open while the agents worked; teardown comes after that and does not count. Whether the agents' plans overlapped, so that a cohort formed at all, is their doing after assignment and is not a condition: every turn-free attempt without a capability failure counts for the arm, and the ledger reports the treatment received and a median over treated attempts beside it. A capability failure while the agents work is the one exclusion after assignment, because #110 forbids reporting it as a turn-free run (AC3); such attempts are listed with their reasons, never dropped silently. Each attempt records its transcript's length and hash when it ends; validity and every Claude measure are read from that prefix (Claude Code may append rows after it exits: a response still being written at teardown is not counted, and that agent's settlement is then unknown), and a prefix that changed counts as unreadable.
|
|
63
|
+
|
|
64
|
+
`scripts/benchmarks/manifest-v2-ablation-106.json` is the #106 ablation: the advisory arm against the same arm with `experiments.stale_notices: "deliver"`, which turns stale-notice dropping off and changes nothing else, on the same release and conditions. Its attempts are counted apart from the four-arm plan.
|
|
65
|
+
|
|
66
|
+
Pass `--repeat <n>` for the n-th repeat of a case (0 for the first): the arm order is row (case index + repeat) of a Williams design (0, 1, n-1, 2, n-2, ... shifted by the row), so over n consecutive rows every arm runs right before every other one once and repeats of one pair change which arm runs last. The manifest's `plan` names the release pilot (case 0, three repeats, 12 attempts, an active-time ceiling of one hour) and the study (ten cases, two repeats, 80 attempts, 6 hours 40 minutes at 300 s each, setup, grading and teardown excluded); `runner.py` refuses a plan whose attempts or ceiling do not follow from its arms, cases and repeats.
|
|
67
|
+
|
|
68
|
+
## Coordination ledger
|
|
69
|
+
|
|
70
|
+
```sh
|
|
71
|
+
python3 scripts/benchmarks/ledger.py --run /private/tmp/ahub-0124-r1 [--run /private/tmp/ahub-0124-r2 ...]
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
It reads each run record, the Claude transcript it names and the fixture's git history, and writes `ledger.json` into the first run directory with a `units` table that names the unit and coverage of every measure. Give `--run` once per directory to pool the repeats of one plan; an attempt a directory's `cohort.json` planned that wrote no record is listed as missing. Per attempt: completion (the runner's end reason and its detail, such as `wall-timeout`, and whether every task has a done); setup and active time; first candidate, completion intents, integration and check times, and each agent's settlement read from its own record (the end of its last native turn after the last done, turns started by late messages included; the last done itself when it did not work after it), and how long the agents could still write after the active time ended (`stopped_s`); usage per agent in task (to its last done on the board) and over the whole attempt: Codex turns, assistant messages, token-usage updates whose running total grew past the total before the window, and the token growth; Claude assistant messages (unique message ids), turns and the tokens their usage records; Claude's main-loop requests by request id (side requests and retries are not in its transcript), and Codex's provider requests unknown (app-server 0.159 does not send its response ids); Codex turns after its done and what started them; late replies, steered or at the next turn; the hub's fact offers, acknowledgements, bytes offered and acknowledged, build times, hook start-up times and steer round trips by path, in the task window; capability readbacks; validity by the grader's gates, and for turn-free the treatment received; held-back messages (`quiet` events) apart from [FYI] messages (which include the final [FYI] the instructions ask for); stale notices; shadow split predictions with their traces; the hooks each agent's records show, by a label that keeps paths and arguments out, with Claude's hook durations and the hub's own timing of every facts hook call; and contributions. Summaries give medians over valid completed attempts and over the (case, repeat) pairs every arm completed validly, totals over the attempts whose tasks were handed out and to which the measure applies (a Codex measure in a Claude-only arm is not counted as unknown), with the number of attempts each was unknown for, and the reasons for the rest. A repeat given twice is refused, and `--plan pilot` (or `study`) lists every planned attempt that wrote no record, whole repeats included. `ledger.json` holds code fragments from the agents' writes and local paths: keep it with the private run data and never commit it; the summary is what a verification record quotes. In-task windows end at different points by arm (a turn-free integration step comes before the done, an advisory completed-change notice turn after it), so arms are compared on whole-attempt usage.
|
|
75
|
+
|
|
76
|
+
Contributions are a heuristic for possible loss, never a certificate: per agent and file, the identifiers and changed fragments its applied writes introduced (only a Claude tool call with a successful result, or a completed Codex patch, counts; a Write replaces the agent's earlier contribution and is not credited with what other agents wrote; a delete removes it; when a file is moved, every agent's contributions to it are checked at its new path) that the final tree lacks, a fragment counting as present anywhere in the file. A same-name overwrite shows as a lost fragment. Shell commands run during the task are not attributed and are counted under `coverage`, with a missing transcript or an unreadable file. Correctness comes from the official grader, for every arm.
|
|
77
|
+
|
|
78
|
+
## Preregistration (v2)
|
|
79
|
+
|
|
80
|
+
Fixed before outcomes are collected (issue #110):
|
|
81
|
+
|
|
82
|
+
- Order of ablations: the stale-notice change first (#106: `manifest-v2-ablation-106.json`, the advisory arm with and without stale-notice dropping on the same release; comparing with 0.12.3's runs would confound it with that release's Codex hooks), then cohort silence with acknowledged facts and the revision-fenced integration (#107 and #108, the turn-free arm). The split rule (#109) is recorded as shadow predictions only; task allocation stays fixed by case index in both joint arms.
|
|
83
|
+
- Primary outcome: the final submitted artifact of each attempt, graded by the official tests with the same controls for every arm. Completed and timed-out attempts are graded (a timeout's final tree is its submission, collected once the agents were stopped, `stopped_s` after the limit); interrupted, infrastructure-failed and invalid attempts are reported unavailable with their reasons. Time to both done is reported only with completion, failures, timeouts and unavailable attempts beside it, and timing comparisons use attempts that completed in every compared arm.
|
|
84
|
+
- Resource outcomes per provider over the whole attempt, in the units the ledger names: Codex turns, usage growth and tokens; Claude assistant messages, main-loop requests and tokens (amended on 2026-10-02, before the study: side requests are not in Claude's transcript); hook and fact bytes and latencies.
|
|
85
|
+
- The turn-free arm is analysed over every attempt without a capability failure (the only exclusion after assignment, required by AC3, and listed), with the treated attempts' median as a secondary result. A 35% time and 30% usage reduction are hypotheses, not targets met; counters of different providers are not added together.
|
|
86
|
+
- Every preregistered attempt is retained, including failed and infrastructure-unavailable ones; no repeat is selected after inspecting results.
|
|
87
|
+
- The release pilot (case 0, three repeats of four arms) is feasibility evidence only. Changing the default from advisory needs paired quality and resource criteria fixed in advance, held-out repository and task instances, and comparison with both solo controls and the current advisory arm; any fact loss, missed integration, stale-generation injection or held-back workflow event is a safety result that blocks it.
|
|
88
|
+
|
|
89
|
+
## Grade and report
|
|
90
|
+
|
|
91
|
+
The versioned adapter invokes CooperBench's official `test_solo`, confirms the empty-base control fails and the official combined gold patch passes, and substitutes the verified image digest for the evaluator's mutable image tag. It rejects a changed source/data tree, removes stale evaluation outputs, and binds each score to the exact diff/evaluation/evaluator/manifest hashes.
|
|
92
|
+
|
|
93
|
+
```sh
|
|
94
|
+
python3 scripts/benchmarks/runner.py grade --run /private/tmp/ahub-0123-case0 \
|
|
95
|
+
--private-inputs /private/tmp/ahub-0123-private-inputs \
|
|
96
|
+
--upstream-root /tmp/agent-hub-cooperbench-upstream \
|
|
97
|
+
--python /tmp/agent-hub-bench-venv/bin/python
|
|
98
|
+
python3 scripts/benchmarks/runner.py report --run /private/tmp/ahub-0123-case0
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
Report rows use qualified identities (`repo:task:feature`) and unavailable attempts remain outside the score denominator. A grade whose manifest, patch, or evaluation hash no longer matches is rejected.
|
|
102
|
+
|
|
103
|
+
Native run integration must follow the pending settlement and telemetry contracts in #100-102: task approval is distinct from delivery settlement, missing usage remains unknown, and execution limits must name their unit. It must also snapshot and restore any temporary Claude trust/read locks on every exit path, verify native cwd/session/model readiness, use a single native sandbox layer, check the canonical Orca root, and stop only actors, hubs and containers it started. Retry is an explicit new attempt over the complete predeclared cohort, never a score-selected subset.
|
|
104
|
+
|
|
105
|
+
The fixed ten-pair sample is a convenience sample, not the full 652-pair suite. The shared checkout is not isolated CooperBench coop, and these artifacts do not establish leaderboard parity or causal coordination superiority.
|
|
106
|
+
|
|
107
|
+
The optional `--codex-bin` selects an absolute native executable when PATH contains several Codex installations. Its reported version must match the manifest. Codex automatic memories and external agent memory import are disabled for the benchmark thread. Native usage totals include the unscored sandbox probe and are labelled as whole-session counters.
|
|
108
|
+
|
|
109
|
+
`--setup-only` runs every selected arm through native readiness and read-denial probes without assigning feature work. Its cohort is marked as calibration and the grader refuses it. A zero-turn Claude session is bound through its verified Orca launch and then checked against native transcript session IDs after the probe. The instance-fenced metadata file enables the daemon's optional usage reader.
|
package/docs/events.md
CHANGED
|
@@ -14,10 +14,20 @@ marked `private: true`, and PII tasks `pii: true`.
|
|
|
14
14
|
|---|---|
|
|
15
15
|
| `envelope` | `id`, `from`, `to` (absent for broadcast), `priority`, `hop`, `kind`, `task` (task id from refs), `bytes` (UTF-8 size of the body; absent on a private envelope, whose size would say something about the PII text), `private`, `dropped` (`hop` or `fyi` when not delivered) |
|
|
16
16
|
| `overflow`, `undeliverable` | `id`, `from`, `peer` |
|
|
17
|
+
| `stale` | `id`, `from`, `peer`, `task`: a notice about the recipient's open task, dropped unsent because, right before it would have been handed over, that task was closed, had another owner or was gone (issue #106) |
|
|
18
|
+
| `quiet` | `id`, `from`, `peers`: an agent message held back from these members of a silent turn-free cohort; its other recipients got it (issue #107) |
|
|
19
|
+
| `fact` | `peer`, `id` (the offer), `files` (files whose diff it carried), `plans`, `unknown` (files whose change it showed with attribution unknown), `bytes` (the injected text), `via` (`hook` for Claude, `steer` for Codex, `done` with an integration request), `ms` (the hub's time to build it), `hookMs` (the hook process's own start-up and connect time), `accepted` (whether app-server took the steer), `unanswered` (app-server did not answer it within 10 s: it may have gone in), `rttMs` (an accepted steer: from sending it to app-server's answer), `probe` (a context check), `coverage` (it named files earlier changes are not covered for): one fact offer (issue #108) |
|
|
20
|
+
| `fact_ack` | `peer`, `id`, `via` (`hook` and `steer`: a readback found the offer in the native session; `done`: the next `hub_task_done`), `ms` (from the offer): an acknowledged offer, the only thing that moves a peer's view |
|
|
21
|
+
| `capability` | `peer`, `state` (`verified` or `lost`), `via`: a peer's context path for facts |
|
|
22
|
+
| `native_turn_end` | `peer`: Claude's Stop hook, the end of its turn (Codex's is its `turn_end`); quiescence evidence for an integration |
|
|
23
|
+
| `hook_stats` | `peer`, `n` (facts hook calls in the turn, its Stop included), `startupMs` and `maxStartupMs` (the hook processes' start-up and connect time, summed and the largest), `hubMs` (the hub's own time for them): at Claude's Stop (issue #108) |
|
|
24
|
+
| `cohort` | `id`, `event` (`formed`, `joined`, `lifted`), `silent`, `tasks`, `owners`: owners of overlapping tasks formed a cohort, it changed membership, or it stopped being silent (issue #107). Recorded in every regime; only a turn-free project's cohorts can be silent |
|
|
25
|
+
| `split` | `task`, `verdict` (`split`, `single`, `unknown`), `single` (the peer that would finish both units alone soonest), `splitS`, `singleS`, `reason` (for `unknown`), `trace` (the inputs and steps: peer names and numbers only): a shadow split prediction where an overlap forms or changes a cohort, routed or named; it never changes the assignment (issue #109) |
|
|
17
26
|
| `state` | `peer`, `state` |
|
|
18
27
|
| `turn_start` | `peer`, `turn` (`<peer>#<hub run>.<n>`, unique across restarts). A turn follows the adapter: pausing a busy peer does not end it |
|
|
19
28
|
| `turn_end` | `peer`, `turn`, `ms`, `tokens` (when the adapter reported any during the turn), `files` and `snapshotMs` (when snapshots are on: how many files the turn changed, and the time both snapshots took) |
|
|
20
29
|
| `tokens` | `peer`, `n` (tokens added since the previous report) |
|
|
30
|
+
| `usage` | `peer`, `source`, opaque `id`, optional `measuredAt` (provider/source time), requested/served model and provider labels, and any provider-reported input/output/cache/total counters. Missing counters stay unknown. |
|
|
21
31
|
| `task` | `id`, `event` (the board history event, e.g. `proposed`, `assigned`, `done`, `check failed`, `blocked`, `ready`), `by`, `state`, `owner`, `reviewer`, `class`, `pii` |
|
|
22
32
|
| `overlap` | `task`, `owner`, `others` (`task`, `owner`, `paths`, and `symbols` when plans name the same symbol; a name that matches a PII pattern is left out, so either list can be empty), the structured twin of the console notice |
|
|
23
33
|
| `quota` | `peer`, `windows` (`id`, `used`, `resetsAt`), `hard`, `measuredAt` (when the reading was taken, if not when it arrived: Claude's numbers come through a file) |
|
|
@@ -30,9 +40,10 @@ Token usage by adapter:
|
|
|
30
40
|
compaction estimates, usage-limit refreshes and the replay to a reattaching connection add nothing. A thread
|
|
31
41
|
started under the hub counts from zero; a resumed thread's first update is its history and only sets the
|
|
32
42
|
baseline. The model call of a compaction itself is real usage and counts.
|
|
33
|
-
- Claude
|
|
34
|
-
|
|
35
|
-
-
|
|
43
|
+
- Claude native transcript usage is optional and keyed by an opaque hash of session and message identity; streamed records with the same message id count once. The status line tee still carries quota percentages only.
|
|
44
|
+
- The local worker records optional counters returned by OmniRoute. Requested route/model and gateway-reported served model/provider are separate fields; an alias is never treated as a served model.
|
|
45
|
+
- `ahub report` deduplicates usage records by peer, source and id. Coverage counts distinguish calls with provider usage from calls where usage was absent. Token counters are provider-reported values; the report never derives a price or treats missing spend as zero. Estimated price and measured provider spend remain unknown unless a future source reports them.
|
|
46
|
+
- Usage telemetry has no prompt, completion, task text, credential, Access header, session id or transcript path.
|
|
36
47
|
|
|
37
48
|
The file is local and never uploaded. It grows without rotation; delete it to start
|
|
38
49
|
over (the hub recreates it).
|
package/docs/operations.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Operations guide
|
|
2
2
|
|
|
3
|
-
This guide describes ahub 0.12.
|
|
3
|
+
This guide describes ahub 0.12.4 and control protocol 13. Live verification
|
|
4
4
|
results and remaining prerequisites are recorded separately in [the smoke ledger](smoke.md).
|
|
5
5
|
|
|
6
6
|
## Install and start
|
|
@@ -344,6 +344,159 @@ are quoted and marked as other agents' text. The hook runs `ahub`, so it has to
|
|
|
344
344
|
be on the PATH Claude Code's hooks see; otherwise every edit shows a hook error
|
|
345
345
|
(it never blocks). `ahub check-path <file>` prints the same list in a terminal.
|
|
346
346
|
|
|
347
|
+
## Turn-free coordination
|
|
348
|
+
|
|
349
|
+
A notice about an open task of its recipient (the completed-change notice and the
|
|
350
|
+
edit-conflict messages above) is checked again for each recipient right before it
|
|
351
|
+
is handed over, after any digest condensation: if that recipient's task has
|
|
352
|
+
closed, changed owner or is gone, its copy is dropped instead of starting a turn,
|
|
353
|
+
the journal records it as `discarded`, `hub.log` has a `STALE` line and
|
|
354
|
+
`events.jsonl` a `stale` event. Other recipients keep their copies. Approvals,
|
|
355
|
+
check results, review requests, assignments and budget, permission and recovery
|
|
356
|
+
messages are never dropped this way, and a dropped notice never resolves another
|
|
357
|
+
delivery. The condition lives in memory: after a restart, or for the oldest of
|
|
358
|
+
more than 1024 notices, the notice is delivered as before.
|
|
359
|
+
|
|
360
|
+
Set `"coordination": "turn-free"` in `.agenthub/config.json` to let owners of
|
|
361
|
+
overlapping open tasks (the overlap rules above) work without messaging each
|
|
362
|
+
other. The default, `"advisory"`, keeps the behaviour described so far, and it
|
|
363
|
+
stays the default until an evaluation says otherwise (`docs/cooperbench.md`).
|
|
364
|
+
|
|
365
|
+
- Cohorts. Owners of overlapping tasks form a cohort when the overlap is found.
|
|
366
|
+
It is silent only if, at that moment, turn-free is on, no PII task is open and
|
|
367
|
+
every owner's context path is verified (below). It never turns silent later;
|
|
368
|
+
it stops being silent, for good, when an owner without a verified path joins,
|
|
369
|
+
a path is lost or a PII task opens, and every member still at work is told at
|
|
370
|
+
once that it may message again, with the completed-change notices the silence
|
|
371
|
+
held.
|
|
372
|
+
- Verified context path. A hub-launched Claude session (`ahub claude`) gets the
|
|
373
|
+
hub's hook before and after every tool call and at the end of each turn, in
|
|
374
|
+
its `--settings` next to the status line tee (a `--settings` of your own turns
|
|
375
|
+
them off). Codex gets context by steer into its running turn. Until a fact or a
|
|
376
|
+
one-line probe has been read back, the peer counts as unverified: for Claude
|
|
377
|
+
the row Claude Code writes in its transcript for the hook's additional
|
|
378
|
+
context, matched by tool use id and the offer's id; for Codex the steered
|
|
379
|
+
input coming back as a user message item of the turn. A new native session,
|
|
380
|
+
the peer going offline, or three offers past a minute without a readback
|
|
381
|
+
(checked at its boundaries and every 30 seconds; facts sent with an
|
|
382
|
+
integration request wait for the next done and do not count, a refused steer
|
|
383
|
+
is no offer, and an unanswered one may still be read back), makes it
|
|
384
|
+
unverified again, and a Claude session whose
|
|
385
|
+
transcript the hub cannot find gets no facts at all and loses a verified
|
|
386
|
+
path. A session whose hooks stop altogether leaves no offer unread, so it
|
|
387
|
+
stays verified; a peer without a verified path that left three
|
|
388
|
+
offers unread gets none until a readback arrives. Kimi, Pi and the local
|
|
389
|
+
worker have no path yet, so a cohort with one of them is never silent.
|
|
390
|
+
- Silence. While a cohort is silent, an agent message from a member to another
|
|
391
|
+
member is held back for that recipient only: other recipients and the console
|
|
392
|
+
get it unchanged, and `events.jsonl` records a `quiet` event. `hub_send`
|
|
393
|
+
answers `not delivered to <peer>: ...` (or `sent to: ...; not delivered to
|
|
394
|
+
...` when some recipients got it); a native turn answer's sender hears it on
|
|
395
|
+
its next delivery. Workflow messages from the hub are never held back. A
|
|
396
|
+
member's messages stay held until its native turn has ended after its task
|
|
397
|
+
closed (Codex's turn completes, or Claude's Stop hook runs; a member already
|
|
398
|
+
between turns when its task closes has stopped; a paused peer may still be in
|
|
399
|
+
its turn, and Claude's channel going offline says nothing about its session),
|
|
400
|
+
so a late answer is still the cohort's. That settlement is recorded when it happens and
|
|
401
|
+
never undone: the member's next turn is new work, and once every member has
|
|
402
|
+
settled the cohort is over (a task reopened after that is outside it). Only a tool call starting counts as activity after
|
|
403
|
+
a turn end. A message held back from all its recipients does not count
|
|
404
|
+
against the sender's limits.
|
|
405
|
+
- Facts. At each tool call (Claude) or completed tool item (Codex) of a cohort
|
|
406
|
+
member with an open task, the hub offers what changed since it last
|
|
407
|
+
acknowledged them in the files every member's task names and in the files it
|
|
408
|
+
touched, with the other members' new plans; the last member still at work
|
|
409
|
+
keeps the others' files after they finish. A directory a task names stands for
|
|
410
|
+
git's changed (staged or not) files against HEAD, new and deleted files under
|
|
411
|
+
it (200 at most; the fact says when more were cut, until it is read back and
|
|
412
|
+
again in a new session). A file the peer has seen there stays covered after git
|
|
413
|
+
stops listing it (put back to HEAD's bytes, or the directory moved away), so
|
|
414
|
+
the way back is shown (200 at most, the newest versions first; the rest are
|
|
415
|
+
named once as no longer tracked). A file there that the peer has neither seen
|
|
416
|
+
nor touched appears once someone changed or created it: it is named without a
|
|
417
|
+
diff (what happened before is never shown) until the peer reads the fact back,
|
|
418
|
+
and until then it counts as a change the peer has not been shown. `.git` directories
|
|
419
|
+
at any depth and what the denylist keeps from every agent
|
|
420
|
+
(`src/local/deny.ts` and `local.deny`) are never read or shown. A file a peer touched before it had a view of it (a
|
|
421
|
+
partial read, say) is compared with what it was then, so a change landing in
|
|
422
|
+
between is shown. A change is credited to an agent
|
|
423
|
+
only with effect evidence: a Claude Edit, MultiEdit or Write whose result is
|
|
424
|
+
exactly its input applied to the file as observed before it, or a Codex patch
|
|
425
|
+
whose diff is exactly what changed. Shell commands, concurrent writers and
|
|
426
|
+
unreported changes are shown with their attribution unknown, never credited by
|
|
427
|
+
elimination; an agent's own verified writes are not shown back to it. A Codex
|
|
428
|
+
read action never counts as having seen a file (it may be partial); a Claude
|
|
429
|
+
Read does when it returns the whole file: no offset or limit, at most 2000
|
|
430
|
+
lines and no line over 2000 characters. A diff that matches a PII pattern is
|
|
431
|
+
not shown (the file is named, to be read), and a changed file whose name
|
|
432
|
+
matches one is counted, not named. Only an acknowledgement (a readback, or the next
|
|
433
|
+
`hub_task_done` for facts sent with an integration request) moves the peer's
|
|
434
|
+
view, so a fact that does not arrive is
|
|
435
|
+
offered again at a later boundary; no turn is ever started for one. An
|
|
436
|
+
acknowledgement says the context reached the native session, not that the
|
|
437
|
+
model read it. 60 changed lines are shown at most, the cut files named; a
|
|
438
|
+
history longer than the hub keeps is shown with its attribution unknown, and a
|
|
439
|
+
file that falls out of the 64 a peer touched is named once. Only regular files
|
|
440
|
+
of 256 KB or less inside the project are read, re-resolved at every read and
|
|
441
|
+
opened without following links; larger ones are named without a diff. Facts never go through the bus or the delivery journal;
|
|
442
|
+
`events.jsonl` records `fact`, `fact_ack` and `capability` events with bytes
|
|
443
|
+
and latencies.
|
|
444
|
+
- Integration. A member's `hub_task_done` is a completion intent. The member
|
|
445
|
+
whose intent completes the set is asked, as its done result, to check its work
|
|
446
|
+
against the others' (their files, signatures and summaries, plus its own facts)
|
|
447
|
+
and to call `hub_task_done` again; nothing is recorded as done yet. The next
|
|
448
|
+
call counts only for the same target: the same owner, cohort revision and
|
|
449
|
+
files (the named ones, and those each member wrote with an edit tool the hub
|
|
450
|
+
saw, Claude's Edit, MultiEdit and Write or a Codex patch, between being handed
|
|
451
|
+
its task and settling, so a symbol-only overlap counts and
|
|
452
|
+
a member settling keeps its files in; reading a file never moves it, a settled
|
|
453
|
+
member's later work does not count, a shell command's writes outside the named
|
|
454
|
+
paths are not seen, and a file the hub reads whole whose bytes equal HEAD's
|
|
455
|
+
is no change, whatever git's stat data says), with every other owner's native turn
|
|
456
|
+
ended after its done (or that owner idle when it finished); the integrating
|
|
457
|
+
owner's own other tasks in the cohort never count as still running. A done of a member by the console counts as its
|
|
458
|
+
intent too. Edits in between, a new member, an owner change, a failed check or
|
|
459
|
+
a reopened review ask again; a done within two seconds of a request is taken as a retry and gets the
|
|
460
|
+
same request again; after three requests the done is recorded with
|
|
461
|
+
`integration unresolved`, never as integrated, and that revision asks nothing
|
|
462
|
+
more (a configured check then counts as usual). A configured check of the integrating member
|
|
463
|
+
counts only for the target it confirmed; when another member reopens its task,
|
|
464
|
+
that member integrates instead and the earlier one's check counts as usual.
|
|
465
|
+
Inside a silent cohort a member's completed-change notice is held, not sent;
|
|
466
|
+
where no integration step runs for the others (the silence was lifted, a PII
|
|
467
|
+
task opened), the held notices go with the lift notice or the done result, and
|
|
468
|
+
for a task the console finishes, to the console. Open tasks outside the cohort
|
|
469
|
+
get their notices as usual. Facts sent with an integration request are
|
|
470
|
+
offered again at the next boundary until the next done acknowledges them.
|
|
471
|
+
Cohorts live in memory: after a hub restart an open
|
|
472
|
+
request is recorded as unresolved, and when a peer first attaches, each of its
|
|
473
|
+
open tasks that overlaps other work hears that overlaps are settled by message
|
|
474
|
+
again, with the completed-change notices of overlapping tasks finished since
|
|
475
|
+
it was handed the task (a notice may come twice; none is lost).
|
|
476
|
+
- While any PII task is open the project behaves as advisory: no facts, no
|
|
477
|
+
silence and no integration step, and the cohorts that were silent stay lifted.
|
|
478
|
+
When it closes, the hub forgets what it had observed, so nothing changed
|
|
479
|
+
meanwhile is shown as a diff; each member is told which of its files to read
|
|
480
|
+
again.
|
|
481
|
+
- `ahub check-path` asks the hub whether the owner of the claimed path shares a
|
|
482
|
+
silent cohort with the caller, and only then leaves out the request to settle
|
|
483
|
+
by message. `templates/claude-hooks.json` holds only the check-path hook; the
|
|
484
|
+
facts hooks need a hub-launched session.
|
|
485
|
+
- Routing does not change. When a task's overlap with another owner's open task
|
|
486
|
+
forms or changes a cohort, routed or named, the hub records a shadow split
|
|
487
|
+
prediction (`split` event; `ahub route explain <id>` shows its trace as it would
|
|
488
|
+
be now): whether splitting two equal units between the two peers
|
|
489
|
+
(`o_s + u_s < o_f + 2u_f`) would finish sooner than the faster one alone, from
|
|
490
|
+
this hub run's recorded task stages. It is unknown unless the units are equal
|
|
491
|
+
and known, both peers are available (the other owner idle; the task's own peer
|
|
492
|
+
idle or busy taking it) with no other open work (an overlapping task its owner
|
|
493
|
+
has started counts), and each has five measured tasks with no more than 30%
|
|
494
|
+
failures and comparable work times. The work stage of a task ends at its first
|
|
495
|
+
`hub_task_done`, and the task itself is never one of its own observations.
|
|
496
|
+
- `"experiments": {"stale_notices": "deliver"}` in `.agenthub/config.json`
|
|
497
|
+
turns the stale-notice drop off, for a controlled comparison only (the #106
|
|
498
|
+
ablation in `docs/cooperbench.md`); the hub logs it at start.
|
|
499
|
+
|
|
347
500
|
## Approvals and pauses
|
|
348
501
|
|
|
349
502
|
Inspect permission requests in the terminal:
|
|
@@ -474,21 +627,21 @@ Rows without a live process are stale registrations; forget them with
|
|
|
474
627
|
|
|
475
628
|
Upgrade running projects with the target release's own coordinator. It accepts
|
|
476
629
|
a running source on control protocol 9 (0.6.x), 10 (0.7.0 through 0.12.0),
|
|
477
|
-
|
|
630
|
+
11 (0.12.1 and 0.12.2) or 12 (0.12.3) and only
|
|
478
631
|
a target on its own protocol, so the target's coordinator fits every supported
|
|
479
632
|
source and carries every recovery fix released up to it. Protocol 8 and older
|
|
480
633
|
(0.5.x and earlier) are refused as `manual-bootstrap-required`. Run from the
|
|
481
634
|
project directory, without replacing the global CLI first:
|
|
482
635
|
|
|
483
636
|
```bash
|
|
484
|
-
bunx --package @staix/agent-hub@0.12.
|
|
485
|
-
bunx --package @staix/agent-hub@0.12.
|
|
637
|
+
bunx --package @staix/agent-hub@0.12.4 ahub upgrade --to 0.12.4 --dry-run
|
|
638
|
+
bunx --package @staix/agent-hub@0.12.4 ahub upgrade --to 0.12.4 --yes
|
|
486
639
|
```
|
|
487
640
|
|
|
488
641
|
| Running now | Coordinator to use |
|
|
489
642
|
| --- | --- |
|
|
490
643
|
| 0.6.x (protocol 9) | the target's, through `bunx` as above |
|
|
491
|
-
| 0.7.0 through 0.12.0 (protocol 10) | the target's, through `bunx` as above |
|
|
644
|
+
| 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12) | the target's, through `bunx` as above |
|
|
492
645
|
| any supported source, with the installed CLI already at the target | `ahub upgrade` below, which is the same coordinator |
|
|
493
646
|
| 0.5.x or earlier (protocol 8 and older) | not supported: bootstrap by hand with the matching CLI |
|
|
494
647
|
|
package/docs/smoke.md
CHANGED
|
@@ -1114,3 +1114,19 @@ work. Times are `hub.log` UTC. Issues #89-#95 come from this run.
|
|
|
1114
1114
|
Retried from the console, the question reached Kimi led by the loss notice ("The hub stopped unexpectedly ...",
|
|
1115
1115
|
with the delivery id), and Kimi answered at 13:19:26.6.
|
|
1116
1116
|
- **Pi's tokens were missing from the report** (#94).
|
|
1117
|
+
|
|
1118
|
+
### Channel settlement and usage (0.12.3, protocol 12)
|
|
1119
|
+
|
|
1120
|
+
With real Codex and Claude sessions in a disposable project, deliver a workflow
|
|
1121
|
+
assignment to Claude, then a directed Codex question. Reply to the question and
|
|
1122
|
+
approve the workflow task. Confirm `ahub status` still lists the workflow delivery
|
|
1123
|
+
as awaiting settlement, without a queue hold, and a later important message arrives.
|
|
1124
|
+
Call `hub_delivery_done` with the workflow channel's delivery ID and generation.
|
|
1125
|
+
Confirm only that row becomes completed. Wrong peer, stale generation and unrelated
|
|
1126
|
+
IDs must be refused. Disconnect during another accepted delivery; confirm
|
|
1127
|
+
`needs_review` survives reconnect and requires explicit inspected queue resolution.
|
|
1128
|
+
|
|
1129
|
+
Compare Claude usage events against its explicit native session transcript, deduped
|
|
1130
|
+
by assistant message ID. Compare local usage events with actual successful provider
|
|
1131
|
+
response counters; absent counters remain unknown and model aliases remain separate
|
|
1132
|
+
from reported served-model provenance. Do not infer dollar spend from token counts.
|
|
@@ -397,6 +397,16 @@ inside a peer.
|
|
|
397
397
|
not paused again, and still gets the resume envelope once it attaches. With no status line
|
|
398
398
|
of the user's own to wrap, the tee prints a short usage line instead of a blank one.
|
|
399
399
|
|
|
400
|
+
### Execution budgets (issue #102)
|
|
401
|
+
|
|
402
|
+
- Execution budgets are opt-in records in `hub.db`, separate from provider quota-window budgets. Configure and inspect them with `ahub budget execution configure <config.json>`, `ahub budget execution status [id]`, and `ahub budget execution disable <id>`. The JSON names a stable `id`, `kind` (`task` or `run`), eligible `peers`, optional `taskId` for task scope, and `limits` keyed by `model_calls`, `tool_calls`, `elapsed_ms`, or `tokens`.
|
|
403
|
+
- A run budget configuration can be saved as JSON and passed to `configure`, for example `{"id":"run:cooperbench-1","kind":"run","peers":["pi","local"],"limits":{"model_calls":40,"tool_calls":32,"elapsed_ms":180000}}`. A task configuration uses `{"id":"task:42","kind":"task","taskId":42,"peers":["pi","local"],"limits":{"model_calls":8,"tool_calls":12}}`. `status` reports scope, eligible peers, limits, measured usage, and the remaining amount or exhaustion reason.
|
|
404
|
+
- Task scope applies only when that task appears in the delivery's `refs.task`; run scope applies across all eligible peers' turns until disabled. If a digest carries several matching tasks, all applicable task scopes and each run scope are admitted atomically. Configuration and counters survive new turns and daemon restarts. Reconfiguring the same id changes limits/eligible peers without resetting consumed usage; a scope id cannot be rebound to a different task or scope kind.
|
|
405
|
+
- `model_calls` counts each admitted provider request. `tool_calls` counts each admitted tool execution (including Pi user shell commands). These reservations happen before requests or effects. The existing `pi.max_steps` remains a per-agent-turn tool execution ceiling; `local.max_steps` remains a per-agent-turn model-loop iteration ceiling. Neither legacy default is reinterpreted or disabled.
|
|
406
|
+
- `elapsed_ms` starts when the budget is configured and is checked at every model/tool admission. `tokens` is explicit only when usage telemetry is available; because a future request's token cost is not known before sending it, a configured token cap without a safe reservation estimate stops the next model request with `unknown_usage`, rather than treating missing usage as zero.
|
|
407
|
+
- Exhaustion ends the delivery without automatic replay. A stop before side effects is reported as `needs_review`; after side effects it reports that partial work may exist and also requires review. Provider quota interruption, budget exhaustion, legacy step caps, and successful native completion remain distinct reasons.
|
|
408
|
+
- Only `pi` and `local` are currently accepted as budgeted peers because they expose a pre-request and pre-tool admission point. Native Claude/Codex/Kimi limits are rejected until their adapters can enforce the same contract. With no execution budget configured, the new meter has no effect.
|
|
409
|
+
|
|
400
410
|
### Shared memory (claude-mem)
|
|
401
411
|
|
|
402
412
|
- Capture. Claude: native plugin hooks. Codex: claude-mem Codex plugin. Kimi: `dot ai
|
|
@@ -861,6 +871,14 @@ session modes (`default`, `plan`, `auto`, `yolo`) but nothing per server or tool
|
|
|
861
871
|
the issue's design listed. `ahub queue list` and hub.log keep them; a later
|
|
862
872
|
schema version can add them without changing the meaning of a field.
|
|
863
873
|
|
|
874
|
+
## Amendment: provider usage provenance (issue #101)
|
|
875
|
+
|
|
876
|
+
- OmniRoute keeps validated optional usage counters from the provider response and sanitized served-model/provider labels. The caller's requested route/model remains separate from the reported served model.
|
|
877
|
+
- Local worker records one usage event per successful provider response, including a response with no usage counters so reports can show missing coverage. Credentials, Access headers, prompts, completions, task text, session ids and transcript paths never enter telemetry.
|
|
878
|
+
- An optional Claude transcript reader accepts explicit session identity and transcript path, but emits only completed assistant-message usage with an opaque session/message hash. Repeated streaming records collapse to the final record for that message; event reports deduplicate repeated polling and resumed transcript reads.
|
|
879
|
+
- Reports sum only provider-reported counters and show per-counter known-record coverage. Missing usage remains unknown. Estimated price and measured provider spend are separate and remain unknown unless sourced; token counts are not prices.
|
|
880
|
+
- `ahub report` coverage describes recorded provider calls; it does not claim complete account or billing coverage.
|
|
881
|
+
|
|
864
882
|
## Amendment: per-turn snapshots and undo (issue #33)
|
|
865
883
|
|
|
866
884
|
- Backend: git tree objects only, written through a copy of the index
|
|
@@ -1151,3 +1169,51 @@ check, then reads current contents after approval and requires the old fragment
|
|
|
1151
1169
|
to occur exactly once. The approved replacement applies to those current bytes,
|
|
1152
1170
|
so another peer's unrelated edits made during the wait are retained. A changed
|
|
1153
1171
|
fragment or an escaping path returns an error without a write.
|
|
1172
|
+
|
|
1173
|
+
## Live channel settlement (issues #100-#104 follow-up)
|
|
1174
|
+
|
|
1175
|
+
Wire protocol 12 separates live notification acceptance from recovery uncertainty.
|
|
1176
|
+
A live `accepted` Claude notification remains visible as `liveAccepted` in status
|
|
1177
|
+
while later notifications can arrive. A correlated reply settles only its own
|
|
1178
|
+
message. Task acceptance, completion and approval are independent of delivery
|
|
1179
|
+
settlement and cannot complete unrelated messages.
|
|
1180
|
+
|
|
1181
|
+
The Claude channel API does not expose a verified native turn-end event. Instead,
|
|
1182
|
+
channel metadata supplies `delivery_id` and a connection `delivery_generation`;
|
|
1183
|
+
`hub_delivery_done` is an explicit acknowledgement of handled work. The daemon
|
|
1184
|
+
requires the current peer socket, matching generation, a delivery actually handed
|
|
1185
|
+
to that socket, and a live accepted journal row. This acknowledgement is not proof
|
|
1186
|
+
of a native turn boundary. Failed notifications, disconnects, replaced sessions
|
|
1187
|
+
and interrupted daemon instances retain `needs_review`, with no automatic replay.
|
|
1188
|
+
A person inspects and resolves uncertain work through the existing revision-fenced
|
|
1189
|
+
`ahub queue` commands. Older source protocols 9, 10 and 11 remain authenticated
|
|
1190
|
+
upgrade sources; ordinary clients must use protocol 12 (13 from 0.12.4, with 12 a source).
|
|
1191
|
+
|
|
1192
|
+
## Amendment: stale overlap notices (issue #106)
|
|
1193
|
+
|
|
1194
|
+
- The completed-change notice (#31) and the edit-conflict notices (#32, #91) are about an open task of their recipient. `Tasks.whileOpen` publishes each to one recipient and remembers the condition (the newest 1024). The bus asks `Tasks.relevant(peer, env)` per recipient when it builds a delivery and again right before it hands the delivery over, after condensation; it is false only for a recorded notice to that peer whose task is gone, has another owner, or left the states it was about: `proposed`, `in_progress`, `changes_requested`, and for conflict notices also `in_review` (#91 reports conflicts with work under review). A dropped copy is a journal row `discarded` with `stale: task #N is no longer open for <peer>`, a `STALE` log line and a `stale` event; other recipients and other deliveries are untouched.
|
|
1195
|
+
- An envelope without a record (another kind, a restart, an evicted record) is delivered as before: its purpose is never guessed from its kind. Assignments, approvals, check results, review requests and budget, permission and recovery messages are never conditional.
|
|
1196
|
+
- The ablation is the advisory arm with and without stale-notice dropping on the same release (`manifest-v2-ablation-106.json`, #110); comparing with 0.12.3's advisory runs would confound it with that release's Codex hooks.
|
|
1197
|
+
|
|
1198
|
+
## Amendment: turn-free cohorts (issue #107)
|
|
1199
|
+
|
|
1200
|
+
- `coordination` in the project config: `"advisory"` (default) or `"turn-free"`; anything else is advisory with a log line. While a PII task is open the project is advisory.
|
|
1201
|
+
- `src/hub/cohorts.ts`: owners of overlapping open tasks (#31) form a cohort when the overlap is found (assignment, an accept with a plan). Members are tasks with their owner and owner generation (ownership events); every change of membership or owner, a withdrawn intent and a lift bump the cohort's revision. A cohort is silent if, when formed, every owner's context path is verified (#108); it never becomes silent later, and it is lifted for good (members still at work told at once, with the held notices) when an unverified owner joins, a path is lost or a PII task opens. Cohorts live in memory: after a restart, each peer's first attach tells its open overlapping tasks that overlaps are settled by message again and replays the completed-change notices of overlapping tasks finished since they were handed over.
|
|
1202
|
+
- Silence is a per-recipient bus policy (`BusOptions.silence`), applied once the audience is final (implicit replies and `digest` resolved): an agent's chat from a member that has not settled to another member is not queued for that member; the other recipients get the envelope unchanged, so correlation, priority caps, dedupe and hops are untouched. A `quiet` bus event goes to the log and `events.jsonl`; `hub_send` answers with the held-back peers (an error when nobody got it); a native turn answer's sender hears it as a note. A member settles once its task left the open states and its native turn has ended after that (Codex's adapter leaving its turn, Claude's Stop hook; a member already between turns when its task closes settles at once; a paused peer counts by its adapter's state, and Claude's channel going offline says nothing about its session). Settlement is recorded as it happens (`Cohorts.turnEnded`, `Cohorts.closed`) and never undone: a settled member's next turn is new work, and a cohort whose members all settled is over. A message held back from all its recipients is not admitted against the sender's limits.
|
|
1203
|
+
- Texts: in a silent cohort, overlap results and notices, task envelopes and the conflict notices name the plans and say not to message; `hub_task_accept` returns the overlapping plans with every accept; `ahub check-path` asks the hub (`silenced` control request). A member's completed-change notice to another member is held (`Cohort.held`), never dropped: it goes with the lift notice when silence is lifted, with the done result of a member for whom no integration step runs (turn-free off, a PII task open), or to the console when the console finishes the task (its done still counts as that member's intent). Notices to open tasks outside the cohort go as usual. Otherwise the advisory texts stay.
|
|
1204
|
+
- Integration: a member's `hub_task_done` records an intent. The member whose intent completes the set is selected in one synchronous step and gets an integration request as its done result (the others' files, signatures and summaries, plus its facts); the request is recorded as `integration requested`, not done. Its next done is accepted (`integrated`, then the usual done) only for the same owner generation, cohort revision and file hash (named files, for a directory git's changed, new and deleted files under it against HEAD, and the files each member wrote with an edit tool the hub saw (Claude's Edit, MultiEdit and Write; Codex's patches) between being handed its task and settling; reads never count, nor shell writes outside the named paths, and a file the hub reads whole whose bytes equal HEAD's is no change there), with its facts acknowledged (the done first acknowledges the facts sent with the request) and every other owner stopped after its intent (a native turn end after it, or idle when it was recorded); the integrating owner's own other tasks never count as running. A cohort whose members have all settled is over, whoever asks first, and is never lifted or reopened after that. A done within two seconds of a request is a retry and gets the same request, with nothing acknowledged or counted. Otherwise it is asked again; past three requests the done is recorded with `integration unresolved`, and that revision asks nothing more (the member's check counts as usual). A failed check or changes requested withdraws that member's intent; a check of the integrating member that passes after its target moved within the same revision counts as finished late; once the revision moved (another member reopened), the former integrating member is an earlier finisher and its check counts. After a restart an open request is recorded as unresolved.
|
|
1205
|
+
|
|
1206
|
+
## Amendment: turn-free facts at tool boundaries (issue #108)
|
|
1207
|
+
|
|
1208
|
+
- `src/hub/facts.ts` keeps three things apart per peer: what the hub observed in the tree (latest version and transitions per file), what it offered (bounded offers with their file versions and plans), and what the peer acknowledged (its view). Only an acknowledgement moves a view; a later boundary offers everything since the view again. Files are the paths every task of each live cohort in which the peer has an open task names (a directory expands to git's changed files against HEAD, staged or not, and its new and deleted files, 200 at most, with a note when more were cut, told until it is read back and again in a new session), contained in the project as real paths and never `.git` or anything the denylist (`src/local/deny.ts`, `local.deny`) keeps from agents, plus the last 64 it touched; a file that falls out of those is named once. A file the peer touched before it had a view of it is compared with its version at that touch, so a change in between is shown, not absorbed into a first look; a file it saw under a named directory stays covered after git stops listing it (put back to HEAD's bytes, or the directory moved away; 200 at most, the rest named once as no longer tracked). A directory file it has neither seen nor touched is named without a diff when it appears (diffing it against HEAD would show what changed while tracking was off, a PII window, and credit unseen changes by elimination), and counts as not yet shown until that fact is read back. Only regular files of 256 KB or less are read, re-resolved at every read and opened without following a final link or blocking. A history longer than the 64 kept transitions makes the change's attribution unknown.
|
|
1209
|
+
- Attribution: a transition is credited only with effect evidence: a Claude Edit, MultiEdit or Write whose result equals its input applied to the file as observed at PreToolUse with no observation in between, or a Codex `fileChange` whose changed lines equal the observed diff. Everything else is shown with its attribution unknown; an agent's own verified writes advance its view and are not shown back. A Claude Read advances its view only when it returned the whole file (no offset or limit, at most 2000 lines, none over 2000 characters); a Codex read action never does (it may be partial). A diff that matches a PII pattern is replaced by a note naming the file, and a changed file whose name matches one is counted without its name.
|
|
1210
|
+
- Acknowledgement and capability: Claude's hook runs before and after every tool and at Stop (`ahub claude` injects it in a turn-free project); the pre answer's text becomes `additionalContext`, and the row Claude Code writes in its transcript for it (matched by tool use id and offer id) is the readback. Codex's offer goes in with `steerText` (bound to the running turn), and the steered input coming back as a `userMessage` item is the readback. A readback marks the peer's path verified; a new session or thread, the peer going offline, or three offers past a minute without one (checked at boundaries and every 30 seconds; facts sent with an integration request wait for the next done and never count, a refused steer is dropped, not unread, and an unanswered one stays pending, since it may have gone in) unverifies it. Until verified, a peer gets a one-line probe at most three times per session, and no offers at all once three are unread; a Claude session whose transcript the hub cannot find gets none at all and loses a verified path.
|
|
1211
|
+
- The control contract is protocol 13: `facts` carries the phase, tool, input, tool use id, session id and transcript path; `silenced` serves check-path; `send` results name held-back recipients. Protocol 12 stays an authenticated recovery source.
|
|
1212
|
+
- While a PII task is open nothing is observed or offered; every board change re-checks this, so tracking switches at once. When tracking resumes everything observed is dropped and each peer is told which files earlier changes are not covered for. Turn ends and tool starts are recorded in every regime.
|
|
1213
|
+
- `events.jsonl`: `fact` (bytes, build time, the hook's own time, the steer round trip), `fact_ack` (latency), `capability`, `native_turn_end`, `hook_stats` (every facts hook call of a Claude turn, at its Stop), `cohort`.
|
|
1214
|
+
|
|
1215
|
+
## Amendment: shadow split prediction (issue #109)
|
|
1216
|
+
|
|
1217
|
+
- `predictSplit(input)` is a pure function next to `assign()`, which no longer takes a split option: assignment never changes. For the pair formed by the peer routing picks and the owner of an overlapping open task, it compares a split (the later of each peer's orientation plus one unit) with the best peer alone (orientation plus two units). It is unknown unless the units are equal and known (one task of the same class is one unit), both peers are available (the other owner idle; the task's own peer idle or busy taking it) with no other open work (an overlapping task its owner has started counts), and each has five measured tasks in this hub run with at most 30% failures and work times whose IQR does not exceed their median; a difference under a tenth of the single time is inconclusive.
|
|
1218
|
+
- `Tasks.splitObservations` reads only tasks handed out in this hub run (one version and hook profile), never the routed task itself, leaves out claims, types failures (failed check, changes requested, escalation, release, decline, unresolved integration), ends the work stage at the first done call (checks and integration are not work) and leaves the stages unknown for an accept recorded by the done itself. The `split` event carries the trace.
|
|
1219
|
+
- An overlap that forms or changes a cohort records a `split` event for the task, routed or named (a benchmark names every owner); `ahub route explain <id>` appends the trace as it would be now. Rerouting needs held-out evidence through #110 first.
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# Agent-hub 0.12.3 verification
|
|
2
|
+
|
|
3
|
+
Implementation and candidate verification for issues #100 through #104, performed on 2026-10-02. Production application is verified separately after publication. The final runtime source revision was `e12f824c5b951ac5e8d2c448e53cf207a8062d60`; subsequent changes only record verification and two scoped quality trade-offs.
|
|
4
|
+
|
|
5
|
+
## Runtime checks
|
|
6
|
+
|
|
7
|
+
- Claude: a live accepted notification does not become a recovery hold. Correlated replies and explicit `hub_delivery_done` settle only the matching delivery on the current connection generation. Task approval remains independent of delivery settlement.
|
|
8
|
+
- Local usage: two real provider requests recorded 5,994 input tokens and 59 output tokens (6,053 total). The requested alias `coding` identified the physical served model `glm-5.3-flash`. Dollar spend was not measured.
|
|
9
|
+
- Execution budgets: a shared two-model-call budget blocked the third Local request before dispatch. A zero-model-call Pi budget blocked the real relay without token growth or tool effects. Typed budget denial retained the Pi task owner and did not escalate or reassign it.
|
|
10
|
+
- Pi self-development: the candidate hub assigned an inspection task, recorded its implementation plan and accepted the resulting source findings. Three unrelated shell permission requests were denied; Pi completed through managed reads.
|
|
11
|
+
- Recovery: a real published 0.12.2/protocol-11 source transitioned to candidate protocol 12, preserving queued envelope IDs, task-state digest and manual pauses.
|
|
12
|
+
- Test process ownership: real owned daemon and detached Bun descendant detection passed; unrelated temporary daemons and a Python command mentioning an agent-hub path were excluded. Ownership uses process birth identity, not text matching.
|
|
13
|
+
|
|
14
|
+
Strict model/tool admission currently covers Local and Pi. Native Codex and Claude expose optional usage and benchmark wall time; they are not presented as instrumented strict execution-budget peers. Unknown counters and spend remain unknown.
|
|
15
|
+
|
|
16
|
+
## Native CooperBench verification
|
|
17
|
+
|
|
18
|
+
The checked-in manifest fixes ten official feature pairs. Release verification repeats the complete three-condition cohort for one pair: `pallets_click_task:2068`, features `1` and `6`. This verifies runner integration and repeatability; it is not a replacement ten-pair performance study.
|
|
19
|
+
|
|
20
|
+
- Official source commit: `63b9d44d9f39a02fccf5bf0052db48a917a011fd`.
|
|
21
|
+
- Codex CLI 0.159.3, `gpt-6.1-sol`, medium effort.
|
|
22
|
+
- Claude Code 2.1.287, `claude-opus-5-5`, medium effort.
|
|
23
|
+
- Conditions: solo Codex; solo Claude; Codex plus Claude sharing one fixture checkout.
|
|
24
|
+
- Each feature-work arm has a 300-second wall limit. Setup and denied hidden-file probes are unscored.
|
|
25
|
+
- The official evaluator runs on the recorded image digest, with empty-base-fail and combined-oracle-pass controls. Qualified feature identities, private input bytes, source tree and submission/evaluation hashes are bound to the cohort.
|
|
26
|
+
|
|
27
|
+
| Cohort | Solo Codex | Solo Claude | Codex + Claude | Official grading |
|
|
28
|
+
|---|---|---|---|---|
|
|
29
|
+
| R1 | Completed, 88.4 s | Completed, 42.7 s | Completed, 94.3 s | 3/3 artifacts passed both features; controls passed |
|
|
30
|
+
| R2 | Completed, 104.5 s | Completed, 44.5 s | Wall timeout, 300.3 s | 2/2 available artifacts passed both features; collaboration unscored; controls passed |
|
|
31
|
+
| R3 | Completed, 107.2 s | Completed, 71.2 s | Completed, 95.4 s | Not graded; runner source changed after preparation |
|
|
32
|
+
| R4, final runtime source | Completed, 120.0 s | Completed, 46.0 s | Completed, 108.8 s | 3/3 artifacts passed both features; controls passed |
|
|
33
|
+
|
|
34
|
+
The first repeat was graded at its recorded earlier candidate revision. The second repeat's collaboration timeout is retained and excluded from the score denominator. The third repeat completed all native arms, but is not graded because a later portability fix changed the pinned runner source. Prepared hashes were not rewritten.
|
|
35
|
+
|
|
36
|
+
Native session counters include setup probes. Raw prompts, conversations, session IDs, provider traces, private inputs and gold solutions remain outside version control. Every completed cohort restored its owned read locks and Claude trust flag and stopped its recorded actors and hub daemons. Orca does not currently expose a repo removal command; inactive fixture registrations remain.
|
|
37
|
+
|
|
38
|
+
This is a shared-workspace convenience sample. No full CooperBench leaderboard, isolated-coop parity, Databricks internal benchmark parity or causal coordination advantage is claimed.
|
|
39
|
+
|
|
40
|
+
## Automated gate and review
|
|
41
|
+
|
|
42
|
+
`scripts/check.sh`: 605 tests passed, 0 failed, 3,062 expectations across 63 files. Typecheck, bundle freshness, npm package contents and process ownership checks passed; final output was `check: OK`. Linux and macOS CI also passed at the reviewed runtime revision.
|
|
43
|
+
|
|
44
|
+
Independent host Codex OCR delegation reviews covered all changed runtime, benchmark and infrastructure source plus manually excluded tests/docs. The generated committed plugin bundle is validated by the build freshness gate. All Important findings were fixed and re-reviewed against the exact final commit before merge.
|