@staix/agent-hub 0.12.16 → 0.12.18
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +14 -0
- package/README.md +4 -4
- package/docs/agent-notes/adapters.md +1 -0
- package/docs/agent-notes/bus.md +2 -0
- package/docs/agent-notes/tasks.md +3 -1
- package/docs/agent-notes/tests.md +1 -0
- package/docs/events.md +37 -3
- package/docs/operations.md +149 -10
- package/docs/quickstart.md +3 -3
- package/docs/security.md +28 -1
- package/docs/smoke.md +93 -0
- package/docs/specs/2026-09-19-agent-hub-design.md +176 -1
- package/docs/verification/2026-10-09-agent-shell-t0.md +123 -0
- package/docs/verified.json +26 -17
- package/package.json +1 -1
- package/plugins/agent-hub/.claude-plugin/plugin.json +1 -1
- package/plugins/agent-hub/server.js +17 -5
- package/src/adapters/acp.ts +2 -2
- package/src/adapters/claude-channel.ts +3 -2
- package/src/adapters/codex-appserver.ts +3 -2
- package/src/adapters/local-worker.ts +16 -12
- package/src/cli/console-state.ts +279 -0
- package/src/cli/console.ts +224 -0
- package/src/cli/facts-hook.ts +8 -1
- package/src/cli/identity-audit.ts +66 -0
- package/src/cli/identity.ts +54 -0
- package/src/cli/launch.ts +15 -2
- package/src/cli/main.ts +75 -32
- package/src/cli/tail-render.ts +17 -0
- package/src/cli/upgrade-runtime.ts +1 -1
- package/src/hub/attribution.ts +20 -0
- package/src/hub/board.ts +6 -2
- package/src/hub/bus.ts +88 -10
- package/src/hub/child-process.ts +11 -0
- package/src/hub/conductor.ts +196 -0
- package/src/hub/control-client.ts +3 -3
- package/src/hub/daemon.ts +388 -51
- package/src/hub/envelope.ts +1 -1
- package/src/hub/events.ts +10 -4
- package/src/hub/hub-tools.ts +13 -0
- package/src/hub/report.ts +162 -4
- package/src/hub/supervision.ts +152 -0
- package/src/hub/tasks.ts +30 -15
- package/src/hub/usage.ts +37 -2
- package/src/pi/launch.ts +2 -1
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,20 @@ Issue and pull request numbers in the entries for 0.7.7 and earlier refer to the
|
|
|
4
4
|
|
|
5
5
|
## Unreleased
|
|
6
6
|
|
|
7
|
+
## 0.12.18
|
|
8
|
+
|
|
9
|
+
- Attribute token increments, provider usage and completed turns to task ids by delivery or a single open task; add `ahub report --by task` with task/class totals, explicit unknown counters, unattributed shares and a separate historical bucket. Preserve PII id-only exports and derive no prices (#200).
|
|
10
|
+
- Add semantic colors to console stream headers, footer and panels with `--color=auto|always|never`; keep plain-text Unicode geometry, sanitize incoming controls before the fixed palette, and restore terminal attributes on exit. Tail, logs, JSON and polling remain unchanged (#201).
|
|
11
|
+
- Control protocol stays 15 and events schema stays 1; a running 0.12.17 hub upgrades with the 0.12.18 coordinator.
|
|
12
|
+
|
|
13
|
+
## 0.12.17
|
|
14
|
+
|
|
15
|
+
- Add an operator console with confirmed approvals, bounded status polling and optional peer, approval, task, queue and event panels; open it from interactive `up` and preserve the tail renderer (#190, #191).
|
|
16
|
+
- Identify agent shell CLI calls as their peer, refuse human-only commands before connecting and record ids-only audit notices without an as-user bypass (#193).
|
|
17
|
+
- Grant steering tools only to one explicit conductor role, preserve the actor on task changes and keep conductor holds separate from human, budget and recovery holds (#194).
|
|
18
|
+
- Batch public task milestones and human-action reminders through the existing digest window, replace repeated pending notices and report completed native supervision turns with unknown usage preserved (#195).
|
|
19
|
+
- Advance the control protocol to 15 for approval lifecycle and supervision metadata, retaining controlled recovery from supported previous sources.
|
|
20
|
+
|
|
7
21
|
## 0.12.16
|
|
8
22
|
|
|
9
23
|
- Record source-verification commits for current operating documents and agent notes; fail missing coverage, invalid stamps and README version drift, and report stale source scopes deterministically (#188).
|
package/README.md
CHANGED
|
@@ -4,7 +4,7 @@ Native multi-agent hub for one developer's machine: Claude Code, Codex, Kimi Cod
|
|
|
4
4
|
hub-owned local-LLM worker collaborate as peers in independent project directories, with
|
|
5
5
|
task-aware model routing (Switchyard) in front of a self-hosted gateway (OmniRoute).
|
|
6
6
|
|
|
7
|
-
Status: 0.12.
|
|
7
|
+
Status: 0.12.18, control protocol 15. Durable delivery records distinguish queued
|
|
8
8
|
work from uncertain execution. The [smoke checklist](docs/smoke.md) records
|
|
9
9
|
verified paths and remaining prerequisites.
|
|
10
10
|
|
|
@@ -22,13 +22,13 @@ Install from npm:
|
|
|
22
22
|
|
|
23
23
|
```bash
|
|
24
24
|
bun add -g @staix/agent-hub && ahub setup
|
|
25
|
-
cd <your project> && ahub init && ahub up
|
|
25
|
+
cd <your project> && ahub init && ahub up
|
|
26
26
|
```
|
|
27
27
|
|
|
28
28
|
Or install the same version from GitHub:
|
|
29
29
|
|
|
30
30
|
```bash
|
|
31
|
-
bun add -g github:STAIxBWLB/agent-hub#v0.12.
|
|
31
|
+
bun add -g github:STAIxBWLB/agent-hub#v0.12.18 && ahub setup
|
|
32
32
|
```
|
|
33
33
|
|
|
34
34
|
The installed commands remain `ahub` and `agent-hub`.
|
|
@@ -100,7 +100,7 @@ only. `--backend dgx` or `--backend mlx` explicitly pins a Pi session's backend.
|
|
|
100
100
|
Pi connects to an authenticated relay with model aliases `dgx/coding`,
|
|
101
101
|
`dgx/fast`, and `mlx/fast`; upstream gateway credentials stay in the hub.
|
|
102
102
|
Its managed tools use the existing path guards, shell sandbox and terminal
|
|
103
|
-
approval flow. Keep `ahub
|
|
103
|
+
approval flow. Keep `ahub console` open to approve write/edit/shell operations.
|
|
104
104
|
Tool-call receipts survive daemon restart; an interrupted operation with an
|
|
105
105
|
unknown outcome must be reconciled before repeating it. Cancelling a Pi turn
|
|
106
106
|
does not automatically hand its task to a cloud peer.
|
|
@@ -15,4 +15,5 @@ Scope: the Codex, ACP and Pi adapters and the Pi extension (turn `generation` an
|
|
|
15
15
|
- Native owner teardown must handle a lost shutdown acknowledgement using verified process identity. Pending launches must be revocable, and active tools/streaming output must count as watchdog activity.
|
|
16
16
|
- Pi owner identity is one shared contract, `processSignature` in `src/pi/process-signature.ts`; the startup claim check, the extension's claim, the TUI owner monitor, `recoveryReady` and the verified teardown all go through it. Its `ps` read pins LC_ALL=C and TZ=UTC (the `processTable` contract) because `lstart` follows the reader's timezone and locale, and a hub on UTC with a Pi child on system time otherwise hashes different strings for the same process and refuses the valid owner (#174). Strict pid + start + command identity stays; never accept PID-only or a fuzzy command match.
|
|
17
17
|
- Return the shared `toolResultFailed` verdict as `/tool`'s `failed`; map only `failed === true` to Pi's `isError`. Preserve text and the older bridge's absent-flag behavior.
|
|
18
|
+
- Spawn a native peer with `peerChildEnv(peer, source)` so its own `AGENTHUB_PEER_ID` replaces inherited agent markers. Sandbox tool environments carry their executing peer too. Preserve the native shell environment policy; Codex's `CODEX_THREAD_ID` supplies the fallback after filtering. Before changing marker detection or launcher environments, read `docs/verification/2026-10-09-agent-shell-t0.md` for pinned source and live-probe limits.
|
|
18
19
|
- Emit Pi's ceiling event at the pre-effect rejection boundary (`src/pi/ceiling.ts`), before `agent_end`. Accept only a valid current-session/current-generation signal during a running turn; retain the first signal and attach it to that turn's failure. Never classify a ceiling from free text.
|
package/docs/agent-notes/bus.md
CHANGED
|
@@ -1,5 +1,7 @@
|
|
|
1
1
|
# Bus, digests and replies
|
|
2
2
|
|
|
3
|
+
- Supervision replaces only pending, never attempted or accepted, envelopes with the same stable key. Preserve the first batch deadline and existing digest/queue ceilings; role/feed revocation withdraws pending notices only. Count supervision from actual transport admission followed by a successful native turn completion, never from enqueue or failure-to-idle.
|
|
4
|
+
|
|
3
5
|
Scope: delivery, digests and condensation, reply addressing, priority and limits. Read before editing `src/hub/bus.ts`, `src/hub/envelope.ts`, `src/hub/limits.ts`, `src/hub/delivery-journal.ts`, `src/hub/inference.ts` (digest condensation), or how an adapter addresses a reply or sets its priority.
|
|
4
6
|
|
|
5
7
|
- A failed digest is retried one envelope at a time, so a poison envelope cannot take its neighbours down with it.
|
|
@@ -1,7 +1,9 @@
|
|
|
1
1
|
# Task board and hub tools
|
|
2
2
|
|
|
3
|
-
Scope: the task flow, assignment, the state machine and the hub's MCP tools. Read before editing `src/hub/tasks.ts`, `src/hub/board.ts`, `src/hub/routing.ts`, `src/hub/task-sweep.ts` or `src/hub/hub-tools.ts`.
|
|
3
|
+
Scope: the task flow, assignment, the state machine and the hub's MCP tools. Read before editing `src/hub/tasks.ts`, `src/hub/board.ts`, `src/hub/routing.ts`, `src/hub/task-sweep.ts`, `src/hub/conductor.ts`, `src/hub/supervision.ts` or `src/hub/hub-tools.ts`.
|
|
4
4
|
|
|
5
5
|
- Tool callers are models: MCP `inputSchema` is not enforced on the way in. Normalize at the boundary (`cleanRefs`) before anything reaches the board, and never throw after a board write.
|
|
6
6
|
- `assign()` never defaults to the task's current owner, or a decline can only come back to the decliner.
|
|
7
7
|
- The state machine allows `in_progress -> approved` only for classes without a reviewer; `review()` checks `in_review` itself.
|
|
8
|
+
- Conductor authority requires the explicit unique role before capability checks. Assignments preserve the acting peer. A peer releases only its own conductor hold; operator release and manual, quota and recovery holds remain separate paths.
|
|
9
|
+
- Supervision reasons use structured enums, never history notes, check output or approval titles. Round signatures survive restart; later summaries count only newly joined tasks.
|
|
@@ -5,4 +5,5 @@ Scope: tests, fakes and the test gate. Read before adding or changing a test or
|
|
|
5
5
|
- Tests that are not about batching build the bus with `batchMs: 0`; with the default 15 s window a lone status envelope looks like a lost message.
|
|
6
6
|
- In Bun 1.3.14 a test timeout that fires while `Bun.spawnSync` runs can start the next test inside spawnSync's own event loop, where another spawnSync then spins at full CPU for good (trigger not pinned down; see README). Bun 1.4.2 still runs the next test inside the outer spawnSync's event loop. `scripts/check.sh` runs tests with a 20 s timeout and under `scripts/hang-watch.sh`; a test that spawns and reads the process table gets a timeout of its own.
|
|
7
7
|
- A child process gets a scrubbed environment, so test knobs for fakes travel in a wrapper script, not in `process.env`.
|
|
8
|
+
- Bind agent identity markers explicitly in native/CLI fixture children. The full gate starts the test runner without inherited agent markers for human CLI fixtures; never weaken production marker detection to make them pass.
|
|
8
9
|
- Run `bun scripts/seeded-check.ts` before reporting the seeded-guard gate complete. Keep the six `scripts/seeds.json` cases sequential; require the same named test green before its seeded assertion fails. Seed rot, survivors, setup/compiler failures, timeout/watchdog termination and process leaks fail the gate. Never run the full seed runner from a unit test or mutate the source checkout; inspect preserved failed fixtures before removing them. CI runs this in the required Linux `seeded guards` job after the ordinary check, separately from `scripts/check.sh`.
|
package/docs/events.md
CHANGED
|
@@ -30,9 +30,9 @@ marked `private: true`, and PII tasks `pii: true`.
|
|
|
30
30
|
| `split` | `task`, `where` (`routing`: routing chose the first owner of a task overlapping another owner's task not started yet, not an escalation, relay or reassignment, the record calibration reads; `cohort`: an overlap formed or changed a cohort), `verdict` (`split`, `single`, `unknown`), `single` (the peer that would finish both units alone soonest), `splitS`, `singleS`, `reason` (for `unknown`), `trace` (the inputs and steps: peer names, their profiles of versions and coordination, and numbers only): a shadow split prediction; it never changes the assignment (issue #109) |
|
|
31
31
|
| `state` | `peer`, `state` |
|
|
32
32
|
| `turn_start` | `peer`, `turn` (`<peer>#<hub run>.<n>`, unique across restarts). A turn follows the adapter: pausing a busy peer does not end it |
|
|
33
|
-
| `turn_end` | `peer`, `turn`, `ms`, `tokens` (when the adapter reported any during the turn), `files` and `snapshotMs` (when snapshots are on: how many files the turn changed, and the time both snapshots took) |
|
|
34
|
-
| `tokens` | `peer`, `n` (tokens added since the previous report) |
|
|
35
|
-
| `usage` | `peer`, `source`, opaque `id`, optional `measuredAt` (provider/source time), requested/served model and provider labels, and any provider-reported input/output/cache/total counters. Missing counters stay unknown. |
|
|
33
|
+
| `turn_end` | `peer`, `turn`, `ms`, `tokens` (when the adapter reported any during the turn), `files` and `snapshotMs` (when snapshots are on: how many files the turn changed, and the time both snapshots took); optional numeric `task`, `attribution` (`delivery`, `single_open`, `unattributed`) frozen at turn start, `pii: true` for an attributed PII task |
|
|
34
|
+
| `tokens` | `peer`, `n` (tokens added since the previous report), optional numeric `task`, `attribution` (`delivery`, `single_open`, `unattributed`), `pii: true` for an attributed PII task |
|
|
35
|
+
| `usage` | `peer`, `source`, opaque `id`, optional `measuredAt` (provider/source time), requested/served model and provider labels, and any provider-reported input/output/cache/total counters; optional numeric `task`, `attribution` (`delivery`, `single_open`, `unattributed`), `pii: true` for an attributed PII task. Missing counters stay unknown. |
|
|
36
36
|
| `task` | `id`, `event` (the board history event, e.g. `proposed`, `assigned`, `done`, `check failed`, `blocked`, `ready`), `by`, `state`, `owner`, `reviewer`, `class`, `pii` |
|
|
37
37
|
| `overlap` | `task`, `owner`, `others` (`task`, `owner`, `paths`, and `symbols` when plans name the same symbol; a name that matches a PII pattern is left out, so either list can be empty), the structured twin of the console notice |
|
|
38
38
|
| `quota` | `peer`, `windows` (`id`, `used`, `resetsAt`), `hard`, `measuredAt` (when the reading was taken, if not when it arrived: Claude's numbers come through a file) |
|
|
@@ -50,6 +50,40 @@ Token usage by adapter:
|
|
|
50
50
|
- `ahub report` deduplicates usage records by peer, source and id. Coverage counts distinguish calls with provider usage from calls where usage was absent. Token counters are provider-reported values; the report never derives a price or treats missing spend as zero. Estimated price and measured provider spend remain unknown unless a future source reports them.
|
|
51
51
|
- Usage telemetry has no prompt, completion, task text, credential, Access header, session id or transcript path.
|
|
52
52
|
|
|
53
|
+
## Per-task usage reports
|
|
54
|
+
|
|
55
|
+
`ahub report --by task` (also `--json` and `--since`) reads only `events.jsonl`.
|
|
56
|
+
It reports task ids, the latest recorded class/outcome, completed logical turns,
|
|
57
|
+
and wall time from the first `in_progress` task event to the first `approved`
|
|
58
|
+
event. Wall time is unknown if either boundary is absent or the latest outcome
|
|
59
|
+
is not approved. Each task and class rollup includes per-peer native token
|
|
60
|
+
increments and provider-reported input/output/cache/total counters. Usage is
|
|
61
|
+
deduplicated by peer, source and id; missing counters are `null` in JSON and
|
|
62
|
+
`unknown` in text, with known-record counts alongside measured subsets. A
|
|
63
|
+
missing task history has unknown class/outcome and belongs to the unknown class
|
|
64
|
+
rollup. PII tasks have ids and a `pii: true` flag, never titles. No prices are derived.
|
|
65
|
+
|
|
66
|
+
At write time, the first applicable attribution rule wins:
|
|
67
|
+
|
|
68
|
+
- `delivery`: the current turn's original delivery names exactly one distinct positive task id, including work by a peer that does not own it.
|
|
69
|
+
- `single_open`: otherwise the peer owns exactly one `in_progress` task.
|
|
70
|
+
- `unattributed`: otherwise no task is assigned to the record.
|
|
71
|
+
|
|
72
|
+
`Bus.onDeliver` observes originals immediately before `peer.deliver` starts the
|
|
73
|
+
turn. Pending task identity is consumed at turn start and cleared on delivery
|
|
74
|
+
admission/failure, so later user-started native turns do not inherit it. Usage
|
|
75
|
+
and token events apply the rule at write time; `turn_end` keeps the start-time
|
|
76
|
+
attribution. Local usage uses its request-bound route policy task when present.
|
|
77
|
+
The existing relay has no task-bearing route usage event; this change adds no
|
|
78
|
+
new usage source. Schema version 1 and the control protocol are unchanged.
|
|
79
|
+
|
|
80
|
+
The unattributed share is always printed for token increments and deduplicated
|
|
81
|
+
usage records, against all recorded increments/usage records. A zero denominator
|
|
82
|
+
has unknown share. Records with no `attribution` field belong to a separate
|
|
83
|
+
`before attribution` bucket displayed beside the share, even if a task field is
|
|
84
|
+
present. Neither bucket is redistributed. `ahub export` preserves these fields
|
|
85
|
+
as raw JSON lines; plain `ahub report` keeps its existing behavior.
|
|
86
|
+
|
|
53
87
|
The file is local and never uploaded. It grows without rotation; delete it to start
|
|
54
88
|
over (the hub recreates it).
|
|
55
89
|
|
package/docs/operations.md
CHANGED
|
@@ -1,8 +1,136 @@
|
|
|
1
1
|
# Operations guide
|
|
2
2
|
|
|
3
|
-
This guide describes ahub 0.12.
|
|
3
|
+
This guide describes ahub 0.12.18 and control protocol 15. Live verification
|
|
4
4
|
results and remaining prerequisites are recorded separately in [the smoke ledger](smoke.md).
|
|
5
5
|
|
|
6
|
+
## Operator console and panels
|
|
7
|
+
|
|
8
|
+
`ahub console` combines the existing tail stream with peer, quota, approval-age
|
|
9
|
+
and input rows. `ahub up` opens it when both input and output are terminals;
|
|
10
|
+
`--no-console` and redirected input/output retain the start-only behavior.
|
|
11
|
+
`ahub tail` remains available with its existing rendering.
|
|
12
|
+
|
|
13
|
+
Allowing an approval requires selection and a separate confirmation. Denying
|
|
14
|
+
does not. Only options the daemon supplied are selectable. Typing a command
|
|
15
|
+
prevents approval shortcuts from interpreting that input as an answer. Full
|
|
16
|
+
approval titles are terminal-only; expiry and answers from another console remove
|
|
17
|
+
the pending item. Approval audit records contain id, peer, option kind, response
|
|
18
|
+
time and answering surface, without the title.
|
|
19
|
+
|
|
20
|
+
Tab toggles stream and panels; `ahub console --panels` starts in panels. Peers,
|
|
21
|
+
Approvals, Tasks, Queue and Events support arrow keys or j/k, Enter for detail,
|
|
22
|
+
Escape to return and `?` for help. `:` enters a command. Assignment, delivery
|
|
23
|
+
resolution and allow decisions require confirmation; delivery resolution requires
|
|
24
|
+
a reason. Tasks use the same public redaction as the board. Panels need at least
|
|
25
|
+
80 columns by 24 rows; smaller terminals stay in stream mode. Task and queue
|
|
26
|
+
polling runs only while the corresponding panel is visible. Leaving restores
|
|
27
|
+
the terminal and returning from panels replays the bounded stream buffer.
|
|
28
|
+
|
|
29
|
+
Console color policy is `--color=auto|always|never`, with `auto` as the default.
|
|
30
|
+
Auto enables color only when both input and output are TTYs, `TERM` is not
|
|
31
|
+
`dumb`, and `NO_COLOR` is empty or absent. Explicit `always` overrides these
|
|
32
|
+
conditions, including redirected stream output; `never` disables styling.
|
|
33
|
+
Neither option changes terminal size requirements, panel decisions, cursor
|
|
34
|
+
management or the final reset. Invalid values fail before connecting.
|
|
35
|
+
|
|
36
|
+
| Meaning | Palette | Examples |
|
|
37
|
+
|---|---|---|
|
|
38
|
+
| Information and navigation | Cyan, bold cyan for active/selected labels | Tabs, section/peer labels and `>` selection marker |
|
|
39
|
+
| Success and availability | Green | Idle peer, approved task |
|
|
40
|
+
| Waiting and attention | Yellow | Busy/paused peer, pending approval, confirmation, review/ready task, important priority |
|
|
41
|
+
| Failure and intervention | Red | Failed/check-failed task, undeliverable/overflow, `needs_review` queue, denial request or expired/cancelled approval |
|
|
42
|
+
| Metadata | Bright black | Ages and remaining approval time |
|
|
43
|
+
|
|
44
|
+
Offline peers, primary titles, action details and message body lines keep the
|
|
45
|
+
terminal's default foreground. Labels and prompts remain readable without
|
|
46
|
+
color. Stream styling applies only to the header, using structured event data;
|
|
47
|
+
message text cannot choose a color. A local denial is shown as requested, not
|
|
48
|
+
as a confirmed receipt; a remote answered closure has no option-kind metadata
|
|
49
|
+
and is not guessed to be a denial. A fixed palette is applied after terminal
|
|
50
|
+
control sanitization. Width, clipping, wrapping and cursor placement use plain
|
|
51
|
+
Unicode text. Each styled span and every interactive exit restores attributes.
|
|
52
|
+
Color adds no polling, timers or extra redraws. `ahub tail`, logs and JSON stay
|
|
53
|
+
unchanged.
|
|
54
|
+
|
|
55
|
+
```sh
|
|
56
|
+
ahub console --panels --color=auto
|
|
57
|
+
NO_COLOR=1 ahub console
|
|
58
|
+
ahub console --color=never
|
|
59
|
+
ahub console --color=always > console-stream.txt
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
Actual light/dark terminal readability and bounded idle-CPU observations are
|
|
63
|
+
recorded separately in the smoke ledger; fake-terminal tests establish policy,
|
|
64
|
+
geometry, sanitization and restoration only.
|
|
65
|
+
|
|
66
|
+
The command input accepts existing status, board, task, review, say, pause,
|
|
67
|
+
resume, budget, queue, permit, ask, remember, route, turns, undo, check-path and
|
|
68
|
+
report operations. It executes an argument vector with closed stdin. Lifecycle,
|
|
69
|
+
launch, nested console, setup, UI, logs and tail commands are refused.
|
|
70
|
+
|
|
71
|
+
## Conducting a team from Claude Code or Codex
|
|
72
|
+
|
|
73
|
+
Start the daemon with `ahub up --no-console`, set exactly one conductor in the
|
|
74
|
+
project configuration, then open `ahub console` in a split terminal for approvals:
|
|
75
|
+
|
|
76
|
+
```json
|
|
77
|
+
{
|
|
78
|
+
"roles": { "claude": ["planner", "reviewer", "conductor"] },
|
|
79
|
+
"conductor": { "feed": "own" }
|
|
80
|
+
}
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
Launch that peer with `ahub claude` for channel pushes, or select `codex` in
|
|
84
|
+
`roles` and launch `ahub codex`. The conductor splits work into owned tasks,
|
|
85
|
+
observes `hub_status`, moves stalled work and obtains review before reporting
|
|
86
|
+
results and open decisions. It does not implement the tasks it handed out.
|
|
87
|
+
`hub_peer_start` starts local, Kimi or headless Pi; requests for native TUIs
|
|
88
|
+
return the `ahub` command for the person to run, after validating it through
|
|
89
|
+
the shared launcher planner. The wrapper plans again at launch when native
|
|
90
|
+
endpoints are available. `hub_peer_hold` and `hub_peer_release`
|
|
91
|
+
manage only holds placed by that conductor. Assignment also requires `assign`
|
|
92
|
+
when the conductor has an explicit capabilities list.
|
|
93
|
+
|
|
94
|
+
Check the returned owner and task state after assignment. Routing skips paused
|
|
95
|
+
peers, so assign work before placing a delivery hold. Hand work out through the
|
|
96
|
+
board and report to the person with `hub_send` addressed to `user`, or a `[FYI]`
|
|
97
|
+
final response in the native TUI. Broadcasting implementation instructions can
|
|
98
|
+
cause an otherwise unassigned owner to claim duplicate work.
|
|
99
|
+
|
|
100
|
+
The person answers approvals in the console, resolves `needs_review` deliveries
|
|
101
|
+
with `ahub queue resolve`, and handles budget overrides and hub lifecycle.
|
|
102
|
+
Running these commands from an agent shell is refused; the CLI retains the
|
|
103
|
+
agent's identity even when invoked through a shell tool.
|
|
104
|
+
|
|
105
|
+
`conductor.feed` is `own` by default, `all` for all tasks, or `off`. Ordinary
|
|
106
|
+
milestones share the existing digest window and collapse repeated queued
|
|
107
|
+
task/kind notices. They wait behind a busy conductor. An aged approval or
|
|
108
|
+
`needs_review` hold is important, contains only peer/tool-age or delivery id,
|
|
109
|
+
and directs the conductor to ask the person. A completed task set emits one
|
|
110
|
+
round notice until a new task joins. Removing the role or switching the feed
|
|
111
|
+
off withdraws pending feed notices.
|
|
112
|
+
|
|
113
|
+
For a Claude conductor, `ahub claude` also observes native session and turn
|
|
114
|
+
boundaries when facts injection and task-idle sweeps are off. Keep the managed
|
|
115
|
+
hooks enabled to measure completion and supervision usage. Passing your own
|
|
116
|
+
`--settings` takes precedence and produces a warning when it replaces that
|
|
117
|
+
observation. The launcher records its private session identity in ordinary
|
|
118
|
+
terminals too; this does not grant terminal-recovery authority.
|
|
119
|
+
Ordinary Claude sessions with turn-free facts or task-idle sweeps enabled get the
|
|
120
|
+
same session/start observation. Other ordinary launches remain non-opt-in.
|
|
121
|
+
|
|
122
|
+
`ahub report` records conductor actions and completed native turns containing
|
|
123
|
+
supervision. Tokens describe the whole measured turn, which may also contain
|
|
124
|
+
other work; they are not a per-notice cost estimate. Missing measurements stay
|
|
125
|
+
unknown. Use the live smoke ledger to assess observed turns and tokens per
|
|
126
|
+
approved task before choosing `all`; an unmeasured run is not a cost benchmark.
|
|
127
|
+
Claude turn counts use authenticated native completion events. Older logical
|
|
128
|
+
state counts are labelled; an idle channel or approved task alone does not prove
|
|
129
|
+
that the native answer finished.
|
|
130
|
+
The Stop hook acknowledgement is not a counted completion. The daemon checks the
|
|
131
|
+
native transcript after the hook can return; missing or changed-session evidence
|
|
132
|
+
stays unknown.
|
|
133
|
+
|
|
6
134
|
## Install and start
|
|
7
135
|
|
|
8
136
|
Use Bun 1.3 or newer. Install the released package and install its Claude
|
|
@@ -14,12 +142,11 @@ ahub setup
|
|
|
14
142
|
cd <project>
|
|
15
143
|
ahub init
|
|
16
144
|
ahub up
|
|
17
|
-
ahub tail
|
|
18
145
|
```
|
|
19
146
|
|
|
20
147
|
`ahub init` writes the project configuration and managed instruction blocks.
|
|
21
148
|
Run it after an upgrade when those blocks need refreshing. `ahub setup`
|
|
22
|
-
updates the shared Claude plugin. Keep `ahub
|
|
149
|
+
updates the shared Claude plugin. Keep `ahub console` open when a local worker
|
|
23
150
|
may request an approval.
|
|
24
151
|
|
|
25
152
|
`.agenthub/config.json` can be committed and shared. The fields that choose
|
|
@@ -256,11 +383,16 @@ token counts, never message bodies or task titles.
|
|
|
256
383
|
```bash
|
|
257
384
|
ahub report --since 7d # turns, busy time and tokens per peer, messages, overlaps, task events
|
|
258
385
|
ahub report --since 7d --json # the same numbers as JSON
|
|
386
|
+
ahub report --by task # tokens, turns and wall time per task and class, with the unattributed share
|
|
259
387
|
ahub export --since 24h # the raw events as JSON lines, for your own analysis
|
|
260
388
|
```
|
|
261
389
|
|
|
262
390
|
`ahub report` counts the same overlap warnings as `scripts/overlaps.ts`, from the
|
|
263
|
-
structured events instead of log lines.
|
|
391
|
+
structured events instead of log lines. `--by task` uses the task each usage and
|
|
392
|
+
token record was attributed to when it was written: the delivery that started the
|
|
393
|
+
turn, otherwise the peer's only `in_progress` task. Everything else is reported as
|
|
394
|
+
unattributed, and records from before 0.12.18 as a separate bucket; neither is
|
|
395
|
+
redistributed (rules: [events](events.md#per-task-usage-reports)).
|
|
264
396
|
|
|
265
397
|
## Turns and undo
|
|
266
398
|
|
|
@@ -521,6 +653,13 @@ stays the default until an evaluation says otherwise (`docs/cooperbench.md`).
|
|
|
521
653
|
|
|
522
654
|
Inspect permission requests in the terminal:
|
|
523
655
|
|
|
656
|
+
```bash
|
|
657
|
+
ahub console
|
|
658
|
+
```
|
|
659
|
+
|
|
660
|
+
Select the requested option in the console and confirm an allow decision.
|
|
661
|
+
For a separate plain terminal, the existing command remains available:
|
|
662
|
+
|
|
524
663
|
```bash
|
|
525
664
|
ahub tail
|
|
526
665
|
ahub permit <request-id> allow
|
|
@@ -647,21 +786,21 @@ Rows without a live process are stale registrations; forget them with
|
|
|
647
786
|
|
|
648
787
|
Upgrade running projects with the target release's own coordinator. It accepts
|
|
649
788
|
a running source on control protocol 9 (0.6.x), 10 (0.7.0 through 0.12.0),
|
|
650
|
-
11 (0.12.1 and 0.12.2), 12 (0.12.3), 13 (0.12.4 through 0.12.15)
|
|
789
|
+
11 (0.12.1 and 0.12.2), 12 (0.12.3), 13 (0.12.4 through 0.12.15) 14 (0.12.16) or 15 (0.12.17 and 0.12.18), and only
|
|
651
790
|
a target on its own protocol, so the target's coordinator fits every supported
|
|
652
791
|
source and carries every recovery fix released up to it. Protocol 8 and older
|
|
653
792
|
(0.5.x and earlier) are refused as `manual-bootstrap-required`. Run from the
|
|
654
793
|
project directory, without replacing the global CLI first:
|
|
655
794
|
|
|
656
795
|
```bash
|
|
657
|
-
bunx --package @staix/agent-hub@0.12.
|
|
658
|
-
bunx --package @staix/agent-hub@0.12.
|
|
796
|
+
bunx --package @staix/agent-hub@0.12.18 ahub upgrade --to 0.12.18 --dry-run
|
|
797
|
+
bunx --package @staix/agent-hub@0.12.18 ahub upgrade --to 0.12.18 --yes
|
|
659
798
|
```
|
|
660
799
|
|
|
661
800
|
| Running now | Coordinator to use |
|
|
662
801
|
| --- | --- |
|
|
663
802
|
| 0.6.x (protocol 9) | the target's, through `bunx` as above |
|
|
664
|
-
| 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.15 (protocol 13), 0.12.16 (protocol 14) | the target's, through `bunx` as above |
|
|
803
|
+
| 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.15 (protocol 13), 0.12.16 (protocol 14), 0.12.17 and 0.12.18 (protocol 15) | the target's, through `bunx` as above |
|
|
665
804
|
| any supported source, with the installed CLI already at the target | `ahub upgrade` below, which is the same coordinator |
|
|
666
805
|
| 0.5.x or earlier (protocol 8 and older) | not supported: bootstrap by hand with the matching CLI |
|
|
667
806
|
|
|
@@ -693,14 +832,14 @@ projects first:
|
|
|
693
832
|
|
|
694
833
|
```bash
|
|
695
834
|
ahub restart --dry-run
|
|
696
|
-
ahub upgrade --to 0.12.
|
|
835
|
+
ahub upgrade --to 0.12.18 --dry-run
|
|
697
836
|
```
|
|
698
837
|
|
|
699
838
|
Apply only after reviewing the plan:
|
|
700
839
|
|
|
701
840
|
```bash
|
|
702
841
|
ahub restart --yes
|
|
703
|
-
ahub upgrade --to 0.12.
|
|
842
|
+
ahub upgrade --to 0.12.18 --yes
|
|
704
843
|
ahub recovery status <operation-id>
|
|
705
844
|
ahub recovery resume <operation-id>
|
|
706
845
|
ahub recovery abort <operation-id>
|
package/docs/quickstart.md
CHANGED
|
@@ -17,7 +17,7 @@ ahub setup # installs the Claude Code channel plu
|
|
|
17
17
|
Or use the matching GitHub release:
|
|
18
18
|
|
|
19
19
|
```bash
|
|
20
|
-
bun add -g github:STAIxBWLB/agent-hub#v0.12.
|
|
20
|
+
bun add -g github:STAIxBWLB/agent-hub#v0.12.18
|
|
21
21
|
ahub setup
|
|
22
22
|
```
|
|
23
23
|
|
|
@@ -34,8 +34,8 @@ In your project directory, one terminal each:
|
|
|
34
34
|
|
|
35
35
|
```bash
|
|
36
36
|
ahub init # .agenthub/config.json, .agenthub/routing.toml, the marker block in AGENTS.md
|
|
37
|
-
ahub up # the daemon
|
|
38
|
-
ahub
|
|
37
|
+
ahub up # starts the daemon and opens the operator console in a terminal
|
|
38
|
+
ahub console # use in another terminal if up used --no-console; Tab opens panels
|
|
39
39
|
ahub kimi # Kimi, headless
|
|
40
40
|
ahub codex # Codex TUI, attached through the hub
|
|
41
41
|
ahub claude # Claude Code with the hub channel
|
package/docs/security.md
CHANGED
|
@@ -8,7 +8,7 @@ agent-hub connects agents that can each run commands. This page says what the hu
|
|
|
8
8
|
- **The control link is loopback plus a secret.** The daemon and the Codex proxy bind 127.0.0.1 only. The control WebSocket requires a per-run token (`.agenthub/state/control-token`, mode 600), and both servers refuse any request that carries an `Origin` header: any web page can open a WebSocket to localhost, and browsers always send `Origin`. External clients cannot claim the console user's id or a hub-managed peer's id.
|
|
9
9
|
- **Permission prompts stay on.** `ahub claude` and `ahub codex` add nothing that weakens the agents' own prompts. Kimi's and `local`'s permission requests are relayed to the console and cancelled after `approvals.timeout_s` (default 120 s) of silence; the macOS notification for a waiting request carries the peer and the tool name only. The one exception is the hub's own tools (`hub_send` and the task tools, matched by exact name): Kimi's requests for them are approved once without a prompt and logged by name, the same trust Codex gets through `approval_mode` in the hub's config. They invoke hub-owned operations rather than arbitrary file or shell tools, and every call passes the hub's own checks. Identity comes from the exact permission title, or from the earlier tool-call title bound to the same call id and resolved against the configured MCP servers when the permission title contains argument JSON (Qwen). Payload text and unrelated display titles never establish identity. A payload longer than the console shows is marked as cut and never offers a session-wide grant. `--unattended` turns prompts off, says so loudly, and is never the default.
|
|
10
10
|
- **A committed config cannot choose launch commands, credential files, data endpoints or a wider sandbox.** The machine-local fields (`kimi_cmd`, `codex_bin`, `pi.cmd`, `checks`, `mlx.bin`, `mlx.runtimeDir`, `mlx.modelPath`, `omniroute.urls`, `omniroute.access_hosts`, the `omniroute` key files, `memory.worker_url`, `local.read_allow`, `local.bash_network`, `local.network_allow`) apply only from a config file git confirms nobody committed: `.agenthub/config.json` or `.agenthub/config.local.json`, matched by file identity so no other spelling the file system accepts slips past, and `.agenthub` itself not a committed symlink or submodule. Without a repository, or when git fails, they keep their defaults; an empty value always means the default. "Untracked" is answered by the repository that contains the project: a checkout copied or extracted into an unrelated repository, or into an ignored directory of one, is trusted like your own files. So a cloned repository cannot choose a launch command, a completion check, a gateway to send a key file to, a memory endpoint, or a wider sandbox. A command in `checks` runs as you, outside the local worker's sandbox, like a git hook. Nothing in `routing.toml` or in task text is ever run. `routing.toml` and the other shared fields still come from the checkout, and they matter: `routing.toml` picks the models the local worker and the hub's inference use at your gateway and can turn the PII constraint off, and roles and budget shape who does what. Review them in a repository you do not trust.
|
|
11
|
-
- **Telemetry holds no bodies.** `.agenthub/state/events.jsonl` (issue #40) records envelope ids, routing and sizes, task ids and states, overlapping paths and token counts. It never records a message body, a task title or detail, and marks private (PII) envelopes and tasks as such. It stays on the machine; `ahub export` only prints it.
|
|
11
|
+
- **Telemetry holds no bodies.** `.agenthub/state/events.jsonl` (issue #40) records envelope ids, routing and sizes, task ids and states, overlapping paths and token counts, and the task id (with a PII flag) that each usage and token record is attributed to. It never records a message body, a task title or detail, and marks private (PII) envelopes and tasks as such. It stays on the machine; `ahub export` only prints it.
|
|
12
12
|
- **Snapshots stay in your repository.** Per-turn snapshots (issue #33) are git objects in the project's own object store, written through a temporary index; nothing is referenced, pushed or copied elsewhere, and `git gc` prunes them. They hold what the work tree held, including untracked files that are not ignored, so keep secrets in ignored files. They carry the repository's own permissions, and nothing caps their disk use but `git gc`. A turn of a peer holding an open PII task is not snapshotted; a PII file left in the project is snapshotted by later turns like any other file. `ahub undo` restores only files whose current content is exactly what the turn left.
|
|
13
13
|
- **The edit hook reads, never decides.** `ahub check-path --hook` (issue #32) reads hub.db and returns context for Claude and a line for you; it sets no permission decision, so your permission rules stay in charge. It names other owners' task ids, titles and states, which then reach Claude's model; PII tasks are left out.
|
|
14
14
|
- **The session record holds identities only.** `.agenthub/state/sessions.json` (issue #37, mode 600) keeps each attached peer's recovery metadata: launch options, session and thread ids, Pi's session file path. No message or task text; loss notices name deliveries by id, sender and public task title.
|
|
@@ -28,6 +28,33 @@ agent-hub connects agents that can each run commands. This page says what the hu
|
|
|
28
28
|
- **Network-level proof for PII.** The hub proves at its own boundaries (tests search every output) that PII text does not leave; packet-level verification of your gateway path is an operations check.
|
|
29
29
|
- **Other operating systems' sandboxes.** Only macOS seatbelt is implemented.
|
|
30
30
|
|
|
31
|
+
## Agent CLI identity and conductor authority
|
|
32
|
+
|
|
33
|
+
Inside an agent session, `ahub` identifies the caller from `AGENTHUB_PEER_ID`,
|
|
34
|
+
or the native Claude/Codex shell marker when the hub marker is absent. It connects
|
|
35
|
+
as that peer in tools mode. Messages and task changes retain that actor;
|
|
36
|
+
`say` defaults to status priority and important messages still require the peer's
|
|
37
|
+
capability. Conflicting or malformed markers fail closed. There is no `--as-user`.
|
|
38
|
+
Human-only operations are refused before connecting, including permission answers,
|
|
39
|
+
queue resolution, budget overrides, lifecycle and recovery operations, and `ask`.
|
|
40
|
+
The operator uses a plain terminal or `ahub console`. A Claude `!` command that
|
|
41
|
+
inherits the agent markers follows the same rule. The console strips control
|
|
42
|
+
sequences from agent and daemon text before it quotes forged headers, and its
|
|
43
|
+
colors come only from a fixed palette applied afterwards, so message text cannot
|
|
44
|
+
set styles or operate the terminal.
|
|
45
|
+
|
|
46
|
+
This is an honest default against accidental impersonation and injected commands.
|
|
47
|
+
It does not make the token inaccessible to an agent with unrestricted project
|
|
48
|
+
shell access. The existing OS sandbox and loopback authentication boundaries apply.
|
|
49
|
+
|
|
50
|
+
Steering tools require an explicit `conductor` role, independent of default-allow
|
|
51
|
+
capabilities. Only one peer may hold that role. A conductor may inspect public
|
|
52
|
+
state, assign or escalate work, start supported headless peers and place its own
|
|
53
|
+
delivery holds. It cannot answer approvals, resolve durable deliveries, override
|
|
54
|
+
budget pauses or release a human hold. Pending approval summaries exclude titles;
|
|
55
|
+
PII tasks remain public stubs. Conduct events contain ids only. Supervision feeds
|
|
56
|
+
use structured reasons and never carry check output or approval bodies.
|
|
57
|
+
|
|
31
58
|
## Reporting
|
|
32
59
|
|
|
33
60
|
Please report vulnerabilities privately through GitHub's "Report a vulnerability" on this repository rather than in a public issue.
|
package/docs/smoke.md
CHANGED
|
@@ -2,6 +2,99 @@
|
|
|
2
2
|
|
|
3
3
|
`scripts/check.sh` covers everything against fakes. The legs below need real accounts and an interactive terminal, so they are run by hand and recorded here.
|
|
4
4
|
|
|
5
|
+
## Operator console and conductor candidate (#190, #191, #193-#195)
|
|
6
|
+
|
|
7
|
+
Candidate 0.12.17, protocol 15, observed on 2026-10-09 KST. The final Claude
|
|
8
|
+
observer runs used immutable source `20ad5e6`; their native completion receipts
|
|
9
|
+
and actual console effects were checked independently. The Kimi new-tool leg
|
|
10
|
+
remains quota-blocked, as recorded below.
|
|
11
|
+
|
|
12
|
+
- Actual Codex 0.146.0 (`gpt-5.5`) and Claude 2.1.295 (Opus 5.5) TUIs started
|
|
13
|
+
real local and headless Pi peers, proposed exactly two tasks owned initially
|
|
14
|
+
by local, reassigned Beta to Pi, and placed and released their own holds.
|
|
15
|
+
The real owners checked their outputs and called `hub_task_done`; the native
|
|
16
|
+
conductor independently read the exact files before approving both tasks.
|
|
17
|
+
Real CLI calls retained the agent actor, and human-only queue/permission
|
|
18
|
+
actions were refused from an agent shell.
|
|
19
|
+
- The person authorized `alpha.txt = ALPHA` (5 bytes) and `beta.txt = BETA`
|
|
20
|
+
(4 bytes), both without a newline. Separate selection and confirmation keys
|
|
21
|
+
reached the actual console PTY and produced allow-once console audit records.
|
|
22
|
+
Out-of-scope directory-list commands were cancelled. The final Claude own
|
|
23
|
+
run used three allow-once answers because Pi combined its exact write and
|
|
24
|
+
bounded byte checks in one approved command; it also recorded two cancelled
|
|
25
|
+
local directory-list requests. No permission or task completion was fabricated.
|
|
26
|
+
- The final Claude own run recorded nine completed native turns and nine
|
|
27
|
+
matching daemon Stop records, 27 unique native usage records, and
|
|
28
|
+
1,888,242 tokens including cached input. Eight completed native turns contained
|
|
29
|
+
supervision, with 1,292,922 recorded tokens. Its final message UUID/id,
|
|
30
|
+
`end_turn`, `turn_duration`, instance/session/launcher and opaque Stop receipt
|
|
31
|
+
matched; native idle and an empty pending delivery queue were verified before
|
|
32
|
+
owned-process cleanup. Local recorded three turns and 72,046 gateway tokens;
|
|
33
|
+
Pi recorded two turns and 47,177 native tokens.
|
|
34
|
+
- A read-only Claude feed-off continuation on the final observer preserved the
|
|
35
|
+
existing approved tasks and exact files. It verified one completed native turn,
|
|
36
|
+
four usage records and 263,413 cached-inclusive tokens, with matching daemon
|
|
37
|
+
Stop, native idle and delivery settlement. An obsolete review hold was first
|
|
38
|
+
discarded through the authenticated public operator API after fresh approved
|
|
39
|
+
task/history and exact BETA readback. That discard is not native execution.
|
|
40
|
+
- A real console with two pending cards and no operator input consumed 0.07 CPU
|
|
41
|
+
seconds over 40.018 wall seconds in stream mode (0.175% of one CPU). A separate
|
|
42
|
+
actual Peers panel, 120x40, with two pending cards and no operator input used
|
|
43
|
+
0.07 CPU seconds over 40.012 wall seconds (0.175%). These are cumulative `ps`
|
|
44
|
+
process CPU samples in distinct modes/runs, not whole-host idle measurements.
|
|
45
|
+
- In a plain Claude TUI without the development-channel flag, one actual
|
|
46
|
+
`hub_status` MCP call succeeded and the daemon audited the Claude status
|
|
47
|
+
action. A unique directed operator push was accepted by the bridge, but no
|
|
48
|
+
pushed native user row or answer appeared in the observed 20.075-second window.
|
|
49
|
+
This proves tool access separately from the bounded negative push observation.
|
|
50
|
+
The real channel-enabled runs above received actual review and supervision
|
|
51
|
+
pushes. The plain probe and cleanup completed in 56.724 seconds, with no model
|
|
52
|
+
file, shell, task or settings actions.
|
|
53
|
+
- Real Kimi Code CLI 2.1.1 completed ACP initialization and `session/new`, then
|
|
54
|
+
rejected the single prompt with HTTP 403 for its weekly account usage limit.
|
|
55
|
+
No requested native tool event or role-refusal result occurred. The reset time
|
|
56
|
+
was not supplied. The owning CLI listed only the managed OAuth Kimi provider
|
|
57
|
+
(four models, default `kimi-code/k3`), so no configured alternative provider was
|
|
58
|
+
found. This is an incomplete external prerequisite, not a new-tool pass. No
|
|
59
|
+
purchase, provider/configuration change or model retry was performed; owned
|
|
60
|
+
processes were stopped after 3.142 seconds.
|
|
61
|
+
|
|
62
|
+
The user explicitly approved release 0.12.17 with only the Kimi native new-tool
|
|
63
|
+
and ordinary-role refusal T0 evidence deferred on 2026-10-09. That native
|
|
64
|
+
prerequisite remains unverified; the quota rejection is not a tool pass. This
|
|
65
|
+
decision does not defer the other native/source gates or authorize a purchase,
|
|
66
|
+
credentials, provider/configuration change or account retry. The same isolated
|
|
67
|
+
read-only Kimi probe remains required before its native coverage is marked verified.
|
|
68
|
+
|
|
69
|
+
Earlier captures remain part of the evidence:
|
|
70
|
+
|
|
71
|
+
- Codex feed-off included an interrupted original and read-only continuation:
|
|
72
|
+
six logical native turns, 49 increments and 1,979,016 tokens. Codex feed-own
|
|
73
|
+
recorded nine turns, 34 increments and 1,399,470 tokens. Their journals were
|
|
74
|
+
independently checked after cleanup. The first baseline interruption and
|
|
75
|
+
operator reconciliation are preserved; neither native task completion nor
|
|
76
|
+
board approval was substituted by the harness.
|
|
77
|
+
- The original Claude runs approved the actual files but an early idle-based
|
|
78
|
+
harness stopped their final review answers. Subsequent `da2ebbc` feed-off
|
|
79
|
+
completed five native turns and 1,451,782 tokens but certified only one daemon
|
|
80
|
+
Stop. A `6034dd9` read-only continuation completed another native turn and
|
|
81
|
+
261,483 tokens, while its Stop remained unavailable during the native hook.
|
|
82
|
+
Both incomplete observer captures are preserved alongside raw and audited
|
|
83
|
+
summaries. The final post-ACK observer above closes the fresh completion proof;
|
|
84
|
+
it does not retroactively certify the earlier missed Stop records.
|
|
85
|
+
|
|
86
|
+
All token totals describe whole native/model turns, including cached input and
|
|
87
|
+
other work. Operator waiting, cancelled requests, marker probes, interrupted
|
|
88
|
+
continuations and differing source revisions make these observations unsuitable
|
|
89
|
+
for a causal feed-overhead or model-efficiency comparison. Unknown measurements
|
|
90
|
+
and the Kimi prerequisite remain explicit.
|
|
91
|
+
|
|
92
|
+
The harness uses real Python PTYs. Operator-file-input forwards only
|
|
93
|
+
chat-authorized keys, removes inherited Orca terminal ownership from fixture
|
|
94
|
+
children, and generates no approval automatically. Private original and
|
|
95
|
+
continuation captures remain separate. The shell-marker probes and vendor
|
|
96
|
+
limits are recorded in [the identity T0 ledger](verification/2026-10-09-agent-shell-t0.md).
|
|
97
|
+
|
|
5
98
|
## Approval race live reproduction and candidate verification (#98)
|
|
6
99
|
|
|
7
100
|
Measured on 2026-10-02 KST with installed 0.12.1 and the correction candidate,
|