@staix/agent-hub 0.12.16 → 0.12.18

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/CHANGELOG.md +14 -0
  2. package/README.md +4 -4
  3. package/docs/agent-notes/adapters.md +1 -0
  4. package/docs/agent-notes/bus.md +2 -0
  5. package/docs/agent-notes/tasks.md +3 -1
  6. package/docs/agent-notes/tests.md +1 -0
  7. package/docs/events.md +37 -3
  8. package/docs/operations.md +149 -10
  9. package/docs/quickstart.md +3 -3
  10. package/docs/security.md +28 -1
  11. package/docs/smoke.md +93 -0
  12. package/docs/specs/2026-09-19-agent-hub-design.md +176 -1
  13. package/docs/verification/2026-10-09-agent-shell-t0.md +123 -0
  14. package/docs/verified.json +26 -17
  15. package/package.json +1 -1
  16. package/plugins/agent-hub/.claude-plugin/plugin.json +1 -1
  17. package/plugins/agent-hub/server.js +17 -5
  18. package/src/adapters/acp.ts +2 -2
  19. package/src/adapters/claude-channel.ts +3 -2
  20. package/src/adapters/codex-appserver.ts +3 -2
  21. package/src/adapters/local-worker.ts +16 -12
  22. package/src/cli/console-state.ts +279 -0
  23. package/src/cli/console.ts +224 -0
  24. package/src/cli/facts-hook.ts +8 -1
  25. package/src/cli/identity-audit.ts +66 -0
  26. package/src/cli/identity.ts +54 -0
  27. package/src/cli/launch.ts +15 -2
  28. package/src/cli/main.ts +75 -32
  29. package/src/cli/tail-render.ts +17 -0
  30. package/src/cli/upgrade-runtime.ts +1 -1
  31. package/src/hub/attribution.ts +20 -0
  32. package/src/hub/board.ts +6 -2
  33. package/src/hub/bus.ts +88 -10
  34. package/src/hub/child-process.ts +11 -0
  35. package/src/hub/conductor.ts +196 -0
  36. package/src/hub/control-client.ts +3 -3
  37. package/src/hub/daemon.ts +388 -51
  38. package/src/hub/envelope.ts +1 -1
  39. package/src/hub/events.ts +10 -4
  40. package/src/hub/hub-tools.ts +13 -0
  41. package/src/hub/report.ts +162 -4
  42. package/src/hub/supervision.ts +152 -0
  43. package/src/hub/tasks.ts +30 -15
  44. package/src/hub/usage.ts +37 -2
  45. package/src/pi/launch.ts +2 -1
package/CHANGELOG.md CHANGED
@@ -4,6 +4,20 @@ Issue and pull request numbers in the entries for 0.7.7 and earlier refer to the
4
4
 
5
5
  ## Unreleased
6
6
 
7
+ ## 0.12.18
8
+
9
+ - Attribute token increments, provider usage and completed turns to task ids by delivery or a single open task; add `ahub report --by task` with task/class totals, explicit unknown counters, unattributed shares and a separate historical bucket. Preserve PII id-only exports and derive no prices (#200).
10
+ - Add semantic colors to console stream headers, footer and panels with `--color=auto|always|never`; keep plain-text Unicode geometry, sanitize incoming controls before the fixed palette, and restore terminal attributes on exit. Tail, logs, JSON and polling remain unchanged (#201).
11
+ - Control protocol stays 15 and events schema stays 1; a running 0.12.17 hub upgrades with the 0.12.18 coordinator.
12
+
13
+ ## 0.12.17
14
+
15
+ - Add an operator console with confirmed approvals, bounded status polling and optional peer, approval, task, queue and event panels; open it from interactive `up` and preserve the tail renderer (#190, #191).
16
+ - Identify agent shell CLI calls as their peer, refuse human-only commands before connecting and record ids-only audit notices without an as-user bypass (#193).
17
+ - Grant steering tools only to one explicit conductor role, preserve the actor on task changes and keep conductor holds separate from human, budget and recovery holds (#194).
18
+ - Batch public task milestones and human-action reminders through the existing digest window, replace repeated pending notices and report completed native supervision turns with unknown usage preserved (#195).
19
+ - Advance the control protocol to 15 for approval lifecycle and supervision metadata, retaining controlled recovery from supported previous sources.
20
+
7
21
  ## 0.12.16
8
22
 
9
23
  - Record source-verification commits for current operating documents and agent notes; fail missing coverage, invalid stamps and README version drift, and report stale source scopes deterministically (#188).
package/README.md CHANGED
@@ -4,7 +4,7 @@ Native multi-agent hub for one developer's machine: Claude Code, Codex, Kimi Cod
4
4
  hub-owned local-LLM worker collaborate as peers in independent project directories, with
5
5
  task-aware model routing (Switchyard) in front of a self-hosted gateway (OmniRoute).
6
6
 
7
- Status: 0.12.16, control protocol 14. Durable delivery records distinguish queued
7
+ Status: 0.12.18, control protocol 15. Durable delivery records distinguish queued
8
8
  work from uncertain execution. The [smoke checklist](docs/smoke.md) records
9
9
  verified paths and remaining prerequisites.
10
10
 
@@ -22,13 +22,13 @@ Install from npm:
22
22
 
23
23
  ```bash
24
24
  bun add -g @staix/agent-hub && ahub setup
25
- cd <your project> && ahub init && ahub up && ahub tail
25
+ cd <your project> && ahub init && ahub up
26
26
  ```
27
27
 
28
28
  Or install the same version from GitHub:
29
29
 
30
30
  ```bash
31
- bun add -g github:STAIxBWLB/agent-hub#v0.12.16 && ahub setup
31
+ bun add -g github:STAIxBWLB/agent-hub#v0.12.18 && ahub setup
32
32
  ```
33
33
 
34
34
  The installed commands remain `ahub` and `agent-hub`.
@@ -100,7 +100,7 @@ only. `--backend dgx` or `--backend mlx` explicitly pins a Pi session's backend.
100
100
  Pi connects to an authenticated relay with model aliases `dgx/coding`,
101
101
  `dgx/fast`, and `mlx/fast`; upstream gateway credentials stay in the hub.
102
102
  Its managed tools use the existing path guards, shell sandbox and terminal
103
- approval flow. Keep `ahub tail` open to approve write/edit/shell operations.
103
+ approval flow. Keep `ahub console` open to approve write/edit/shell operations.
104
104
  Tool-call receipts survive daemon restart; an interrupted operation with an
105
105
  unknown outcome must be reconciled before repeating it. Cancelling a Pi turn
106
106
  does not automatically hand its task to a cloud peer.
@@ -15,4 +15,5 @@ Scope: the Codex, ACP and Pi adapters and the Pi extension (turn `generation` an
15
15
  - Native owner teardown must handle a lost shutdown acknowledgement using verified process identity. Pending launches must be revocable, and active tools/streaming output must count as watchdog activity.
16
16
  - Pi owner identity is one shared contract, `processSignature` in `src/pi/process-signature.ts`; the startup claim check, the extension's claim, the TUI owner monitor, `recoveryReady` and the verified teardown all go through it. Its `ps` read pins LC_ALL=C and TZ=UTC (the `processTable` contract) because `lstart` follows the reader's timezone and locale, and a hub on UTC with a Pi child on system time otherwise hashes different strings for the same process and refuses the valid owner (#174). Strict pid + start + command identity stays; never accept PID-only or a fuzzy command match.
17
17
  - Return the shared `toolResultFailed` verdict as `/tool`'s `failed`; map only `failed === true` to Pi's `isError`. Preserve text and the older bridge's absent-flag behavior.
18
+ - Spawn a native peer with `peerChildEnv(peer, source)` so its own `AGENTHUB_PEER_ID` replaces inherited agent markers. Sandbox tool environments carry their executing peer too. Preserve the native shell environment policy; Codex's `CODEX_THREAD_ID` supplies the fallback after filtering. Before changing marker detection or launcher environments, read `docs/verification/2026-10-09-agent-shell-t0.md` for pinned source and live-probe limits.
18
19
  - Emit Pi's ceiling event at the pre-effect rejection boundary (`src/pi/ceiling.ts`), before `agent_end`. Accept only a valid current-session/current-generation signal during a running turn; retain the first signal and attach it to that turn's failure. Never classify a ceiling from free text.
@@ -1,5 +1,7 @@
1
1
  # Bus, digests and replies
2
2
 
3
+ - Supervision replaces only pending, never attempted or accepted, envelopes with the same stable key. Preserve the first batch deadline and existing digest/queue ceilings; role/feed revocation withdraws pending notices only. Count supervision from actual transport admission followed by a successful native turn completion, never from enqueue or failure-to-idle.
4
+
3
5
  Scope: delivery, digests and condensation, reply addressing, priority and limits. Read before editing `src/hub/bus.ts`, `src/hub/envelope.ts`, `src/hub/limits.ts`, `src/hub/delivery-journal.ts`, `src/hub/inference.ts` (digest condensation), or how an adapter addresses a reply or sets its priority.
4
6
 
5
7
  - A failed digest is retried one envelope at a time, so a poison envelope cannot take its neighbours down with it.
@@ -1,7 +1,9 @@
1
1
  # Task board and hub tools
2
2
 
3
- Scope: the task flow, assignment, the state machine and the hub's MCP tools. Read before editing `src/hub/tasks.ts`, `src/hub/board.ts`, `src/hub/routing.ts`, `src/hub/task-sweep.ts` or `src/hub/hub-tools.ts`.
3
+ Scope: the task flow, assignment, the state machine and the hub's MCP tools. Read before editing `src/hub/tasks.ts`, `src/hub/board.ts`, `src/hub/routing.ts`, `src/hub/task-sweep.ts`, `src/hub/conductor.ts`, `src/hub/supervision.ts` or `src/hub/hub-tools.ts`.
4
4
 
5
5
  - Tool callers are models: MCP `inputSchema` is not enforced on the way in. Normalize at the boundary (`cleanRefs`) before anything reaches the board, and never throw after a board write.
6
6
  - `assign()` never defaults to the task's current owner, or a decline can only come back to the decliner.
7
7
  - The state machine allows `in_progress -> approved` only for classes without a reviewer; `review()` checks `in_review` itself.
8
+ - Conductor authority requires the explicit unique role before capability checks. Assignments preserve the acting peer. A peer releases only its own conductor hold; operator release and manual, quota and recovery holds remain separate paths.
9
+ - Supervision reasons use structured enums, never history notes, check output or approval titles. Round signatures survive restart; later summaries count only newly joined tasks.
@@ -5,4 +5,5 @@ Scope: tests, fakes and the test gate. Read before adding or changing a test or
5
5
  - Tests that are not about batching build the bus with `batchMs: 0`; with the default 15 s window a lone status envelope looks like a lost message.
6
6
  - In Bun 1.3.14 a test timeout that fires while `Bun.spawnSync` runs can start the next test inside spawnSync's own event loop, where another spawnSync then spins at full CPU for good (trigger not pinned down; see README). Bun 1.4.2 still runs the next test inside the outer spawnSync's event loop. `scripts/check.sh` runs tests with a 20 s timeout and under `scripts/hang-watch.sh`; a test that spawns and reads the process table gets a timeout of its own.
7
7
  - A child process gets a scrubbed environment, so test knobs for fakes travel in a wrapper script, not in `process.env`.
8
+ - Bind agent identity markers explicitly in native/CLI fixture children. The full gate starts the test runner without inherited agent markers for human CLI fixtures; never weaken production marker detection to make them pass.
8
9
  - Run `bun scripts/seeded-check.ts` before reporting the seeded-guard gate complete. Keep the six `scripts/seeds.json` cases sequential; require the same named test green before its seeded assertion fails. Seed rot, survivors, setup/compiler failures, timeout/watchdog termination and process leaks fail the gate. Never run the full seed runner from a unit test or mutate the source checkout; inspect preserved failed fixtures before removing them. CI runs this in the required Linux `seeded guards` job after the ordinary check, separately from `scripts/check.sh`.
package/docs/events.md CHANGED
@@ -30,9 +30,9 @@ marked `private: true`, and PII tasks `pii: true`.
30
30
  | `split` | `task`, `where` (`routing`: routing chose the first owner of a task overlapping another owner's task not started yet, not an escalation, relay or reassignment, the record calibration reads; `cohort`: an overlap formed or changed a cohort), `verdict` (`split`, `single`, `unknown`), `single` (the peer that would finish both units alone soonest), `splitS`, `singleS`, `reason` (for `unknown`), `trace` (the inputs and steps: peer names, their profiles of versions and coordination, and numbers only): a shadow split prediction; it never changes the assignment (issue #109) |
31
31
  | `state` | `peer`, `state` |
32
32
  | `turn_start` | `peer`, `turn` (`<peer>#<hub run>.<n>`, unique across restarts). A turn follows the adapter: pausing a busy peer does not end it |
33
- | `turn_end` | `peer`, `turn`, `ms`, `tokens` (when the adapter reported any during the turn), `files` and `snapshotMs` (when snapshots are on: how many files the turn changed, and the time both snapshots took) |
34
- | `tokens` | `peer`, `n` (tokens added since the previous report) |
35
- | `usage` | `peer`, `source`, opaque `id`, optional `measuredAt` (provider/source time), requested/served model and provider labels, and any provider-reported input/output/cache/total counters. Missing counters stay unknown. |
33
+ | `turn_end` | `peer`, `turn`, `ms`, `tokens` (when the adapter reported any during the turn), `files` and `snapshotMs` (when snapshots are on: how many files the turn changed, and the time both snapshots took); optional numeric `task`, `attribution` (`delivery`, `single_open`, `unattributed`) frozen at turn start, `pii: true` for an attributed PII task |
34
+ | `tokens` | `peer`, `n` (tokens added since the previous report), optional numeric `task`, `attribution` (`delivery`, `single_open`, `unattributed`), `pii: true` for an attributed PII task |
35
+ | `usage` | `peer`, `source`, opaque `id`, optional `measuredAt` (provider/source time), requested/served model and provider labels, and any provider-reported input/output/cache/total counters; optional numeric `task`, `attribution` (`delivery`, `single_open`, `unattributed`), `pii: true` for an attributed PII task. Missing counters stay unknown. |
36
36
  | `task` | `id`, `event` (the board history event, e.g. `proposed`, `assigned`, `done`, `check failed`, `blocked`, `ready`), `by`, `state`, `owner`, `reviewer`, `class`, `pii` |
37
37
  | `overlap` | `task`, `owner`, `others` (`task`, `owner`, `paths`, and `symbols` when plans name the same symbol; a name that matches a PII pattern is left out, so either list can be empty), the structured twin of the console notice |
38
38
  | `quota` | `peer`, `windows` (`id`, `used`, `resetsAt`), `hard`, `measuredAt` (when the reading was taken, if not when it arrived: Claude's numbers come through a file) |
@@ -50,6 +50,40 @@ Token usage by adapter:
50
50
  - `ahub report` deduplicates usage records by peer, source and id. Coverage counts distinguish calls with provider usage from calls where usage was absent. Token counters are provider-reported values; the report never derives a price or treats missing spend as zero. Estimated price and measured provider spend remain unknown unless a future source reports them.
51
51
  - Usage telemetry has no prompt, completion, task text, credential, Access header, session id or transcript path.
52
52
 
53
+ ## Per-task usage reports
54
+
55
+ `ahub report --by task` (also `--json` and `--since`) reads only `events.jsonl`.
56
+ It reports task ids, the latest recorded class/outcome, completed logical turns,
57
+ and wall time from the first `in_progress` task event to the first `approved`
58
+ event. Wall time is unknown if either boundary is absent or the latest outcome
59
+ is not approved. Each task and class rollup includes per-peer native token
60
+ increments and provider-reported input/output/cache/total counters. Usage is
61
+ deduplicated by peer, source and id; missing counters are `null` in JSON and
62
+ `unknown` in text, with known-record counts alongside measured subsets. A
63
+ missing task history has unknown class/outcome and belongs to the unknown class
64
+ rollup. PII tasks have ids and a `pii: true` flag, never titles. No prices are derived.
65
+
66
+ At write time, the first applicable attribution rule wins:
67
+
68
+ - `delivery`: the current turn's original delivery names exactly one distinct positive task id, including work by a peer that does not own it.
69
+ - `single_open`: otherwise the peer owns exactly one `in_progress` task.
70
+ - `unattributed`: otherwise no task is assigned to the record.
71
+
72
+ `Bus.onDeliver` observes originals immediately before `peer.deliver` starts the
73
+ turn. Pending task identity is consumed at turn start and cleared on delivery
74
+ admission/failure, so later user-started native turns do not inherit it. Usage
75
+ and token events apply the rule at write time; `turn_end` keeps the start-time
76
+ attribution. Local usage uses its request-bound route policy task when present.
77
+ The existing relay has no task-bearing route usage event; this change adds no
78
+ new usage source. Schema version 1 and the control protocol are unchanged.
79
+
80
+ The unattributed share is always printed for token increments and deduplicated
81
+ usage records, against all recorded increments/usage records. A zero denominator
82
+ has unknown share. Records with no `attribution` field belong to a separate
83
+ `before attribution` bucket displayed beside the share, even if a task field is
84
+ present. Neither bucket is redistributed. `ahub export` preserves these fields
85
+ as raw JSON lines; plain `ahub report` keeps its existing behavior.
86
+
53
87
  The file is local and never uploaded. It grows without rotation; delete it to start
54
88
  over (the hub recreates it).
55
89
 
@@ -1,8 +1,136 @@
1
1
  # Operations guide
2
2
 
3
- This guide describes ahub 0.12.16 and control protocol 14. Live verification
3
+ This guide describes ahub 0.12.18 and control protocol 15. Live verification
4
4
  results and remaining prerequisites are recorded separately in [the smoke ledger](smoke.md).
5
5
 
6
+ ## Operator console and panels
7
+
8
+ `ahub console` combines the existing tail stream with peer, quota, approval-age
9
+ and input rows. `ahub up` opens it when both input and output are terminals;
10
+ `--no-console` and redirected input/output retain the start-only behavior.
11
+ `ahub tail` remains available with its existing rendering.
12
+
13
+ Allowing an approval requires selection and a separate confirmation. Denying
14
+ does not. Only options the daemon supplied are selectable. Typing a command
15
+ prevents approval shortcuts from interpreting that input as an answer. Full
16
+ approval titles are terminal-only; expiry and answers from another console remove
17
+ the pending item. Approval audit records contain id, peer, option kind, response
18
+ time and answering surface, without the title.
19
+
20
+ Tab toggles stream and panels; `ahub console --panels` starts in panels. Peers,
21
+ Approvals, Tasks, Queue and Events support arrow keys or j/k, Enter for detail,
22
+ Escape to return and `?` for help. `:` enters a command. Assignment, delivery
23
+ resolution and allow decisions require confirmation; delivery resolution requires
24
+ a reason. Tasks use the same public redaction as the board. Panels need at least
25
+ 80 columns by 24 rows; smaller terminals stay in stream mode. Task and queue
26
+ polling runs only while the corresponding panel is visible. Leaving restores
27
+ the terminal and returning from panels replays the bounded stream buffer.
28
+
29
+ Console color policy is `--color=auto|always|never`, with `auto` as the default.
30
+ Auto enables color only when both input and output are TTYs, `TERM` is not
31
+ `dumb`, and `NO_COLOR` is empty or absent. Explicit `always` overrides these
32
+ conditions, including redirected stream output; `never` disables styling.
33
+ Neither option changes terminal size requirements, panel decisions, cursor
34
+ management or the final reset. Invalid values fail before connecting.
35
+
36
+ | Meaning | Palette | Examples |
37
+ |---|---|---|
38
+ | Information and navigation | Cyan, bold cyan for active/selected labels | Tabs, section/peer labels and `>` selection marker |
39
+ | Success and availability | Green | Idle peer, approved task |
40
+ | Waiting and attention | Yellow | Busy/paused peer, pending approval, confirmation, review/ready task, important priority |
41
+ | Failure and intervention | Red | Failed/check-failed task, undeliverable/overflow, `needs_review` queue, denial request or expired/cancelled approval |
42
+ | Metadata | Bright black | Ages and remaining approval time |
43
+
44
+ Offline peers, primary titles, action details and message body lines keep the
45
+ terminal's default foreground. Labels and prompts remain readable without
46
+ color. Stream styling applies only to the header, using structured event data;
47
+ message text cannot choose a color. A local denial is shown as requested, not
48
+ as a confirmed receipt; a remote answered closure has no option-kind metadata
49
+ and is not guessed to be a denial. A fixed palette is applied after terminal
50
+ control sanitization. Width, clipping, wrapping and cursor placement use plain
51
+ Unicode text. Each styled span and every interactive exit restores attributes.
52
+ Color adds no polling, timers or extra redraws. `ahub tail`, logs and JSON stay
53
+ unchanged.
54
+
55
+ ```sh
56
+ ahub console --panels --color=auto
57
+ NO_COLOR=1 ahub console
58
+ ahub console --color=never
59
+ ahub console --color=always > console-stream.txt
60
+ ```
61
+
62
+ Actual light/dark terminal readability and bounded idle-CPU observations are
63
+ recorded separately in the smoke ledger; fake-terminal tests establish policy,
64
+ geometry, sanitization and restoration only.
65
+
66
+ The command input accepts existing status, board, task, review, say, pause,
67
+ resume, budget, queue, permit, ask, remember, route, turns, undo, check-path and
68
+ report operations. It executes an argument vector with closed stdin. Lifecycle,
69
+ launch, nested console, setup, UI, logs and tail commands are refused.
70
+
71
+ ## Conducting a team from Claude Code or Codex
72
+
73
+ Start the daemon with `ahub up --no-console`, set exactly one conductor in the
74
+ project configuration, then open `ahub console` in a split terminal for approvals:
75
+
76
+ ```json
77
+ {
78
+ "roles": { "claude": ["planner", "reviewer", "conductor"] },
79
+ "conductor": { "feed": "own" }
80
+ }
81
+ ```
82
+
83
+ Launch that peer with `ahub claude` for channel pushes, or select `codex` in
84
+ `roles` and launch `ahub codex`. The conductor splits work into owned tasks,
85
+ observes `hub_status`, moves stalled work and obtains review before reporting
86
+ results and open decisions. It does not implement the tasks it handed out.
87
+ `hub_peer_start` starts local, Kimi or headless Pi; requests for native TUIs
88
+ return the `ahub` command for the person to run, after validating it through
89
+ the shared launcher planner. The wrapper plans again at launch when native
90
+ endpoints are available. `hub_peer_hold` and `hub_peer_release`
91
+ manage only holds placed by that conductor. Assignment also requires `assign`
92
+ when the conductor has an explicit capabilities list.
93
+
94
+ Check the returned owner and task state after assignment. Routing skips paused
95
+ peers, so assign work before placing a delivery hold. Hand work out through the
96
+ board and report to the person with `hub_send` addressed to `user`, or a `[FYI]`
97
+ final response in the native TUI. Broadcasting implementation instructions can
98
+ cause an otherwise unassigned owner to claim duplicate work.
99
+
100
+ The person answers approvals in the console, resolves `needs_review` deliveries
101
+ with `ahub queue resolve`, and handles budget overrides and hub lifecycle.
102
+ Running these commands from an agent shell is refused; the CLI retains the
103
+ agent's identity even when invoked through a shell tool.
104
+
105
+ `conductor.feed` is `own` by default, `all` for all tasks, or `off`. Ordinary
106
+ milestones share the existing digest window and collapse repeated queued
107
+ task/kind notices. They wait behind a busy conductor. An aged approval or
108
+ `needs_review` hold is important, contains only peer/tool-age or delivery id,
109
+ and directs the conductor to ask the person. A completed task set emits one
110
+ round notice until a new task joins. Removing the role or switching the feed
111
+ off withdraws pending feed notices.
112
+
113
+ For a Claude conductor, `ahub claude` also observes native session and turn
114
+ boundaries when facts injection and task-idle sweeps are off. Keep the managed
115
+ hooks enabled to measure completion and supervision usage. Passing your own
116
+ `--settings` takes precedence and produces a warning when it replaces that
117
+ observation. The launcher records its private session identity in ordinary
118
+ terminals too; this does not grant terminal-recovery authority.
119
+ Ordinary Claude sessions with turn-free facts or task-idle sweeps enabled get the
120
+ same session/start observation. Other ordinary launches remain non-opt-in.
121
+
122
+ `ahub report` records conductor actions and completed native turns containing
123
+ supervision. Tokens describe the whole measured turn, which may also contain
124
+ other work; they are not a per-notice cost estimate. Missing measurements stay
125
+ unknown. Use the live smoke ledger to assess observed turns and tokens per
126
+ approved task before choosing `all`; an unmeasured run is not a cost benchmark.
127
+ Claude turn counts use authenticated native completion events. Older logical
128
+ state counts are labelled; an idle channel or approved task alone does not prove
129
+ that the native answer finished.
130
+ The Stop hook acknowledgement is not a counted completion. The daemon checks the
131
+ native transcript after the hook can return; missing or changed-session evidence
132
+ stays unknown.
133
+
6
134
  ## Install and start
7
135
 
8
136
  Use Bun 1.3 or newer. Install the released package and install its Claude
@@ -14,12 +142,11 @@ ahub setup
14
142
  cd <project>
15
143
  ahub init
16
144
  ahub up
17
- ahub tail
18
145
  ```
19
146
 
20
147
  `ahub init` writes the project configuration and managed instruction blocks.
21
148
  Run it after an upgrade when those blocks need refreshing. `ahub setup`
22
- updates the shared Claude plugin. Keep `ahub tail` open when a local worker
149
+ updates the shared Claude plugin. Keep `ahub console` open when a local worker
23
150
  may request an approval.
24
151
 
25
152
  `.agenthub/config.json` can be committed and shared. The fields that choose
@@ -256,11 +383,16 @@ token counts, never message bodies or task titles.
256
383
  ```bash
257
384
  ahub report --since 7d # turns, busy time and tokens per peer, messages, overlaps, task events
258
385
  ahub report --since 7d --json # the same numbers as JSON
386
+ ahub report --by task # tokens, turns and wall time per task and class, with the unattributed share
259
387
  ahub export --since 24h # the raw events as JSON lines, for your own analysis
260
388
  ```
261
389
 
262
390
  `ahub report` counts the same overlap warnings as `scripts/overlaps.ts`, from the
263
- structured events instead of log lines.
391
+ structured events instead of log lines. `--by task` uses the task each usage and
392
+ token record was attributed to when it was written: the delivery that started the
393
+ turn, otherwise the peer's only `in_progress` task. Everything else is reported as
394
+ unattributed, and records from before 0.12.18 as a separate bucket; neither is
395
+ redistributed (rules: [events](events.md#per-task-usage-reports)).
264
396
 
265
397
  ## Turns and undo
266
398
 
@@ -521,6 +653,13 @@ stays the default until an evaluation says otherwise (`docs/cooperbench.md`).
521
653
 
522
654
  Inspect permission requests in the terminal:
523
655
 
656
+ ```bash
657
+ ahub console
658
+ ```
659
+
660
+ Select the requested option in the console and confirm an allow decision.
661
+ For a separate plain terminal, the existing command remains available:
662
+
524
663
  ```bash
525
664
  ahub tail
526
665
  ahub permit <request-id> allow
@@ -647,21 +786,21 @@ Rows without a live process are stale registrations; forget them with
647
786
 
648
787
  Upgrade running projects with the target release's own coordinator. It accepts
649
788
  a running source on control protocol 9 (0.6.x), 10 (0.7.0 through 0.12.0),
650
- 11 (0.12.1 and 0.12.2), 12 (0.12.3), 13 (0.12.4 through 0.12.15) or 14 (0.12.16), and only
789
+ 11 (0.12.1 and 0.12.2), 12 (0.12.3), 13 (0.12.4 through 0.12.15) 14 (0.12.16) or 15 (0.12.17 and 0.12.18), and only
651
790
  a target on its own protocol, so the target's coordinator fits every supported
652
791
  source and carries every recovery fix released up to it. Protocol 8 and older
653
792
  (0.5.x and earlier) are refused as `manual-bootstrap-required`. Run from the
654
793
  project directory, without replacing the global CLI first:
655
794
 
656
795
  ```bash
657
- bunx --package @staix/agent-hub@0.12.16 ahub upgrade --to 0.12.16 --dry-run
658
- bunx --package @staix/agent-hub@0.12.16 ahub upgrade --to 0.12.16 --yes
796
+ bunx --package @staix/agent-hub@0.12.18 ahub upgrade --to 0.12.18 --dry-run
797
+ bunx --package @staix/agent-hub@0.12.18 ahub upgrade --to 0.12.18 --yes
659
798
  ```
660
799
 
661
800
  | Running now | Coordinator to use |
662
801
  | --- | --- |
663
802
  | 0.6.x (protocol 9) | the target's, through `bunx` as above |
664
- | 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.15 (protocol 13), 0.12.16 (protocol 14) | the target's, through `bunx` as above |
803
+ | 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.15 (protocol 13), 0.12.16 (protocol 14), 0.12.17 and 0.12.18 (protocol 15) | the target's, through `bunx` as above |
665
804
  | any supported source, with the installed CLI already at the target | `ahub upgrade` below, which is the same coordinator |
666
805
  | 0.5.x or earlier (protocol 8 and older) | not supported: bootstrap by hand with the matching CLI |
667
806
 
@@ -693,14 +832,14 @@ projects first:
693
832
 
694
833
  ```bash
695
834
  ahub restart --dry-run
696
- ahub upgrade --to 0.12.16 --dry-run
835
+ ahub upgrade --to 0.12.18 --dry-run
697
836
  ```
698
837
 
699
838
  Apply only after reviewing the plan:
700
839
 
701
840
  ```bash
702
841
  ahub restart --yes
703
- ahub upgrade --to 0.12.16 --yes
842
+ ahub upgrade --to 0.12.18 --yes
704
843
  ahub recovery status <operation-id>
705
844
  ahub recovery resume <operation-id>
706
845
  ahub recovery abort <operation-id>
@@ -17,7 +17,7 @@ ahub setup # installs the Claude Code channel plu
17
17
  Or use the matching GitHub release:
18
18
 
19
19
  ```bash
20
- bun add -g github:STAIxBWLB/agent-hub#v0.12.16
20
+ bun add -g github:STAIxBWLB/agent-hub#v0.12.18
21
21
  ahub setup
22
22
  ```
23
23
 
@@ -34,8 +34,8 @@ In your project directory, one terminal each:
34
34
 
35
35
  ```bash
36
36
  ahub init # .agenthub/config.json, .agenthub/routing.toml, the marker block in AGENTS.md
37
- ahub up # the daemon for this directory (loopback only)
38
- ahub tail # keep open: the conversation, state changes, permission requests
37
+ ahub up # starts the daemon and opens the operator console in a terminal
38
+ ahub console # use in another terminal if up used --no-console; Tab opens panels
39
39
  ahub kimi # Kimi, headless
40
40
  ahub codex # Codex TUI, attached through the hub
41
41
  ahub claude # Claude Code with the hub channel
package/docs/security.md CHANGED
@@ -8,7 +8,7 @@ agent-hub connects agents that can each run commands. This page says what the hu
8
8
  - **The control link is loopback plus a secret.** The daemon and the Codex proxy bind 127.0.0.1 only. The control WebSocket requires a per-run token (`.agenthub/state/control-token`, mode 600), and both servers refuse any request that carries an `Origin` header: any web page can open a WebSocket to localhost, and browsers always send `Origin`. External clients cannot claim the console user's id or a hub-managed peer's id.
9
9
  - **Permission prompts stay on.** `ahub claude` and `ahub codex` add nothing that weakens the agents' own prompts. Kimi's and `local`'s permission requests are relayed to the console and cancelled after `approvals.timeout_s` (default 120 s) of silence; the macOS notification for a waiting request carries the peer and the tool name only. The one exception is the hub's own tools (`hub_send` and the task tools, matched by exact name): Kimi's requests for them are approved once without a prompt and logged by name, the same trust Codex gets through `approval_mode` in the hub's config. They invoke hub-owned operations rather than arbitrary file or shell tools, and every call passes the hub's own checks. Identity comes from the exact permission title, or from the earlier tool-call title bound to the same call id and resolved against the configured MCP servers when the permission title contains argument JSON (Qwen). Payload text and unrelated display titles never establish identity. A payload longer than the console shows is marked as cut and never offers a session-wide grant. `--unattended` turns prompts off, says so loudly, and is never the default.
10
10
  - **A committed config cannot choose launch commands, credential files, data endpoints or a wider sandbox.** The machine-local fields (`kimi_cmd`, `codex_bin`, `pi.cmd`, `checks`, `mlx.bin`, `mlx.runtimeDir`, `mlx.modelPath`, `omniroute.urls`, `omniroute.access_hosts`, the `omniroute` key files, `memory.worker_url`, `local.read_allow`, `local.bash_network`, `local.network_allow`) apply only from a config file git confirms nobody committed: `.agenthub/config.json` or `.agenthub/config.local.json`, matched by file identity so no other spelling the file system accepts slips past, and `.agenthub` itself not a committed symlink or submodule. Without a repository, or when git fails, they keep their defaults; an empty value always means the default. "Untracked" is answered by the repository that contains the project: a checkout copied or extracted into an unrelated repository, or into an ignored directory of one, is trusted like your own files. So a cloned repository cannot choose a launch command, a completion check, a gateway to send a key file to, a memory endpoint, or a wider sandbox. A command in `checks` runs as you, outside the local worker's sandbox, like a git hook. Nothing in `routing.toml` or in task text is ever run. `routing.toml` and the other shared fields still come from the checkout, and they matter: `routing.toml` picks the models the local worker and the hub's inference use at your gateway and can turn the PII constraint off, and roles and budget shape who does what. Review them in a repository you do not trust.
11
- - **Telemetry holds no bodies.** `.agenthub/state/events.jsonl` (issue #40) records envelope ids, routing and sizes, task ids and states, overlapping paths and token counts. It never records a message body, a task title or detail, and marks private (PII) envelopes and tasks as such. It stays on the machine; `ahub export` only prints it.
11
+ - **Telemetry holds no bodies.** `.agenthub/state/events.jsonl` (issue #40) records envelope ids, routing and sizes, task ids and states, overlapping paths and token counts, and the task id (with a PII flag) that each usage and token record is attributed to. It never records a message body, a task title or detail, and marks private (PII) envelopes and tasks as such. It stays on the machine; `ahub export` only prints it.
12
12
  - **Snapshots stay in your repository.** Per-turn snapshots (issue #33) are git objects in the project's own object store, written through a temporary index; nothing is referenced, pushed or copied elsewhere, and `git gc` prunes them. They hold what the work tree held, including untracked files that are not ignored, so keep secrets in ignored files. They carry the repository's own permissions, and nothing caps their disk use but `git gc`. A turn of a peer holding an open PII task is not snapshotted; a PII file left in the project is snapshotted by later turns like any other file. `ahub undo` restores only files whose current content is exactly what the turn left.
13
13
  - **The edit hook reads, never decides.** `ahub check-path --hook` (issue #32) reads hub.db and returns context for Claude and a line for you; it sets no permission decision, so your permission rules stay in charge. It names other owners' task ids, titles and states, which then reach Claude's model; PII tasks are left out.
14
14
  - **The session record holds identities only.** `.agenthub/state/sessions.json` (issue #37, mode 600) keeps each attached peer's recovery metadata: launch options, session and thread ids, Pi's session file path. No message or task text; loss notices name deliveries by id, sender and public task title.
@@ -28,6 +28,33 @@ agent-hub connects agents that can each run commands. This page says what the hu
28
28
  - **Network-level proof for PII.** The hub proves at its own boundaries (tests search every output) that PII text does not leave; packet-level verification of your gateway path is an operations check.
29
29
  - **Other operating systems' sandboxes.** Only macOS seatbelt is implemented.
30
30
 
31
+ ## Agent CLI identity and conductor authority
32
+
33
+ Inside an agent session, `ahub` identifies the caller from `AGENTHUB_PEER_ID`,
34
+ or the native Claude/Codex shell marker when the hub marker is absent. It connects
35
+ as that peer in tools mode. Messages and task changes retain that actor;
36
+ `say` defaults to status priority and important messages still require the peer's
37
+ capability. Conflicting or malformed markers fail closed. There is no `--as-user`.
38
+ Human-only operations are refused before connecting, including permission answers,
39
+ queue resolution, budget overrides, lifecycle and recovery operations, and `ask`.
40
+ The operator uses a plain terminal or `ahub console`. A Claude `!` command that
41
+ inherits the agent markers follows the same rule. The console strips control
42
+ sequences from agent and daemon text before it quotes forged headers, and its
43
+ colors come only from a fixed palette applied afterwards, so message text cannot
44
+ set styles or operate the terminal.
45
+
46
+ This is an honest default against accidental impersonation and injected commands.
47
+ It does not make the token inaccessible to an agent with unrestricted project
48
+ shell access. The existing OS sandbox and loopback authentication boundaries apply.
49
+
50
+ Steering tools require an explicit `conductor` role, independent of default-allow
51
+ capabilities. Only one peer may hold that role. A conductor may inspect public
52
+ state, assign or escalate work, start supported headless peers and place its own
53
+ delivery holds. It cannot answer approvals, resolve durable deliveries, override
54
+ budget pauses or release a human hold. Pending approval summaries exclude titles;
55
+ PII tasks remain public stubs. Conduct events contain ids only. Supervision feeds
56
+ use structured reasons and never carry check output or approval bodies.
57
+
31
58
  ## Reporting
32
59
 
33
60
  Please report vulnerabilities privately through GitHub's "Report a vulnerability" on this repository rather than in a public issue.
package/docs/smoke.md CHANGED
@@ -2,6 +2,99 @@
2
2
 
3
3
  `scripts/check.sh` covers everything against fakes. The legs below need real accounts and an interactive terminal, so they are run by hand and recorded here.
4
4
 
5
+ ## Operator console and conductor candidate (#190, #191, #193-#195)
6
+
7
+ Candidate 0.12.17, protocol 15, observed on 2026-10-09 KST. The final Claude
8
+ observer runs used immutable source `20ad5e6`; their native completion receipts
9
+ and actual console effects were checked independently. The Kimi new-tool leg
10
+ remains quota-blocked, as recorded below.
11
+
12
+ - Actual Codex 0.146.0 (`gpt-5.5`) and Claude 2.1.295 (Opus 5.5) TUIs started
13
+ real local and headless Pi peers, proposed exactly two tasks owned initially
14
+ by local, reassigned Beta to Pi, and placed and released their own holds.
15
+ The real owners checked their outputs and called `hub_task_done`; the native
16
+ conductor independently read the exact files before approving both tasks.
17
+ Real CLI calls retained the agent actor, and human-only queue/permission
18
+ actions were refused from an agent shell.
19
+ - The person authorized `alpha.txt = ALPHA` (5 bytes) and `beta.txt = BETA`
20
+ (4 bytes), both without a newline. Separate selection and confirmation keys
21
+ reached the actual console PTY and produced allow-once console audit records.
22
+ Out-of-scope directory-list commands were cancelled. The final Claude own
23
+ run used three allow-once answers because Pi combined its exact write and
24
+ bounded byte checks in one approved command; it also recorded two cancelled
25
+ local directory-list requests. No permission or task completion was fabricated.
26
+ - The final Claude own run recorded nine completed native turns and nine
27
+ matching daemon Stop records, 27 unique native usage records, and
28
+ 1,888,242 tokens including cached input. Eight completed native turns contained
29
+ supervision, with 1,292,922 recorded tokens. Its final message UUID/id,
30
+ `end_turn`, `turn_duration`, instance/session/launcher and opaque Stop receipt
31
+ matched; native idle and an empty pending delivery queue were verified before
32
+ owned-process cleanup. Local recorded three turns and 72,046 gateway tokens;
33
+ Pi recorded two turns and 47,177 native tokens.
34
+ - A read-only Claude feed-off continuation on the final observer preserved the
35
+ existing approved tasks and exact files. It verified one completed native turn,
36
+ four usage records and 263,413 cached-inclusive tokens, with matching daemon
37
+ Stop, native idle and delivery settlement. An obsolete review hold was first
38
+ discarded through the authenticated public operator API after fresh approved
39
+ task/history and exact BETA readback. That discard is not native execution.
40
+ - A real console with two pending cards and no operator input consumed 0.07 CPU
41
+ seconds over 40.018 wall seconds in stream mode (0.175% of one CPU). A separate
42
+ actual Peers panel, 120x40, with two pending cards and no operator input used
43
+ 0.07 CPU seconds over 40.012 wall seconds (0.175%). These are cumulative `ps`
44
+ process CPU samples in distinct modes/runs, not whole-host idle measurements.
45
+ - In a plain Claude TUI without the development-channel flag, one actual
46
+ `hub_status` MCP call succeeded and the daemon audited the Claude status
47
+ action. A unique directed operator push was accepted by the bridge, but no
48
+ pushed native user row or answer appeared in the observed 20.075-second window.
49
+ This proves tool access separately from the bounded negative push observation.
50
+ The real channel-enabled runs above received actual review and supervision
51
+ pushes. The plain probe and cleanup completed in 56.724 seconds, with no model
52
+ file, shell, task or settings actions.
53
+ - Real Kimi Code CLI 2.1.1 completed ACP initialization and `session/new`, then
54
+ rejected the single prompt with HTTP 403 for its weekly account usage limit.
55
+ No requested native tool event or role-refusal result occurred. The reset time
56
+ was not supplied. The owning CLI listed only the managed OAuth Kimi provider
57
+ (four models, default `kimi-code/k3`), so no configured alternative provider was
58
+ found. This is an incomplete external prerequisite, not a new-tool pass. No
59
+ purchase, provider/configuration change or model retry was performed; owned
60
+ processes were stopped after 3.142 seconds.
61
+
62
+ The user explicitly approved release 0.12.17 with only the Kimi native new-tool
63
+ and ordinary-role refusal T0 evidence deferred on 2026-10-09. That native
64
+ prerequisite remains unverified; the quota rejection is not a tool pass. This
65
+ decision does not defer the other native/source gates or authorize a purchase,
66
+ credentials, provider/configuration change or account retry. The same isolated
67
+ read-only Kimi probe remains required before its native coverage is marked verified.
68
+
69
+ Earlier captures remain part of the evidence:
70
+
71
+ - Codex feed-off included an interrupted original and read-only continuation:
72
+ six logical native turns, 49 increments and 1,979,016 tokens. Codex feed-own
73
+ recorded nine turns, 34 increments and 1,399,470 tokens. Their journals were
74
+ independently checked after cleanup. The first baseline interruption and
75
+ operator reconciliation are preserved; neither native task completion nor
76
+ board approval was substituted by the harness.
77
+ - The original Claude runs approved the actual files but an early idle-based
78
+ harness stopped their final review answers. Subsequent `da2ebbc` feed-off
79
+ completed five native turns and 1,451,782 tokens but certified only one daemon
80
+ Stop. A `6034dd9` read-only continuation completed another native turn and
81
+ 261,483 tokens, while its Stop remained unavailable during the native hook.
82
+ Both incomplete observer captures are preserved alongside raw and audited
83
+ summaries. The final post-ACK observer above closes the fresh completion proof;
84
+ it does not retroactively certify the earlier missed Stop records.
85
+
86
+ All token totals describe whole native/model turns, including cached input and
87
+ other work. Operator waiting, cancelled requests, marker probes, interrupted
88
+ continuations and differing source revisions make these observations unsuitable
89
+ for a causal feed-overhead or model-efficiency comparison. Unknown measurements
90
+ and the Kimi prerequisite remain explicit.
91
+
92
+ The harness uses real Python PTYs. Operator-file-input forwards only
93
+ chat-authorized keys, removes inherited Orca terminal ownership from fixture
94
+ children, and generates no approval automatically. Private original and
95
+ continuation captures remain separate. The shell-marker probes and vendor
96
+ limits are recorded in [the identity T0 ledger](verification/2026-10-09-agent-shell-t0.md).
97
+
5
98
  ## Approval race live reproduction and candidate verification (#98)
6
99
 
7
100
  Measured on 2026-10-02 KST with installed 0.12.1 and the correction candidate,