@staix/agent-hub 0.12.16 → 0.12.17

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (43) hide show
  1. package/CHANGELOG.md +8 -0
  2. package/README.md +4 -4
  3. package/docs/agent-notes/adapters.md +1 -0
  4. package/docs/agent-notes/bus.md +2 -0
  5. package/docs/agent-notes/tasks.md +3 -1
  6. package/docs/agent-notes/tests.md +1 -0
  7. package/docs/operations.md +106 -9
  8. package/docs/quickstart.md +3 -3
  9. package/docs/security.md +24 -0
  10. package/docs/smoke.md +93 -0
  11. package/docs/specs/2026-09-19-agent-hub-design.md +120 -1
  12. package/docs/verification/2026-10-09-agent-shell-t0.md +123 -0
  13. package/docs/verified.json +26 -17
  14. package/package.json +1 -1
  15. package/plugins/agent-hub/.claude-plugin/plugin.json +1 -1
  16. package/plugins/agent-hub/server.js +17 -5
  17. package/src/adapters/acp.ts +2 -2
  18. package/src/adapters/claude-channel.ts +3 -2
  19. package/src/adapters/codex-appserver.ts +3 -2
  20. package/src/adapters/local-worker.ts +5 -5
  21. package/src/cli/console-state.ts +213 -0
  22. package/src/cli/console.ts +201 -0
  23. package/src/cli/facts-hook.ts +8 -1
  24. package/src/cli/identity-audit.ts +66 -0
  25. package/src/cli/identity.ts +54 -0
  26. package/src/cli/launch.ts +15 -2
  27. package/src/cli/main.ts +62 -29
  28. package/src/cli/tail-render.ts +17 -0
  29. package/src/cli/upgrade-runtime.ts +1 -1
  30. package/src/hub/board.ts +6 -2
  31. package/src/hub/bus.ts +82 -10
  32. package/src/hub/child-process.ts +11 -0
  33. package/src/hub/conductor.ts +196 -0
  34. package/src/hub/control-client.ts +3 -3
  35. package/src/hub/daemon.ts +367 -46
  36. package/src/hub/envelope.ts +1 -1
  37. package/src/hub/events.ts +6 -1
  38. package/src/hub/hub-tools.ts +13 -0
  39. package/src/hub/report.ts +53 -4
  40. package/src/hub/supervision.ts +152 -0
  41. package/src/hub/tasks.ts +30 -15
  42. package/src/hub/usage.ts +37 -2
  43. package/src/pi/launch.ts +2 -1
package/CHANGELOG.md CHANGED
@@ -4,6 +4,14 @@ Issue and pull request numbers in the entries for 0.7.7 and earlier refer to the
4
4
 
5
5
  ## Unreleased
6
6
 
7
+ ## 0.12.17
8
+
9
+ - Add an operator console with confirmed approvals, bounded status polling and optional peer, approval, task, queue and event panels; open it from interactive `up` and preserve the tail renderer (#190, #191).
10
+ - Identify agent shell CLI calls as their peer, refuse human-only commands before connecting and record ids-only audit notices without an as-user bypass (#193).
11
+ - Grant steering tools only to one explicit conductor role, preserve the actor on task changes and keep conductor holds separate from human, budget and recovery holds (#194).
12
+ - Batch public task milestones and human-action reminders through the existing digest window, replace repeated pending notices and report completed native supervision turns with unknown usage preserved (#195).
13
+ - Advance the control protocol to 15 for approval lifecycle and supervision metadata, retaining controlled recovery from supported previous sources.
14
+
7
15
  ## 0.12.16
8
16
 
9
17
  - Record source-verification commits for current operating documents and agent notes; fail missing coverage, invalid stamps and README version drift, and report stale source scopes deterministically (#188).
package/README.md CHANGED
@@ -4,7 +4,7 @@ Native multi-agent hub for one developer's machine: Claude Code, Codex, Kimi Cod
4
4
  hub-owned local-LLM worker collaborate as peers in independent project directories, with
5
5
  task-aware model routing (Switchyard) in front of a self-hosted gateway (OmniRoute).
6
6
 
7
- Status: 0.12.16, control protocol 14. Durable delivery records distinguish queued
7
+ Status: 0.12.17, control protocol 15. Durable delivery records distinguish queued
8
8
  work from uncertain execution. The [smoke checklist](docs/smoke.md) records
9
9
  verified paths and remaining prerequisites.
10
10
 
@@ -22,13 +22,13 @@ Install from npm:
22
22
 
23
23
  ```bash
24
24
  bun add -g @staix/agent-hub && ahub setup
25
- cd <your project> && ahub init && ahub up && ahub tail
25
+ cd <your project> && ahub init && ahub up
26
26
  ```
27
27
 
28
28
  Or install the same version from GitHub:
29
29
 
30
30
  ```bash
31
- bun add -g github:STAIxBWLB/agent-hub#v0.12.16 && ahub setup
31
+ bun add -g github:STAIxBWLB/agent-hub#v0.12.17 && ahub setup
32
32
  ```
33
33
 
34
34
  The installed commands remain `ahub` and `agent-hub`.
@@ -100,7 +100,7 @@ only. `--backend dgx` or `--backend mlx` explicitly pins a Pi session's backend.
100
100
  Pi connects to an authenticated relay with model aliases `dgx/coding`,
101
101
  `dgx/fast`, and `mlx/fast`; upstream gateway credentials stay in the hub.
102
102
  Its managed tools use the existing path guards, shell sandbox and terminal
103
- approval flow. Keep `ahub tail` open to approve write/edit/shell operations.
103
+ approval flow. Keep `ahub console` open to approve write/edit/shell operations.
104
104
  Tool-call receipts survive daemon restart; an interrupted operation with an
105
105
  unknown outcome must be reconciled before repeating it. Cancelling a Pi turn
106
106
  does not automatically hand its task to a cloud peer.
@@ -15,4 +15,5 @@ Scope: the Codex, ACP and Pi adapters and the Pi extension (turn `generation` an
15
15
  - Native owner teardown must handle a lost shutdown acknowledgement using verified process identity. Pending launches must be revocable, and active tools/streaming output must count as watchdog activity.
16
16
  - Pi owner identity is one shared contract, `processSignature` in `src/pi/process-signature.ts`; the startup claim check, the extension's claim, the TUI owner monitor, `recoveryReady` and the verified teardown all go through it. Its `ps` read pins LC_ALL=C and TZ=UTC (the `processTable` contract) because `lstart` follows the reader's timezone and locale, and a hub on UTC with a Pi child on system time otherwise hashes different strings for the same process and refuses the valid owner (#174). Strict pid + start + command identity stays; never accept PID-only or a fuzzy command match.
17
17
  - Return the shared `toolResultFailed` verdict as `/tool`'s `failed`; map only `failed === true` to Pi's `isError`. Preserve text and the older bridge's absent-flag behavior.
18
+ - Spawn a native peer with `peerChildEnv(peer, source)` so its own `AGENTHUB_PEER_ID` replaces inherited agent markers. Sandbox tool environments carry their executing peer too. Preserve the native shell environment policy; Codex's `CODEX_THREAD_ID` supplies the fallback after filtering. Before changing marker detection or launcher environments, read `docs/verification/2026-10-09-agent-shell-t0.md` for pinned source and live-probe limits.
18
19
  - Emit Pi's ceiling event at the pre-effect rejection boundary (`src/pi/ceiling.ts`), before `agent_end`. Accept only a valid current-session/current-generation signal during a running turn; retain the first signal and attach it to that turn's failure. Never classify a ceiling from free text.
@@ -1,5 +1,7 @@
1
1
  # Bus, digests and replies
2
2
 
3
+ - Supervision replaces only pending, never attempted or accepted, envelopes with the same stable key. Preserve the first batch deadline and existing digest/queue ceilings; role/feed revocation withdraws pending notices only. Count supervision from actual transport admission followed by a successful native turn completion, never from enqueue or failure-to-idle.
4
+
3
5
  Scope: delivery, digests and condensation, reply addressing, priority and limits. Read before editing `src/hub/bus.ts`, `src/hub/envelope.ts`, `src/hub/limits.ts`, `src/hub/delivery-journal.ts`, `src/hub/inference.ts` (digest condensation), or how an adapter addresses a reply or sets its priority.
4
6
 
5
7
  - A failed digest is retried one envelope at a time, so a poison envelope cannot take its neighbours down with it.
@@ -1,7 +1,9 @@
1
1
  # Task board and hub tools
2
2
 
3
- Scope: the task flow, assignment, the state machine and the hub's MCP tools. Read before editing `src/hub/tasks.ts`, `src/hub/board.ts`, `src/hub/routing.ts`, `src/hub/task-sweep.ts` or `src/hub/hub-tools.ts`.
3
+ Scope: the task flow, assignment, the state machine and the hub's MCP tools. Read before editing `src/hub/tasks.ts`, `src/hub/board.ts`, `src/hub/routing.ts`, `src/hub/task-sweep.ts`, `src/hub/conductor.ts`, `src/hub/supervision.ts` or `src/hub/hub-tools.ts`.
4
4
 
5
5
  - Tool callers are models: MCP `inputSchema` is not enforced on the way in. Normalize at the boundary (`cleanRefs`) before anything reaches the board, and never throw after a board write.
6
6
  - `assign()` never defaults to the task's current owner, or a decline can only come back to the decliner.
7
7
  - The state machine allows `in_progress -> approved` only for classes without a reviewer; `review()` checks `in_review` itself.
8
+ - Conductor authority requires the explicit unique role before capability checks. Assignments preserve the acting peer. A peer releases only its own conductor hold; operator release and manual, quota and recovery holds remain separate paths.
9
+ - Supervision reasons use structured enums, never history notes, check output or approval titles. Round signatures survive restart; later summaries count only newly joined tasks.
@@ -5,4 +5,5 @@ Scope: tests, fakes and the test gate. Read before adding or changing a test or
5
5
  - Tests that are not about batching build the bus with `batchMs: 0`; with the default 15 s window a lone status envelope looks like a lost message.
6
6
  - In Bun 1.3.14 a test timeout that fires while `Bun.spawnSync` runs can start the next test inside spawnSync's own event loop, where another spawnSync then spins at full CPU for good (trigger not pinned down; see README). Bun 1.4.2 still runs the next test inside the outer spawnSync's event loop. `scripts/check.sh` runs tests with a 20 s timeout and under `scripts/hang-watch.sh`; a test that spawns and reads the process table gets a timeout of its own.
7
7
  - A child process gets a scrubbed environment, so test knobs for fakes travel in a wrapper script, not in `process.env`.
8
+ - Bind agent identity markers explicitly in native/CLI fixture children. The full gate starts the test runner without inherited agent markers for human CLI fixtures; never weaken production marker detection to make them pass.
8
9
  - Run `bun scripts/seeded-check.ts` before reporting the seeded-guard gate complete. Keep the six `scripts/seeds.json` cases sequential; require the same named test green before its seeded assertion fails. Seed rot, survivors, setup/compiler failures, timeout/watchdog termination and process leaks fail the gate. Never run the full seed runner from a unit test or mutate the source checkout; inspect preserved failed fixtures before removing them. CI runs this in the required Linux `seeded guards` job after the ordinary check, separately from `scripts/check.sh`.
@@ -1,8 +1,99 @@
1
1
  # Operations guide
2
2
 
3
- This guide describes ahub 0.12.16 and control protocol 14. Live verification
3
+ This guide describes ahub 0.12.17 and control protocol 15. Live verification
4
4
  results and remaining prerequisites are recorded separately in [the smoke ledger](smoke.md).
5
5
 
6
+ ## Operator console and panels
7
+
8
+ `ahub console` combines the existing tail stream with peer, quota, approval-age
9
+ and input rows. `ahub up` opens it when both input and output are terminals;
10
+ `--no-console` and redirected input/output retain the start-only behavior.
11
+ `ahub tail` remains available with its existing rendering.
12
+
13
+ Allowing an approval requires selection and a separate confirmation. Denying
14
+ does not. Only options the daemon supplied are selectable. Typing a command
15
+ prevents approval shortcuts from interpreting that input as an answer. Full
16
+ approval titles are terminal-only; expiry and answers from another console remove
17
+ the pending item. Approval audit records contain id, peer, option kind, response
18
+ time and answering surface, without the title.
19
+
20
+ Tab toggles stream and panels; `ahub console --panels` starts in panels. Peers,
21
+ Approvals, Tasks, Queue and Events support arrow keys or j/k, Enter for detail,
22
+ Escape to return and `?` for help. `:` enters a command. Assignment, delivery
23
+ resolution and allow decisions require confirmation; delivery resolution requires
24
+ a reason. Tasks use the same public redaction as the board. Panels need at least
25
+ 80 columns by 24 rows; smaller terminals stay in stream mode. Task and queue
26
+ polling runs only while the corresponding panel is visible. Leaving restores
27
+ the terminal and returning from panels replays the bounded stream buffer.
28
+
29
+ The command input accepts existing status, board, task, review, say, pause,
30
+ resume, budget, queue, permit, ask, remember, route, turns, undo, check-path and
31
+ report operations. It executes an argument vector with closed stdin. Lifecycle,
32
+ launch, nested console, setup, UI, logs and tail commands are refused.
33
+
34
+ ## Conducting a team from Claude Code or Codex
35
+
36
+ Start the daemon with `ahub up --no-console`, set exactly one conductor in the
37
+ project configuration, then open `ahub console` in a split terminal for approvals:
38
+
39
+ ```json
40
+ {
41
+ "roles": { "claude": ["planner", "reviewer", "conductor"] },
42
+ "conductor": { "feed": "own" }
43
+ }
44
+ ```
45
+
46
+ Launch that peer with `ahub claude` for channel pushes, or select `codex` in
47
+ `roles` and launch `ahub codex`. The conductor splits work into owned tasks,
48
+ observes `hub_status`, moves stalled work and obtains review before reporting
49
+ results and open decisions. It does not implement the tasks it handed out.
50
+ `hub_peer_start` starts local, Kimi or headless Pi; requests for native TUIs
51
+ return the `ahub` command for the person to run, after validating it through
52
+ the shared launcher planner. The wrapper plans again at launch when native
53
+ endpoints are available. `hub_peer_hold` and `hub_peer_release`
54
+ manage only holds placed by that conductor. Assignment also requires `assign`
55
+ when the conductor has an explicit capabilities list.
56
+
57
+ Check the returned owner and task state after assignment. Routing skips paused
58
+ peers, so assign work before placing a delivery hold. Hand work out through the
59
+ board and report to the person with `hub_send` addressed to `user`, or a `[FYI]`
60
+ final response in the native TUI. Broadcasting implementation instructions can
61
+ cause an otherwise unassigned owner to claim duplicate work.
62
+
63
+ The person answers approvals in the console, resolves `needs_review` deliveries
64
+ with `ahub queue resolve`, and handles budget overrides and hub lifecycle.
65
+ Running these commands from an agent shell is refused; the CLI retains the
66
+ agent's identity even when invoked through a shell tool.
67
+
68
+ `conductor.feed` is `own` by default, `all` for all tasks, or `off`. Ordinary
69
+ milestones share the existing digest window and collapse repeated queued
70
+ task/kind notices. They wait behind a busy conductor. An aged approval or
71
+ `needs_review` hold is important, contains only peer/tool-age or delivery id,
72
+ and directs the conductor to ask the person. A completed task set emits one
73
+ round notice until a new task joins. Removing the role or switching the feed
74
+ off withdraws pending feed notices.
75
+
76
+ For a Claude conductor, `ahub claude` also observes native session and turn
77
+ boundaries when facts injection and task-idle sweeps are off. Keep the managed
78
+ hooks enabled to measure completion and supervision usage. Passing your own
79
+ `--settings` takes precedence and produces a warning when it replaces that
80
+ observation. The launcher records its private session identity in ordinary
81
+ terminals too; this does not grant terminal-recovery authority.
82
+ Ordinary Claude sessions with turn-free facts or task-idle sweeps enabled get the
83
+ same session/start observation. Other ordinary launches remain non-opt-in.
84
+
85
+ `ahub report` records conductor actions and completed native turns containing
86
+ supervision. Tokens describe the whole measured turn, which may also contain
87
+ other work; they are not a per-notice cost estimate. Missing measurements stay
88
+ unknown. Use the live smoke ledger to assess observed turns and tokens per
89
+ approved task before choosing `all`; an unmeasured run is not a cost benchmark.
90
+ Claude turn counts use authenticated native completion events. Older logical
91
+ state counts are labelled; an idle channel or approved task alone does not prove
92
+ that the native answer finished.
93
+ The Stop hook acknowledgement is not a counted completion. The daemon checks the
94
+ native transcript after the hook can return; missing or changed-session evidence
95
+ stays unknown.
96
+
6
97
  ## Install and start
7
98
 
8
99
  Use Bun 1.3 or newer. Install the released package and install its Claude
@@ -14,12 +105,11 @@ ahub setup
14
105
  cd <project>
15
106
  ahub init
16
107
  ahub up
17
- ahub tail
18
108
  ```
19
109
 
20
110
  `ahub init` writes the project configuration and managed instruction blocks.
21
111
  Run it after an upgrade when those blocks need refreshing. `ahub setup`
22
- updates the shared Claude plugin. Keep `ahub tail` open when a local worker
112
+ updates the shared Claude plugin. Keep `ahub console` open when a local worker
23
113
  may request an approval.
24
114
 
25
115
  `.agenthub/config.json` can be committed and shared. The fields that choose
@@ -521,6 +611,13 @@ stays the default until an evaluation says otherwise (`docs/cooperbench.md`).
521
611
 
522
612
  Inspect permission requests in the terminal:
523
613
 
614
+ ```bash
615
+ ahub console
616
+ ```
617
+
618
+ Select the requested option in the console and confirm an allow decision.
619
+ For a separate plain terminal, the existing command remains available:
620
+
524
621
  ```bash
525
622
  ahub tail
526
623
  ahub permit <request-id> allow
@@ -647,21 +744,21 @@ Rows without a live process are stale registrations; forget them with
647
744
 
648
745
  Upgrade running projects with the target release's own coordinator. It accepts
649
746
  a running source on control protocol 9 (0.6.x), 10 (0.7.0 through 0.12.0),
650
- 11 (0.12.1 and 0.12.2), 12 (0.12.3), 13 (0.12.4 through 0.12.15) or 14 (0.12.16), and only
747
+ 11 (0.12.1 and 0.12.2), 12 (0.12.3), 13 (0.12.4 through 0.12.15) 14 (0.12.16) or 15 (0.12.17), and only
651
748
  a target on its own protocol, so the target's coordinator fits every supported
652
749
  source and carries every recovery fix released up to it. Protocol 8 and older
653
750
  (0.5.x and earlier) are refused as `manual-bootstrap-required`. Run from the
654
751
  project directory, without replacing the global CLI first:
655
752
 
656
753
  ```bash
657
- bunx --package @staix/agent-hub@0.12.16 ahub upgrade --to 0.12.16 --dry-run
658
- bunx --package @staix/agent-hub@0.12.16 ahub upgrade --to 0.12.16 --yes
754
+ bunx --package @staix/agent-hub@0.12.17 ahub upgrade --to 0.12.17 --dry-run
755
+ bunx --package @staix/agent-hub@0.12.17 ahub upgrade --to 0.12.17 --yes
659
756
  ```
660
757
 
661
758
  | Running now | Coordinator to use |
662
759
  | --- | --- |
663
760
  | 0.6.x (protocol 9) | the target's, through `bunx` as above |
664
- | 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.15 (protocol 13), 0.12.16 (protocol 14) | the target's, through `bunx` as above |
761
+ | 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.15 (protocol 13), 0.12.16 (protocol 14), 0.12.17 (protocol 15) | the target's, through `bunx` as above |
665
762
  | any supported source, with the installed CLI already at the target | `ahub upgrade` below, which is the same coordinator |
666
763
  | 0.5.x or earlier (protocol 8 and older) | not supported: bootstrap by hand with the matching CLI |
667
764
 
@@ -693,14 +790,14 @@ projects first:
693
790
 
694
791
  ```bash
695
792
  ahub restart --dry-run
696
- ahub upgrade --to 0.12.16 --dry-run
793
+ ahub upgrade --to 0.12.17 --dry-run
697
794
  ```
698
795
 
699
796
  Apply only after reviewing the plan:
700
797
 
701
798
  ```bash
702
799
  ahub restart --yes
703
- ahub upgrade --to 0.12.16 --yes
800
+ ahub upgrade --to 0.12.17 --yes
704
801
  ahub recovery status <operation-id>
705
802
  ahub recovery resume <operation-id>
706
803
  ahub recovery abort <operation-id>
@@ -17,7 +17,7 @@ ahub setup # installs the Claude Code channel plu
17
17
  Or use the matching GitHub release:
18
18
 
19
19
  ```bash
20
- bun add -g github:STAIxBWLB/agent-hub#v0.12.16
20
+ bun add -g github:STAIxBWLB/agent-hub#v0.12.17
21
21
  ahub setup
22
22
  ```
23
23
 
@@ -34,8 +34,8 @@ In your project directory, one terminal each:
34
34
 
35
35
  ```bash
36
36
  ahub init # .agenthub/config.json, .agenthub/routing.toml, the marker block in AGENTS.md
37
- ahub up # the daemon for this directory (loopback only)
38
- ahub tail # keep open: the conversation, state changes, permission requests
37
+ ahub up # starts the daemon and opens the operator console in a terminal
38
+ ahub console # use in another terminal if up used --no-console; Tab opens panels
39
39
  ahub kimi # Kimi, headless
40
40
  ahub codex # Codex TUI, attached through the hub
41
41
  ahub claude # Claude Code with the hub channel
package/docs/security.md CHANGED
@@ -28,6 +28,30 @@ agent-hub connects agents that can each run commands. This page says what the hu
28
28
  - **Network-level proof for PII.** The hub proves at its own boundaries (tests search every output) that PII text does not leave; packet-level verification of your gateway path is an operations check.
29
29
  - **Other operating systems' sandboxes.** Only macOS seatbelt is implemented.
30
30
 
31
+ ## Agent CLI identity and conductor authority
32
+
33
+ Inside an agent session, `ahub` identifies the caller from `AGENTHUB_PEER_ID`,
34
+ or the native Claude/Codex shell marker when the hub marker is absent. It connects
35
+ as that peer in tools mode. Messages and task changes retain that actor;
36
+ `say` defaults to status priority and important messages still require the peer's
37
+ capability. Conflicting or malformed markers fail closed. There is no `--as-user`.
38
+ Human-only operations are refused before connecting, including permission answers,
39
+ queue resolution, budget overrides, lifecycle and recovery operations, and `ask`.
40
+ The operator uses a plain terminal or `ahub console`. A Claude `!` command that
41
+ inherits the agent markers follows the same rule.
42
+
43
+ This is an honest default against accidental impersonation and injected commands.
44
+ It does not make the token inaccessible to an agent with unrestricted project
45
+ shell access. The existing OS sandbox and loopback authentication boundaries apply.
46
+
47
+ Steering tools require an explicit `conductor` role, independent of default-allow
48
+ capabilities. Only one peer may hold that role. A conductor may inspect public
49
+ state, assign or escalate work, start supported headless peers and place its own
50
+ delivery holds. It cannot answer approvals, resolve durable deliveries, override
51
+ budget pauses or release a human hold. Pending approval summaries exclude titles;
52
+ PII tasks remain public stubs. Conduct events contain ids only. Supervision feeds
53
+ use structured reasons and never carry check output or approval bodies.
54
+
31
55
  ## Reporting
32
56
 
33
57
  Please report vulnerabilities privately through GitHub's "Report a vulnerability" on this repository rather than in a public issue.
package/docs/smoke.md CHANGED
@@ -2,6 +2,99 @@
2
2
 
3
3
  `scripts/check.sh` covers everything against fakes. The legs below need real accounts and an interactive terminal, so they are run by hand and recorded here.
4
4
 
5
+ ## Operator console and conductor candidate (#190, #191, #193-#195)
6
+
7
+ Candidate 0.12.17, protocol 15, observed on 2026-10-09 KST. The final Claude
8
+ observer runs used immutable source `20ad5e6`; their native completion receipts
9
+ and actual console effects were checked independently. The Kimi new-tool leg
10
+ remains quota-blocked, as recorded below.
11
+
12
+ - Actual Codex 0.146.0 (`gpt-5.5`) and Claude 2.1.295 (Opus 5.5) TUIs started
13
+ real local and headless Pi peers, proposed exactly two tasks owned initially
14
+ by local, reassigned Beta to Pi, and placed and released their own holds.
15
+ The real owners checked their outputs and called `hub_task_done`; the native
16
+ conductor independently read the exact files before approving both tasks.
17
+ Real CLI calls retained the agent actor, and human-only queue/permission
18
+ actions were refused from an agent shell.
19
+ - The person authorized `alpha.txt = ALPHA` (5 bytes) and `beta.txt = BETA`
20
+ (4 bytes), both without a newline. Separate selection and confirmation keys
21
+ reached the actual console PTY and produced allow-once console audit records.
22
+ Out-of-scope directory-list commands were cancelled. The final Claude own
23
+ run used three allow-once answers because Pi combined its exact write and
24
+ bounded byte checks in one approved command; it also recorded two cancelled
25
+ local directory-list requests. No permission or task completion was fabricated.
26
+ - The final Claude own run recorded nine completed native turns and nine
27
+ matching daemon Stop records, 27 unique native usage records, and
28
+ 1,888,242 tokens including cached input. Eight completed native turns contained
29
+ supervision, with 1,292,922 recorded tokens. Its final message UUID/id,
30
+ `end_turn`, `turn_duration`, instance/session/launcher and opaque Stop receipt
31
+ matched; native idle and an empty pending delivery queue were verified before
32
+ owned-process cleanup. Local recorded three turns and 72,046 gateway tokens;
33
+ Pi recorded two turns and 47,177 native tokens.
34
+ - A read-only Claude feed-off continuation on the final observer preserved the
35
+ existing approved tasks and exact files. It verified one completed native turn,
36
+ four usage records and 263,413 cached-inclusive tokens, with matching daemon
37
+ Stop, native idle and delivery settlement. An obsolete review hold was first
38
+ discarded through the authenticated public operator API after fresh approved
39
+ task/history and exact BETA readback. That discard is not native execution.
40
+ - A real console with two pending cards and no operator input consumed 0.07 CPU
41
+ seconds over 40.018 wall seconds in stream mode (0.175% of one CPU). A separate
42
+ actual Peers panel, 120x40, with two pending cards and no operator input used
43
+ 0.07 CPU seconds over 40.012 wall seconds (0.175%). These are cumulative `ps`
44
+ process CPU samples in distinct modes/runs, not whole-host idle measurements.
45
+ - In a plain Claude TUI without the development-channel flag, one actual
46
+ `hub_status` MCP call succeeded and the daemon audited the Claude status
47
+ action. A unique directed operator push was accepted by the bridge, but no
48
+ pushed native user row or answer appeared in the observed 20.075-second window.
49
+ This proves tool access separately from the bounded negative push observation.
50
+ The real channel-enabled runs above received actual review and supervision
51
+ pushes. The plain probe and cleanup completed in 56.724 seconds, with no model
52
+ file, shell, task or settings actions.
53
+ - Real Kimi Code CLI 2.1.1 completed ACP initialization and `session/new`, then
54
+ rejected the single prompt with HTTP 403 for its weekly account usage limit.
55
+ No requested native tool event or role-refusal result occurred. The reset time
56
+ was not supplied. The owning CLI listed only the managed OAuth Kimi provider
57
+ (four models, default `kimi-code/k3`), so no configured alternative provider was
58
+ found. This is an incomplete external prerequisite, not a new-tool pass. No
59
+ purchase, provider/configuration change or model retry was performed; owned
60
+ processes were stopped after 3.142 seconds.
61
+
62
+ The user explicitly approved release 0.12.17 with only the Kimi native new-tool
63
+ and ordinary-role refusal T0 evidence deferred on 2026-10-09. That native
64
+ prerequisite remains unverified; the quota rejection is not a tool pass. This
65
+ decision does not defer the other native/source gates or authorize a purchase,
66
+ credentials, provider/configuration change or account retry. The same isolated
67
+ read-only Kimi probe remains required before its native coverage is marked verified.
68
+
69
+ Earlier captures remain part of the evidence:
70
+
71
+ - Codex feed-off included an interrupted original and read-only continuation:
72
+ six logical native turns, 49 increments and 1,979,016 tokens. Codex feed-own
73
+ recorded nine turns, 34 increments and 1,399,470 tokens. Their journals were
74
+ independently checked after cleanup. The first baseline interruption and
75
+ operator reconciliation are preserved; neither native task completion nor
76
+ board approval was substituted by the harness.
77
+ - The original Claude runs approved the actual files but an early idle-based
78
+ harness stopped their final review answers. Subsequent `da2ebbc` feed-off
79
+ completed five native turns and 1,451,782 tokens but certified only one daemon
80
+ Stop. A `6034dd9` read-only continuation completed another native turn and
81
+ 261,483 tokens, while its Stop remained unavailable during the native hook.
82
+ Both incomplete observer captures are preserved alongside raw and audited
83
+ summaries. The final post-ACK observer above closes the fresh completion proof;
84
+ it does not retroactively certify the earlier missed Stop records.
85
+
86
+ All token totals describe whole native/model turns, including cached input and
87
+ other work. Operator waiting, cancelled requests, marker probes, interrupted
88
+ continuations and differing source revisions make these observations unsuitable
89
+ for a causal feed-overhead or model-efficiency comparison. Unknown measurements
90
+ and the Kimi prerequisite remain explicit.
91
+
92
+ The harness uses real Python PTYs. Operator-file-input forwards only
93
+ chat-authorized keys, removes inherited Orca terminal ownership from fixture
94
+ children, and generates no approval automatically. Private original and
95
+ continuation captures remain separate. The shell-marker probes and vendor
96
+ limits are recorded in [the identity T0 ledger](verification/2026-10-09-agent-shell-t0.md).
97
+
5
98
  ## Approval race live reproduction and candidate verification (#98)
6
99
 
7
100
  Measured on 2026-10-02 KST with installed 0.12.1 and the correction candidate,
@@ -1343,6 +1343,125 @@ whether to continue the native session or restart; this change provides no
1343
1343
  automatic session replacement. Status, tail and dashboard expose readings with
1344
1344
  source, measurement time and freshness beside quota information.
1345
1345
 
1346
- The control contract is protocol 14. Recovery sources 9 through 13 remain
1346
+ Release 0.12.16 uses control protocol 14. Recovery sources 9 through 13 remain
1347
1347
  supported; protocol 13 identifies releases 0.12.4 through 0.12.15, while
1348
1348
  0.12.16 uses protocol 14.
1349
+
1350
+ ## Operator console and conductor (issues #190, #191, #193, #194, #195)
1351
+
1352
+ `ahub console` owns one authenticated console connection. It shares tail's
1353
+ renderer, renders a DECSTBM stream and footer, and provides optional alternate
1354
+ screen panels. `up` opens it only with terminal input and output, unless
1355
+ `--no-console` is given. Tail remains a plain stream. Console commands use an
1356
+ allowlisted argv with stdin closed; lifecycle and native launches are excluded.
1357
+ Allow options require confirmation and never interpret nonempty input as an
1358
+ approval. Deny is direct. Pending requests include expiry and are withdrawn by
1359
+ `permission_closed` on any answer or cancellation. Answer audits have no title.
1360
+ Panels expose Peers, Approvals, Tasks, Queue and Events with bounded polling and
1361
+ Unicode cell widths, fall back below 80x24, and restore terminal state on exit.
1362
+
1363
+ Agent shell CLI calls connect as the detected peer in tools mode. Hub launches
1364
+ set `AGENTHUB_PEER_ID`; the pinned installed Codex shell injects
1365
+ `CODEX_THREAD_ID` after environment filtering, and Claude uses `CLAUDECODE`.
1366
+ Malformed or conflicting markers fail closed. Console-only commands are denied
1367
+ before connecting, with no as-user escape. Refusals enter a bounded ids-only
1368
+ local audit spool which a running daemon consumes; this preserves the no-connect
1369
+ rule while making refusals visible on the console and in the log. A stopped
1370
+ daemon cannot show a live notice; it consumes remaining records on startup.
1371
+ This does not establish a security boundary against an unrestricted shell
1372
+ which can read the token. Native source evidence and live probe results remain
1373
+ separate in the smoke ledger.
1374
+
1375
+ Exactly one explicit conductor role may be configured. Default-allow
1376
+ capabilities do not grant it; role authority is checked on each operation and
1377
+ refreshed on a subsequent connection. Shared MCP tools provide public status,
1378
+ public task history, actor-preserving assign/escalate, headless local/Kimi/Pi
1379
+ start, and conductor-owned persistent holds. TUI starts return launch commands.
1380
+ Assign/escalate need `assign` if the peer has a capabilities entry. Approval
1381
+ answers, durable queue resolution, budget overrides and lifecycle remain human
1382
+ operations. A conductor cannot release human or budget holds. Audit events are
1383
+ ids-only and report counts their actions.
1384
+
1385
+ The conductor feed defaults to own tasks, also supports all and off, and uses
1386
+ the existing bus digest window. Repeated queued task/kind milestones replace
1387
+ the earlier pending notice; accepted or in-flight deliveries are preserved.
1388
+ Structured milestones carry public titles or PII stubs, never raw history
1389
+ notes or check output. Only aged approval summaries and needs-review delivery
1390
+ ids are important. Role/feed revocation withdraws pending feed notices. A
1391
+ completed task set produces one round notice until a new task joins; subsequent
1392
+ rounds count only newly joined tasks. Pure status supervision queues wait for
1393
+ the existing digest deadline rather than flushing at the ordinary batch-count
1394
+ threshold. Mixed traffic retains the ordinary admission behavior. The existing
1395
+ 10-original digest ceiling and 200-entry queue ceiling still bound admission;
1396
+ larger windows may require multiple bounded deliveries.
1397
+
1398
+ Supervision cost counts completed native turns that received feed notices,
1399
+ with whole-turn token readings when available. Other work may share a turn;
1400
+ these readings are not per-notice token attribution. Missing measurements are
1401
+ unknown. Native TUI conductor, sandbox and feed-on smoke results must be recorded
1402
+ as observed outcomes, separately from unit/fake protocol tests.
1403
+
1404
+ Managed Claude launches with turn-free facts, task-idle sweeps or a conductor
1405
+ role install native observation hooks. A conductor receives them even when facts
1406
+ injection and task-idle sweeps are disabled; ordinary non-opt-in launches remain
1407
+ unchanged. SessionStart registers the native
1408
+ session and UserPromptSubmit starts observation; PreToolUse keeps a tool turn
1409
+ active and Stop closes it. Explicit caller settings remain authoritative and
1410
+ produce a warning when they replace these hooks. An ordinary terminal records a
1411
+ private launcher identity without creating an Orca terminal-recovery record.
1412
+ The facts control request accepts session/start/pre/post/stop phases and carries
1413
+ nativeInstanceId and nativeLaunchId alongside sessionId and transcriptPath.
1414
+ The command hook forwards launcher identity only for its matching state directory
1415
+ and peer; library calls targeting another hub do not inherit that identity.
1416
+ The daemon fences observations to its current instance and launcher, validates
1417
+ the session transcript, and rejects stale stop/post observations. Native prompt
1418
+ text is never included in those requests or events.
1419
+
1420
+ A genuine native Stop requires the current daemon/launcher binding and an actual
1421
+ assistant end_turn transcript message at or after the native turn's first start,
1422
+ distinct from the completed-message baseline recorded at that start. Later tool
1423
+ activity does not move this turn-start boundary. Missing start evidence or a
1424
+ start from another session, launch or channel claim leaves completion unknown.
1425
+ Accepted completion consumes that start before the peer becomes idle, so another
1426
+ message cannot reuse it. A valid bound Stop request receives one prompt
1427
+ acknowledgement with pending=true, within the existing two-second hook deadline.
1428
+ The acknowledgement records observation only, never completion. After the hook
1429
+ can return, a deferred observer waits up to 1200 monotonic milliseconds for its
1430
+ transcript append to become visible. Each read and final consumption revalidate
1431
+ the captured start, session,
1432
+ launch, peer and claim. A seen previous-turn baseline still waits while a new
1433
+ current start exists; a consumed duplicate remains a no-op. Timeout, malformed or
1434
+ oversized evidence and superseded context leave completion unknown. This wait
1435
+ does not retry a user action or relax authority.
1436
+ The observer sends no second request reply and never changes global hook settings.
1437
+ The live harness also binds its final receipt to the current private launch id.
1438
+ Its opaque
1439
+ deduplication id binds session, launch and message; transport replacement or new
1440
+ activity cannot turn a replay into another completion. Unbound or idless legacy
1441
+ Stop events cannot certify completion or finish supervision. A current private
1442
+ launcher marker takes precedence over an older terminal-recovery launch id for
1443
+ native observation, while preserving the recovery records and their authority.
1444
+ Claude turn reports
1445
+ prefer unique native Stop events; historical logical-state counts are labelled
1446
+ as such, and an unobserved completion remains unknown. Channel idle, watchdog
1447
+ expiry, board approval and a tool reply do not independently establish native
1448
+ completion. The live harness waits for the current fixture/instance/session's
1449
+ final transcript end_turn and turn_duration after its last review before
1450
+ reporting a completed case or terminating its native TUI. It also requires the
1451
+ matching opaque native completion event, an idle peer and settled delivery rows.
1452
+
1453
+ These additions use control protocol 15. Supported recovery sources include
1454
+ protocol 14 (0.12.16) alongside the previous source protocols.
1455
+
1456
+ For release 0.12.17 only, the user explicitly accepted deferral of real Kimi ACP
1457
+ new-tool invocation and ordinary-role refusal T0 evidence on 2026-10-09. The
1458
+ managed OAuth account rejected its prompt with HTTP 403 for the weekly quota
1459
+ before any tool event; reset time is unknown and no configured native alternative
1460
+ was found. This prerequisite remains unverified. Actual Claude/Codex conductor
1461
+ workflows, Claude shell/channel observations, shared MCP/fake-adapter authority
1462
+ checks, real Kimi ACP initialization/session creation and full source gates are
1463
+ separately qualified. The deferral authorizes no purchase, credentials,
1464
+ provider/configuration change, account retry or authority relaxation. Once quota
1465
+ is available under an authorized account, the same isolated read-only status-tool
1466
+ and ordinary-role refusal probe must record a native tool event and daemon result
1467
+ before this Kimi prerequisite can be marked verified.