claude-slack-channel-bots 0.8.2 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -6,15 +6,23 @@ A single HTTP MCP server that holds one Slack Socket Mode connection and routes
6
6
 
7
7
  ## Quick Start
8
8
 
9
+ **[Bun](https://bun.sh) is required** — it is CSCB's runtime and the interpreter its install and start scripts run under. Install bun before you install CSCB. `npm install -g claude-slack-channel-bots` on a box without bun fails during postinstall (the postinstall script is a bun script).
10
+
9
11
  1. **Install globally via bun:**
10
12
 
11
13
  ```sh
12
14
  bun install -g claude-slack-channel-bots
13
15
  ```
14
16
 
15
- The postinstall script creates skeleton config files in `~/.claude/channels/slack/`.
17
+ 2. **Trust the package so postinstall runs:**
18
+
19
+ ```sh
20
+ bun pm -g trust claude-slack-channel-bots
21
+ ```
22
+
23
+ Bun blocks the lifecycle scripts of untrusted packages, so the postinstall does not run on the plain `install` above — you must trust the package for it to fire. The `-g` flag targets the global install; without it `bun pm trust` looks for a `package.json` in the current directory and errors with `No package.json was found`. (Run `bun pm -g untrusted` to confirm it is listed first.) The postinstall then creates skeleton config files in `~/.claude/channels/slack/`. Skip this step and a later `start` fails with `missing prerequisite: config.json`.
16
24
 
17
- 2. **Run the setup skill:**
25
+ 3. **Run the setup skill:**
18
26
 
19
27
  The package includes a Claude Code skill at `skills/setup-slack-channel-bots/` that walks you through the entire configuration. Copy or symlink it into `~/.claude/skills/`, then run:
20
28
 
@@ -24,7 +32,7 @@ A single HTTP MCP server that holds one Slack Socket Mode connection and routes
24
32
 
25
33
  It handles Slack app creation, tokens, routing, access control, hooks, and validation — and skips anything already configured.
26
34
 
27
- 3. **Start the server:**
35
+ 4. **Start the server:**
28
36
 
29
37
  ```sh
30
38
  claude-slack-channel-bots start
@@ -136,6 +144,10 @@ Tokens and runtime options are read from environment variables. There is no `.en
136
144
  | `SLACK_STATE_DIR` | Override the directory where `config.json`, `access.json`, and runtime state are stored. Defaults to `~/.claude/channels/slack`. |
137
145
  | `SLACK_ACCESS_MODE` | Set to `static` to load `access.json` once at startup and cache it for the lifetime of the process rather than re-reading it on every event. Useful in high-throughput environments where disk reads are a concern. |
138
146
  | `SLACK_DRY_RUN` | Set to `1` to start the server without Slack credentials. Token validation is skipped, Socket Mode and `web.auth.test()` are not called, and MCP tool calls (`reply`, `react`, etc.) are logged instead of sent. Useful for integration testing. |
147
+ | `CSCB_LOG_MAX_BYTES` | Rotate `server.log` / `clean_restart.log` when the active file reaches this many bytes. Defaults to `10485760` (10 MiB). Values `<= 0` or non-numeric are ignored. |
148
+ | `CSCB_LOG_KEEP` | Number of rotated generations to retain (`server.log.1` … `server.log.N`). Defaults to `5`. Set to `0` to keep none (the log is truncated instead of rolled). Values `< 0` or non-numeric are ignored. |
149
+ | `CSCB_AD_VERBOSE` | Set to a truthy value (`1`, `true`, `yes`, `on`) to restore the agent-director library's per-poll `SubprocessClient: <verb> ok` success dumps in `server.log`. Off by default — these routine dumps are dropped so the log stays readable. Failures and warnings from agent-director always pass through regardless of this flag. Read once at server startup, so it takes effect on server restart. |
150
+ | `CSCB_HTTP_VERBOSE` | Set to a truthy value (`1`, `true`, `yes`, `on`) to restore the per-request MCP access line (`HTTP <method> <path> session=…`) in `server.log`. Off by default — the `/mcp` endpoint is hit on every client poll and SSE open, so these routine lines are dropped to keep the log readable. Session connect/disconnect, route mismatches, and errors are logged unconditionally regardless of this flag. Checked per request, so it takes effect without a restart. |
139
151
 
140
152
  Shell profile example:
141
153
 
@@ -185,7 +197,7 @@ A skeleton file is created by postinstall. Populate it before running `start`.
185
197
  | `routes` | object | required | Map of Slack channel ID → route entry. Each entry requires a `cwd` field: the working directory for that session. Used to identify sessions via `roots/list` after MCP handshake. `~` is expanded. Each `cwd` must be unique across all routes. May also include an optional `claude_config_dir` string (see below). |
186
198
  | `default_route` | string | — | CWD path to use when a message arrives on a channel with no explicit entry in `routes`. Must match an existing route `cwd`. Channels that are in `routes` but whose session is not yet registered have their messages dropped — they do not fall back to `default_route`. |
187
199
  | `default_dm_session` | string | — | CWD path of the session that handles direct messages. Must match an existing route `cwd`. |
188
- | `bind` | string | `"127.0.0.1"` | Interface the HTTP server binds to. Use `"0.0.0.0"` to expose on all interfaces. |
200
+ | `bind` | string | `"127.0.0.1"` | Interface the HTTP server binds to. Use `"0.0.0.0"` to expose on all interfaces. The in-process cron scheduler delivers via `127.0.0.1`, so `bind` must include loopback (the default, or `0.0.0.0`) for scheduled fires to work. |
189
201
  | `port` | number | `3100` | Port the HTTP server listens on. |
190
202
  | `session_restart_delay` | number | `60` | Seconds to wait before auto-restarting a dead session. Set to `0` to disable auto-restart. Must be non-negative. |
191
203
  | `health_check_interval` | number | `120` | Seconds between periodic liveness polls. Set to `0` to disable. Must be non-negative. |
@@ -197,8 +209,12 @@ A skeleton file is created by postinstall. Populate it before running `start`.
197
209
  | `cozempic_prescription` | string | `"standard"` | Cozempic cleaning intensity before resume. Valid values: `gentle`, `standard`, `aggressive`. Has no effect if cozempic is not installed. |
198
210
  | `message_archive_db` | string | — | Path to a SQLite DB where every inbound Slack message is archived in real time. Parent directories are created if missing; schema is initialized on first open. Compatible with the `archive-messages.py` backfill script — both can write concurrently. Feature is disabled when absent. |
199
211
  | `claude_config_dir` | string | — | Path to a Claude on-disk config directory. When set, managed sessions launch with `CLAUDE_CONFIG_DIR='<resolved-path>'` so the bot authenticates against a specific account. `~` is expanded and the path is resolved to absolute. Per-route `routes[id].claude_config_dir` overrides this top-level value for individual channels. When neither is set, Claude's own default applies. Must be non-empty when set. |
200
- | `resume_enabled` | boolean | `true` | When `false`, the session manager always performs a fresh Claude session launch instead of resuming, both on startup and on runtime auto-restart, even when a stored session ID exists. Disabling this skips the `--resume` flag entirely. Use this as a workaround if your Claude Code version crashes with "sandbox required but unavailable" on `--resume` (a known regression in v2.1.120). |
212
+ | `resume_enabled` | boolean | `true` | When `true` (default), a bot whose session died — including after a host reboot or pod resume — comes back with its prior conversation history intact instead of starting fresh. When `false`, the session manager always performs a fresh launch instead of resuming, both on startup and on runtime auto-restart, even when a stored session exists. Set `false` as a workaround if your Claude Code version crashes with "sandbox required but unavailable" on resume (a known regression in v2.1.120). Requires a system-installed `agent-director` ≥ 0.8.0 for reboot recovery to actually restore history. |
201
213
  | `agent_director_poll_interval_ms` | number | `1000` | Poll interval (ms) for the agent-director permission relay tick. Must be a positive integer in `[200, 3_600_000]`. Replaces the pre-rename `claude_director_poll_interval_ms` — the old name is rejected at startup. Unknown top-level config fields are also rejected to surface stale configs after the rename. |
214
+ | `stop_hook_bootstrap` | boolean | `true` | Controls whether the server installs the CSCB-managed Slack Reply Guard Stop hook into `<claude_config_dir>/settings.json` at boot (see [Slack Reply Guard (Stop hook)](#slack-reply-guard-stop-hook)). Set to `false` to disable installation for every route and to actively remove any previously-installed managed entry. Per-route `routes[id].stop_hook_bootstrap` overrides this top-level value. Non-boolean values are rejected by config validation at startup. |
215
+ | `cron_table_path` | string | `<config dir>/crontab` | Path to the crontable for the built-in cron scheduler (`cscb_cron`). Defaults to `crontab` in the directory of the loaded `config.json`. `~` is expanded like other path keys. The resolved path is exported into every managed session as `CSCB_CRONTABLE_PATH` so bots can find the crontable and self-schedule (see [Scheduled Prompts](#scheduled-prompts-cscb_cron)). Must be a non-empty string when set. Changing it requires a server restart. |
216
+ | `cron_log_path` | string | `<config dir>/cron.log` | Path to the `cscb_cron` log file. Defaults to `cron.log` in the directory of the loaded `config.json`. `~` is expanded like other path keys. Must be a non-empty string when set. Changing it requires a server restart. |
217
+ | `cron_log_max_bytes` | number | — | Size cap in bytes for the cron log. Must be a positive integer when set. Cron-log pruning is disabled when absent. Changing it requires a server restart. |
202
218
 
203
219
  #### Per-route `claude_config_dir` override
204
220
 
@@ -221,6 +237,28 @@ When you want different bot sessions to authenticate as different Claude account
221
237
 
222
238
  `C_PERSONAL` launches with the Max account; `C_CORPORATE` falls through to the top-level value and uses the corporate account. Use `claude auth login --claudeai` (or `--console`) with `CLAUDE_CONFIG_DIR` set to the same directory to populate each config dir before starting the server.
223
239
 
240
+ #### Per-route `stop_hook_bootstrap` override
241
+
242
+ Set `stop_hook_bootstrap` on an individual route to override the top-level default for that one bot. Per-route values win over the top-level value; routes without their own value inherit the top-level default (which is itself `true` when absent).
243
+
244
+ ```json
245
+ {
246
+ "routes": {
247
+ "C_EDIT_ONLY_BOT": {
248
+ "cwd": "~/projects/gamma",
249
+ "claude_config_dir": "~/.claude-gamma",
250
+ "stop_hook_bootstrap": false
251
+ },
252
+ "C_NORMAL_BOT": {
253
+ "cwd": "~/projects/delta",
254
+ "claude_config_dir": "~/.claude-delta"
255
+ }
256
+ }
257
+ }
258
+ ```
259
+
260
+ A per-route opt-out only *fully* disables the guard for that bot when the route owns a **dedicated** `claude_config_dir` — see [Shared-dir aggregation](#shared-dir-aggregation) for the interaction when routes share a dir.
261
+
224
262
  ---
225
263
 
226
264
  ### Access Control (access.json)
@@ -317,6 +355,16 @@ Behavior by case:
317
355
  - **Stale PID file** (process no longer running): removes the PID file, prints `server is not running (removed stale PID file)`, exits 0.
318
356
  - **Live process:** sends `SIGTERM`, polls for exit for up to `stop_timeout` seconds (default 30s). Prints `[slack] Server stopped.` on clean exit. Escalates to `SIGKILL` if the process does not exit within `stop_timeout`.
319
357
 
358
+ Plain `stop` leaves the managed bots running — they are meant to survive a server restart. Pass `--stop-bots` to gracefully exit the bots first:
359
+
360
+ ```sh
361
+ claude-slack-channel-bots stop --stop-bots
362
+ ```
363
+
364
+ This mirrors `clean_restart`'s order: the server is stopped **first**, then the bot teardown runs for each route — pause the bot, poll until it exits (or up to `exit_timeout` seconds), then force-kill on timeout. Teardown kills but never deletes each row, preserving its `claude_session_id` so the bots can resume their conversation history on the next start. Stopping the server first prevents its `onsessionclosed`/`scheduleRestart` handler from respawning a just-exited bot mid-teardown (which would delete its `ended` row and history). Use it when you want a clean, flushed shutdown of the bots (for example before a host reboot).
365
+
366
+ If agent-director is unreachable, the teardown **fails loudly** — the command prints the error and exits non-zero rather than silently reporting a clean stop. (A missing config is best-effort: teardown is skipped but the server stop still succeeds, since the server is already down.)
367
+
320
368
  ### `claude-slack-channel-bots clean_restart`
321
369
 
322
370
  Gracefully exits all managed Claude Code sessions, then stops and starts the server.
@@ -325,12 +373,15 @@ Gracefully exits all managed Claude Code sessions, then stops and starts the ser
325
373
  claude-slack-channel-bots clean_restart
326
374
  ```
327
375
 
328
- For each configured route, calls `client.pause({claude_instance_id})` via agent-director and polls `client.status(...)` until the spawn transitions to `ended` / `missing` (or `client.list(...)` returns no row). If the spawn does not exit within `exit_timeout` seconds (default 120s), the spawn is force-killed via `client.kill(...)`. All routes are processed in parallel. Individual session errors are logged and do not abort the restart. After the server restarts, the SR-1.4 collision-then-act dispatcher decides resume-vs-fresh per route — agent-director owns Claude session-id state, not CSCB.
376
+ For each configured route, calls `client.pause({claude_instance_id})` via agent-director and polls `client.status(...)` until the spawn transitions to `ended` / `missing` (or `client.list(...)` returns no row). If the spawn does not exit within `exit_timeout` seconds (default 120s), the spawn is force-killed via `client.kill(...)`. Teardown kills but never deletes each row, preserving its `claude_session_id` so bots resume their conversation history on the next start. All routes are processed in parallel. After the server restarts, the SR-1.4 collision-then-act dispatcher decides resume-vs-fresh per route — agent-director owns Claude session-id state, not CSCB.
377
+
378
+ A benign kill outcome — the row already being gone — is tolerated per-route and does not abort the restart. Any other per-route teardown failure, including a pause failure that escalates to a kill which then fails to reach agent-director, is fatal: it fails loudly and aborts the restart (non-zero exit).
329
379
 
330
380
  Behavior by case:
331
381
 
332
382
  - **No configured routes:** skips the shutdown phase and proceeds directly to stop/start.
333
383
  - **Server already stopped:** `stop` reports `server is not running`; `start` then brings up a fresh server.
384
+ - **agent-director unreachable:** teardown fails loudly and the restart is aborted (non-zero exit); no new server is started. The `no spawn row` message appears only when a route genuinely has no spawn, never when the client failed to reach agent-director.
334
385
 
335
386
  ### PID file
336
387
 
@@ -423,6 +474,8 @@ On success, returns HTTP 200:
423
474
 
424
475
  ### Example: crontab reminder
425
476
 
477
+ For recurring prompts, prefer the built-in scheduler (see [Scheduled Prompts](#scheduled-prompts-cscb_cron)) — it needs no host cron and delivers straight into a bot channel. The host-crontab example below is an alternative when you already run `cron`:
478
+
426
479
  ```sh
427
480
  # crontab -e
428
481
  0 9 * * 1 curl -s -X POST http://localhost:3100/interject \
@@ -432,6 +485,124 @@ On success, returns HTTP 200:
432
485
 
433
486
  ---
434
487
 
488
+ ## Scheduled Prompts (cscb_cron)
489
+
490
+ The server fires scheduled prompts into bot channels once per minute, reading them from a crontable. Each fire is delivered as an `/interject` message into the target channel, exactly as if a script had POSTed it.
491
+
492
+ ### The crontable
493
+
494
+ Schedules live in the crontable file at `cron_table_path` (default `<config dir>/crontab`, where `<config dir>` is the directory of your loaded `config.json`; override it with the `cron_table_path` key in `config.json`). The server creates the file on first boot if it is absent, with a self-documenting comment header describing the line format. See [Crontable format](#crontable-format) below for the full reference. With the default config location the crontable is at `~/.claude/channels/slack/crontab`:
495
+
496
+ ```sh
497
+ cat ~/.claude/channels/slack/crontab
498
+ ```
499
+
500
+ ### Crontable format
501
+
502
+ On first boot the server auto-creates the crontable with this self-documenting header:
503
+
504
+ ```
505
+ # CSCB crontable — scheduled prompts for the Slack channel bots.
506
+ #
507
+ # One schedule per line. Fields are positional and whitespace-delimited:
508
+ #
509
+ # <min> <hour> <dom> <mon> <dow> <prompt-path> [<channel-id>[,<channel-id>...]]
510
+ #
511
+ # tokens 1-5 : a standard 5-field cron expression (minute hour day-of-month
512
+ # month day-of-week).
513
+ # token 6 : path to the prompt file to run. It must contain NO spaces — a
514
+ # line with more than 7 whitespace-delimited tokens is a parse
515
+ # error (a path with spaces is unrepresentable). The prompt
516
+ # file's content is capped at 32KB (enforced when the job fires).
517
+ # token 7 : OPTIONAL comma-separated list of Slack channel IDs to target.
518
+ # Omit it entirely to target ALL bots — that omission IS the
519
+ # all-bots form. There is NO all-bots wildcard: a literal '*' in
520
+ # the channel position is a parse error, not "all channels".
521
+ #
522
+ # Lines beginning with '#' and blank lines are ignored. A malformed line is
523
+ # skipped on its own; sibling lines still schedule.
524
+ #
525
+ # Example (every day at 09:00, run grooming-tick.md, target two channels):
526
+ # 0 9 * * * /home/horde/prompts/grooming-tick.md C0123ABC,C0456DEF
527
+ #
528
+ # Example (every hour on the hour, run standup.md, target all bots):
529
+ # 0 * * * * /home/horde/prompts/standup.md
530
+ ```
531
+
532
+ Each schedule is one line of **exactly 5 cron fields**, then the prompt-file path, then an optional comma-separated channel list:
533
+
534
+ ```
535
+ 0 9 * * 1 /home/horde/prompts/weekly-report.md C0123ABC,C0456DEF
536
+ ```
537
+
538
+ Rules:
539
+
540
+ - **Exactly 5 cron fields** (minute hour day-of-month month day-of-week). Croner's 6-field (seconds-precision) and `@macro` forms are **not** supported.
541
+ - **Omit the channel list to target ALL bots** — the omission itself is the all-bots form. (All-bots delivery is currently deferred — see [Delivery semantics](#delivery-semantics).)
542
+ - **No `*` wildcard in the channel position.** A literal `*` where a channel ID belongs is a parse error, not "all channels".
543
+ - **Prompt paths cannot contain spaces.** A path with spaces is unrepresentable; the extra tokens make the line a parse error and it is skipped.
544
+ - **`#` comments and blank lines are allowed** and ignored.
545
+ - **A bad line is skipped and logged** (as `parse-error` in the cron log), never fatal — sibling lines still schedule.
546
+
547
+ Because the count is positional, a **6-field line silently mis-parses instead of erroring.** For example:
548
+
549
+ ```
550
+ 0 0 1 1 1 1 ~/prompts/p.md
551
+ ```
552
+
553
+ Here the 6th field (`1`) is taken as the prompt path and the real path (`~/prompts/p.md`) is taken as the channel list. No error is raised — every slot is filled with something syntactically acceptable — so the schedule fires on a nonsense cadence against a nonsense path. Keep expressions to exactly 5 fields.
554
+
555
+ ### Path resolution
556
+
557
+ The prompt-file path resolves as follows:
558
+
559
+ - A leading `~` expands to the home directory.
560
+ - A **relative** path resolves against the **crontable's own directory** — not `$HOME`. This is a deliberate divergence from system cron's convention, so you can keep a `prompts/` directory alongside the crontable and reference it as `prompts/standup.md`.
561
+ - An **absolute** path is used as-is.
562
+
563
+ ### How fires appear
564
+
565
+ A scheduled fire arrives in the channel as an `/interject` message whose `sender` label is `cscb-cron:<prompt-file-basename>` — for a prompt file `standup.md` the sender is `cscb-cron:standup`. This distinguishes a cron tick from a human and from peer-bot traffic. Each schedule delivers its own message independently, so when several schedules match the same minute for the same channel each one arrives as its own `/interject` message.
566
+
567
+ ### The cron log
568
+
569
+ Every fire outcome is recorded in the cron log at `cron_log_path` (default `<config dir>/cron.log`; override with the `cron_log_path` key). Each attempt writes one line per target channel plus a per-fire summary line carrying `delivered=N failed=M` counts. The log is plain text, so `grep no-session cron.log` yields readable lines.
570
+
571
+ The `outcome` field of each line is one of these classes:
572
+
573
+ | Outcome | What happened | What to do |
574
+ |---|---|---|
575
+ | `delivered` | The prompt reached the target channel's session. | Nothing — success. |
576
+ | `no-session` | The channel is routed but no live session is connected, so the message was dropped. | Bring the session up. Failed fires are **never** retried or queued (see below). |
577
+ | `unknown-channel` | The target channel is not in `config.json → routes`. | Fix the channel ID in the crontable, or add the route. |
578
+ | `prompt-missing` | The prompt file did not exist at fire time. | Create the file or correct its path in the crontable. |
579
+ | `prompt-unreadable` | The prompt file existed but could not be read (see the `errno`). | Fix file permissions or the path. |
580
+ | `prompt-oversize` | The prompt exceeds the 32KB `/interject` cap and was skipped, never truncated. | Shorten the prompt file. |
581
+ | `parse-error` | The crontable line could not be parsed. | Fix the line — see the crontable header for the format. |
582
+ | `http-error` | The localhost POST hit an unexpected HTTP status or a network failure. | Check that the server is listening on loopback (see the `bind` note below) and inspect the `errno`/`status` in the line. |
583
+ | `fanout-deferred` | A channel-less (all-bots) line was matched but not delivered. | None — all-bots fan-out is not yet enabled; give the line an explicit channel to deliver it today. |
584
+
585
+ ### Delivery semantics
586
+
587
+ - **No retry.** A failed fire is logged and dropped — never queued or re-sent. A channel with no live session fails every fire until its session is running again; the server does not queue the missed prompts.
588
+ - **Missed fires are skipped, not caught up.** While the server is down, no scheduled prompts fire, and they are not replayed on restart. The `scheduler started, N schedules loaded` line in the cron log marks when scheduling resumed, bounding the outage window.
589
+ - **Edits take effect after a server restart.** The crontable is read and parsed once, at scheduler start. Editing the file — by hand or by a bot appending a line — does nothing until the server restarts and reloads the table.
590
+ - **Server-local time.** Cron expressions are evaluated in the server's local timezone.
591
+ - **Channel-less lines are deferred.** A line with no channel is currently matched but logged `fanout-deferred` and not delivered. Give a line an explicit channel to have it fire.
592
+ - **`bind` must include loopback.** The scheduler delivers via `127.0.0.1`, so a `bind` set to a single non-loopback interface makes every fire fail with `http-error`. Use the default `127.0.0.1` or `0.0.0.0`.
593
+
594
+ ### Bot self-scheduling
595
+
596
+ Every managed session carries the resolved crontable path in the `CSCB_CRONTABLE_PATH` environment variable, so a bot can schedule its own prompts without being told where the crontable lives. Discover it from inside a session:
597
+
598
+ ```sh
599
+ echo $CSCB_CRONTABLE_PATH
600
+ ```
601
+
602
+ The crontable is the single source of truth for schedules. When adding a schedule, **append** a new line — never rewrite, reorder, or delete other lines. An appended line takes effect after the next server restart (the table is read once at scheduler start).
603
+
604
+ ---
605
+
435
606
  ## Permission Relay
436
607
 
437
608
  When Claude Code requires tool approval, the permission relay surfaces an interactive Slack message with **Allow** and **Deny** buttons instead of blocking the TUI. Architecture is polling-based on the `agent-director` library — there are **no hook scripts to install** and no HTTP long-poll loops.
@@ -452,6 +623,70 @@ The Slack app must have **interactivity enabled** with **Socket Mode** as the de
452
623
 
453
624
  The `AskUserQuestion` tool is denied for every CSCB-spawned bot via the agent-director template (`deny: ['AskUserQuestion']`). Bots respond to operator questions via the Slack `reply` MCP tool instead. There is no `ask-relay.sh` hook and no `/ask` HTTP route.
454
625
 
626
+ ### Memory-directory reads
627
+
628
+ The template also pre-allows each bot to read its own persistent-memory directory, so those reads don't surface a permission prompt to a human. One `Read(//<config-dir>/projects/*/memory/**)` rule is derived per distinct Claude config directory in your routing config (a route's `claude_config_dir`, the top-level `claude_config_dir`, or the `~/.claude` default). The rule is scoped to `projects/*/memory/**` only — never the config-dir root, which holds live credentials — so it never pre-authorizes credential reads.
629
+
630
+ ---
631
+
632
+ ## Slack Reply Guard (Stop hook)
633
+
634
+ CSCB ships a Claude Code Stop hook that enforces a simple rule for every bot session it manages: **when the most recent real user message on the turn came from Slack, the assistant must call the `mcp__slack-channel-router__reply` tool before ending the turn**. If it does not, the Stop hook exits `2`, and Claude Code shows the assistant the reminder `Slack user is waiting for a reply. You must respond by calling the mcp__slack-channel-router__reply tool before ending your turn.` and re-runs it once. The retry sets `stop_hook_active=true`, which short-circuits the guard, so exactly one forced retry occurs per turn — never an infinite block loop.
635
+
636
+ ### What the server writes, and where
637
+
638
+ CSCB owns installing the hook on your behalf. On every server boot, alongside the trust-folder bootstrap, the server walks every route, groups them by effective `claude_config_dir` (per-route override falls back to the top-level value), and patches `<claude_config_dir>/settings.json` in place. For each dir it ensures **exactly one** managed Stop-hook group of the shape:
639
+
640
+ ```jsonc
641
+ {
642
+ "hooks": {
643
+ "Stop": [
644
+ { "hooks": [ { "type": "command", "command": "<absolute path>/stop-hooks/slack-reply-guard.sh" } ] }
645
+ ]
646
+ }
647
+ }
648
+ ```
649
+
650
+ The command is an absolute path to the script inside CSCB's installed package tree. There is no `matcher` field — Stop is not a tool-scoped event. Any other Stop hooks you have configured, and every other key in `settings.json`, are preserved. Writes are atomic (`.tmp` + `rename`). A missing `settings.json` is created with just this group; a malformed `settings.json` is left untouched and a startup error is recorded.
651
+
652
+ **Recognition rule.** The server treats *any* Stop-hook `command` string containing the substring `slack-reply-guard.sh` as CSCB-managed. Duplicates from prior boots are collapsed to one canonical entry; stale entries (from an older install path) are rewritten to the current absolute path — this is the self-heal path across upgrades.
653
+
654
+ ### Timing: on-disk at every boot, effective at next Claude process start
655
+
656
+ The bootstrap rewrites `settings.json` on **every** CSCB boot, so the on-disk entry always reflects the currently-installed release's absolute path. Claude Code, however, only reads hook configuration when a Claude process starts. On a CSCB restart, live sessions are reconnected and keep their already-running Claude processes — they will not pick up an updated hook path until the next fresh spawn or the next resume of a dead/missing session for that route.
657
+
658
+ ### Shared-dir aggregation
659
+
660
+ The install/remove decision is per **directory**, not per route. If two routes resolve to the same `claude_config_dir`, the managed entry is installed when at least one of them has the guard enabled, and removed only when all of them have it disabled. Consequence: a per-route opt-out fully disables the guard for a bot only when that route owns a *dedicated* `claude_config_dir`. A route that shares a dir with any enabled route still gets the guard on that shared dir.
661
+
662
+ ### Personal-dir refusal
663
+
664
+ The bootstrap refuses to touch the operator's own `~/.claude` directory. If a route's effective `claude_config_dir` resolves (via `realpathSync`, with a lexical fallback for paths that do not exist on disk) to your home `.claude` dir, nothing is written and a startup error is recorded. This prevents CSCB from ever installing a bot-oriented Stop hook into your interactive Claude Code config.
665
+
666
+ ### Bots without a `claude_config_dir`
667
+
668
+ Routes with no effective `claude_config_dir` — neither per-route nor top-level — are skipped. Empty or whitespace-only values are treated as absent (so `resolve("")` never lands in the process cwd). If you want the guard on a bot, give its route a real `claude_config_dir`.
669
+
670
+ ### v1 limitations — opt these bots out
671
+
672
+ The v1 guard only recognises a reply via `mcp__slack-channel-router__reply`. Bots whose only Slack surface is `edit_message` or `react` will end their turn without producing a matching `tool_use`, and the guard will block them and force one useless retry every turn. **Opt these bots out** by setting `stop_hook_bootstrap: false` on the route (see the field reference below), and give the route a dedicated `claude_config_dir` — see [Shared-dir aggregation](#shared-dir-aggregation).
673
+
674
+ ### Opting out
675
+
676
+ The `stop_hook_bootstrap` boolean lives on the top level of `config.json` and on individual routes. It defaults to `true`. Set it to `false` at the top level to disable the bootstrap for every route; set it on an individual route to override the top-level default for one bot. See the [Field reference](#field-reference) and [Per-route `stop_hook_bootstrap` override](#per-route-stop_hook_bootstrap-override) below for the field details and the per-route-vs-shared-dir interaction.
677
+
678
+ ### Tag drift — fail-open, verify after upgrades
679
+
680
+ The guard's Slack-origination predicate is a substring match on the prefix `<channel source="slack` in the transcript entry Claude Code writes for every Slack-delivered turn. CSCB only sends `{content, meta}` over MCP; the `<channel source="…">` wrapper is rendered by the **Claude Code harness itself** when it serialises the MCP tool result into the transcript, and the `source` attribute is the MCP server name (e.g. `slack-channel-router`). That tag is therefore an **external, harness-owned contract** — a future Claude Code release can rename it or restructure the wrapper without touching CSCB, and the guard's predicate would silently stop matching. Because the contract sits outside CSCB, the guard is designed to fail open on drift, and a post-upgrade verification recipe (below) exists so operators catch a silent-dark guard the next time the harness changes the tag.
681
+
682
+ The guard is **fail-open by design**: any error, missing transcript, missing `jq`, or absence of the tag results in `exit 0` (turn allowed). This means a future rename of the `<channel>` tag will silently disable the guard rather than break the bot. After every CSCB or Claude Code upgrade, verify the guard end-to-end:
683
+
684
+ 1. Send the bot a Slack message that requires a reply.
685
+ 2. Confirm the reply lands in Slack.
686
+ 3. In the bot's transcript file (`<claude_config_dir>/projects/<slug>/*.jsonl` — guard-covered bots always run with a dedicated `claude_config_dir`, since the bootstrap refuses the operator's personal `~/.claude`), grep for `<channel source="slack` on the triggering message and for a subsequent assistant entry containing `"name":"mcp__slack-channel-router__reply"` in a `tool_use` block.
687
+
688
+ If the tag prefix no longer appears, the guard is dark — file an issue.
689
+
455
690
  ---
456
691
 
457
692
  ## Troubleshooting
@@ -474,12 +709,32 @@ Messages to channels not listed in `access.json → channels` and not present in
474
709
  **Permission relay not working**
475
710
  Check that the Slack app has interactivity enabled (Interactivity & Shortcuts → toggle on). Verify the bot is in `check_permission` state via `agent-director list --state check_permission --label service=cscb` (operator CLI). Inspect `server.log` for `permission-poller:` lines — skipped-tick WARNs at 5+ consecutive skips signal that the poll interval is too tight; increase `agent_director_poll_interval_ms` in `config.json`.
476
711
 
712
+ **Bot appears dead / posts a "blocked on a native permission prompt" warning**
713
+ The bot is wedged in `check_permission` on a native Claude Code TUI prompt that never reached Slack (a permission decision AD recorded but could not deliver). The bot stops responding, and after ~90 s the poller posts a one-shot channel warning. Recover by inspecting the native prompt with `agent-director read-pane --claude-instance-id <id>`, then killing and respawning the session (`agent-director kill <id>` or tmux-kill, then let the server restart it or `claude-slack-channel-bots stop && claude-slack-channel-bots start`). Do **not** use `send-keys` — agent-director hard-rejects it while the spawn is in this relayed permission state. The warning fires once per wedge episode; the detector re-arms if the bot later wedges again.
714
+
477
715
  **Session not restarting after crash**
478
- After 3 consecutive launch failures for a route, auto-restart is suspended until the server is restarted. Restart the server with `claude-slack-channel-bots stop && claude-slack-channel-bots start`. To disable auto-restart entirely, set `session_restart_delay` to `0` in `config.json`.
716
+ Auto-restart backs off exponentially on repeated launch failures — the delay doubles from `session_restart_delay` (default 60s) on each consecutive failure, up to a 15-minute ceiling. After 5 consecutive failures the route hits a cap: a `SpawnCapReached` message is posted to the channel and automatic restarts stop.
717
+
718
+ Sending a message in a channel whose session is dead but not yet capped triggers a fast recovery: the restart is scheduled immediately (the backoff delay is clamped down to 5 seconds for an explicit human trigger, never raised), and the sender is told the session is starting and to retry in a moment. The dropped message itself is **not** delivered or replayed — recovery only starts the session; you must resend after it comes up. This human trigger still counts each failed launch toward the backoff/cap, and a restart already pending or active is not stacked.
719
+
720
+ A **capped** route (or one with auto-restart disabled via `session_restart_delay: 0`) does **not** recover on an inbound message — firing another launch there would only burn a spawn attempt against a route that cannot come up. The sender is told plainly that the channel will not self-recover and an operator must restart the server. To clear the cap and retry, restart the server with `claude-slack-channel-bots stop && claude-slack-channel-bots start`; the failure counter is in-process and cleared on restart, giving each route a fresh attempt. To disable auto-restart entirely, set `session_restart_delay` to `0` in `config.json`.
721
+
722
+ **Bot alive but silently unresponsive (MCP disconnected)**
723
+ A bot can stay running yet lose its MCP connection to the server — the process is alive but no longer reachable, so it stops responding without ever emitting a disconnect event. The periodic health-check recovers this automatically: once a channel is seen alive-but-disconnected on two consecutive ticks, the health-check schedules a reconnect (or a relaunch if the process has since died), so a stranded channel comes back with no inbound message and no server restart. The recovery lands within roughly two `health_check_interval` periods (default 120 s each) plus the restart backoff delay (default `session_restart_delay` 60 s) before the reconnect runs — about 3–5 minutes with default settings. A bot mid-turn (`working` state) is deliberately left alone and reconnected on a later tick once its turn settles.
724
+
725
+ A closely related symptom is a bot that still *looks* connected but silently drops every inbound message — its underlying message stream went away without the connection registering as closed. The same health-check path recovers this on the same two-consecutive-tick cadence, so no inbound message or server restart is needed. If a message does arrive on such a channel before recovery lands, the sender is told the message was not delivered and to retry in a moment, rather than getting silence.
726
+
727
+ To verify recovery in the field, tail `server.log` for a stranded channel and confirm a tick-driven recovery lands — look for a `[slack] Scheduling restart for channel=<id> in <N>s (backoff)` line and a `[slack] Session alive but disconnected — reconnecting MCP for channel=<id>` line naming that channel (and, when the bot was mid-turn, a `deferring /mcp reconnect to a later tick (b.9a7/b.rmy)` line first). A **capped** route is exempt: the health-check skips it entirely, including reconnects, so a capped channel still requires a server restart (see below).
479
728
 
480
729
  **Session stuck during clean_restart**
481
730
  If a session does not exit within `exit_timeout` seconds (default 120s), `clean_restart` force-kills the spawn via `agent-director kill` and proceeds. To manually recover, run `agent-director list --label service=cscb` to find lingering spawns and `agent-director kill <claude_instance_id>` to clear them, then `claude-slack-channel-bots stop && claude-slack-channel-bots start`.
482
731
 
732
+ **`clean_restart` or `stop --stop-bots` exits non-zero with an agent-director teardown error**
733
+ This is intentional: when agent-director is unreachable, the teardown cannot run, so the command fails loudly rather than silently no-op'ing and (for `clean_restart`) restarting on top of bots it never touched. Confirm agent-director is installed and responsive with `agent-director version`, then re-run the command. Teardown kills but never deletes rows on any failure path, so it is always safe to retry once agent-director is reachable.
734
+
735
+ **Bots come back with no memory of the prior conversation after a reboot**
736
+ With `resume_enabled: true`, a bot whose host rebooted (or pod resumed) should return with its conversation history. If it comes back amnesiac, confirm the system-installed `agent-director` is **≥ 0.8.0** (`agent-director version`) — reboot recovery relies on capabilities added in that release. Note that `bun run install-check` does **not** confirm this: its client floor is `0.7.0`, lower than the reboot-recovery requirement, so install-check passes on a `0.7.x` binary that still yields amnesiac bots. Verify the resume requirement directly with `agent-director version`. Note: legacy sessions created before upgrading to 0.8.0 may lose history exactly once on their first post-upgrade recovery, then resume cleanly thereafter.
737
+
483
738
  **Session crashes on resume with "sandbox required but unavailable"**
484
739
  This is a known regression in certain Claude Code releases (e.g. v2.1.120) where `--resume` triggers a sandbox check that fails in headless environments. Set `resume_enabled: false` in `config.json` to disable `--resume` entirely — the bot will always start a fresh Claude session instead of resuming a prior conversation, both on startup and on runtime auto-restart:
485
740
 
@@ -492,9 +747,17 @@ This is a known regression in certain Claude Code releases (e.g. v2.1.120) where
492
747
 
493
748
  ---
494
749
 
750
+ ## Server log rotation
751
+
752
+ The server daemon writes runtime output to `~/.claude/channels/slack/server.log` (and `clean_restart.log` for the `clean_restart` subcommand). Rotation is built into CSCB, so it applies on every machine with no per-host logrotate config: when the active file crosses `CSCB_LOG_MAX_BYTES` it is rolled to `server.log.1`, the previous `.1` → `.2`, and so on up to `CSCB_LOG_KEEP` generations; the oldest is discarded. Rotated generations are not compressed. See the environment-variable table above for the size, retention, and verbosity settings that tune this behavior.
753
+
754
+ Note this covers only `server.log` / `clean_restart.log`. `startup-errors.log` and `permission-trail.jsonl` are separate append-only files with their own retention (see below).
755
+
756
+ ---
757
+
495
758
  ## Startup errors
496
759
 
497
- CSCB writes fatal startup errors to `~/.claude/channels/slack/startup-errors.log` (override the directory with `SLACK_STATE_DIR`) in addition to stderr. Each entry is a single timestamped line. The file is append-only and never rotated by CSCB — copy `docs/logrotate-startup-errors.conf` into `/etc/logrotate.d/` if you want host-level rotation.
760
+ CSCB writes startup errors to `~/.claude/channels/slack/startup-errors.log` (override the directory with `SLACK_STATE_DIR`) in addition to stderr. Each entry is a single timestamped line. The file is append-only and never rotated by CSCB — copy `docs/logrotate-startup-errors.conf` into `/etc/logrotate.d/` if you want host-level rotation. Most classes below are fatal (the process exits non-zero); the JSONL-persistence warnings at the end are non-fatal and do not block startup.
498
761
 
499
762
  Classes you may see:
500
763
 
@@ -512,6 +775,15 @@ Classes you may see:
512
775
  - `ad-same-user-stat` — Non-ENOENT stat error on the state DB (permissions, I/O). Investigate the file before re-launching.
513
776
  - `ad-template-install` — `client.makeTemplate(...)` rejected the boot-time refresh of the `slack-channel-bot` template. The line includes the agent-director `errName`.
514
777
 
778
+ The following classes are **non-fatal warnings** about conversation-memory loss. They are recorded to the same log but never exit the process or block startup. The first four are written by the JSONL-persistence safeguard, which runs *before* the resume path to warn about an *impending* loss; the last two (`jsonl-transcript-lost-on-resume`, `jsonl-diagnosis-inconclusive`) are written *by the resume path itself* when it tried to resume a row and either confirmed a wipe or could not determine whether one occurred:
779
+
780
+ - `jsonl-non-persistent` — a session-transcript storage root (`<claude_config_dir>/projects`) is on a `tmpfs`/`ramfs` mount, so nothing there survives a reboot and session resume is structurally impossible on this host. A warning is also posted to the affected channels — those whose transcript storage root is the flagged mount, not every routed channel. Move the config dir to a persistent filesystem.
781
+ - `jsonl-persistence-check-warning` — the safeguard could not determine the filesystem type of a transcript storage root (unreadable/unparseable `/proc/self/mountinfo`, or an unresolvable path), so persistence is unverified. No Slack post is made. Investigate the mount before relying on resume.
782
+ - `jsonl-transcript-stale-path` — a channel's saved transcript exists on disk at the resolved fallback path, but agent-director's recorded `jsonl_path` points elsewhere (missing/empty). On the next restart the resume path would treat it as missing and wipe the channel's memory. A warning is posted to that channel; an operator should reconcile the path before restarting.
783
+ - `jsonl-transcript-lost` — a channel's transcript is gone from disk (neither the recorded nor the fallback path exists), yet the message archive shows messages in that channel since the bot spawned. Conversation history has been lost and resume will start the bot fresh. A warning is posted to that channel. Requires `message_archive_db` to be configured for the archive evidence.
784
+ - `jsonl-transcript-lost-on-resume` — resume actually threw `ErrJsonlMissing` for a channel, the bot was delete+fresh-spawned, and the message archive shows messages in that channel since it spawned — so conversation history was destroyed by this recovery, not merely at risk. A warning is also posted to that channel. This is the resume path's own after-the-fact report (distinct from the pre-resume `jsonl-transcript-lost` warning above); the log line names every transcript path tried and whether each came from agent-director or was computed locally. A missing transcript on a channel that was *idle since spawn* — the archive was consulted and shows zero messages since spawn — is expected (the transcript is created lazily on first message) and produces a quiet log line only, no error class and no channel post. Requires `message_archive_db` for the archive evidence; when the archive cannot be consulted the case is instead reported as `jsonl-diagnosis-inconclusive` below.
785
+ - `jsonl-diagnosis-inconclusive` — resume threw `ErrJsonlMissing` and the bot was delete+fresh-spawned, but the diagnosis could not determine whether conversation history was lost: the agent-director row could not be fetched, `started_at` was unparseable, the message archive could not be read, or `message_archive_db` is not configured at all. Because "inconclusive" correlates with the same storage problems that cause real loss, this is surfaced (not silently downgraded to a benign never-created): the line records *why* the diagnosis failed, and a warning is posted to the channel worded as uncertainty ("restarted fresh; could not determine whether prior history was preserved") rather than as a confirmed loss. When the reason is an unconfigured archive, the message notes that diagnosis is impossible without `message_archive_db` and suggests enabling it. These channels are counted separately in the startup summary as `fresh-after-inconclusive-amnesia` (distinct from the `fresh-after-amnesia` count).
786
+
515
787
  ---
516
788
 
517
789
  ## Release process
@@ -528,14 +800,15 @@ The bump kind is **required** — there is no default. The skill exits with a us
528
800
 
529
801
  ### Preflight gates
530
802
 
531
- Before any side-effecting step runs, `/publish` enforces six fail-fast gates. Any failure aborts before the version is bumped, the tarball is packed, or anything is committed:
803
+ Before any side-effecting step runs, `/publish` enforces seven fail-fast gates. Any failure aborts before the version is bumped, the tarball is packed, or anything is committed:
532
804
 
533
805
  1. **Clean working tree on `main` in sync with origin/main.** No uncommitted changes; HEAD branch is `main`; `main` is exactly equal to `origin/main` after `git fetch origin`.
534
806
  2. **Tests exist and pass.** At least one `*.test.ts` file under `tests/` and `bun test` exits zero.
535
807
  3. **Typecheck passes.** `bun run typecheck` exits zero.
536
808
  4. **npm authenticated.** `npm whoami` exits zero (run `npm login` first if not).
537
809
  5. **Next version not already published.** `npm view claude-slack-channel-bots@<next-version> version` must report nothing.
538
- 6. **`/ci` integration suite passes.** The full Docker-based integration test suite is run via the `/ci` skill and must report PASS. **`/ci` is mandatory and has no opt-out flag** — release without an unbroken integration run is not possible through this skill.
810
+ 6. **No stranded finished work.** `scripts/audit-finished-tickets.sh` must exit zero. It flags any `finished` ticket whose fix is neither on `main` nor explicitly closed (ops-only / no-repro / superseded / documents-only outside any git repo), any unmerged branch tied to a finished ticket, and any unmerged branch that references no known ticket at all. A release cannot ship while a fix is silently stranded on a dead branch. The gate is read-only — it never mutates tickets or git. This gate runs in Phase 1 preflight, **before** the `/ci` gate below, so a stranded-work failure aborts the release before the Docker suite ever runs. The gate splits its failure by the audit's exit code: audit exit 1 (stranded work) → preflight exit 16; audit exit 2 (setup failure — the Bugs/Plans/Ideas hives, a git repo, or a `main` ref were not locatable from this checkout, e.g. a throwaway `/tmp` clone) → preflight exit 17, whose fix is to rerun `/publish` from the canonical checkout that sits beside the hives. **This gate is green today — it exits zero with no findings.** There is no known expected debt in either class, so a non-zero result is a new, real finding to investigate before releasing. Do not re-list findings here — the audit's own output is the inventory.
811
+ 7. **`/ci` integration suite passes.** The full Docker-based integration test suite is run via the `/ci` skill and must report PASS. **`/ci` is mandatory and has no opt-out flag** — release without an unbroken integration run is not possible through this skill.
539
812
 
540
813
  ### What happens during a release
541
814
 
@@ -570,7 +843,7 @@ For operators upgrading from a pre-`agent-director` install:
570
843
 
571
844
  1. **Install the new CSCB**: `bun remove claude-director` (if present) and `bun install -g claude-slack-channel-bots@^<new>`. The `agent-director` library is pulled in transitively — no separate install step.
572
845
  2. **Delete any old relay hooks** — see [Upgrading from pre-Epic-2 (v0.5.x → v0.6.x)](#upgrading-from-pre-epic-2-v05x--v06x) for the cleanup commands.
573
- 3. **Configure agent-director's `find-missing` sweep**. CSCB does NOT call `client.findMissing(...)`; reconciling stuck rows is the operator's responsibility. Add a cron entry (or systemd timer) that runs `agent-director find-missing` on a cadence that matches your tolerance — e.g. every minute on a busy host:
846
+ 3. **(Optional) Configure an agent-director `find-missing` sweep**. CSCB itself runs `find-missing` once before a resume during dead-session recovery, so a bot that died on reboot comes back with its history intact. A standalone periodic sweep is no longer required for CSCB recovery, but remains useful if you want stuck rows from non-CSCB spawns reconciled on a cadence. To add one, use a cron entry (or systemd timer):
574
847
  ```cron
575
848
  * * * * * /usr/local/bin/agent-director find-missing --timeout 30s
576
849
  ```
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "claude-slack-channel-bots",
3
- "version": "0.8.2",
3
+ "version": "0.10.0",
4
4
  "description": "Multi-session Slack-to-Claude bridge — run multiple Claude Code bots across Slack channels via Socket Mode",
5
5
  "type": "module",
6
6
  "bin": {
@@ -13,10 +13,11 @@
13
13
  "scripts/install-check.ts",
14
14
  "slack-app-manifest.yml",
15
15
  "README.md",
16
- "skills/"
16
+ "skills/",
17
+ "stop-hooks/"
17
18
  ],
18
19
  "engines": {
19
- "bun": ">=1.0.0"
20
+ "bun": ">=1.0.21"
20
21
  },
21
22
  "scripts": {
22
23
  "postinstall": "bun src/postinstall.ts && bun scripts/fixup-bun-cache.ts",
@@ -31,7 +32,8 @@
31
32
  "@modelcontextprotocol/sdk": "^1.0.0",
32
33
  "@slack/socket-mode": "^2.0.0",
33
34
  "@slack/web-api": "^7.0.0",
34
- "agent-director": "^0.7.8",
35
+ "agent-director": "^0.10.0",
36
+ "croner": "^10.0.1",
35
37
  "semver": "^7.6.0"
36
38
  },
37
39
  "devDependencies": {
@@ -130,25 +130,25 @@ fallthrough.
130
130
  inspection (`ls -la <binary_path>`) and removal of the bad entry,
131
131
  then re-installation via the AD install command. Re-run Step 1.
132
132
 
133
- 3. **`probe-timeout`** — `agent-director --version` did not return within
133
+ 3. **`probe-timeout`** — `agent-director version` did not return within
134
134
  the probe window. Likely the binary is hanging on startup
135
135
  (corrupted, mismatched architecture, missing shared library).
136
- Recommend a manual `<binary_path> --version` invocation to confirm,
136
+ Recommend a manual `<binary_path> version` invocation to confirm,
137
137
  then reinstall via the AD install command. Re-run Step 1.
138
138
 
139
- 4. **`probe-nonzero-exit`** — `agent-director --version` exited with a
139
+ 4. **`probe-nonzero-exit`** — `agent-director version` exited with a
140
140
  non-zero code. The stderr block surfaces `exitCode` and any
141
141
  `diagnostic` from AD. Show the user the values and recommend a
142
- manual reproduction (`<binary_path> --version`), then reinstall.
142
+ manual reproduction (`<binary_path> version`), then reinstall.
143
143
  Re-run Step 1.
144
144
 
145
- 5. **`probe-killed-by-signal`** — `agent-director --version` was killed
145
+ 5. **`probe-killed-by-signal`** — `agent-director version` was killed
146
146
  by a signal (SIGSEGV, SIGBUS, etc.). The stderr block surfaces
147
147
  `signal`. The binary is likely corrupted or built for a different
148
148
  architecture. Recommend a full reinstall via the AD install command.
149
149
  Re-run Step 1.
150
150
 
151
- 6. **`unparseable-version`** — `agent-director --version` returned but
151
+ 6. **`unparseable-version`** — `agent-director version` returned but
152
152
  the output could not be parsed as semver. The stderr's `diagnostic`
153
153
  field carries the raw output. Likely the installed binary is from
154
154
  a pre-release line that uses non-semver tags (e.g. `v0.6.3-dev`
@@ -353,7 +353,7 @@ Check that agent-director is installed and the `slack-channel-bot` template
353
353
  is registered:
354
354
 
355
355
  ```bash
356
- agent-director --version
356
+ agent-director version
357
357
  ```
358
358
 
359
359
  The `slack-channel-bot` template is registered automatically at CSCB startup
@@ -12,8 +12,8 @@
12
12
  * Construction: AD 0.7.0+ exposes an async `Client.create()` factory; the
13
13
  * subprocess constructor is protected and `new Client(...)` is a compile-time
14
14
  * TS error. The SR-5.1 startup gate (src/agent-director-startup.ts) awaits
15
- * `Client.create({ storePath, createIfMissing: true, logger: console })` and
16
- * then hands the resolved instance to `setClient()`. Verb call sites pull the
15
+ * `Client.create({ storePath, createIfMissing: true, logger: makeFilteredAdLogger(console) })`
16
+ * and then hands the resolved instance to `setClient()`. Verb call sites pull the
17
17
  * installed singleton through `getClient()`; calling `getClient()` before the
18
18
  * gate has installed a Client throws an internal bug-marker error (it
19
19
  * indicates a caller-site bug, not a runtime condition).
@@ -35,9 +35,28 @@
35
35
  * own pause timeout via CSCB-side polling and never relies on the library's
36
36
  * pause budget, so the class never reaches a CSCB handler.
37
37
  *
38
+ * CSCB-synthetic subclasses (NOT emitted by the agent-director library — minted
39
+ * inside CSCB and dispatched through the same instanceof-branching convention so
40
+ * SR-0.2 holds for them too):
41
+ * - ErrSpawnCapReached (restart backoff / consecutive-failure cap latch)
42
+ *
38
43
  * SPDX-License-Identifier: MIT
39
44
  */
40
45
 
46
+ import { AgentDirectorError } from 'agent-director'
47
+
48
+ /**
49
+ * CSCB-synthetic error — never emitted by the agent-director library. Minted by
50
+ * the restart backoff cap path (server.ts onCapReached) when a channel hits the
51
+ * consecutive session-launch failure cap and automatic restarts are suspended.
52
+ * Branched on via `instanceof` in remediationHint (SR-0.2 — no string matching).
53
+ */
54
+ export class ErrSpawnCapReached extends AgentDirectorError {
55
+ constructor(description: string) {
56
+ super('spawn', 'SpawnCapReached', description)
57
+ }
58
+ }
59
+
41
60
  export {
42
61
  AgentDirectorError,
43
62
  ErrClientClosed,