claude-spotter 1.4.17 → 1.4.19

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,82 @@
1
1
  # Changelog
2
2
 
3
+ ## 1.4.19
4
+
5
+ 親セッションの暴走を誘発できたHook出力の信頼境界を修正する。監査用AIは内部で構造化判定を返すだけとし、
6
+ その自由文やprovider出力を親モデルへ渡さない。親向け出力はSpotterプログラムが検証済みtool IDから
7
+ 決定論的に生成する、採否を親が独立判断できる非命令形の助言へ限定する。
8
+
9
+ ### 変更点
10
+
11
+ - **共通parent-output projector**: Claude / Codexの両Hookが同じprojectorを使う。tool IDは
12
+ ASCII grammar、160文字、5件、重複排除、安定sort、助言全体2,000文字の上限を持ち、改行・制御文字・
13
+ backtick・Markdown・超長大IDを拒否する。監査用AIの`reason` / `raw`はAPI入力に持たない。
14
+ - **非命令形のUserPrompt助言**: `additionalContext`はcatalog照合済みtool IDと固定テンプレートだけから
15
+ 作る。「使え」「補正せよ」「ユーザーへ伝えよ」といった命令や監査用AIの説明文を含めない。
16
+ - **failure分離**: backend codeをallow-list済みの固定状態へ写像し、固定`systemMessage`・固定stderr・
17
+ 構造Hook eventへ出す。backend message、provider stdout / stderr、未知code本文を反射しない。この保証は
18
+ Hook出力生成までで、全Codex App/background面でのUI可視性は未保証。
19
+ - **Stop持ち越し廃止**: finding / failureを`.spotter/pending/`へ新規保存せず、次の無関係な
20
+ UserPromptSubmitへ配送しない。findingは構造Hook event、failureは固定診断として当該Stopで完結する。
21
+ 旧same-session pendingは内容を読まずにunlinkし、`ENOENT`以外の失敗も固定診断だけにする。
22
+ - **安全性維持**: `decision:"block"`、Stop継続、UserPromptSubmit exit 2の範囲拡大は行わず、既存の
23
+ 再帰Hook / daemon proliferation防止、marker、short prompt、non-blocking failure契約を維持する。
24
+
25
+ ### 検証
26
+
27
+ projector / Claude / Codexのtargeted testとfull suiteを実行し、447件中445 pass・0 fail・既存2 skip。
28
+ AI/backend/provider sentinel、unsafe ID、legacy pending非読取unlink、Stop次turn非配送、catalog/transcript
29
+ 読込例外の固定degradationを確認した。敵対的再監査は初回BLOCKER 2件(Codex読込例外の境界外、
30
+ UI可視性の過剰主張)を検出・修正し、再監査BLOCKER 0。`npm pack --dry-run`は62 files、global
31
+ `spotter 1.4.19`へlocal installし、projector smoke・Hook diagnostics・global UserPromptSubmit smokeを
32
+ 確認した。公開前release gateとしてfull test、pack、秘密混入、CI、registry tarballを再検証し、
33
+ `v1.4.19` tag・npm `latest`・GitHub Release・global installを同一versionへ揃える。
34
+
35
+ ## 1.4.18
36
+
37
+ auditor model の更新を model 名の場当たり的な置換から切り離し、versioned policy と再現可能な比較 eval を
38
+ 導入する。反復評価24/24 exactの `gpt-5.6-terra × medium` をowner裁定でproductionへ昇格した。
39
+ `latest` aliasやCodex CLIの暗黙既定、失敗時fallback、eval artifactからの自動昇格は使わない。
40
+
41
+ ### 変更点
42
+
43
+ - **model policy**: 意味論的な auditor role、production selection、`gpt-5.6-luna × low` /
44
+ `gpt-5.6-terra × low` / `gpt-5.6-terra × medium` の評価 profile、policy version、検証状態を
45
+ 単一 module に集約した。medium追加時にpolicy versionを2、production昇格時に3へ上げた。
46
+ - **backend / diagnostics**: model と reasoning effort を backend 生成時に一度だけ解決し、成功・失敗の
47
+ structured result と diagnostics に effective selection、選択元、policy version、検証状態を残す。
48
+ model invocation が失敗しても別 model へ retry しない。
49
+ - **比較 eval**: `spotter auditor model-matrix` を追加した。同じ versioned fixture を固定順序で実行し、
50
+ fixture hash、Codex CLI version、schema / exact match、false positive / negative、p50 / p95、timeout、
51
+ catalog 外 name と anomaly を bounded artifact に記録する。Codex JSONL `turn.completed.usage` から
52
+ token数だけを抽出し、raw event本文は保存しない。ChatGPTプランの金額costはAPI価格で代用せず
53
+ `not-available-chatgpt-plan` とし、artifact自身がmodel昇格を許可することはない。
54
+ - **usage-limit diagnostics**: Codex CLIが実測済みの利用上限文言で非ゼロ終了した場合をgenericな
55
+ `E_CODEX_CLI_EXIT` から `E_CODEX_CLI_USAGE_LIMIT` へ分離した。認証失効を先に判定し、一般的な429は
56
+ 誤分類しない。Hookはリセット時刻まで待つかCodexプランを確認する復旧案内を表示し、fallbackは行わない。
57
+ - **model unavailable diagnostics**: Codex CLIが指定modelをChatGPTアカウントで利用できないと返した場合を
58
+ `E_CODEX_CLI_MODEL_UNAVAILABLE`へ分離した。認証・利用上限を優先し、providerのstdout/stderrはredact、
59
+ effective selectionとexit codeは診断用に保持する。別modelへのfallbackは行わない。
60
+ - **公式model更新監視**: OpenAI公式のlatest-model / ChatGPT pricing Markdownの両方に揃った完全3種familyだけを
61
+ 数値版比較し、週次workflowが評価提案Issueを重複なく作る。policy書換え・model呼出・自動昇格は行わない。
62
+ - **運用SLO / Stop実測**: UserPromptSubmit / Stopのp50・p95・timeout率と品質gateを日本語で正本化した。
63
+ CLIのStop continuationはmax-1を確認した一方、App background/app-server taskはStop非発火だったため、
64
+ active Appのblock挙動が未確認のまま既存契約を変えず、pending deliveryを維持する。
65
+
66
+ ### 検証
67
+
68
+ `auditor-model-matrix-cmd` / `auditor-cmd` / `cli` の targeted test 29 / 29 pass、全体
69
+ 439 / 437 pass / 0 fail / 2 skip。
70
+ 循環・非文字列 backend 応答、model selection 不一致、version 出力の秘密混入、dirty fixture、filtered
71
+ hallucination を敵対的に再検証し、blocker 0 を確認した。代表 fixture の repeat=1 operational smoke は
72
+ baseline / Luna / Terra の全12件が Codex CLI usage limit で `E_CODEX_CLI_EXIT` となったため、model 品質・
73
+ latency・availability の比較には使わなかった。Pro20回復後のrepeat=3を2回実行し、Terra lowは合算
74
+ 23/24 exactでbaseline 18/24、Luna low 17/24より最良。usage対応runではtoken usageを36/36取得した。
75
+ Terra lowにも見逃しが1件出たためmediumを追加評価した。Terra mediumは2回のrepeat=3で合計24/24 exact、
76
+ FP/FN 0、timeout 0となり、owner裁定でproductionへ採用した。ChatGPTプランの金額costは取得不能として
77
+ 明示した。採用runは全体p95 4.361秒で、stage別SLOとtimeout/workload感度も
78
+ [`docs/04_operational-slo.md`](https://github.com/kitepon-rgb/Spotter/blob/main/docs/04_operational-slo.md)へ固定した。
79
+
3
80
  ## 1.4.17
4
81
 
5
82
  v1.4.16 で実装済みだった stale Unix socket recovery を、v1.4.16 tag を改変せずこの patch の実配布候補に含める。
@@ -25,8 +102,10 @@ daemon 異常死後の orphan socket を次回起動前に安全に除去する
25
102
 
26
103
  `1c67698` の clean worktree から npm pack / temp prefix install / CLI version / Hook install・reinstall を
27
104
  smoke し、`node --test` は 383 / 381 pass / 0 fail / 2 skip。targeted test と adversarial review も
28
- blocker 0。OS CI matrix は macOS / Linux / Windows × Node 22.5.0 / 22.x の全6件が green。
29
- publish 後の npm / GitHub Release / fresh global install 三者一致確認は残る。
105
+ blocker 0。v1.4.17-only README / CHANGELOG を載せた local release candidate は `6ea6a2b`。
106
+ CLI help、58-entry pack、同じ full suite を再確認した。OS CI matrix はmacOS / Linux / Windows ×
107
+ Node 22.5.0 / 22.xの全6件がgreen。最終 SHA `7987f2a`をtag / npm / GitHub Releaseへ公開し、
108
+ npm `latest`とfresh global installの三者一致を確認した。
30
109
 
31
110
  ## 1.4.16
32
111
 
package/README.ja.md CHANGED
@@ -15,7 +15,7 @@
15
15
 
16
16
  Claude には「使えるツールがあるのに、使うべきタイミングで使わない」という構造的な弱点があります。記録すべき決定を memory / caveat MCP に残さない、docs lookup MCP を呼ばずに古い知識で応答する、ブラウザ自動化 MCP で確認せず UI 状態を推測する — **「分からないと自覚できない」から、ツールを取りに行けない**。
17
17
 
18
- Spotter はツールカタログを完全に把握した別の監査エージェントで、ユーザー入力と主役 AI の応答を並走監査します。自動選択では Claude host は Codex CLI があればそれを、なければ session-scoped Haiku を選び、Codex host は Codex CLI を既定にします。明示 backend override は host より優先しますが、runtime failure で別 backend へ黙って切り替えません。見落としは透明化された指摘として次に利用できる文脈へ届けます。**主役 AI が自覚して自己監査する**設計は本プロダクトの存在意義を破壊するため、hook 経由でその意思と独立に検出します。
18
+ Spotter はツールカタログを完全に把握した別の監査エージェントで、ユーザー入力と主役 AI の応答を並走監査します。自動選択では Claude host は Codex CLI があればそれを、なければ session-scoped Haiku を選び、Codex host は Codex CLI を既定にします。明示 backend override は host より優先しますが、runtime failure で別 backend へ黙って切り替えません。応答前は検証済みtool IDだけを固定・非命令形の助言へ変換でき、応答後のfindingは構造eventに留めて後続turnへ注入しません。監査用AIの自由文が親セッションへ入ることはありません。**主役 AI が自覚して自己監査する**設計は本プロダクトの存在意義を破壊するため、hook 経由でその意思と独立に検出します。
19
19
 
20
20
  <p align="center">
21
21
  <img src=".github/concept.svg" alt="Claude が答え、Spotter が見ている" width="80%">
@@ -62,6 +62,7 @@ Codex 側では現行の `[features].hooks = true` を有効化し、互換の
62
62
  Spotter が所有する Codex handler は現行の同期 command schema で生成します。install / upgrade 後は `/hooks` で review して新しい Codex session を開いてください。`spotter codex-hook diagnostics` は登録と readiness を診断しますが、trust を内部状態から推測しません。
63
63
 
64
64
  Spotter を upgrade した後、release note で hook 設定変更が案内されている場合は、各 install 済みプロジェクトで `spotter install` を再実行してください。global package update でコード経路は変わりますが、既存 `.claude/settings.json` の timeout 値は自動では書き換わりません。
65
+ `v1.4.19`はruntimeの出力変換だけを変更するため、install済みprojectで`spotter install`をやり直す必要はありません。global packageを更新し、新しいClaude/Codexセッションを開いてください。
65
66
 
66
67
  ```bash
67
68
  spotter uninstall # このプロジェクトの hook 登録を解除
@@ -88,20 +89,19 @@ spotter codex-hook install
88
89
 
89
90
  ### 1 ターンの監査フロー
90
91
 
91
- Claude Code と Codex では `Stop` の受け口が違います。下の図は Claude host の流れです。
92
- Codex native `Stop` は遅延配送で、不足ツールの指摘を queue し、次の same-session
93
- `UserPromptSubmit` で提示します。
92
+ Claude Code と Codex は同じ安全なparent-output projectorを使います。監査用AIの自由文は内部に留め、
93
+ `UserPromptSubmit`では検証済みtool IDだけを固定・非命令形の助言へ変換します。`Stop` findingは
94
+ 構造eventに記録し、後のturnへ注入しません。
94
95
 
95
96
  ```mermaid
96
97
  flowchart TD
97
98
  U([User 発話]) --> UPH[UserPromptSubmit hook<br/>Spotter が発話とカタログから一次判定]
98
- UPH --> BT[Claude Thinking<br/>Spotter の推奨を<br/>additionalContext で受信]
99
+ UPH --> BT[Host model<br/>検証済みtool IDの<br/>固定助言を受け取る場合がある]
99
100
  BT --> BA([Claude の最初の応答])
100
101
  BA --> SH[Stop hook<br/>応答と使用済みツールから最終チェック]
101
102
  SH --> DEC{見落とし<br/>あり?}
102
103
  DEC -->|なし| DONE([完了])
103
- DEC -->|あり| SB[.spotter/pending/ に積む<br/>v1.4.8 deferred delivery]
104
- SB --> NEXT([次の UserPromptSubmit で<br/>additionalContext として配信])
104
+ DEC -->|あり| EVT[構造Hook eventへ記録<br/>次turnへは注入しない]
105
105
  ```
106
106
 
107
107
  ### カタログの収集経路
@@ -172,6 +172,8 @@ spotter codex-hook install
172
172
  # Codex native hooks の修復 / 明示登録 (通常は spotter install が実行)
173
173
  spotter codex-hook diagnostics
174
174
  # Codex hook の登録/readiness を診断。trust は /hooks で review
175
+ spotter auditor model-matrix --fixtures test/fixtures/auditor-model-matrix.v1.json
176
+ # pinned auditor model profile を再現可能に比較する experimental eval
175
177
  spotter uninstall # hook 登録を解除 (~/.spotter は残す)
176
178
  ```
177
179
 
@@ -190,9 +192,11 @@ Primary auditor backend policy: Claude hooks の auto selection は PATH に Cod
190
192
  `SPOTTER_AUDITOR_BACKEND` の明示 override はどちらの host でも優先し、runtime failure では別 backend へ
191
193
  hidden fallback しません。
192
194
  Codex 側の SessionStart hook は `.spotter/tool-db.codex.json` を bg refresh し、Claude DB には触れません。
193
- Codex CLI auditor の子プロセスは、hook 判定を安く速く保つため既定で `gpt-5.4-mini` と
194
- `model_reasoning_effort="low"` を明示指定します。実測や制御された実験では
195
- `SPOTTER_CODEX_CLI_MODEL` / `SPOTTER_CODEX_CLI_REASONING_EFFORT` で上書きできます。
195
+ Codex CLI auditor は versioned product policy を使い、production は反復 fixture 評価を通過した
196
+ `gpt-5.6-terra × medium`。`gpt-5.6-luna × low` / `gpt-5.6-terra × low` は比較 profile として残し、
197
+ profile から production へ自動昇格しません。`latest` alias や
198
+ 親 Codex の default を暗黙継承せず、失敗時に別 model へ retry しません。制御された実験では
199
+ `SPOTTER_CODEX_CLI_MODEL` / `SPOTTER_CODEX_CLI_REASONING_EFFORT` で上書きでき、diagnostics は unverified と表示します。
196
200
  明示 smoke には `SPOTTER_AUDITOR_BACKEND=codex-sidecar` も使えます。
197
201
 
198
202
  ## 設計ドキュメント
@@ -205,15 +209,16 @@ Codex CLI auditor の子プロセスは、hook 判定を安く速く保つため
205
209
 
206
210
  ## 既知の制約
207
211
 
208
- - v1.4.8 以降、Claude / Codex 両 host で `Stop` hook は **遅延配送 (deferred delivery)** に統一されています。`Stop` で見落としツールを検出した場合、Spotter は `<projectRoot>/.spotter/pending/<sessionId>.json` に指摘を積み、次の same-session `UserPromptSubmit` で `additionalContext` として配信します。当ターンの最初の応答は transcript にそのまま残ります
209
- - pending ファイルは Claude / Codex が同じパス (`.spotter/pending/`) を共有します。host-neutral 設計です
210
- - **Haiku の JSON スキーマ違反は v0.5.0 以降「想定済み異常」として session renew + `role_collapse_reset` で回復**します。**v1.4.15 以降、auditor/daemon の失敗はプロンプトをブロックしません**: `UserPromptSubmit` は `[Spotter からの警告]` を出して exit 0。v1.4.17 では `Stop` 失敗も warning pending に積み、次の same-session prompt で1回配信します。直後に session が終わる場合だけ、配送先となる次 prompt がありません
212
+ - `Stop` hookは最初の応答がstream済みになった後で発火します。v1.4.19以降は応答の書換え・継続強制・次turnへの監査文配送をせず、findingを構造Hook eventへ記録します。pre-responseの`UserPromptSubmit`精度が引き続き主軸です
213
+ - `UserPromptSubmit.additionalContext`は受動的metadataではなくモデル可視contextです。v1.4.19以降はcatalog一致・grammar検証済みtool IDだけから決定論的に生成し、監査用AIのreason、backend message、provider stdout/stderrを反射しません
214
+ - **Haiku の JSON スキーマ違反は v0.5.0 以降「想定済み異常」として session renew + `role_collapse_reset` で回復**します。**v1.4.19以降、auditor/daemonの失敗はモデルcontextにせずnon-blockingを維持**します。allow-list済み固定`systemMessage`・固定stderr・構造Hook eventだけを出してexit 0にします
211
215
 
212
216
  <details>
213
217
  <summary><strong>📋 最近のハイライト</strong></summary>
214
218
 
219
+ - **親出力をルールベース化** (v1.4.19) — 監査用AIの自由文は親セッションへ入りません。検証済みtool IDだけがUserPromptSubmitの任意助言になり、Stop findingは無関係な次turnへ持ち越されません
215
220
  - **daemon は異常死しても復活する** (v1.4.16) — daemon が graceful shutdown を経ず死んでも (マシンスリープ / 強制終了 / `SessionEnd` 前の crash)、残った Unix socket が以後の起動を塞がなくなった。`startDaemon` が bind 前に orphan socket を除去するので、次の `UserPromptSubmit` の auto-resurrect が `EADDRINUSE` で crash-loop せずに成功し、「そのセッションが永久に未監査」になる事態を防ぐ
216
- - **失敗は声に出して縮退、host を固めない** (v1.4.15) — auditor backend が失敗したとき (例: codex のログイン失効) も、`UserPromptSubmit` hook はプロンプトを黙って消さずに `[Spotter からの警告]` を出して通す。codex ログイン失効時は直し方 (`codex login`) を明示する
221
+ - **失敗は声に出して縮退、hostを固めない** (v1.4.15) — この版でbackend failureによるpromptのsilent消去を止めた。v1.4.19以降もnon-blocking挙動は維持し、旧model可視警告文は固定`systemMessage`・stderr・構造event診断へ置換した
217
222
  - **プラグイン形式の MCP サーバー対応** — `plugin:everything-claude-code:context7` のように名前に内部コロンを含むサーバーを正しくパースし、配下のツールをカタログに取り込めるようになった (旧版はこの形式のサーバーをすべて単一の `"plugin"` に潰して、Claude の監査から silent に脱落させていた)
218
223
  - **プロジェクト単位の監査隔離** — daemon が監査に使うのはローカル DB のみ。グローバル DB は description 再利用キャッシュに役割限定。**他プロジェクト**でインストールしたツールが現プロジェクトの監査に混入することはない
219
224
  - **手放しでカタログ維持** — `spotter install` が Claude DB を自動 seed、Claude / Codex それぞれの SessionStart が host-local DB を bg refresh する。手書き管理は一切不要
package/README.md CHANGED
@@ -15,7 +15,7 @@
15
15
 
16
16
  Claude has a structural blind spot: **it can't reach for a tool it doesn't realize it needs**. It may skip a project memory MCP when a decision should be recorded, answer from stale memory instead of a docs-lookup MCP, or reason about UI state without a browser-automation MCP. The model can't always tell when it doesn't know — so the tool stays unused.
17
17
 
18
- Spotter runs a separate auditor with the full tool catalog and checks both the user's prompt and the primary agent's reply. Automatic selection uses Codex CLI on a Claude host when available, otherwise the session-scoped Haiku path; on a Codex host it defaults to Codex CLI. An explicit backend override takes precedence, but a runtime failure never silently switches backend. When Spotter finds a missed tool, it injects a transparent recommendation into the next available context. **The primary agent is never asked to self-audit** — that would defeat the premise. Detection happens through hooks, independent of the primary agent's intent.
18
+ Spotter runs a separate auditor with the full tool catalog and checks both the user's prompt and the primary agent's reply. Automatic selection uses Codex CLI on a Claude host when available, otherwise the session-scoped Haiku path; on a Codex host it defaults to Codex CLI. An explicit backend override takes precedence, but a runtime failure never silently switches backend. Before the primary reply, validated tool IDs may become fixed, non-directive advice; after the reply, findings remain structured events and are not injected into a later turn. Auditor prose never enters the parent session. **The primary agent is never asked to self-audit** — that would defeat the premise. Detection happens through hooks, independent of the primary agent's intent.
19
19
 
20
20
  <p align="center">
21
21
  <img src=".github/concept.svg" alt="Claude answers · Spotter watches" width="80%">
@@ -62,6 +62,7 @@ For Codex, install enables the current `[features].hooks = true` flag and still
62
62
  Installer-owned Codex handlers use the current synchronous command schema. After install or upgrade, review them with `/hooks`, then open a fresh Codex session; `spotter codex-hook diagnostics` reports registration/readiness but does not guess hook trust.
63
63
 
64
64
  After upgrading Spotter, re-run `spotter install` in each installed project when release notes mention hook setting changes. The global package update changes the code path, but existing `.claude/settings.json` timeout values are not rewritten automatically.
65
+ `v1.4.19` changes runtime output projection only, so already installed projects do not need another `spotter install`; update the global package and open a fresh Claude/Codex session.
65
66
 
66
67
  ```bash
67
68
  spotter uninstall # remove hooks from this project
@@ -88,20 +89,19 @@ spotter codex-hook install
88
89
 
89
90
  ### Audit flow per turn
90
91
 
91
- Claude Code and Codex have different `Stop` surfaces. The diagram below is the Claude
92
- host flow. Codex native `Stop` uses deferred delivery: findings are queued and shown on
93
- the next same-session `UserPromptSubmit`.
92
+ Claude Code and Codex share the same safe parent-output projector. Auditor prose stays
93
+ internal; only validated tool IDs can become fixed, non-imperative advice on
94
+ `UserPromptSubmit`. `Stop` records structured findings without injecting them into a later turn.
94
95
 
95
96
  ```mermaid
96
97
  flowchart TD
97
98
  U([User prompt]) --> UPH[UserPromptSubmit hook<br/>Spotter audits prompt against catalog]
98
- UPH --> BT[Claude thinking<br/>receives Spotter's recommendations<br/>as additionalContext]
99
+ UPH --> BT[Host model<br/>may receive fixed advisory<br/>with validated tool IDs]
99
100
  BT --> BA([Claude's first answer])
100
101
  BA --> SH[Stop hook<br/>Spotter re-audits answer + tools used]
101
102
  SH --> DEC{Missed<br/>tool?}
102
103
  DEC -->|No| DONE([Done])
103
- DEC -->|Yes| SB[Queue finding to .spotter/pending/<br/>v1.4.8 deferred delivery]
104
- SB --> NEXT([Surfaces as additionalContext<br/>on next UserPromptSubmit])
104
+ DEC -->|Yes| EVT[Record structured Hook event<br/>no next-turn injection]
105
105
  ```
106
106
 
107
107
  ### Catalog discovery
@@ -173,6 +173,8 @@ spotter codex-hook install
173
173
  # repair / explicitly register Codex native hooks (normally handled by spotter install)
174
174
  spotter codex-hook diagnostics
175
175
  # check Codex hook registration/readiness; trust is reviewed with /hooks
176
+ spotter auditor model-matrix --fixtures test/fixtures/auditor-model-matrix.v1.json
177
+ # experimental reproducible comparison of pinned auditor model profiles
176
178
  spotter uninstall # remove hooks from this project (leaves ~/.spotter intact)
177
179
  ```
178
180
 
@@ -191,10 +193,12 @@ otherwise the Haiku-compatible path. Codex native hooks automatically select Cod
191
193
  `SPOTTER_AUDITOR_BACKEND` override wins on either host; runtime failure never triggers a hidden fallback.
192
194
  The Codex SessionStart hook refreshes `.spotter/tool-db.codex.json` in the background
193
195
  without touching the Claude DB.
194
- Codex CLI auditor child processes explicitly use `gpt-5.4-mini` with
195
- `model_reasoning_effort="low"` by default so hook checks stay cheap and fast;
196
+ Codex CLI auditor child processes use a versioned product policy. The production selection is
197
+ `gpt-5.6-terra × medium`, promoted after repeated fixture evaluation. `gpt-5.6-luna × low` and
198
+ `gpt-5.6-terra × low` remain comparison profiles; profiles never trigger automatic upgrades.
199
+ Spotter does not inherit a `latest` alias or the parent Codex default, and an invocation failure never retries another model.
196
200
  `SPOTTER_CODEX_CLI_MODEL` and `SPOTTER_CODEX_CLI_REASONING_EFFORT` can override
197
- those values for smoke tests or controlled experiments.
201
+ the production values for controlled experiments; diagnostics mark overrides as unverified.
198
202
  `SPOTTER_AUDITOR_BACKEND=codex-sidecar` is available for explicit sidecar auditor smoke.
199
203
 
200
204
  ## Design docs
@@ -207,15 +211,16 @@ those values for smoke tests or controlled experiments.
207
211
 
208
212
  ## Known limitations
209
213
 
210
- - The `Stop` hook fires **after** the first answer has already been streamed. Spotter therefore queues a finding for the next same-session prompt instead of rewriting that answer. Detection accuracy in `UserPromptSubmit` (the *pre-response* stage) remains the primary quality axis
211
- - `Stop` hook is **deferred** for both Claude and Codex hosts as of v1.4.8. When Spotter finds a missed tool at `Stop`, it appends the finding to `<projectRoot>/.spotter/pending/<sessionId>.json` and surfaces it on the next same-session `UserPromptSubmit` as `additionalContext`. The original assistant message stays as the turn's final transcript entry — no `decision:"block"` re-generation cycle. The same pending file is shared by Claude and Codex (host-neutral path)
212
- - **Since v0.5.0, JSON schema violations from Haiku are treated as expected anomalies** (session renew + `role_collapse_reset`). **Since v1.4.15, an auditor/daemon failure no longer blocks the prompt**: `UserPromptSubmit` emits a loud `[Spotter からの警告]` and exits 0. In v1.4.17, a `Stop` failure is queued as the same kind of warning and delivered once on the next same-session prompt. If the session ends immediately, no later prompt exists and that final warning cannot be surfaced
214
+ - The `Stop` hook fires **after** the first answer has already been streamed. Spotter records a structured finding but does not rewrite the answer, force a continuation, or inject auditor text into the next prompt. Detection accuracy in `UserPromptSubmit` (the *pre-response* stage) remains the primary quality axis
215
+ - `UserPromptSubmit.additionalContext` is model-visible context, not passive metadata. Since v1.4.19 it is generated only by a deterministic projector from catalog-matched, grammar-checked tool IDs. Auditor reasons, backend messages, and provider stdout/stderr are never reflected into it
216
+ - **Since v0.5.0, JSON schema violations from Haiku are treated as expected anomalies** (session renew + `role_collapse_reset`). **Since v1.4.19, an auditor/daemon failure remains non-blocking without becoming model context**: Claude and Codex emit only an allow-listed fixed `systemMessage`, fixed stderr, and a structured Hook event, then exit 0
213
217
 
214
218
  <details>
215
219
  <summary><strong>📋 Recent highlights</strong></summary>
216
220
 
221
+ - **Rule-based parent output boundary** (v1.4.19) — auditor AI prose cannot enter parent-session Hook output. Validated tool IDs become optional fixed advice on `UserPromptSubmit`; `Stop` findings stay structured and are never carried into an unrelated next turn
217
222
  - **Daemon recovers after an ungraceful death** (v1.4.16) — if the daemon dies without graceful shutdown (machine sleep, force-quit, crash before `SessionEnd`), the Unix socket it leaves behind no longer bricks every restart. `startDaemon` removes the orphaned socket before binding, so the next `UserPromptSubmit` auto-resurrect succeeds instead of crash-looping on `EADDRINUSE` and leaving the session permanently unaudited
218
- - **Failures degrade loudly, never freeze the host** (v1.4.15) — when the auditor backend fails (e.g. codex login expired), the `UserPromptSubmit` hook surfaces a `[Spotter からの警告]` and lets your prompt through instead of silently erasing it. codex login expiry names the one-line fix (`codex login`)
223
+ - **Failures degrade loudly, never freeze the host** (v1.4.15) — this release stopped backend failure from silently erasing a prompt. Since v1.4.19, the non-blocking behavior remains but the old model-visible warning text is replaced by fixed `systemMessage`, stderr, and structured event diagnostics
219
224
  - **Plugin-scoped MCP servers** — names like `plugin:everything-claude-code:context7` (with internal colons) are now parsed correctly and their tools enter the catalog. Earlier versions silently collapsed all plugin MCP servers into a single literal `"plugin"`, dropping their tools from Claude's audit
220
225
  - **Per-project / per-host audit isolation** — the daemon audits against the local DB only; global DBs are host-specific description caches. Tools discovered in *other* projects or another host can never bleed into this project's audit set
221
226
  - **Zero-touch catalog** — `spotter install` seeds the Claude DB automatically; Claude and Codex SessionStart hooks keep their host-local DBs fresh in the background. You never have to maintain the tool list by hand
package/bin/spotter.mjs CHANGED
@@ -50,6 +50,8 @@ Usage:
50
50
  (experimental) run primary auditor backend once
51
51
  spotter auditor matrix --stage STAGE --input FILE
52
52
  (experimental) compare primary auditor backend matrix
53
+ spotter auditor model-matrix --fixtures FILE
54
+ (experimental) evaluate pinned Codex auditor profiles
53
55
  spotter daemon start --session-id ID (internal) run session daemon
54
56
  spotter hook <event> (internal) hook dispatch
55
57
  events: session-start | user-prompt |
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "claude-spotter",
3
- "version": "1.4.17",
3
+ "version": "1.4.19",
4
4
  "description": "Audit agent running alongside Claude Code that catches missed tool calls — 気づく役と実行する役の分離",
5
5
  "type": "module",
6
6
  "bin": {
@@ -0,0 +1,151 @@
1
+ import { CODEX_AUDITOR_MODEL_POLICY } from '../src/core/codex-auditor-model-policy.mjs';
2
+
3
+ export const LATEST_MODELS_URL = 'https://developers.openai.com/api/docs/guides/latest-model.md';
4
+ export const PRICING_URL = 'https://learn.chatgpt.com/docs/pricing.md';
5
+ export const MAX_BODY_BYTES = 1024 * 1024;
6
+ export const FETCH_TIMEOUT_MS = 15_000;
7
+
8
+ const ROLES = ['sol', 'terra', 'luna'];
9
+ const PROPOSAL = '同じfixtureでLuna low/Terra lowを比較し、品質不足時のみTerra mediumを評価する。productionの自動昇格・書換えは行わない。';
10
+
11
+ export class ModelPolicyCheckError extends Error {
12
+ constructor(code) {
13
+ super(code);
14
+ this.name = 'ModelPolicyCheckError';
15
+ this.code = code;
16
+ }
17
+ }
18
+
19
+ function fail(code) { throw new ModelPolicyCheckError(code); }
20
+
21
+ function versionKey(major, minor) { return `${Number(major)}.${Number(minor)}`; }
22
+
23
+ export function compareVersions(left, right) {
24
+ const [leftMajor, leftMinor] = left.split('.').map(Number);
25
+ const [rightMajor, rightMinor] = right.split('.').map(Number);
26
+ return leftMajor - rightMajor || leftMinor - rightMinor;
27
+ }
28
+
29
+ export function extractCompleteFamilies(text, source) {
30
+ const regex = source === 'latest'
31
+ ? /\bgpt-(\d+)\.(\d+)-(sol|terra|luna)\b/gi
32
+ : /\bGPT-(\d+)\.(\d+)\s+(Sol|Terra|Luna)\b/g;
33
+ const byVersion = new Map();
34
+ for (const match of text.matchAll(regex)) {
35
+ const version = versionKey(match[1], match[2]);
36
+ const roles = byVersion.get(version) ?? new Set();
37
+ roles.add(match[3].toLowerCase());
38
+ byVersion.set(version, roles);
39
+ }
40
+ return [...byVersion.entries()]
41
+ .filter(([, roles]) => ROLES.every((role) => roles.has(role)))
42
+ .map(([version]) => ({ version, models: ROLES.map((role) => `gpt-${version}-${role}`) }))
43
+ .sort((a, b) => compareVersions(a.version, b.version));
44
+ }
45
+
46
+ async function readBoundedBody(response) {
47
+ const declaredLength = Number(response.headers?.get?.('content-length'));
48
+ if (Number.isFinite(declaredLength) && declaredLength > MAX_BODY_BYTES) fail('E_MODEL_POLICY_CHECK_BODY_TOO_LARGE');
49
+ if (!response.body?.getReader) {
50
+ const text = await response.text();
51
+ if (Buffer.byteLength(text) > MAX_BODY_BYTES) fail('E_MODEL_POLICY_CHECK_BODY_TOO_LARGE');
52
+ return text;
53
+ }
54
+ const reader = response.body.getReader();
55
+ const chunks = [];
56
+ let size = 0;
57
+ try {
58
+ for (;;) {
59
+ const { done, value } = await reader.read();
60
+ if (done) break;
61
+ size += value.byteLength;
62
+ if (size > MAX_BODY_BYTES) {
63
+ await reader.cancel();
64
+ fail('E_MODEL_POLICY_CHECK_BODY_TOO_LARGE');
65
+ }
66
+ chunks.push(value);
67
+ }
68
+ } finally {
69
+ reader.releaseLock?.();
70
+ }
71
+ return new TextDecoder().decode(Buffer.concat(chunks));
72
+ }
73
+
74
+ export async function fetchMarkdown(url, { fetchFn = fetch, timeoutMs = FETCH_TIMEOUT_MS } = {}) {
75
+ const controller = new AbortController();
76
+ const timeout = setTimeout(() => controller.abort(), timeoutMs);
77
+ try {
78
+ let response;
79
+ try {
80
+ response = await fetchFn(url, { signal: controller.signal, headers: { accept: 'text/markdown,text/plain;q=0.9' } });
81
+ } catch {
82
+ fail('E_MODEL_POLICY_CHECK_FETCH_FAILED');
83
+ }
84
+ if (!response || response.status !== 200) fail('E_MODEL_POLICY_CHECK_HTTP_STATUS');
85
+ return await readBoundedBody(response);
86
+ } finally {
87
+ clearTimeout(timeout);
88
+ }
89
+ }
90
+
91
+ function policyProduction(policy) {
92
+ const model = policy?.production?.model;
93
+ const match = typeof model === 'string' && /^gpt-(\d+)\.(\d+)-(sol|terra|luna)$/.exec(model);
94
+ if (!match) fail('E_MODEL_POLICY_CHECK_INVALID_POLICY');
95
+ return { model, version: versionKey(match[1], match[2]) };
96
+ }
97
+
98
+ export async function runModelPolicyCheck({
99
+ fetchFn = fetch,
100
+ now = () => new Date(),
101
+ policy = CODEX_AUDITOR_MODEL_POLICY,
102
+ timeoutMs = FETCH_TIMEOUT_MS,
103
+ } = {}) {
104
+ const production = policyProduction(policy);
105
+ const [latestMarkdown, pricingMarkdown] = await Promise.all([
106
+ fetchMarkdown(LATEST_MODELS_URL, { fetchFn, timeoutMs }),
107
+ fetchMarkdown(PRICING_URL, { fetchFn, timeoutMs }),
108
+ ]);
109
+ const latestFamilies = extractCompleteFamilies(latestMarkdown, 'latest');
110
+ const pricingFamilies = extractCompleteFamilies(pricingMarkdown, 'pricing');
111
+ if (latestFamilies.length === 0 || pricingFamilies.length === 0) fail('E_MODEL_POLICY_CHECK_REQUIRED_FAMILY_MISSING');
112
+ const latestDetected = latestFamilies.at(-1);
113
+ const pricingDetected = pricingFamilies.at(-1);
114
+ if (latestDetected.version !== pricingDetected.version) fail('E_MODEL_POLICY_CHECK_SOURCE_MISMATCH');
115
+ const pricingVersionSet = new Set(pricingFamilies.map(({ version }) => version));
116
+ const commonFamilies = latestFamilies.filter(({ version }) => pricingVersionSet.has(version));
117
+ const detectedFamily = latestDetected;
118
+ const status = compareVersions(detectedFamily.version, production.version) > 0 ? 'update-available' : 'current';
119
+ const diagnostics = status === 'current'
120
+ && commonFamilies.some(({ version }) => version !== production.version)
121
+ ? ['E_MODEL_POLICY_CHECK_NON_PRODUCTION_CANDIDATE_PRESENT']
122
+ : [];
123
+ return {
124
+ schema: 'spotter.codex_model_update_check.v1',
125
+ checkedAt: now().toISOString(),
126
+ policy: { version: policy.policyVersion, production: policy.production.model },
127
+ sources: [LATEST_MODELS_URL, PRICING_URL],
128
+ detectedFamily,
129
+ candidates: commonFamilies,
130
+ status,
131
+ proposal: status === 'update-available' ? PROPOSAL : null,
132
+ diagnostics,
133
+ };
134
+ }
135
+
136
+ export async function main({ stdout = process.stdout, stderr = process.stderr, run = runModelPolicyCheck } = {}) {
137
+ try {
138
+ const artifact = await run();
139
+ stdout.write(`${JSON.stringify(artifact)}\n`);
140
+ return 0;
141
+ } catch (error) {
142
+ const code = error instanceof ModelPolicyCheckError ? error.code : 'E_MODEL_POLICY_CHECK_UNEXPECTED';
143
+ stderr.write(`${code}\n`);
144
+ return 1;
145
+ }
146
+ }
147
+
148
+ if (import.meta.url === new URL(process.argv[1], 'file:').href) {
149
+ const exitCode = await main();
150
+ process.exitCode = exitCode;
151
+ }
@@ -2,6 +2,7 @@ import { readFile } from 'node:fs/promises';
2
2
  import { resolve } from 'node:path';
3
3
  import { createAuditorBackend } from '../core/auditor-backend.mjs';
4
4
  import { readLocal } from '../tool-db/refresh.mjs';
5
+ import { runAuditorModelMatrixCommand } from './auditor-model-matrix-cmd.mjs';
5
6
 
6
7
  const AUDITOR_USAGE = `spotter auditor — experimental primary auditor smoke commands
7
8
 
@@ -10,6 +11,8 @@ Usage:
10
11
  [--project DIR] [--host-agent claude|codex|automation|unknown]
11
12
  [--backend haiku|codex-cli|codex-sidecar|auto]
12
13
  spotter auditor matrix --stage user_input|turn_end --input FILE [--project DIR]
14
+ spotter auditor model-matrix --fixtures FILE [--profile baseline|luna|terra|terra-medium]...
15
+ [--repeat N] [--project DIR] [--output FILE]
13
16
 
14
17
  Input JSON:
15
18
  user_input: {"user_input":"..."} or {"userInput":"..."}
@@ -32,6 +35,14 @@ export async function runAuditorCommand({ argv = process.argv.slice(2) } = {}) {
32
35
  await runAuditorMatrixCommand({ argv: argv.slice(1) });
33
36
  return;
34
37
  }
38
+ if (sub === 'model-matrix') {
39
+ if (argv.slice(1).includes('--help') || argv.slice(1).includes('-h')) {
40
+ process.stdout.write(`Usage: spotter auditor model-matrix --fixtures FILE [--profile baseline|luna|terra|terra-medium]...\n [--repeat N] [--project DIR] [--output FILE]\n`);
41
+ return;
42
+ }
43
+ await runAuditorModelMatrixCommand({ argv: argv.slice(1) });
44
+ return;
45
+ }
35
46
  process.stderr.write(`unknown auditor subcommand: ${sub}\n${AUDITOR_USAGE}`);
36
47
  process.exit(2);
37
48
  }