claude-spotter 1.4.27 → 1.4.29
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +34 -0
- package/README.ja.md +16 -0
- package/README.md +16 -0
- package/bin/spotter.mjs +8 -0
- package/package.json +7 -2
- package/scripts/verify-release-commit.mjs +50 -0
- package/scripts/verify-release-commit.test.mjs +48 -0
- package/src/cli/codex-hook-cmd.mjs +186 -2
- package/src/cli/diagnostics-cmd.mjs +1 -1
- package/src/cli/evaluation-cmd.mjs +129 -0
- package/src/core/codex-transcript.mjs +13 -3
- package/src/core/evaluation-context.mjs +184 -0
- package/src/core/evaluation-report.mjs +84 -0
- package/src/core/evaluation-store.mjs +461 -0
- package/src/core/evaluation-tool-id.mjs +382 -0
- package/src/daemon/daemon.mjs +93 -4
- package/src/hooks/pre-tool-use.mjs +57 -1
- package/src/hooks/user-prompt.mjs +83 -2
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,39 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 1.4.29 — 2026-08-04
|
|
4
|
+
|
|
5
|
+
- **提案率とtool採用率の端末内評価を追加。** Claude / Codexの成功UserPromptSubmitを母数に、
|
|
6
|
+
safe projector後に実際に提示したtool itemと、同じturnでのcanonical tool利用を
|
|
7
|
+
`~/.spotter/evaluation.db`へ記録する。採用率の母数はoutcome確定itemだけとし、利用記録不全は
|
|
8
|
+
`outcome_missing`へ分離する。組込みtoolは評価対象外で、既存turn-end auditorのraw usedToolsは変更しない。
|
|
9
|
+
- **提案時の別文脈を保存。** proposal確定時刻を記録し、既存Throughline `observer-read`を一度だけ呼ぶ。
|
|
10
|
+
Spotterがauditorへ渡したcontextとは別fieldに保存し、失敗時はcontextだけunavailableとしてretryしない。
|
|
11
|
+
返却されたhost / thread hashを提案元sessionと照合し、同じprojectの並行sessionを取り違えた場合も
|
|
12
|
+
snapshotを保存しない。
|
|
13
|
+
- **欠測とtool identityを率から分離。** Claudeの利用記録失敗、Codex transcript不全、Stopが30分以上
|
|
14
|
+
欠落したopen proposalは`outcome_missing`として投影し、非採用へ混ぜない。Codex Skillは提案済みの
|
|
15
|
+
正規`SKILL.md`を実際にreadしたcallだけを採用として認識する。現行Codexのouter `exec`内で実行される
|
|
16
|
+
nested MCP callと、同じturn内の複数`exec`も個別の利用入力として保持する。
|
|
17
|
+
- **初回同時openを直列化。** 2 projectが未作成の評価DBを同時に開いてもWAL設定前に
|
|
18
|
+
`database is locked`とならないよう、bounded busy timeoutをschema初期化より先に有効化する。
|
|
19
|
+
retry、lockfile、background回収は追加しない。
|
|
20
|
+
- **評価CLIを追加。** `spotter evaluation report`で全project / project / tool / host別の
|
|
21
|
+
`S/P/I/C/A/M`、`P/S`、`A/C`を表示し、`cases` / `case`で非採用caseのrequest、2種類の文脈、
|
|
22
|
+
提案ID、利用ID、outcomeを確認できる。reportは保存済みSQLiteだけを読み、JSONはWindows
|
|
23
|
+
PowerShell 5.1でも壊れないASCII-safe形式で出力する。
|
|
24
|
+
- **検証・公開。** 固定fixture、Claude / Codex hook、Throughline束縛、2 process同時writeを含む
|
|
25
|
+
568 tests(macOS / Linuxは566 pass / 2 skip、Windowsは556 pass / 12 platform skip)、
|
|
26
|
+
macOS / Linux / Windows × Node 22.13 / 22.xのCI run `30913375991`、75-file packを通過した。
|
|
27
|
+
別projectの実Codex turnではThroughline文脈付きproposalから
|
|
28
|
+
`S=1 P=1 I=1 C=1 A=1 M=0`を記録し、Claude実turnのpass観測と両hostのproposal fixtureも確認した。
|
|
29
|
+
Windowsではouter `exec`内の絶対pathを壊さずSkill採用へ変換する回帰も固定した。
|
|
30
|
+
|
|
31
|
+
## 1.4.28 — 2026-07-21
|
|
32
|
+
|
|
33
|
+
- **Windows Codex SessionStartの誤timeoutを修正。** detached refreshを起動するSpotter所有hookの
|
|
34
|
+
host側上限を5秒から30秒へ変更し、Windows nativeのNode起動・project discoveryが5秒を超えても
|
|
35
|
+
SessionStart失敗として切られないようにした。再installは旧5秒設定をcanonical 30秒へ正規化する。
|
|
36
|
+
|
|
3
37
|
## 1.4.27 — 2026-07-21
|
|
4
38
|
|
|
5
39
|
- **Windows PowerShell 5.1向け診断JSONをASCII安全化。** `spotter diagnostics logs --json`は
|
package/README.ja.md
CHANGED
|
@@ -13,6 +13,16 @@
|
|
|
13
13
|
|
|
14
14
|
> **気づく役と実行する役を分離する。** Claude Code の横で並走し、主役の Claude が**ツールを呼び忘れたとき**だけ静かに指摘する監査役。
|
|
15
15
|
|
|
16
|
+
[kitepon.dev](https://kitepon.dev/)を運営する[クオ(@QLyun35332)](https://x.com/QLyun35332)が
|
|
17
|
+
開発・メンテナンスしています。
|
|
18
|
+
|
|
19
|
+
## 所有境界
|
|
20
|
+
|
|
21
|
+
本repositoryは監査挙動、project marker、installer、release、diagnosticsを所有します。
|
|
22
|
+
製品横断のcatalogとhost統合は、kitepon.devの製品開発を支える内部基盤
|
|
23
|
+
[dotagents](https://github.com/kitepon-rgb/dotagents)が担当します。
|
|
24
|
+
MarkItDownは別区分の第三者CLIです。
|
|
25
|
+
|
|
16
26
|
Claude には「使えるツールがあるのに、使うべきタイミングで使わない」という構造的な弱点があります。記録すべき決定を memory / caveat MCP に残さない、docs lookup MCP を呼ばずに古い知識で応答する、ブラウザ自動化 MCP で確認せず UI 状態を推測する — **「分からないと自覚できない」から、ツールを取りに行けない**。
|
|
17
27
|
|
|
18
28
|
Spotter はツールカタログを完全に把握した別の監査エージェントで、ユーザー入力と主役 AI の応答を並走監査します。自動選択では Claude host は Codex CLI があればそれを、なければ session-scoped Haiku を選び、Codex host は Codex CLI を既定にします。明示 backend override は host より優先しますが、runtime failure で別 backend へ黙って切り替えません。応答前は検証済みtool IDだけを固定・非命令形の助言へ変換でき、応答後のfindingは構造eventに留めて後続turnへ注入しません。監査用AIの自由文が親セッションへ入ることはありません。**主役 AI が自覚して自己監査する**設計は本プロダクトの存在意義を破壊するため、hook 経由でその意思と独立に検出します。
|
|
@@ -207,6 +217,12 @@ spotter db rebuild # Claude local + Claude global DB を両方消してか
|
|
|
207
217
|
spotter status # 稼働中の daemon 一覧
|
|
208
218
|
spotter doctor # 環境診断 (Node / claude CLI / Codex readiness / tool-db 整合性)
|
|
209
219
|
spotter diagnostics logs # daemon log から pass=false / backend latency / anomaly signal を集計
|
|
220
|
+
spotter evaluation report
|
|
221
|
+
# 端末内DBからproject横断の提案率・tool採用率を集計
|
|
222
|
+
spotter evaluation cases --outcome not-adopted
|
|
223
|
+
# 提案されたが同じturnで使われなかったtool itemを一覧
|
|
224
|
+
spotter evaluation case <observation-id>
|
|
225
|
+
# request、2種類の文脈、提案、利用、outcomeを確認
|
|
210
226
|
spotter codex risk-check --findings findings.json --host-agent claude
|
|
211
227
|
# Spotter finding を codex-sidecar に渡して read-only risk analysis
|
|
212
228
|
spotter codex review|explore|opinion --findings findings.json --host-agent claude
|
package/README.md
CHANGED
|
@@ -13,6 +13,16 @@
|
|
|
13
13
|
|
|
14
14
|
> **Separate the spotter from the doer.** Spotter runs alongside Claude Code and quietly flags the moments when your primary Claude **forgets to use a tool it has access to**.
|
|
15
15
|
|
|
16
|
+
Built and maintained by [Quo](https://x.com/QLyun35332) at [kitepon.dev](https://kitepon.dev/en).
|
|
17
|
+
|
|
18
|
+
## Ownership boundary
|
|
19
|
+
|
|
20
|
+
This repository owns auditor behavior, project markers, installers, release,
|
|
21
|
+
and diagnostics. The cross-product catalog and host integration are handled by
|
|
22
|
+
[dotagents](https://github.com/kitepon-rgb/dotagents), the internal development
|
|
23
|
+
toolchain behind kitepon.dev's products.
|
|
24
|
+
MarkItDown is managed separately as a third-party CLI.
|
|
25
|
+
|
|
16
26
|
Claude has a structural blind spot: **it can't reach for a tool it doesn't realize it needs**. It may skip a project memory MCP when a decision should be recorded, answer from stale memory instead of a docs-lookup MCP, or reason about UI state without a browser-automation MCP. The model can't always tell when it doesn't know — so the tool stays unused.
|
|
17
27
|
|
|
18
28
|
Spotter runs a separate auditor with the full tool catalog and checks both the user's prompt and the primary agent's reply. Automatic selection uses Codex CLI on a Claude host when available, otherwise the session-scoped Haiku path; on a Codex host it defaults to Codex CLI. An explicit backend override takes precedence, but a runtime failure never silently switches backend. Before the primary reply, validated tool IDs may become fixed, non-directive advice; after the reply, findings remain structured events and are not injected into a later turn. Auditor prose never enters the parent session. **The primary agent is never asked to self-audit** — that would defeat the premise. Detection happens through hooks, independent of the primary agent's intent.
|
|
@@ -216,6 +226,12 @@ spotter doctor # environment check (Node / claude CLI / Codex readines
|
|
|
216
226
|
spotter diagnostics logs # summarize daemon logs for pass=false / backend latency / anomaly signals
|
|
217
227
|
spotter diagnostics runtime-errors
|
|
218
228
|
# print the local allow-listed runtime-error aggregate snapshot (no network)
|
|
229
|
+
spotter evaluation report
|
|
230
|
+
# show cross-project proposal/adoption counts and rates from the local DB
|
|
231
|
+
spotter evaluation cases --outcome not-adopted
|
|
232
|
+
# list proposed tools that were not used in the same turn
|
|
233
|
+
spotter evaluation case <observation-id>
|
|
234
|
+
# inspect request, Spotter context, Throughline snapshot, proposal, usage, and outcome
|
|
219
235
|
spotter codex risk-check --findings findings.json --host-agent claude
|
|
220
236
|
# run read-only codex-sidecar risk analysis for Spotter findings
|
|
221
237
|
spotter codex review|explore|opinion --findings findings.json --host-agent claude
|
package/bin/spotter.mjs
CHANGED
|
@@ -11,6 +11,7 @@ import { runCodexCommand } from '../src/cli/codex-cmd.mjs';
|
|
|
11
11
|
import { runCodexHookCommand } from '../src/cli/codex-hook-cmd.mjs';
|
|
12
12
|
import { runAuditorCommand } from '../src/cli/auditor-cmd.mjs';
|
|
13
13
|
import { runDiagnosticsCommand } from '../src/cli/diagnostics-cmd.mjs';
|
|
14
|
+
import { runEvaluationCommand } from '../src/cli/evaluation-cmd.mjs';
|
|
14
15
|
import { runFactoryDiagnostics } from '../src/cli/factory-diagnostics.mjs';
|
|
15
16
|
import { runDaemonStart } from '../src/cli/daemon-cmd.mjs';
|
|
16
17
|
import { runSessionStart } from '../src/hooks/session-start.mjs';
|
|
@@ -45,6 +46,10 @@ Usage:
|
|
|
45
46
|
spotter diagnostics logs [--json] summarize daemon logs for precision diagnostics
|
|
46
47
|
spotter diagnostics runtime-errors [snapshot|ack|resolve|reopen|compact]
|
|
47
48
|
consume the local allow-listed aggregate store
|
|
49
|
+
spotter evaluation report [filters] [--json]
|
|
50
|
+
spotter evaluation cases --outcome OUTCOME [filters] [--json]
|
|
51
|
+
spotter evaluation case OBSERVATION_ID [--json]
|
|
52
|
+
read saved proposal-adoption observations
|
|
48
53
|
spotter codex risk-check --findings FILE
|
|
49
54
|
run read-only codex-sidecar risk analysis
|
|
50
55
|
spotter codex review|explore|opinion --findings FILE
|
|
@@ -124,6 +129,9 @@ async function main() {
|
|
|
124
129
|
}
|
|
125
130
|
await runDiagnosticsCommand({ argv: rest });
|
|
126
131
|
return;
|
|
132
|
+
case 'evaluation':
|
|
133
|
+
await runEvaluationCommand({ argv: rest });
|
|
134
|
+
return;
|
|
127
135
|
case 'daemon': {
|
|
128
136
|
const sub = rest[0];
|
|
129
137
|
if (sub === 'start') { await runDaemonStart({ argv: rest.slice(1) }); return; }
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "claude-spotter",
|
|
3
|
-
"version": "1.4.
|
|
3
|
+
"version": "1.4.29",
|
|
4
4
|
"description": "Audit agent running alongside Claude Code that catches missed tool calls — 気づく役と実行する役の分離",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
@@ -11,6 +11,8 @@
|
|
|
11
11
|
},
|
|
12
12
|
"scripts": {
|
|
13
13
|
"test": "node --test",
|
|
14
|
+
"verify:release-commit": "node scripts/verify-release-commit.mjs",
|
|
15
|
+
"prepublishOnly": "npm run verify:release-commit && npm test",
|
|
14
16
|
"postinstall": "node scripts/postinstall.mjs",
|
|
15
17
|
"preuninstall": "node scripts/preuninstall.mjs"
|
|
16
18
|
},
|
|
@@ -23,7 +25,10 @@
|
|
|
23
25
|
"audit",
|
|
24
26
|
"spotter"
|
|
25
27
|
],
|
|
26
|
-
"author":
|
|
28
|
+
"author": {
|
|
29
|
+
"name": "Quo / クオ at kitepon.dev",
|
|
30
|
+
"url": "https://kitepon.dev/"
|
|
31
|
+
},
|
|
27
32
|
"license": "MIT",
|
|
28
33
|
"repository": {
|
|
29
34
|
"type": "git",
|
|
@@ -0,0 +1,50 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
|
|
3
|
+
// publishする成果が、既定ブランチへ着地済みのcommitから作られていることを保証する。
|
|
4
|
+
// 着地していないbranchからpublishすると、そのbranchが取り残された時点で公開物が
|
|
5
|
+
// 後続releaseから消え、統合契約だけが存在しない面を指し続ける(0.3.0の孤児releaseが実例)。
|
|
6
|
+
|
|
7
|
+
import assert from "node:assert/strict";
|
|
8
|
+
import { spawnSync } from "node:child_process";
|
|
9
|
+
import path from "node:path";
|
|
10
|
+
import { fileURLToPath } from "node:url";
|
|
11
|
+
|
|
12
|
+
const scriptPath = fileURLToPath(import.meta.url);
|
|
13
|
+
const defaultProjectDirectory = path.dirname(path.dirname(scriptPath));
|
|
14
|
+
|
|
15
|
+
export function verifyReleaseCommit({
|
|
16
|
+
projectDirectory = defaultProjectDirectory,
|
|
17
|
+
spawn = spawnSync,
|
|
18
|
+
} = {}) {
|
|
19
|
+
const git = (...args) => {
|
|
20
|
+
const result = spawn("git", args, { cwd: projectDirectory, encoding: "utf8" });
|
|
21
|
+
return { ok: result.status === 0, stdout: (result.stdout ?? "").trim() };
|
|
22
|
+
};
|
|
23
|
+
|
|
24
|
+
const head = git("rev-parse", "HEAD");
|
|
25
|
+
assert.ok(head.ok, "git HEAD を解決できません");
|
|
26
|
+
const originHead = git("symbolic-ref", "--quiet", "refs/remotes/origin/HEAD");
|
|
27
|
+
const defaultRef = originHead.ok && originHead.stdout
|
|
28
|
+
? originHead.stdout.replace("refs/remotes/", "")
|
|
29
|
+
: "origin/main";
|
|
30
|
+
const defaultResolved = git("rev-parse", "--verify", `${defaultRef}^{commit}`);
|
|
31
|
+
assert.ok(defaultResolved.ok, `既定ブランチ ${defaultRef} を解決できません。git fetch origin を先に実行してください。`);
|
|
32
|
+
const isAncestor = git("merge-base", "--is-ancestor", head.stdout, defaultRef);
|
|
33
|
+
assert.ok(
|
|
34
|
+
isAncestor.ok,
|
|
35
|
+
`publish対象 ${head.stdout.slice(0, 12)} が ${defaultRef} の祖先ではありません。`
|
|
36
|
+
+ " 先に既定ブランチへ着地させてpushしてから publish してください。"
|
|
37
|
+
);
|
|
38
|
+
const dirty = git("status", "--porcelain", "--untracked-files=normal");
|
|
39
|
+
assert.ok(dirty.ok, "working treeの状態を確認できません");
|
|
40
|
+
assert.equal(
|
|
41
|
+
dirty.stdout,
|
|
42
|
+
"",
|
|
43
|
+
`working treeに未commitの変更があります。publish対象commitとpayloadが一致しません:\n${dirty.stdout}`
|
|
44
|
+
);
|
|
45
|
+
return `release commit ${head.stdout.slice(0, 12)} is landed on ${defaultRef}.`;
|
|
46
|
+
}
|
|
47
|
+
|
|
48
|
+
if (process.argv[1] && path.resolve(process.argv[1]) === scriptPath) {
|
|
49
|
+
console.log(verifyReleaseCommit());
|
|
50
|
+
}
|
|
@@ -0,0 +1,48 @@
|
|
|
1
|
+
import assert from "node:assert/strict";
|
|
2
|
+
import { execFileSync } from "node:child_process";
|
|
3
|
+
import { mkdtemp, mkdir, rm, writeFile } from "node:fs/promises";
|
|
4
|
+
import { tmpdir } from "node:os";
|
|
5
|
+
import path from "node:path";
|
|
6
|
+
import test from "node:test";
|
|
7
|
+
|
|
8
|
+
import { verifyReleaseCommit } from "./verify-release-commit.mjs";
|
|
9
|
+
|
|
10
|
+
function git(cwd, ...args) {
|
|
11
|
+
return execFileSync("git", args, { cwd, encoding: "utf8", stdio: ["ignore", "pipe", "pipe"] });
|
|
12
|
+
}
|
|
13
|
+
|
|
14
|
+
async function createReleaseFixture(t) {
|
|
15
|
+
const root = await mkdtemp(path.join(tmpdir(), "spotter-release-gate-"));
|
|
16
|
+
const remote = path.join(root, "origin.git");
|
|
17
|
+
const work = path.join(root, "work");
|
|
18
|
+
t.after(async () => rm(root, { recursive: true, force: true }));
|
|
19
|
+
await mkdir(work);
|
|
20
|
+
git(root, "init", "--bare", remote);
|
|
21
|
+
git(work, "init", "--initial-branch=main");
|
|
22
|
+
git(work, "config", "user.name", "Spotter release gate test");
|
|
23
|
+
git(work, "config", "user.email", "spotter-release-gate@example.invalid");
|
|
24
|
+
await writeFile(path.join(work, ".gitignore"), "ignored.tmp\n", "utf8");
|
|
25
|
+
await writeFile(path.join(work, "tracked.txt"), "tracked\n", "utf8");
|
|
26
|
+
git(work, "add", ".gitignore", "tracked.txt");
|
|
27
|
+
git(work, "commit", "-m", "fixture");
|
|
28
|
+
git(work, "remote", "add", "origin", remote);
|
|
29
|
+
git(work, "push", "--set-upstream", "origin", "main");
|
|
30
|
+
return work;
|
|
31
|
+
}
|
|
32
|
+
|
|
33
|
+
test("release gate passes for a commit landed on main", async (t) => {
|
|
34
|
+
const work = await createReleaseFixture(t);
|
|
35
|
+
assert.match(verifyReleaseCommit({ projectDirectory: work }), /landed on origin\/main/);
|
|
36
|
+
});
|
|
37
|
+
|
|
38
|
+
test("release gate rejects an untracked publish payload", async (t) => {
|
|
39
|
+
const work = await createReleaseFixture(t);
|
|
40
|
+
await writeFile(path.join(work, "untracked.txt"), "must block publish\n", "utf8");
|
|
41
|
+
assert.throws(() => verifyReleaseCommit({ projectDirectory: work }), /working treeに未commitの変更があります/);
|
|
42
|
+
});
|
|
43
|
+
|
|
44
|
+
test("release gate permits ignored generated output", async (t) => {
|
|
45
|
+
const work = await createReleaseFixture(t);
|
|
46
|
+
await writeFile(path.join(work, "ignored.tmp"), "generated\n", "utf8");
|
|
47
|
+
assert.match(verifyReleaseCommit({ projectDirectory: work }), /landed on origin\/main/);
|
|
48
|
+
});
|
|
@@ -1,4 +1,5 @@
|
|
|
1
1
|
import { spawnSync } from 'node:child_process';
|
|
2
|
+
import { randomUUID } from 'node:crypto';
|
|
2
3
|
import { existsSync, realpathSync } from 'node:fs';
|
|
3
4
|
import { mkdir, readFile, writeFile } from 'node:fs/promises';
|
|
4
5
|
import { homedir } from 'node:os';
|
|
@@ -34,11 +35,21 @@ import {
|
|
|
34
35
|
readProjectAuditorContextConfig,
|
|
35
36
|
} from '../core/auditor-context.mjs';
|
|
36
37
|
import { observeRuntimeErrorIsolatedSafe } from '../core/runtime-error-store.mjs';
|
|
38
|
+
import { createEvaluationStore } from '../core/evaluation-store.mjs';
|
|
39
|
+
import { loadEvaluationObserverContext } from '../core/evaluation-context.mjs';
|
|
40
|
+
import {
|
|
41
|
+
canonicalizeCodexNestedMcpToolIds,
|
|
42
|
+
canonicalizeCodexSkillReadToolIds,
|
|
43
|
+
canonicalizeProposedToolIds,
|
|
44
|
+
canonicalizeUsedToolIds,
|
|
45
|
+
} from '../core/evaluation-tool-id.mjs';
|
|
46
|
+
import { version as SPOTTER_VERSION } from '../version.mjs';
|
|
37
47
|
|
|
38
48
|
const HERE = dirname(fileURLToPath(import.meta.url));
|
|
39
49
|
const PACKAGE_ROOT = resolve(HERE, '..', '..');
|
|
40
50
|
const SPOTTER_BIN = join(PACKAGE_ROOT, 'bin', 'spotter.mjs');
|
|
41
51
|
const CODEX_HOOK_TIMEOUT_SEC = 60;
|
|
52
|
+
const CODEX_SESSION_START_TIMEOUT_SEC = 30;
|
|
42
53
|
const DEFAULT_CODEX_HOOK_AUDITOR_TIMEOUT_MS = 20_000;
|
|
43
54
|
const DEFAULT_CODEX_STOP_SHORT_FINAL_MAX_CHARS = 120;
|
|
44
55
|
const CODEX_HOOK_FEATURE_NAMES = ['hooks', 'codex_hooks'];
|
|
@@ -124,22 +135,77 @@ export async function runCodexUserPromptSubmitHook({
|
|
|
124
135
|
writeOutput = (text) => process.stdout.write(text),
|
|
125
136
|
writeError = (text) => process.stderr.write(text),
|
|
126
137
|
runtimeErrorObserver = async () => ({ collected: false, reason: 'observer_not_configured' }),
|
|
138
|
+
createEvaluationStoreFn = createEvaluationStore,
|
|
139
|
+
loadEvaluationObserverContextFn = loadEvaluationObserverContext,
|
|
140
|
+
randomUUIDFn = randomUUID,
|
|
141
|
+
now = () => Date.now(),
|
|
127
142
|
} = {}) {
|
|
128
143
|
if (isChildCall()) return;
|
|
129
144
|
const input = await readInput();
|
|
145
|
+
const observationId = randomUUIDFn();
|
|
146
|
+
const recordedAtMs = now();
|
|
130
147
|
const projectRoot = findSpotterMarker(input.cwd);
|
|
131
148
|
if (!projectRoot) return;
|
|
132
149
|
const startedAt = Date.now();
|
|
133
150
|
let contextDurationMs = null;
|
|
151
|
+
let auditorContextConfig;
|
|
134
152
|
|
|
135
153
|
const prompt = requireString(input, 'prompt');
|
|
154
|
+
const sessionId = codexSessionId(input);
|
|
155
|
+
const recordEvaluation = async ({ auditStatus, proposedToolIds = [], backend = null, model = null, auditorSeenContext = null, config = undefined }) => {
|
|
156
|
+
if (!sessionId) {
|
|
157
|
+
reportEvaluationFailure(writeError);
|
|
158
|
+
return;
|
|
159
|
+
}
|
|
160
|
+
const proposals = canonicalizeProposedToolIds(proposedToolIds);
|
|
161
|
+
let observerContext = { status: 'not_requested', snapshot: null };
|
|
162
|
+
let proposedAtMs = recordedAtMs;
|
|
163
|
+
try {
|
|
164
|
+
if (auditStatus === 'success' && proposals.resolvedToolIds.length > 0) {
|
|
165
|
+
proposedAtMs = now();
|
|
166
|
+
observerContext = await loadEvaluationObserverContextFn({
|
|
167
|
+
projectRoot,
|
|
168
|
+
host: 'codex',
|
|
169
|
+
sessionId,
|
|
170
|
+
config,
|
|
171
|
+
recordedAtMs: proposedAtMs,
|
|
172
|
+
});
|
|
173
|
+
}
|
|
174
|
+
const store = createEvaluationStoreFn();
|
|
175
|
+
try {
|
|
176
|
+
store.recordTurn({
|
|
177
|
+
observationId,
|
|
178
|
+
recordedAtMs: proposals.resolvedToolIds.length > 0 ? proposedAtMs : recordedAtMs,
|
|
179
|
+
proposedAtMs,
|
|
180
|
+
projectPath: projectRoot,
|
|
181
|
+
host: 'codex',
|
|
182
|
+
sessionId,
|
|
183
|
+
auditStatus,
|
|
184
|
+
requestText: prompt,
|
|
185
|
+
auditorSeenContext,
|
|
186
|
+
observerContextStatus: observerContext.status,
|
|
187
|
+
observerSnapshot: observerContext.snapshot,
|
|
188
|
+
proposedToolIds: proposals.resolvedToolIds,
|
|
189
|
+
backend,
|
|
190
|
+
model,
|
|
191
|
+
spotterVersion: SPOTTER_VERSION,
|
|
192
|
+
});
|
|
193
|
+
} finally {
|
|
194
|
+
store.close();
|
|
195
|
+
}
|
|
196
|
+
} catch {
|
|
197
|
+
reportEvaluationFailure(writeError);
|
|
198
|
+
}
|
|
199
|
+
};
|
|
136
200
|
const legacyPending = await discardLegacyPendingFn({ projectRoot, sessionId: codexSessionId(input) });
|
|
137
201
|
let context;
|
|
138
202
|
const contextStartedAt = Date.now();
|
|
139
203
|
try {
|
|
140
204
|
const config = await readAuditorContextConfigFn(projectRoot);
|
|
205
|
+
auditorContextConfig = config;
|
|
141
206
|
contextDurationMs = Date.now() - contextStartedAt;
|
|
142
207
|
if (config.mode === 'disabled') {
|
|
208
|
+
await recordEvaluation({ auditStatus: 'skipped' });
|
|
143
209
|
await recordCodexHookEventSafe(recordHookEventFn, {
|
|
144
210
|
projectRoot,
|
|
145
211
|
event: {
|
|
@@ -163,6 +229,7 @@ export async function runCodexUserPromptSubmitHook({
|
|
|
163
229
|
});
|
|
164
230
|
contextDurationMs = Date.now() - contextStartedAt;
|
|
165
231
|
} catch (err) {
|
|
232
|
+
await recordEvaluation({ auditStatus: 'error' });
|
|
166
233
|
await observeRuntimeFailure(runtimeErrorObserver, 'auditor_unavailable');
|
|
167
234
|
contextDurationMs = Date.now() - contextStartedAt;
|
|
168
235
|
const failure = projectBackendFailure(err?.code);
|
|
@@ -183,6 +250,7 @@ export async function runCodexUserPromptSubmitHook({
|
|
|
183
250
|
return;
|
|
184
251
|
}
|
|
185
252
|
if (context.status === 'unavailable' || context.status === 'schema_mismatch') {
|
|
253
|
+
await recordEvaluation({ auditStatus: 'error' });
|
|
186
254
|
await observeRuntimeFailure(runtimeErrorObserver, 'auditor_unavailable');
|
|
187
255
|
const failure = projectBackendFailure(context.status === 'unavailable'
|
|
188
256
|
? 'E_AUDITOR_CONTEXT_UNAVAILABLE'
|
|
@@ -204,6 +272,7 @@ export async function runCodexUserPromptSubmitHook({
|
|
|
204
272
|
return;
|
|
205
273
|
}
|
|
206
274
|
if (context.status !== 'fresh') {
|
|
275
|
+
await recordEvaluation({ auditStatus: 'skipped' });
|
|
207
276
|
await recordCodexHookEventSafe(recordHookEventFn, {
|
|
208
277
|
projectRoot,
|
|
209
278
|
event: {
|
|
@@ -236,6 +305,7 @@ export async function runCodexUserPromptSubmitHook({
|
|
|
236
305
|
contextStatus: 'fresh',
|
|
237
306
|
});
|
|
238
307
|
} catch (err) {
|
|
308
|
+
await recordEvaluation({ auditStatus: 'error', backend: err?.backend ?? null, model: err?.diagnostics?.modelSelection?.effectiveModel ?? null });
|
|
239
309
|
if (enteredAuditorBoundary) await observeRuntimeFailure(runtimeErrorObserver, 'auditor_unavailable');
|
|
240
310
|
const failure = projectBackendFailure(err?.code);
|
|
241
311
|
safeWriteError(writeError, failure.stderr);
|
|
@@ -273,10 +343,23 @@ export async function runCodexUserPromptSubmitHook({
|
|
|
273
343
|
},
|
|
274
344
|
}, writeError);
|
|
275
345
|
if (judgment.pass === true) {
|
|
346
|
+
await recordEvaluation({
|
|
347
|
+
auditStatus: 'success',
|
|
348
|
+
backend: judgment.meta?.backend ?? backend.name ?? 'unknown',
|
|
349
|
+
model: judgment.meta?.modelSelection?.effectiveModel ?? null,
|
|
350
|
+
});
|
|
276
351
|
return;
|
|
277
352
|
}
|
|
278
353
|
|
|
279
354
|
const toolIds = projectToolIds(judgment.findings.map((finding) => finding.toolName));
|
|
355
|
+
await recordEvaluation({
|
|
356
|
+
auditStatus: 'success',
|
|
357
|
+
proposedToolIds: toolIds,
|
|
358
|
+
backend: judgment.meta?.backend ?? backend.name ?? 'unknown',
|
|
359
|
+
model: judgment.meta?.modelSelection?.effectiveModel ?? null,
|
|
360
|
+
auditorSeenContext: JSON.stringify(context.turns),
|
|
361
|
+
config: auditorContextConfig,
|
|
362
|
+
});
|
|
280
363
|
const advice = projectParentAdvice(toolIds);
|
|
281
364
|
if (advice) writeCodexUserPromptContexts({ contexts: [advice], writeOutput });
|
|
282
365
|
}
|
|
@@ -290,20 +373,35 @@ export async function runCodexStopHook({
|
|
|
290
373
|
writeOutput = (text) => process.stdout.write(text),
|
|
291
374
|
writeError = (text) => process.stderr.write(text),
|
|
292
375
|
runtimeErrorObserver = async () => ({ collected: false, reason: 'observer_not_configured' }),
|
|
376
|
+
createEvaluationStoreFn = createEvaluationStore,
|
|
377
|
+
codexHome = process.env.CODEX_HOME || join(homedir(), '.codex'),
|
|
378
|
+
now = () => Date.now(),
|
|
293
379
|
} = {}) {
|
|
294
380
|
if (isChildCall()) return;
|
|
295
381
|
const input = await readInput();
|
|
296
|
-
if (input.stop_hook_active === true) return;
|
|
297
382
|
const projectRoot = findSpotterMarker(input.cwd);
|
|
298
383
|
if (!projectRoot) return;
|
|
299
384
|
const startedAt = Date.now();
|
|
300
385
|
const reportError = (text) => safeWriteError(writeError, text);
|
|
386
|
+
if (input.stop_hook_active === true) {
|
|
387
|
+
await closeCodexEvaluationTurn({
|
|
388
|
+
createEvaluationStoreFn,
|
|
389
|
+
sessionId: codexSessionId(input),
|
|
390
|
+
usageStatus: 'incomplete',
|
|
391
|
+
completedAtMs: now(),
|
|
392
|
+
writeError: reportError,
|
|
393
|
+
});
|
|
394
|
+
return;
|
|
395
|
+
}
|
|
301
396
|
const transcriptPath = requireString(input, 'transcript_path');
|
|
302
397
|
const finalResponse = codexLastAssistantMessage(input) ?? '(no final response available)';
|
|
303
398
|
let toolUsage;
|
|
304
399
|
try {
|
|
305
400
|
toolUsage = await readCodexToolUsageFn(transcriptPath);
|
|
306
401
|
} catch (err) {
|
|
402
|
+
await closeCodexEvaluationTurn({
|
|
403
|
+
createEvaluationStoreFn, sessionId: codexSessionId(input), usageStatus: 'incomplete', completedAtMs: now(), writeError: reportError,
|
|
404
|
+
});
|
|
307
405
|
const failure = projectBackendFailure(err?.code);
|
|
308
406
|
reportError(failure.stderr);
|
|
309
407
|
writeCodexSystemMessage({ systemMessage: failure.systemMessage, writeOutput });
|
|
@@ -321,6 +419,15 @@ export async function runCodexStopHook({
|
|
|
321
419
|
return;
|
|
322
420
|
}
|
|
323
421
|
const usedTools = Array.isArray(toolUsage?.usedTools) ? toolUsage.usedTools : [];
|
|
422
|
+
await closeCodexEvaluationTurn({
|
|
423
|
+
createEvaluationStoreFn,
|
|
424
|
+
sessionId: codexSessionId(input),
|
|
425
|
+
toolUsage,
|
|
426
|
+
projectRoot,
|
|
427
|
+
codexHome,
|
|
428
|
+
completedAtMs: now(),
|
|
429
|
+
writeError: reportError,
|
|
430
|
+
});
|
|
324
431
|
const toolUsageEvent = compactCodexToolUsageForEvent(toolUsage);
|
|
325
432
|
if (toolUsageEvent.toolUsageAnomalyCount === 0
|
|
326
433
|
&& shouldSkipShortCodexStop({ finalResponse, usedTools, env: process.env })) {
|
|
@@ -623,6 +730,7 @@ function validateSpotterCodexHookEvent(settings, event) {
|
|
|
623
730
|
const candidates = allEntries.filter((hook) => isSpotterCodexCommand(String(hook?.command ?? '')));
|
|
624
731
|
const expectedCandidates = candidates.filter((hook) => isSpotterCodexCommand(String(hook.command ?? ''), event));
|
|
625
732
|
const expected = expectedCandidates.filter((hook) => hook?.type === 'command');
|
|
733
|
+
const expectedTimeout = event === 'SessionStart' ? CODEX_SESSION_START_TIMEOUT_SEC : CODEX_HOOK_TIMEOUT_SEC;
|
|
626
734
|
const issues = [];
|
|
627
735
|
if (expectedCandidates.length === 0) issues.push('missing');
|
|
628
736
|
if (expectedCandidates.length > 1) issues.push('duplicate');
|
|
@@ -632,6 +740,7 @@ function validateSpotterCodexHookEvent(settings, event) {
|
|
|
632
740
|
if (hook.async === true) issues.push('async:true');
|
|
633
741
|
if (!Object.hasOwn(hook, 'timeout')) issues.push('timeout:missing');
|
|
634
742
|
else if (!Number.isFinite(hook.timeout) || hook.timeout <= 0) issues.push('timeout-invalid');
|
|
743
|
+
else if (hook.timeout !== expectedTimeout) issues.push(`timeout!=${expectedTimeout}`);
|
|
635
744
|
if (Object.hasOwn(hook, 'timeoutSec')) issues.push('timeoutSec');
|
|
636
745
|
if (hook.async === false) issues.push('async:false');
|
|
637
746
|
if (hook.statusMessage === null) issues.push('statusMessage:null');
|
|
@@ -658,6 +767,8 @@ function isIncompatibleSpotterHookIssue(issue) {
|
|
|
658
767
|
'type!=command',
|
|
659
768
|
'async:true',
|
|
660
769
|
'timeout-invalid',
|
|
770
|
+
'timeout!=30',
|
|
771
|
+
'timeout!=60',
|
|
661
772
|
'timeoutSec',
|
|
662
773
|
'commandWindows-invalid',
|
|
663
774
|
].includes(issue);
|
|
@@ -772,6 +883,79 @@ function safeWriteError(writeError, text) {
|
|
|
772
883
|
}
|
|
773
884
|
}
|
|
774
885
|
|
|
886
|
+
function reportEvaluationFailure(writeError) {
|
|
887
|
+
safeWriteError(writeError, 'Spotter の評価記録に失敗しました。\n');
|
|
888
|
+
}
|
|
889
|
+
|
|
890
|
+
function isCompleteCodexUsage(toolUsage, canonicalUsage) {
|
|
891
|
+
if (!toolUsage || !Array.isArray(toolUsage.usedTools) || !Array.isArray(toolUsage.toolCalls)) return false;
|
|
892
|
+
if (toolUsage.scope !== 'current-turn') return false;
|
|
893
|
+
if (Array.isArray(toolUsage.anomalies) && toolUsage.anomalies.length > 0) return false;
|
|
894
|
+
return canonicalUsage.missingCount === 0;
|
|
895
|
+
}
|
|
896
|
+
|
|
897
|
+
function isCodexEvaluationUsage(toolName) {
|
|
898
|
+
return typeof toolName === 'string' && (
|
|
899
|
+
/^mcp__[A-Za-z0-9_-]+__/u.test(toolName)
|
|
900
|
+
|| toolName === 'Skill'
|
|
901
|
+
|| toolName === 'skill'
|
|
902
|
+
|| /^(?:skills?|Skill)(?:__|\.)/u.test(toolName)
|
|
903
|
+
);
|
|
904
|
+
}
|
|
905
|
+
|
|
906
|
+
async function closeCodexEvaluationTurn({
|
|
907
|
+
createEvaluationStoreFn,
|
|
908
|
+
sessionId,
|
|
909
|
+
usedToolIds = [],
|
|
910
|
+
usageStatus,
|
|
911
|
+
toolUsage = null,
|
|
912
|
+
projectRoot = null,
|
|
913
|
+
codexHome = process.env.CODEX_HOME || join(homedir(), '.codex'),
|
|
914
|
+
completedAtMs,
|
|
915
|
+
writeError,
|
|
916
|
+
}) {
|
|
917
|
+
if (!sessionId) {
|
|
918
|
+
reportEvaluationFailure(writeError);
|
|
919
|
+
return;
|
|
920
|
+
}
|
|
921
|
+
try {
|
|
922
|
+
const store = createEvaluationStoreFn();
|
|
923
|
+
try {
|
|
924
|
+
const row = store.database.prepare(`
|
|
925
|
+
SELECT observation_id FROM evaluation_turns
|
|
926
|
+
WHERE session_id = ? AND completed_at_ms IS NULL
|
|
927
|
+
ORDER BY recorded_at_ms DESC LIMIT 1
|
|
928
|
+
`).get(sessionId);
|
|
929
|
+
if (!row) return;
|
|
930
|
+
if (toolUsage) {
|
|
931
|
+
const proposals = store.database.prepare(`
|
|
932
|
+
SELECT tool_id FROM evaluation_items WHERE observation_id = ? ORDER BY tool_id
|
|
933
|
+
`).all(row.observation_id).map((item) => item.tool_id);
|
|
934
|
+
const allUsages = toolUsage.toolCalls ?? toolUsage.usedTools.map((toolName) => ({ toolName }));
|
|
935
|
+
const evaluationUsages = allUsages.filter((usage) => isCodexEvaluationUsage(usage?.toolName));
|
|
936
|
+
const canonicalUsage = canonicalizeUsedToolIds(evaluationUsages, { host: 'codex' });
|
|
937
|
+
const skillReads = await canonicalizeCodexSkillReadToolIds(allUsages, {
|
|
938
|
+
proposedToolIds: proposals,
|
|
939
|
+
projectRoot,
|
|
940
|
+
codexHome,
|
|
941
|
+
});
|
|
942
|
+
const nestedMcpCalls = canonicalizeCodexNestedMcpToolIds(allUsages);
|
|
943
|
+
usedToolIds = [...new Set([
|
|
944
|
+
...canonicalUsage.resolvedToolIds,
|
|
945
|
+
...nestedMcpCalls,
|
|
946
|
+
...skillReads,
|
|
947
|
+
])];
|
|
948
|
+
usageStatus = isCompleteCodexUsage(toolUsage, canonicalUsage) ? 'complete' : 'incomplete';
|
|
949
|
+
}
|
|
950
|
+
store.closeTurn({ observationId: row.observation_id, usedToolIds, usageStatus, completedAtMs });
|
|
951
|
+
} finally {
|
|
952
|
+
store.close();
|
|
953
|
+
}
|
|
954
|
+
} catch {
|
|
955
|
+
reportEvaluationFailure(writeError);
|
|
956
|
+
}
|
|
957
|
+
}
|
|
958
|
+
|
|
775
959
|
// Phase D (hook parity, 2026-05-08): Codex `--project` diagnostics now read the host-neutral
|
|
776
960
|
// `<projectRoot>/.spotter/hook-events.jsonl` and filter to `host:"codex"` so the existing
|
|
777
961
|
// `codex-hook diagnostics` shape (counts of just Codex events) stays intact.
|
|
@@ -884,7 +1068,7 @@ function mergeCodexHooks(current, { nodePath, spotterBin, platform = process.pla
|
|
|
884
1068
|
next.hooks = next.hooks ?? {};
|
|
885
1069
|
const prefix = platform === 'win32' ? '& ' : '';
|
|
886
1070
|
addCodexHook(next, 'SessionStart', `${prefix}${quoteArg(nodePath)} ${quoteArg(spotterBin)} codex-hook session-start`, {
|
|
887
|
-
timeout:
|
|
1071
|
+
timeout: CODEX_SESSION_START_TIMEOUT_SEC,
|
|
888
1072
|
});
|
|
889
1073
|
addCodexHook(next, 'UserPromptSubmit', `${prefix}${quoteArg(nodePath)} ${quoteArg(spotterBin)} codex-hook user-prompt-submit`);
|
|
890
1074
|
addCodexHook(next, 'Stop', `${prefix}${quoteArg(nodePath)} ${quoteArg(spotterBin)} codex-hook stop`);
|
|
@@ -59,7 +59,7 @@ export async function runDiagnosticsLogsCommand({
|
|
|
59
59
|
writeOutput(formatDaemonLogSummary(merged));
|
|
60
60
|
}
|
|
61
61
|
|
|
62
|
-
function stringifyAsciiJson(value) {
|
|
62
|
+
export function stringifyAsciiJson(value) {
|
|
63
63
|
return JSON.stringify(value, null, 2).replace(/[\u007f-\uffff]/g, (character) =>
|
|
64
64
|
`\\u${character.charCodeAt(0).toString(16).padStart(4, '0')}`
|
|
65
65
|
);
|