kairos-chain 3.87.0 → 3.88.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +71 -0
- data/lib/kairos_mcp/version.rb +1 -1
- data/templates/knowledge/multi_llm_review_workflow/multi_llm_review_workflow.md +76 -31
- data/templates/knowledge/multi_llm_reviewer_evaluation/multi_llm_reviewer_evaluation.md +23 -1
- data/templates/skillsets/account_manager/plugin/agents/bookkeeper.md +2 -0
- data/templates/skillsets/agent/plugin/SKILL.md +1 -1
- data/templates/skillsets/agent/plugin/agents/monitor.md +2 -1
- data/templates/skillsets/multi_llm_review/config/multi_llm_review.yml +10 -8
- data/templates/skillsets/multi_llm_review/skillset.json +2 -2
- data/templates/skillsets/multi_llm_review/test/test_multi_llm_review.rb +70 -0
- data/templates/skillsets/multi_llm_review/tools/multi_llm_review.rb +76 -3
- data/templates/skillsets/project_manager/plugin/agents/secretary.md +2 -0
- data/templates/skillsets/skillset_exchange/plugin/agents/reviewer.md +2 -1
- metadata +1 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 6ab07b57ae1574d18326409b3db66f61dabd697c9bc8d52e6ba20b00ee078bb6
|
|
4
|
+
data.tar.gz: ad59a3b0ddaca3913faffcd102cf7a361ad4772b90dcd0950c703884f6a8d0b9
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: bed9eb85e62df8c52384e0ead71655d1c44b375abd79cf4dc85797e910644ce8360743073572d12edd52d515c1bae9eecdf524c22811a44c444cca0766839a1c
|
|
7
|
+
data.tar.gz: 39d2c4d7b801d8cfb9f73e3f2075e87178d7ea2d2bbbb57dd967b24dbf69f72a3812a4f9782c793bebd7a27525793f144226fb7a00ae6712851ef276ceefee6e
|
data/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,77 @@ All notable changes to the `kairos-chain` gem will be documented in this file.
|
|
|
4
4
|
|
|
5
5
|
This project follows [Semantic Versioning](https://semver.org/).
|
|
6
6
|
|
|
7
|
+
## [3.88.0] - 2026-09-30
|
|
8
|
+
|
|
9
|
+
### Added — `multi_llm_review` 0.11.1: an artifact named by path and sha256
|
|
10
|
+
|
|
11
|
+
`multi_llm_review` takes an optional `artifact_sha256`, and `artifact_content`
|
|
12
|
+
is no longer required. When `artifact_content` is empty, the tool reads
|
|
13
|
+
`artifact_path` inside its working directory and refuses unless the bytes hash
|
|
14
|
+
to the pin; a path outside the working directory, a file that is not a regular
|
|
15
|
+
file, a file that changes between the check and the read, and non-UTF-8 content
|
|
16
|
+
are refused, all before any seat is dispatched. When `artifact_content` is
|
|
17
|
+
carried together with a pin, the content must hash to the pin. Callers that pass
|
|
18
|
+
`artifact_content` without a pin are unaffected.
|
|
19
|
+
|
|
20
|
+
Why: the agent SkillSet's DECIDE phase writes tool arguments itself, within an
|
|
21
|
+
output budget, and `autoexec_run` passes them as written. On 2026-09-30 a
|
|
22
|
+
101,549-byte artifact reached the plan as a placeholder string, and approving
|
|
23
|
+
the plan would have sent the placeholder to every seat. With the pin, the plan
|
|
24
|
+
names the file and its hash, and the plan's own hash binds the content.
|
|
25
|
+
|
|
26
|
+
Claude Code keeps a connected server's tool schema until the session restarts,
|
|
27
|
+
so a client started before the upgrade still sees `artifact_content` as required.
|
|
28
|
+
|
|
29
|
+
### Changed — Opus 4.6 retired from the review roster; Sonnet 5.5 takes its seat
|
|
30
|
+
|
|
31
|
+
Operator decision, 2026-09-30: Opus 4.6 is retired from every role it held —
|
|
32
|
+
the Claude CLI review seat and the sub-author role for self-referential
|
|
33
|
+
passages — and Sonnet 5.5 at effort high takes both. The roster entry becomes
|
|
34
|
+
`claude-sonnet-5-5` / `claude_cli_sonnet5.5`; the roster is still four seats.
|
|
35
|
+
|
|
36
|
+
This is the operator's decision, not a like-for-like succession. Opus 4.6 held
|
|
37
|
+
both roles for a documented bias profile (ambiguity-preserving,
|
|
38
|
+
self-reference-friendly); Sonnet 5.5 is uncalibrated, with two review rounds on
|
|
39
|
+
2026-09-30 and nothing else, so no bias rationale is written for it.
|
|
40
|
+
|
|
41
|
+
- L1 `multi_llm_review_workflow` 3.15.0 → 3.15.1 and
|
|
42
|
+
`multi_llm_reviewer_evaluation` 1.6 → 1.7: text a reader follows as current
|
|
43
|
+
procedure now names Sonnet 5.5. Measurements, seat profiles, dated
|
|
44
|
+
diagnoses and changelog entries that name Opus 4.6 stay as written, because
|
|
45
|
+
they are facts about that model. The dev-repo copies of these two entries and
|
|
46
|
+
of `design_to_implementation_workflow` had drifted behind the shipped ones
|
|
47
|
+
(1.5 and 1.1 against 1.6 and 1.2) and are synced.
|
|
48
|
+
- The agent SkillSet's plugin `SKILL.md` names Sonnet 5.5 as the sub-author.
|
|
49
|
+
|
|
50
|
+
Upgrading: `kairos-chain upgrade` never overwrites an instance's
|
|
51
|
+
`config/multi_llm_review.yml`, so an instance keeps Opus 4.6 until its own
|
|
52
|
+
roster entry is edited. The config's comment on the entry says how.
|
|
53
|
+
|
|
54
|
+
## [3.87.1] - 2026-09-24
|
|
55
|
+
|
|
56
|
+
### Changed — shipped subagents name their model and effort in full
|
|
57
|
+
|
|
58
|
+
The four subagents shipped as plugin artifacts — `bookkeeper` (account_manager),
|
|
59
|
+
`secretary` (project_manager), `monitor` (agent) and `reviewer`
|
|
60
|
+
(skillset_exchange) — now pin `model: claude-opus-5-5` and `effort: high`.
|
|
61
|
+
Before, two had no `model` line (they ran on the main conversation's model) and
|
|
62
|
+
two said `sonnet`, and none set `effort`, so they ran at the session's effort.
|
|
63
|
+
|
|
64
|
+
Why: an alias floats and differs by provider — per the Claude Code model-config
|
|
65
|
+
docs, `sonnet` is Sonnet 5 on the Anthropic API but Sonnet 4.5 on Bedrock, Vertex
|
|
66
|
+
and Foundry — and an instance that pinned full IDs locally was reverted to the
|
|
67
|
+
aliases by `kairos-chain upgrade` (same-version content overwrite, 2026-09-24).
|
|
68
|
+
With the template and the instance identical, the upgrade has nothing to revert.
|
|
69
|
+
The cost is a manual edit when a newer model ships; Claude Code warns when a
|
|
70
|
+
requested model has a scheduled retirement date or has been remapped.
|
|
71
|
+
|
|
72
|
+
On Bedrock, Vertex or Foundry, model IDs are provider-specific: map this ID
|
|
73
|
+
to your provider's ID with the `modelOverrides` setting. To run these agents on
|
|
74
|
+
another model or effort, edit the instance copy under
|
|
75
|
+
`.kairos/skillsets/<name>/plugin/agents/` and re-project; `upgrade` overwrites that
|
|
76
|
+
copy, so the edit has to be repeated after each upgrade.
|
|
77
|
+
|
|
7
78
|
## [3.87.0] - 2026-09-24
|
|
8
79
|
|
|
9
80
|
### Added — `model_provenance` SkillSet: which model actually answered
|
data/lib/kairos_mcp/version.rb
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: multi_llm_review_workflow
|
|
3
3
|
description: "Multi-LLM review methodology and execution — workflow pattern, CLI tooling, consensus analysis, Persona Assembly. Applicable to design, implementation, documentation, or any artifact."
|
|
4
|
-
version: "3.15.
|
|
4
|
+
version: "3.15.1"
|
|
5
5
|
tags:
|
|
6
6
|
- workflow
|
|
7
7
|
- review
|
|
@@ -141,7 +141,7 @@ these numbers.
|
|
|
141
141
|
|
|
142
142
|
| Seat | Runs | APPROVE rate design / impl / doc | Median wall s | Median output chars | (c) share, labelled only |
|
|
143
143
|
|---|---|---|---|---|---|
|
|
144
|
-
| `claude_cli_opus4.6` | 134 | 64% (31/48) / 71% (27/38) / 86% (20/23) | 54 | 4,536 | 70% (19/27) |
|
|
144
|
+
| `claude_cli_opus4.6` (retired 2026-09-30) | 134 | 64% (31/48) / 71% (27/38) / 86% (20/23) | 54 | 4,536 | 70% (19/27) |
|
|
145
145
|
| `cursor_composer2.5` | 138 | 13% (7/51) / 48% (24/50) / 25% (6/24) | 133 | 3,795 | 100% (2/2) |
|
|
146
146
|
| `codex_gpt5.6-sol` (retired 2026-09-05) | 138 | 1% (1/54) / 11% (7/60) / 4% (1/24) | 113 | 2,038 | 3% (5/140) |
|
|
147
147
|
| `claude_team_opus-5` (persona) | 131 | 0% (0/51) / 12% (7/56) / 0% (0/22) | not measured per seat | 11,512 | 27% (374/1362) |
|
|
@@ -155,16 +155,24 @@ every row from 2026-09-05 onward for a second reason — the corpus was gathered
|
|
|
155
155
|
with reviewers at medium effort, and reviewers now run at high (see § Thinking
|
|
156
156
|
Effort Configuration), so post-swap rounds are not directly comparable to it.
|
|
157
157
|
|
|
158
|
+
The opus4.6 row is likewise history. Opus 4.6 was retired on 2026-09-30
|
|
159
|
+
(operator decision) and `claude_cli_sonnet5.5` — Sonnet 5.5 at effort high —
|
|
160
|
+
took the seat. Sonnet 5.5 had held it for two rounds by then and has no
|
|
161
|
+
profile, so nothing in the opus4.6 row, including its lenient approval rate,
|
|
162
|
+
may be read as a description of the new occupant.
|
|
163
|
+
|
|
158
164
|
Selection consequences, each tied to the number above it:
|
|
159
165
|
|
|
160
|
-
- **
|
|
161
|
-
least.** Codex
|
|
162
|
-
advisory rate in the corpus (5 of 140 labelled
|
|
163
|
-
document reviews in 23 while 19 of its 27
|
|
164
|
-
|
|
165
|
-
with opus4.6's vote already inside it
|
|
166
|
-
|
|
167
|
-
|
|
166
|
+
- **In this corpus a `codex` APPROVE carried the most information and an
|
|
167
|
+
`opus4.6` APPROVE the least.** Codex approved 1 design review in 54 and paired
|
|
168
|
+
that with the lowest advisory rate in the corpus (5 of 140 labelled
|
|
169
|
+
findings). Opus4.6 approved 20 document reviews in 23 while 19 of its 27
|
|
170
|
+
labelled findings were advisory, so the `3/4 APPROVE` threshold was met, in
|
|
171
|
+
practice, with opus4.6's vote already inside it. Both seats have since
|
|
172
|
+
changed occupant (codex 2026-09-05, opus4.6 2026-09-30), and neither
|
|
173
|
+
successor is calibrated: until they are, do not assume which seat casts the
|
|
174
|
+
lenient vote — read each APPROVE for what it says. (That threshold is a
|
|
175
|
+
reference figure, not a gate — see § Convergence Rules.)
|
|
168
176
|
- **Volume anti-correlates with signal.** The persona seat raises 3,249 of the
|
|
169
177
|
5,286 findings (61%) and 27% of its labelled ones are advisory. Seat personas
|
|
170
178
|
when breadth is wanted; do not seat them to obtain a verdict.
|
|
@@ -497,10 +505,11 @@ they disagree, the config is right and this section is stale.
|
|
|
497
505
|
- [ ] Agent Team Personas model: = orchestrator model (NOT a different model),
|
|
498
506
|
unless you are running personas on a different model on purpose — then
|
|
499
507
|
it is whatever you declare as persona_model (see § Persona execution model)
|
|
500
|
-
- [ ] Subprocess CLI model: Opus 4.6
|
|
501
|
-
|
|
502
|
-
|
|
503
|
-
|
|
508
|
+
- [ ] Subprocess CLI model: Sonnet 5.5 (`claude-sonnet-5-5`; Opus 4.6 held this
|
|
509
|
+
slot until its retirement on 2026-09-30). The other Claude roster slot is
|
|
510
|
+
Opus 5.5, which under the default "delegate" strategy is taken by your
|
|
511
|
+
persona team rather than spawned — so when you are Opus 5.5, Sonnet 5.5 is
|
|
512
|
+
the only Claude CLI subprocess
|
|
504
513
|
- [ ] Codex model: gpt-6-astra, with -m. One codex slot since gpt-5.5 was
|
|
505
514
|
retired 2026-09-05 — do not add a second codex entry expecting the old
|
|
506
515
|
cross-generation pairing
|
|
@@ -853,8 +862,8 @@ outside this repository — see the incident recorded in § Pre-flight checklist
|
|
|
853
862
|
| **Codex** | `codex exec -m <model> -c model_reasoning_effort=high` | stdin pipe: `cat prompt.md \| codex exec -m <model> -` | `-o /path/output.md` | gpt-6-astra — one slot since gpt-5.5 was retired 2026-09-05 |
|
|
854
863
|
| **Cursor Agent** | `agent -p --model composer-2.5` | File reference (stdin NOT supported) | stdout redirect: `> output.md` | composer-2.5, passed explicitly — never relying on the CLI default |
|
|
855
864
|
| **Claude Code** | Agent tool (internal) | Direct prompt string | Write to workspace file | Orchestrator model, or the declared `persona_model` when personas run elsewhere |
|
|
856
|
-
| **Claude CLI (
|
|
857
|
-
| **Claude CLI (frontier)** | `claude -p --model claude-opus-5-5` | stdin pipe: `cat prompt.md \| claude -p --model claude-opus-5-5` | stdout redirect: `> output.md` | Runs only when Opus 5.5 is *not* the orchestrator; when it is, that slot is taken by the persona team |
|
|
865
|
+
| **Claude CLI (Sonnet 5.5)** | `claude -p --model claude-sonnet-5-5 --effort high` | stdin pipe: `cat prompt.md \| claude -p --model claude-sonnet-5-5 --effort high` | stdout redirect: `> output.md` | Sonnet 5.5 — took Opus 4.6's slot on 2026-09-30 by operator decision; uncalibrated |
|
|
866
|
+
| **Claude CLI (frontier)** | `claude -p --model claude-opus-5-5 --effort high` | stdin pipe: `cat prompt.md \| claude -p --model claude-opus-5-5 --effort high` | stdout redirect: `> output.md` | Runs only when Opus 5.5 is *not* the orchestrator; when it is, that slot is taken by the persona team |
|
|
858
867
|
|
|
859
868
|
`--bare` must NOT be passed (established 2026-07-23): it skips credential
|
|
860
869
|
loading and the subprocess fails with "Not logged in". The project-instruction
|
|
@@ -874,7 +883,7 @@ Based on cross-evaluation experiment (7 models × 4 tasks + Nomic, 518 CLI calls
|
|
|
874
883
|
|------|-------|-------------|-----------|
|
|
875
884
|
| **Primary (orchestrator)** | session default | Claude Code's own setting | Not set by this config — record the level actually in effect (see the 2026-09-24 note) |
|
|
876
885
|
| **Reviewer: Agent Team** | = orchestrator, or the declared `persona_model` | the session's level, unless the persona's agent file sets `effort:` | Personas inherit whichever model — and effort — actually runs them |
|
|
877
|
-
| **Reviewer: Claude CLI** |
|
|
886
|
+
| **Reviewer: Claude CLI** | Sonnet 5.5, plus any frontier roster slot the orchestrator is not | `--effort high` (config `effort: high`) | Operator instruction 2026-09-05; supersedes the 2026-04-29 default-effort policy — see the note below the table |
|
|
878
887
|
| **Coding sub-agent** | Opus 5.5 | `--effort high` | Operator default for KairosChain (2026-09-24); not measured here (see note) |
|
|
879
888
|
| **Design sub-agent** | Opus 5.5 | `--effort high` | Operator default for KairosChain (2026-09-24); not measured here (see note) |
|
|
880
889
|
| **Codex** | GPT-6-astra / GPT-5.5 | `-c model_reasoning_effort=high` | Same operator instruction. The earlier "(no flag) / fixed effort" entry was wrong: codex_adapter has always emitted this flag when the roster set `effort` |
|
|
@@ -904,9 +913,19 @@ Note (2026-07-26): the coding / design sub-agent rows moved from Opus 4.7
|
|
|
904
913
|
(retired 2026-06-10) to Opus 5. Their effort values are Anthropic's published
|
|
905
914
|
starting points for Opus 5, not measurements from this project — treat them as
|
|
906
915
|
a starting point to sweep down from, not a validated setting. Separately, the
|
|
907
|
-
**sub-author** role for self-referential passages
|
|
916
|
+
**sub-author** role for self-referential passages stayed Opus 4.6: it was chosen
|
|
908
917
|
for its ambiguity-preserving bias, not for capability, so a frontier successor
|
|
909
|
-
|
|
918
|
+
did not replace it. (Superseded 2026-09-30 — see the next note.)
|
|
919
|
+
|
|
920
|
+
Note (2026-09-30): Opus 4.6 is retired everywhere, the **sub-author** role
|
|
921
|
+
included, by operator decision; Sonnet 5.5 (`claude-sonnet-5-5`, effort high)
|
|
922
|
+
takes it. This is not a like-for-like succession. The reason Opus 4.6 held the
|
|
923
|
+
role was a documented bias profile, and Sonnet 5.5 has none — it had held the
|
|
924
|
+
review seat for two rounds when the decision was made, and nothing else. So the
|
|
925
|
+
sub-author choice rests on the operator's decision alone: do not cite an
|
|
926
|
+
ambiguity-preserving or other bias rationale for Sonnet 5.5 until a measurement
|
|
927
|
+
supports one, and record its sub-author outputs so a profile can accumulate in
|
|
928
|
+
`multi_llm_reviewer_evaluation`.
|
|
910
929
|
|
|
911
930
|
Note (2026-09-24): the frontier rows moved from Opus 5 to Opus 5.5 (operator
|
|
912
931
|
instruction — 5.5 is now Claude Code's default orchestrator and takes over
|
|
@@ -960,11 +979,12 @@ record no longer claims a model that did not answer.
|
|
|
960
979
|
orchestrating LLM MUST pass its own model identifier as `orchestrator_model`.
|
|
961
980
|
|
|
962
981
|
**Rationale**: The reviewer roster contains more than one Claude entry (Opus 5.5
|
|
963
|
-
and
|
|
982
|
+
and Sonnet 5.5 as of 2026-09-30; Opus 4.6 held the second entry until then, and
|
|
983
|
+
Opus 5 held the frontier entry from 2026-07-26). A frontier slot sits in the roster on purpose:
|
|
964
984
|
when it is the current session model it matches `orchestrator_model` and becomes
|
|
965
985
|
the persona-team slot; when it is not, it is dispatched as a `claude -p`
|
|
966
986
|
subprocess. With the current two-entry Claude side, an Opus 5.5 orchestrator leaves
|
|
967
|
-
|
|
987
|
+
Sonnet 5.5 as the only Claude CLI subprocess. No per-orchestrator config branch is
|
|
968
988
|
needed either way.
|
|
969
989
|
To avoid the orchestrator reviewing its
|
|
970
990
|
own output (no independent signal), the dispatcher excludes or delegates the
|
|
@@ -1025,7 +1045,7 @@ multi_llm_review(
|
|
|
1025
1045
|
- If `orchestrator_model` is `null` or unmatched, full roster runs (back-compat).
|
|
1026
1046
|
|
|
1027
1047
|
**Manual-mode equivalent**: When orchestrating by hand, do not assign yourself
|
|
1028
|
-
as a subprocess reviewer. Run the Claude CLI subprocess reviewers (
|
|
1048
|
+
as a subprocess reviewer. Run the Claude CLI subprocess reviewers (Sonnet 5.5, plus
|
|
1029
1049
|
any frontier roster slot you are not); if your own model matches one of them,
|
|
1030
1050
|
skip that entry and use the after-exclusion convergence rule.
|
|
1031
1051
|
|
|
@@ -1040,7 +1060,7 @@ artifact, use `escalate: true` — see § Reserve observers below.
|
|
|
1040
1060
|
The `delegate` strategy lets the orchestrator perform persona-based "Agent Team"
|
|
1041
1061
|
review in its own context — preserving inherited project context that a fresh
|
|
1042
1062
|
`claude -p` subprocess loses. Subprocess reviewers (codex, cursor, Claude CLI
|
|
1043
|
-
|
|
1063
|
+
Sonnet 5.5 and the non-orchestrator frontier model) remain single-LLM.
|
|
1044
1064
|
|
|
1045
1065
|
**Why**: The orchestrator already holds the artifact in context with full project
|
|
1046
1066
|
awareness. Re-shipping it to a sandboxed subprocess discards that context. Same-
|
|
@@ -1349,8 +1369,8 @@ readable until GC. Read them directly and synthesize manually, then re-run
|
|
|
1349
1369
|
before, silently swapping the model behind an unchanged role label
|
|
1350
1370
|
- **Codex workspace**: `-C /path/to/workspace` to set working directory
|
|
1351
1371
|
- **Claude Agent paths**: Write within workspace (e.g., `log/`), not `/tmp`
|
|
1352
|
-
- **Claude CLI (
|
|
1353
|
-
- **Claude CLI parallelism**: Agent tool (internal, orchestrator model) + Bash `claude -p` (external,
|
|
1372
|
+
- **Claude CLI (Sonnet 5.5 / any non-orchestrator frontier slot)**: `claude -p --model claude-sonnet-5-5` (likewise `claude-opus-5-5`) runs as external process. Uses stdin pipe (like Codex). Do NOT pass `--bare` — it skips credential loading and the subprocess dies with "Not logged in" (established 2026-07-23). Project-instruction bias is suppressed via `review_context: independent`, not via `--bare`
|
|
1373
|
+
- **Claude CLI parallelism**: Agent tool (internal, orchestrator model) + Bash `claude -p` (external, Sonnet 5.5 and any other Claude roster slot the orchestrator is not) run truly in parallel as separate processes
|
|
1354
1374
|
- **Claude CLI file access**: a plain `claude -p` review subprocess should not need file access. Ensure the review prompt includes all artifact content inline (rule #6). Use `--add-dir` + `--allowedTools "Read,Glob,Grep"` if file access is genuinely needed, and accept that CLAUDE.md is loaded (the old `--bare` workaround is unusable — see above)
|
|
1355
1375
|
|
|
1356
1376
|
## Prompt Generation Rules
|
|
@@ -1494,7 +1514,7 @@ Step 2: Detect environment, and check the roster against config
|
|
|
1494
1514
|
and treat them as the roster. Detection only tells you whether a default has
|
|
1495
1515
|
drifted; the model each slot runs is named on the command line.
|
|
1496
1516
|
- Report: "Auto mode: Codex (gpt-6-astra), Cursor (composer-2.5),
|
|
1497
|
-
Claude Team (orchestrator model), Claude CLI (
|
|
1517
|
+
Claude Team (orchestrator model), Claude CLI (sonnet-5.5)"
|
|
1498
1518
|
|
|
1499
1519
|
Step 3: Execute the configured roster in parallel (currently 4 slots, one of
|
|
1500
1520
|
which is your own persona team)
|
|
@@ -1502,9 +1522,9 @@ Step 3: Execute the configured roster in parallel (currently 4 slots, one of
|
|
|
1502
1522
|
- Bash(background): agent -p --trust --model composer-2.5 "Read prompt and review..." > log/review_cursor.md
|
|
1503
1523
|
(no effort flag — Cursor has no effort control)
|
|
1504
1524
|
- Agent(background): Claude Team (orchestrator model, e.g. Opus 5.5) → write to log/review_claude_team_opus5.5.md
|
|
1505
|
-
- Bash(background): cat prompt.md | claude -p --model claude-
|
|
1506
|
-
(add a line per further Claude roster slot you are not; with the 2026-09-
|
|
1507
|
-
roster an Opus 5.5 orchestrator has none, so
|
|
1525
|
+
- Bash(background): cat prompt.md | claude -p --model claude-sonnet-5-5 --effort high > log/review_claude_sonnet5.5.md 2>log/review_claude_sonnet5.5.stderr.log
|
|
1526
|
+
(add a line per further Claude roster slot you are not; with the 2026-09-30
|
|
1527
|
+
roster an Opus 5.5 orchestrator has none, so sonnet-5.5 is the only one)
|
|
1508
1528
|
|
|
1509
1529
|
Step 4: Collect and validate
|
|
1510
1530
|
- Wait for all to complete (background task notifications)
|
|
@@ -1541,7 +1561,7 @@ log/{artifact}_review{N}_{llm_id}_{date}.md # Individual reviews
|
|
|
1541
1561
|
log/{artifact}_review{N}_consensus_{date}.md # Consensus analysis
|
|
1542
1562
|
```
|
|
1543
1563
|
|
|
1544
|
-
LLM identifiers: `claude_cli_opus5.5`, `
|
|
1564
|
+
LLM identifiers: `claude_cli_opus5.5`, `claude_cli_sonnet5.5`,
|
|
1545
1565
|
`codex_gpt6-astra`, `cursor_composer2.5`, `cursor_gpt5.4`,
|
|
1546
1566
|
`cursor_premium`. The delegated slot is reported as `claude_team_<model>`
|
|
1547
1567
|
(e.g. `claude_team_claude-opus-5-5`), assembled at collect time — the roster's
|
|
@@ -1553,7 +1573,8 @@ retired 2026-07-26: `claude_cli_fable5` — five consecutive non-substantive
|
|
|
1553
1573
|
returns, 85-128 characters in 5-7 seconds, no findings and no verdict text;
|
|
1554
1574
|
retired 2026-09-05: `codex_gpt5.6-sol`, replaced by `codex_gpt6-astra`, and
|
|
1555
1575
|
`codex_gpt5.5`, not replaced; retired 2026-09-24: `claude_cli_opus5`, succeeded in
|
|
1556
|
-
the same slot by `claude_cli_opus5.5
|
|
1576
|
+
the same slot by `claude_cli_opus5.5`; retired 2026-09-30: `claude_cli_opus4.6`,
|
|
1577
|
+
succeeded in the same slot by `claude_cli_sonnet5.5`. Runs recorded under a retired identifier keep it —
|
|
1557
1578
|
the label names the model that answered, so renaming old records would attribute
|
|
1558
1579
|
one model's findings to another)
|
|
1559
1580
|
|
|
@@ -1948,6 +1969,30 @@ Compression ratio: parallel agent raw → Assembly ≈ 2:1
|
|
|
1948
1969
|
§ Thinking Effort Configuration, including that the persona seat follows the
|
|
1949
1970
|
session's level while CLI seats follow the config.
|
|
1950
1971
|
|
|
1972
|
+
- Opus 4.6 retired → Sonnet 5.5 (v3.15.1, 2026-09-30, operator decision:
|
|
1973
|
+
「全部Opus4.6は引退させて、Sonnet5.5に置き換えて下さい」). Opus 4.6 leaves every
|
|
1974
|
+
role it held — the Claude CLI review seat and the sub-author role — and
|
|
1975
|
+
Sonnet 5.5 at effort high takes both. The seat's label becomes
|
|
1976
|
+
`claude_cli_sonnet5.5`; nothing else about the roster changes (still 4 seats:
|
|
1977
|
+
Opus 5.5 as orchestrator and persona team, Sonnet 5.5 CLI, codex gpt-6-astra,
|
|
1978
|
+
cursor composer-2.5). Only text that prescribes current behaviour was
|
|
1979
|
+
rewritten; two dated passages were annotated rather than rewritten — the
|
|
1980
|
+
Step 0.1 row label gains "(retired 2026-09-30)" with its numbers untouched,
|
|
1981
|
+
and the 2026-07-26 sub-author note is put in the past tense and pointed at
|
|
1982
|
+
the note that supersedes it. The two Claude CLI rows of the Tool Matrix
|
|
1983
|
+
also gain `--effort high`, which they had lacked while the table below them
|
|
1984
|
+
required it. The `claude_cli_opus4.6` row in § Step 0.1, the effort measurements
|
|
1985
|
+
taken on the Opus 4.6 / 4.7 generation, the 2026-08-06 no_verdict diagnosis
|
|
1986
|
+
and every earlier changelog entry stay as written, because they are facts
|
|
1987
|
+
about Opus 4.6. Recorded as the operator's decision rather than as a
|
|
1988
|
+
succession: Opus 4.6 was kept in both roles for a documented bias profile
|
|
1989
|
+
(ambiguity-preserving, self-reference-friendly), while Sonnet 5.5 is
|
|
1990
|
+
uncalibrated — two review rounds on 2026-09-30 and nothing else — so no bias
|
|
1991
|
+
rationale is written for it. Config: `multi_llm_review` 0.11.1 roster entry
|
|
1992
|
+
`claude-sonnet-5-5`; the entry's comment says how to restore Opus 4.6.
|
|
1993
|
+
Upgrading does not rewrite an instance's own `config/multi_llm_review.yml`,
|
|
1994
|
+
so an existing instance keeps Opus 4.6 until that entry is edited by hand.
|
|
1995
|
+
|
|
1951
1996
|
**Key insight**: Design reviews and implementation reviews find
|
|
1952
1997
|
**categorically different bugs**. Both phases are necessary. The corollary that
|
|
1953
1998
|
cost eight rounds to learn: **which of the two you are in is readable from the
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: multi_llm_reviewer_evaluation
|
|
3
3
|
description: "Multi-LLM reviewer performance evaluation — strengths, weaknesses, value-system biases, and recommended workflows. Based on 185+ reviews (Phase 1, 2026-02 to 03) + Phase 2 Case A 4-round Codex bias study (2026-05-04)."
|
|
4
|
-
version: "1.
|
|
4
|
+
version: "1.7"
|
|
5
5
|
tags:
|
|
6
6
|
- multi-llm
|
|
7
7
|
- review
|
|
@@ -47,6 +47,11 @@ Based on 185+ review files across KairosChain development (2026-02-24 to 2026-03
|
|
|
47
47
|
|
|
48
48
|
## Per-Reviewer Profiles
|
|
49
49
|
|
|
50
|
+
> Opus 4.6 was retired on 2026-09-30 (operator decision); Sonnet 5.5 took its
|
|
51
|
+
> review seat (`claude_cli_sonnet5.5`) and the sub-author role. The Opus 4.6
|
|
52
|
+
> profiles below are that model's record and do not describe Sonnet 5.5, which
|
|
53
|
+
> had no profile when the decision was made — two review rounds, and nothing else.
|
|
54
|
+
|
|
50
55
|
### Claude Opus 4.6 (Primary Designer)
|
|
51
56
|
|
|
52
57
|
- **Strength**: Security-level threat modeling, novel architectural alternatives (e.g., discovering `register_gate` as zero-L0 solution)
|
|
@@ -285,6 +290,14 @@ Normative statement and the carryover/new split: L1 `multi_llm_review_workflow`
|
|
|
285
290
|
|
|
286
291
|
> Note: Workflows updated for 4-reviewer default (Opus 4.7 added 2026-04-19). Opus 4.7 profile is provisional pending evaluation data.
|
|
287
292
|
|
|
293
|
+
> Roster note (2026-09-30): the model lines below are the 2026-04-19
|
|
294
|
+
> recommendation (last edited 2026-05-26) and are history, not a roster to launch. Every Claude and
|
|
295
|
+
> Codex model they name has since retired — Opus 4.7 on 2026-06-10, GPT-5.4 on
|
|
296
|
+
> 2026-07-23, Opus 4.6 on 2026-09-30. The roster in use lives in
|
|
297
|
+
> `multi_llm_review/config/multi_llm_review.yml` and nowhere else. What still
|
|
298
|
+
> reads from the lines below is the emphasis per phase, which the Strength
|
|
299
|
+
> Matrix above supports.
|
|
300
|
+
|
|
288
301
|
```
|
|
289
302
|
Design phase: Claude Opus 4.6 + Claude CLI Opus 4.7 + Codex GPT-5.4 + Composer-2.5
|
|
290
303
|
Implementation: Codex GPT-5.4 + Composer-2.5 + Claude Opus 4.6 + Claude CLI Opus 4.7
|
|
@@ -353,6 +366,15 @@ which is decidable by the orchestrator and resistant to value-divergence stallin
|
|
|
353
366
|
|
|
354
367
|
## Changelog
|
|
355
368
|
|
|
369
|
+
- **v1.7 (2026-09-30)**: Opus 4.6 retired by operator decision; Sonnet 5.5
|
|
370
|
+
takes its review seat and the sub-author role. Two notes added, nothing
|
|
371
|
+
deleted: § Per-Reviewer Profiles says the Opus 4.6 profiles are that model's
|
|
372
|
+
record and that Sonnet 5.5 has none, and § Recommended Workflow says its model
|
|
373
|
+
lines are the 2026-04-19 recommendation — every Claude and Codex model in them
|
|
374
|
+
is now retired — with the roster living only in the config. The statistics,
|
|
375
|
+
Strength Matrix, Cost-Benefit and One-Line Summary rows for Opus 4.6 stay:
|
|
376
|
+
they are measurements of Opus 4.6. No profile is written for Sonnet 5.5,
|
|
377
|
+
because no measurement supports one yet.
|
|
356
378
|
- **v1.6 (2026-09-06)**: § Convergence Rule (Updated) rewritten. It had stated
|
|
357
379
|
`3/4 APPROVE = proceed to next step` and `4/4 APPROVE = merge-ready` with no
|
|
358
380
|
note that the ratio is a reference value — while L1 `multi_llm_review_workflow`
|
|
@@ -8,6 +8,8 @@ description: >
|
|
|
8
8
|
the way out. Cannot post, confirm a join, discard, close, or bind evidence — those are the
|
|
9
9
|
operator's, and it hands them back.
|
|
10
10
|
tools: mcp__kairos-chain__am_import, mcp__kairos-chain__am_query, mcp__kairos-chain__am_report
|
|
11
|
+
model: claude-opus-5-5
|
|
12
|
+
effort: high
|
|
11
13
|
---
|
|
12
14
|
|
|
13
15
|
You are the bookkeeper for this operator's ledger.
|
|
@@ -31,7 +31,7 @@ which spawns external LLMs as subprocesses via adapter classes:
|
|
|
31
31
|
|
|
32
32
|
| Adapter | Subprocess command | Use case |
|
|
33
33
|
|---------|-------------------|----------|
|
|
34
|
-
| `ClaudeCodeAdapter` | `claude -p --output-format json` | Sub-author (
|
|
34
|
+
| `ClaudeCodeAdapter` | `claude -p --output-format json` | Sub-author (Sonnet 5.5), persona reviewers |
|
|
35
35
|
| `CodexAdapter` | `codex exec --sandbox read-only` | Codex review |
|
|
36
36
|
| `CursorAdapter` | `agent -p` | Cursor review |
|
|
37
37
|
| `AnthropicAdapter` | Direct API (no subprocess) | Anthropic API calls |
|
|
@@ -3,7 +3,8 @@ name: agent-monitor
|
|
|
3
3
|
description: >
|
|
4
4
|
Post-session review agent for KairosChain cognitive agents.
|
|
5
5
|
Reviews agent status, OODA cycle history, and blockchain records to flag anomalies.
|
|
6
|
-
model:
|
|
6
|
+
model: claude-opus-5-5
|
|
7
|
+
effort: high
|
|
7
8
|
disallowedTools: Write, Edit, Bash
|
|
8
9
|
---
|
|
9
10
|
|
|
@@ -13,7 +13,7 @@
|
|
|
13
13
|
# 2026-08 review threads closed that way without ever reaching their ratio. See
|
|
14
14
|
# L1 multi_llm_review_workflow § Convergence Rules.
|
|
15
15
|
#
|
|
16
|
-
# Roster has 4 reviewers (claude_cli_opus5.5,
|
|
16
|
+
# Roster has 4 reviewers (claude_cli_opus5.5, claude_cli_sonnet5.5,
|
|
17
17
|
# codex_gpt6-astra, cursor_composer2.5).
|
|
18
18
|
# Rules are ratio-based (parser interprets "N/M" as N/M fraction applied
|
|
19
19
|
# to successful count), so the literal numerator/denominator is
|
|
@@ -125,7 +125,7 @@ reviewers:
|
|
|
125
125
|
# "delegate" strategy, is replaced by that orchestrator's own persona
|
|
126
126
|
# Agent Team review; the other is dispatched as a `claude -p` subprocess.
|
|
127
127
|
# So the Claude side is always {orchestrator persona} + {other frontier
|
|
128
|
-
# model via CLI} + {
|
|
128
|
+
# model via CLI} + {Sonnet 5.5 via CLI} — no config branching needed.
|
|
129
129
|
# role_labels are provider-neutral because either entry can end up on
|
|
130
130
|
# either path; the delegated slot is relabelled claude_team_<model> by
|
|
131
131
|
# PersonaAssembly at collect time.
|
|
@@ -141,17 +141,19 @@ reviewers:
|
|
|
141
141
|
# contributing signal. To restore, re-add an entry with model
|
|
142
142
|
# claude-fable-5-1 (Fable 5's successor; Fable 5 itself is not the one to
|
|
143
143
|
# bring back) and widen convergence_rule back to a 6-reviewer basis.
|
|
144
|
-
# Consequence of the removal: the Claude CLI side is
|
|
144
|
+
# Consequence of the removal: the Claude CLI side is sonnet5.5 only when the
|
|
145
145
|
# orchestrator is opus5.5, since the orchestrator's own slot is delegated to
|
|
146
146
|
# the persona team rather than spawned as a subprocess.
|
|
147
147
|
|
|
148
|
-
# Opus 4.6
|
|
149
|
-
#
|
|
150
|
-
# anchor on the Claude side
|
|
148
|
+
# Opus 4.6 retired 2026-09-30 (operator instruction); Sonnet 5.5 at effort
|
|
149
|
+
# high takes the seat. Opus 4.6 had been kept as the calibrated, deliberately
|
|
150
|
+
# non-frontier anchor on the Claude side. Sonnet 5.5 had held the seat for
|
|
151
|
+
# two rounds when it was chosen and is otherwise uncalibrated. To restore, set
|
|
152
|
+
# this entry back to model claude-opus-4-6 and role_label claude_cli_opus4.6.
|
|
151
153
|
- provider: claude_code
|
|
152
|
-
model: claude-
|
|
154
|
+
model: claude-sonnet-5-5
|
|
153
155
|
effort: high
|
|
154
|
-
role_label:
|
|
156
|
+
role_label: claude_cli_sonnet5.5
|
|
155
157
|
|
|
156
158
|
# Opus 4.8 retired 2026-07-25 when Opus 5 entered the roster (same
|
|
157
159
|
# succession logic that retired 4.7 on 2026-06-10: the strictness /
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "multi_llm_review",
|
|
3
|
-
"version": "0.11.
|
|
4
|
-
"description": "v0.10.1 (2026-08-17): the returned block of vote counts is named vote_tally, not convergence. It holds approve_count, reject_count, skip_count, successful_count, threshold and rule -- vote arithmetic -- while the closing condition the L1 workflow states is 'new (a)+(b) P0 = 0', which this SkillSet does not compute and no returned field carries. INV-R2 had already demoted the ratio to a reference value in a code comment while leaving the field named convergence, so every round handed the orchestrator a block whose name claimed what its contents could not answer. Renaming rather than adding a field is deliberate: with nothing in the payload called convergence, the criterion has to be fetched from the findings. Consumers reading the old name get nil; agent_step reads vote_tally with a fallback to convergence for records written earlier. v0.10.0 (finding weight axis + worker-death recovery, 2026-08-06): findings gain a consequence axis beside the severity axis — the prompt contract requires a [consequence: who is harmed, and how, if this is never fixed] clause on every P0, and aggregation records a P0 without one at P2, keeping the stated severity and the demotion reason on the row (severity_stated / severity_demoted: consequence_missing). Only PRESENCE is checked, mechanically; whether a stated consequence is real or trivial stays the orchestrator's call. The dedup key strips the clause so two reviewers naming one defect still merge, and a group where one member states the harm carries it for the row. Motivating measurement (project_orientation_report checker, R5): 3 of 7 P0s were factually correct findings that cost nobody anything. Reviewer prompts also stop naming the round number — telling a reviewer which round it is in (like telling it its counts are compared) selects for finding-production over finding-weight; the L1 workflow v3.10.0 Reviewer incentive rule states the orchestrator-side half. And a worker death no longer discards completed seats: the worker persists each reply as it arrives (partial_results.json), and collect's crash/timeout branches recover the finished seats, entering every unreached seat in the denominator as a skip row (worker_crashed_seat_lost) with the death named in payload.worker_failure — R3 2026-08-06 lost three completed external seats to one stale heartbeat because the only exit was total loss. v0.9.1 (seat access, frozen 2026-08-06 after a one-round review): reviewer prompts carry a <seat_access> block beside the inline artifact — a seat that cannot read the repository must not attempt tool calls and must not open by saying it will read files; it reviews the artifact text alone, marks unverifiable claims [INFERRED], and still opens with its verdict line. The wording is conditional because seats differ (codex runs --sandbox read-only and can read the repository), and the block is not emitted for by_reference delivery, whose existing cannot-read instruction it would contradict. Root cause fixed: the claude subprocess seat runs with tools disabled in an empty working directory, and on implementation artifacts citing file paths it opened with pseudo-tool-call markup or a cannot-access preamble instead of its verdict line, leaving four consecutive rounds as no_verdict while counting every round on design artifacts. v0.9.0 (evidence fidelity, frozen 2026-08-06 after six review rounds): a finding reaches the record whole. It used to be cut at 201 bytes by an inclusive Range in aggregation and bounded again by the 500-character display limit, so a downstream instance measured 18 of 21 findings arriving at exactly 201 bytes with reviews[].raw_text empty on every row. Findings are now bounded in BYTES at FINDING_RECORD_MAX_LEN = 8000 for the record, with DEFAULT_MAX_LEN = 500 still applied by every path that takes a finding into a prompt. Deduplication still keys on the first 80 characters — widening it would stop two reviewers describing one defect from merging, moving the finding count and the convergence denominator — but the surviving text is no longer arbitrary: issue comes from a member whose severity equals the merged severity, distinct texts survive in issue_variants (capped at MAX_ISSUE_VARIANTS = 8, with issue_variants_omitted naming what the cap dropped). Every row now carries raw_text_excerpt unconditionally (4096 bytes) and the reviewer's reply in raw_text on request (include_raw_text, 65536 bytes). Both are SANITISED TRANSCRIPTIONS, not verbatim records, and both tool schemas say so: the text is byte-clamped, NFKC-normalised, stripped of invisible characters, tag-escaped, then byte-clamped again. No field states whether they hold the whole reply, because nothing in this SkillSet can know — a completeness flag was implemented, measured wrong in both directions, and removed. On delegated runs the pending-state record keeps each subprocess reply as it arrived; on single-phase runs the returned payload is the only form there is. Findings carried into a later round's prompt are sanitised and folded to one line, a path that previously took reviewer text into a prompt with no sanitisation at all. v0.8.0 (v0.7 record schema, design frozen 2026-08-01): the verdict vocabulary is the three canonical words plus tense forms only (INV-R1); the ratio and threshold are recorded reference values, not the run's conclusion — the top-level verdict field became reference_verdict and the run is closed by the operator's declaration outside the record (INV-R2); the persona team occupies one seat, its derivation rule is recorded, and a submission smaller than convened — including empty — is accepted with the shortfall on the record (INV-R3/R4); every run writes an existence marker at dispatch, completed records are never garbage-collected, and expired runs are reduced to a minimal trace instead of erased (INV-R4); a divergence-excluded tally is carried beside the main one (INV-R5); the record names its pre-declared spec and carries transport diagnostics as state tags (INV-R6); artifact delivery is a per-seat attribute (inline | by_reference) and an unreachable delivery is refused rather than dispatched (INV-R7). Parallel multi-LLM review orchestration. Dispatches review prompts to N LLM backends via llm_client, collects verdicts, and computes consensus. v0.6.0: reserve observers (escalate) and a declarable persona execution model; the observer set is built in one pass with explicit precedence (ObserverSet); every slot must name its model and role_label, and duplicate names — including the persona team's own — are refused. A reply's verdict is no longer inferred from its prose: it is read from a declared field, from the header the reply opens with when that header carries a verdict name and nothing else, or not at all, in which case the reply leaves the denominator with no_verdict recorded beside its name. The record says why every observer did or did not count (denominator_composition, five skip_reason values, observers_reporting), and the per-reviewer row is written by one mapping rather than two. v0.5.2: the cursor reviewer pins model composer-2.5 instead of inheriting the cursor CLI default, which is operator-editable and had silently become an Anthropic model. v0.5.1: Fable 5 retired from the roster (five consecutive silent returns), convergence 3/5; orchestrator_model description now states the bare-ID rule so a caller does not review its own output. v0.5.0: adds multi_llm_review_wait (Phase 1.5) for explicit subprocess completion gating with next_action recovery hints, and Path A/B doc disambiguation. v0.4.0 (Phase 12): feedback_text + schema_version, sanitization contract for prompt-injection defense, and multi_llm_review_bundle tool for human-handoff paths without dispatch.",
|
|
3
|
+
"version": "0.11.1",
|
|
4
|
+
"description": "v0.11.1 (2026-09-30): an artifact can be named by artifact_path + artifact_sha256 instead of carried in artifact_content, for callers such as an agent plan that cannot carry the text. The tool reads only a regular file inside its working directory, re-checks after reading that the path still names the file it opened, and refuses unless the bytes hash to the pin; carried text given beside a pin must hash to it too. Every refusal happens before any seat is dispatched. v0.10.1 (2026-08-17): the returned block of vote counts is named vote_tally, not convergence. It holds approve_count, reject_count, skip_count, successful_count, threshold and rule -- vote arithmetic -- while the closing condition the L1 workflow states is 'new (a)+(b) P0 = 0', which this SkillSet does not compute and no returned field carries. INV-R2 had already demoted the ratio to a reference value in a code comment while leaving the field named convergence, so every round handed the orchestrator a block whose name claimed what its contents could not answer. Renaming rather than adding a field is deliberate: with nothing in the payload called convergence, the criterion has to be fetched from the findings. Consumers reading the old name get nil; agent_step reads vote_tally with a fallback to convergence for records written earlier. v0.10.0 (finding weight axis + worker-death recovery, 2026-08-06): findings gain a consequence axis beside the severity axis — the prompt contract requires a [consequence: who is harmed, and how, if this is never fixed] clause on every P0, and aggregation records a P0 without one at P2, keeping the stated severity and the demotion reason on the row (severity_stated / severity_demoted: consequence_missing). Only PRESENCE is checked, mechanically; whether a stated consequence is real or trivial stays the orchestrator's call. The dedup key strips the clause so two reviewers naming one defect still merge, and a group where one member states the harm carries it for the row. Motivating measurement (project_orientation_report checker, R5): 3 of 7 P0s were factually correct findings that cost nobody anything. Reviewer prompts also stop naming the round number — telling a reviewer which round it is in (like telling it its counts are compared) selects for finding-production over finding-weight; the L1 workflow v3.10.0 Reviewer incentive rule states the orchestrator-side half. And a worker death no longer discards completed seats: the worker persists each reply as it arrives (partial_results.json), and collect's crash/timeout branches recover the finished seats, entering every unreached seat in the denominator as a skip row (worker_crashed_seat_lost) with the death named in payload.worker_failure — R3 2026-08-06 lost three completed external seats to one stale heartbeat because the only exit was total loss. v0.9.1 (seat access, frozen 2026-08-06 after a one-round review): reviewer prompts carry a <seat_access> block beside the inline artifact — a seat that cannot read the repository must not attempt tool calls and must not open by saying it will read files; it reviews the artifact text alone, marks unverifiable claims [INFERRED], and still opens with its verdict line. The wording is conditional because seats differ (codex runs --sandbox read-only and can read the repository), and the block is not emitted for by_reference delivery, whose existing cannot-read instruction it would contradict. Root cause fixed: the claude subprocess seat runs with tools disabled in an empty working directory, and on implementation artifacts citing file paths it opened with pseudo-tool-call markup or a cannot-access preamble instead of its verdict line, leaving four consecutive rounds as no_verdict while counting every round on design artifacts. v0.9.0 (evidence fidelity, frozen 2026-08-06 after six review rounds): a finding reaches the record whole. It used to be cut at 201 bytes by an inclusive Range in aggregation and bounded again by the 500-character display limit, so a downstream instance measured 18 of 21 findings arriving at exactly 201 bytes with reviews[].raw_text empty on every row. Findings are now bounded in BYTES at FINDING_RECORD_MAX_LEN = 8000 for the record, with DEFAULT_MAX_LEN = 500 still applied by every path that takes a finding into a prompt. Deduplication still keys on the first 80 characters — widening it would stop two reviewers describing one defect from merging, moving the finding count and the convergence denominator — but the surviving text is no longer arbitrary: issue comes from a member whose severity equals the merged severity, distinct texts survive in issue_variants (capped at MAX_ISSUE_VARIANTS = 8, with issue_variants_omitted naming what the cap dropped). Every row now carries raw_text_excerpt unconditionally (4096 bytes) and the reviewer's reply in raw_text on request (include_raw_text, 65536 bytes). Both are SANITISED TRANSCRIPTIONS, not verbatim records, and both tool schemas say so: the text is byte-clamped, NFKC-normalised, stripped of invisible characters, tag-escaped, then byte-clamped again. No field states whether they hold the whole reply, because nothing in this SkillSet can know — a completeness flag was implemented, measured wrong in both directions, and removed. On delegated runs the pending-state record keeps each subprocess reply as it arrived; on single-phase runs the returned payload is the only form there is. Findings carried into a later round's prompt are sanitised and folded to one line, a path that previously took reviewer text into a prompt with no sanitisation at all. v0.8.0 (v0.7 record schema, design frozen 2026-08-01): the verdict vocabulary is the three canonical words plus tense forms only (INV-R1); the ratio and threshold are recorded reference values, not the run's conclusion — the top-level verdict field became reference_verdict and the run is closed by the operator's declaration outside the record (INV-R2); the persona team occupies one seat, its derivation rule is recorded, and a submission smaller than convened — including empty — is accepted with the shortfall on the record (INV-R3/R4); every run writes an existence marker at dispatch, completed records are never garbage-collected, and expired runs are reduced to a minimal trace instead of erased (INV-R4); a divergence-excluded tally is carried beside the main one (INV-R5); the record names its pre-declared spec and carries transport diagnostics as state tags (INV-R6); artifact delivery is a per-seat attribute (inline | by_reference) and an unreachable delivery is refused rather than dispatched (INV-R7). Parallel multi-LLM review orchestration. Dispatches review prompts to N LLM backends via llm_client, collects verdicts, and computes consensus. v0.6.0: reserve observers (escalate) and a declarable persona execution model; the observer set is built in one pass with explicit precedence (ObserverSet); every slot must name its model and role_label, and duplicate names — including the persona team's own — are refused. A reply's verdict is no longer inferred from its prose: it is read from a declared field, from the header the reply opens with when that header carries a verdict name and nothing else, or not at all, in which case the reply leaves the denominator with no_verdict recorded beside its name. The record says why every observer did or did not count (denominator_composition, five skip_reason values, observers_reporting), and the per-reviewer row is written by one mapping rather than two. v0.5.2: the cursor reviewer pins model composer-2.5 instead of inheriting the cursor CLI default, which is operator-editable and had silently become an Anthropic model. v0.5.1: Fable 5 retired from the roster (five consecutive silent returns), convergence 3/5; orchestrator_model description now states the bare-ID rule so a caller does not review its own output. v0.5.0: adds multi_llm_review_wait (Phase 1.5) for explicit subprocess completion gating with next_action recovery hints, and Path A/B doc disambiguation. v0.4.0 (Phase 12): feedback_text + schema_version, sanitization contract for prompt-injection defense, and multi_llm_review_bundle tool for human-handoff paths without dispatch.",
|
|
5
5
|
"author": "Masaomi Hatakeyama",
|
|
6
6
|
"layer": "L1",
|
|
7
7
|
"depends_on": [
|
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
3
|
require 'minitest/autorun'
|
|
4
|
+
require 'minitest/mock'
|
|
4
5
|
require 'json'
|
|
5
6
|
require_relative '../lib/multi_llm_review/consensus'
|
|
6
7
|
require_relative '../lib/multi_llm_review/observer_set'
|
|
@@ -1155,6 +1156,75 @@ module KairosMcp
|
|
|
1155
1156
|
# against the observer set rather than against a helper that only saw
|
|
1156
1157
|
# half of it.
|
|
1157
1158
|
|
|
1159
|
+
# An agent plan names the artifact by path + sha256 instead of
|
|
1160
|
+
# carrying it; the tool reads exactly those bytes or refuses.
|
|
1161
|
+
def test_pinned_artifact_is_read_only_when_it_matches_its_pin
|
|
1162
|
+
File.write('artifact.md', "reviewed text\n")
|
|
1163
|
+
pin = Digest::SHA256.hexdigest("reviewed text\n")
|
|
1164
|
+
|
|
1165
|
+
assert_equal ["reviewed text\n", nil], @tool.send(:read_pinned_artifact, 'artifact.md', pin)
|
|
1166
|
+
assert_match(/not the pinned/, @tool.send(:read_pinned_artifact, 'artifact.md', 'f' * 64)[1])
|
|
1167
|
+
assert_match(/inside/, @tool.send(:read_pinned_artifact, '/etc/hosts', pin)[1])
|
|
1168
|
+
refused = JSON.parse(@tool.call({ 'artifact_name' => 'a', 'review_type' => 'design',
|
|
1169
|
+
'artifact_path' => 'artifact.md' }).first[:text])
|
|
1170
|
+
assert_match(/artifact_sha256 are both given/, refused['error'])
|
|
1171
|
+
end
|
|
1172
|
+
|
|
1173
|
+
# A pin binds whatever is reviewed. Carried text that does not hash to
|
|
1174
|
+
# it is refused before roster resolution (a plan can write a
|
|
1175
|
+
# placeholder beside the pin); carried text that does, and the file's
|
|
1176
|
+
# own bytes, go on to be reviewed.
|
|
1177
|
+
def test_pin_binds_carried_content_as_well_as_the_file
|
|
1178
|
+
seen = nil
|
|
1179
|
+
@tool.define_singleton_method(:resolve_reviewers) do |args, _config|
|
|
1180
|
+
seen = args['artifact_content']
|
|
1181
|
+
raise ObserverSet::RosterError, 'stopped before dispatch'
|
|
1182
|
+
end
|
|
1183
|
+
File.write('artifact.md', "reviewed text\n")
|
|
1184
|
+
pin = Digest::SHA256.hexdigest("reviewed text\n")
|
|
1185
|
+
base = { 'artifact_name' => 'a', 'review_type' => 'design', 'artifact_path' => 'artifact.md' }
|
|
1186
|
+
|
|
1187
|
+
refused = JSON.parse(@tool.call(base.merge('artifact_content' => '<<see artifact_path>>',
|
|
1188
|
+
'artifact_sha256' => pin)).first[:text])
|
|
1189
|
+
assert_match(/does not hash to artifact_sha256/, refused['error'])
|
|
1190
|
+
assert_nil seen, 'a refused call must not reach roster resolution'
|
|
1191
|
+
|
|
1192
|
+
@tool.call(base.merge('artifact_content' => "reviewed text\n", 'artifact_sha256' => " #{pin.upcase} "))
|
|
1193
|
+
assert_equal "reviewed text\n", seen
|
|
1194
|
+
|
|
1195
|
+
seen = nil
|
|
1196
|
+
@tool.call(base.merge('artifact_sha256' => pin))
|
|
1197
|
+
assert_equal "reviewed text\n", seen, 'the file bytes become the reviewed content'
|
|
1198
|
+
end
|
|
1199
|
+
|
|
1200
|
+
# The read is confined to a regular file inside the working directory,
|
|
1201
|
+
# and to the file the checked path named when it was opened.
|
|
1202
|
+
def test_pinned_read_refuses_escapes_non_files_and_a_swap
|
|
1203
|
+
outside = Dir.mktmpdir('mlr-outside-')
|
|
1204
|
+
File.write(File.join(outside, 'secret.md'), "secret\n")
|
|
1205
|
+
secret_pin = Digest::SHA256.hexdigest("secret\n")
|
|
1206
|
+
File.symlink(File.join(outside, 'secret.md'), 'link.md')
|
|
1207
|
+
assert_match(/inside/, @tool.send(:read_pinned_artifact, 'link.md', secret_pin)[1])
|
|
1208
|
+
|
|
1209
|
+
Dir.mkdir('adir')
|
|
1210
|
+
File.mkfifo('afifo')
|
|
1211
|
+
assert_match(/not a regular file/, @tool.send(:read_pinned_artifact, 'adir', secret_pin)[1])
|
|
1212
|
+
assert_match(/not a regular file/, @tool.send(:read_pinned_artifact, 'afifo', secret_pin)[1])
|
|
1213
|
+
assert_match(/unreadable/, @tool.send(:read_pinned_artifact, "a\0b", secret_pin)[1])
|
|
1214
|
+
|
|
1215
|
+
# A swap between open and re-check shows up as a different inode
|
|
1216
|
+
# behind the checked path.
|
|
1217
|
+
File.write('artifact.md', "reviewed text\n")
|
|
1218
|
+
pin = Digest::SHA256.hexdigest("reviewed text\n")
|
|
1219
|
+
elsewhere = File.stat(File.join(outside, 'secret.md'))
|
|
1220
|
+
File.stub(:stat, elsewhere) do
|
|
1221
|
+
assert_match(/changed while/, @tool.send(:read_pinned_artifact, 'artifact.md', pin)[1])
|
|
1222
|
+
end
|
|
1223
|
+
assert_equal ["reviewed text\n", nil], @tool.send(:read_pinned_artifact, 'artifact.md', pin)
|
|
1224
|
+
ensure
|
|
1225
|
+
FileUtils.rm_rf(outside) if outside
|
|
1226
|
+
end
|
|
1227
|
+
|
|
1158
1228
|
def test_delegate_response_writes_pending_state
|
|
1159
1229
|
subprocess_results = [
|
|
1160
1230
|
{ role_label: 'codex', provider: 'codex', model: 'codex-default',
|
|
@@ -80,7 +80,18 @@ module KairosMcp
|
|
|
80
80
|
properties: {
|
|
81
81
|
artifact_content: {
|
|
82
82
|
type: 'string',
|
|
83
|
-
description: '
|
|
83
|
+
description: 'Omit when artifact_path + artifact_sha256 name the file; ' \
|
|
84
|
+
'otherwise the full text of the artifact to review. If a pin is ' \
|
|
85
|
+
'also given, this text must hash to it or the call is refused.'
|
|
86
|
+
},
|
|
87
|
+
artifact_sha256: {
|
|
88
|
+
type: 'string',
|
|
89
|
+
description: 'Pin: sha256 of the bytes to review. With artifact_path and ' \
|
|
90
|
+
'no artifact_content the tool reads the file at artifact_path (a ' \
|
|
91
|
+
'regular file inside its working directory) and refuses unless the ' \
|
|
92
|
+
'bytes hash to this value; with artifact_content, that text must ' \
|
|
93
|
+
'hash to it. For callers that cannot carry the text, such as an ' \
|
|
94
|
+
'agent plan.'
|
|
84
95
|
},
|
|
85
96
|
artifact_name: {
|
|
86
97
|
type: 'string',
|
|
@@ -92,7 +103,9 @@ module KairosMcp
|
|
|
92
103
|
'Required for any roster slot configured with ' \
|
|
93
104
|
'artifact_delivery: by_reference — those slots receive this path ' \
|
|
94
105
|
'plus a sha256 computed over artifact_content instead of the ' \
|
|
95
|
-
'inlined body. Slots delivered inline ignore it
|
|
106
|
+
'inlined body. Slots delivered inline ignore it, except that ' \
|
|
107
|
+
'with artifact_sha256 and no artifact_content it is the file ' \
|
|
108
|
+
'the tool reads for every slot.'
|
|
96
109
|
},
|
|
97
110
|
review_spec: {
|
|
98
111
|
type: 'object',
|
|
@@ -240,7 +253,7 @@ module KairosMcp
|
|
|
240
253
|
'assembly, and no other copy exists anywhere.'
|
|
241
254
|
}
|
|
242
255
|
},
|
|
243
|
-
required: %w[
|
|
256
|
+
required: %w[artifact_name review_type]
|
|
244
257
|
}
|
|
245
258
|
end
|
|
246
259
|
|
|
@@ -266,6 +279,26 @@ module KairosMcp
|
|
|
266
279
|
}))
|
|
267
280
|
end
|
|
268
281
|
|
|
282
|
+
# A caller that cannot carry the text names it by path and sha256.
|
|
283
|
+
# An agent plan is why: its tool arguments are written by a model
|
|
284
|
+
# within an output budget, and a large artifact reached the seats
|
|
285
|
+
# as a placeholder string (2026-09-30). A pin binds whatever is
|
|
286
|
+
# reviewed, so carried text that does not hash to it is refused
|
|
287
|
+
# too — a plan can write a placeholder beside the pin.
|
|
288
|
+
pin = arguments['artifact_sha256'].to_s.strip.downcase
|
|
289
|
+
if arguments['artifact_content'].to_s.empty?
|
|
290
|
+
content, error = read_pinned_artifact(arguments['artifact_path'], pin)
|
|
291
|
+
return text_content(JSON.generate({ 'status' => 'error', 'error' => error })) if error
|
|
292
|
+
|
|
293
|
+
arguments = arguments.merge('artifact_content' => content)
|
|
294
|
+
elsif !pin.empty? && Digest::SHA256.hexdigest(arguments['artifact_content'].to_s) != pin
|
|
295
|
+
return text_content(JSON.generate({
|
|
296
|
+
'status' => 'error',
|
|
297
|
+
'error' => 'artifact_content does not hash to artifact_sha256. Omit ' \
|
|
298
|
+
'artifact_content to have the tool read artifact_path, or drop the pin.'
|
|
299
|
+
}))
|
|
300
|
+
end
|
|
301
|
+
|
|
269
302
|
begin
|
|
270
303
|
reviewers = resolve_reviewers(arguments, config)
|
|
271
304
|
rescue ObserverSet::RosterError => e
|
|
@@ -625,6 +658,46 @@ module KairosMcp
|
|
|
625
658
|
|
|
626
659
|
private
|
|
627
660
|
|
|
661
|
+
# Reads only inside the working directory (the root pin_resolver
|
|
662
|
+
# uses) and only bytes that hash to the caller's pin, so the plan
|
|
663
|
+
# that named the file, not the file's later state, decides what is
|
|
664
|
+
# reviewed. Returns [text, nil] or [nil, error].
|
|
665
|
+
def read_pinned_artifact(path, sha256)
|
|
666
|
+
if path.to_s.strip.empty? || sha256.to_s.strip.empty?
|
|
667
|
+
return [nil, 'artifact_content is required, unless artifact_path and artifact_sha256 are both given']
|
|
668
|
+
end
|
|
669
|
+
|
|
670
|
+
root = File.realpath(Dir.pwd)
|
|
671
|
+
full = File.realpath(File.expand_path(path.to_s, root))
|
|
672
|
+
return [nil, "artifact_path must lie inside #{root}"] unless full.start_with?("#{root}/")
|
|
673
|
+
|
|
674
|
+
# One handle, opened without blocking (a FIFO would otherwise hang
|
|
675
|
+
# the call) and read only if it is a regular file. Afterwards the
|
|
676
|
+
# checked path must still resolve to the file that handle opened,
|
|
677
|
+
# so a symlink swapped in between the check and the read cannot
|
|
678
|
+
# pull bytes from outside the root.
|
|
679
|
+
bytes = nil
|
|
680
|
+
opened = File.open(full, File::RDONLY | File::NONBLOCK, binmode: true) do |f|
|
|
681
|
+
st = f.stat
|
|
682
|
+
bytes = f.read if st.file?
|
|
683
|
+
st
|
|
684
|
+
end
|
|
685
|
+
return [nil, 'artifact_path is not a regular file'] unless opened.file?
|
|
686
|
+
|
|
687
|
+
now = File.stat(full)
|
|
688
|
+
unless File.realpath(full) == full && [now.dev, now.ino] == [opened.dev, opened.ino]
|
|
689
|
+
return [nil, 'artifact_path changed while it was being read']
|
|
690
|
+
end
|
|
691
|
+
|
|
692
|
+
actual = Digest::SHA256.hexdigest(bytes)
|
|
693
|
+
return [nil, "artifact_path hashes to #{actual}, not the pinned #{sha256}"] unless actual == sha256.to_s.strip.downcase
|
|
694
|
+
|
|
695
|
+
text = bytes.force_encoding(Encoding::UTF_8)
|
|
696
|
+
text.valid_encoding? ? [text, nil] : [nil, 'artifact_path is not valid UTF-8']
|
|
697
|
+
rescue SystemCallError, ArgumentError => e
|
|
698
|
+
[nil, "artifact_path unreadable: #{e.message}"]
|
|
699
|
+
end
|
|
700
|
+
|
|
628
701
|
# Phase 1.5 — articulate which fallback_chain path actually ran.
|
|
629
702
|
# Returns Hash with path_taken/tier_actually_used/target_harness/acknowledgment.
|
|
630
703
|
# Keyed on whether a persona was actually convened, not on the
|
|
@@ -8,6 +8,8 @@ description: >
|
|
|
8
8
|
project_manager work-item store. Cannot record irreversible project actions and cannot change or
|
|
9
9
|
read project records — hands those back for the operator.
|
|
10
10
|
tools: mcp__kairos-chain__pm_digest, mcp__kairos-chain__pm_query, mcp__kairos-chain__pm_item
|
|
11
|
+
model: claude-opus-5-5
|
|
12
|
+
effort: high
|
|
11
13
|
---
|
|
12
14
|
|
|
13
15
|
You are the secretary for this operator's project store.
|
|
@@ -3,7 +3,8 @@ name: exchange-reviewer
|
|
|
3
3
|
description: >
|
|
4
4
|
Reviews SkillSet exchange operations for safety and compatibility.
|
|
5
5
|
Checks blockchain integrity, SkillSet health, and skill freshness.
|
|
6
|
-
model:
|
|
6
|
+
model: claude-opus-5-5
|
|
7
|
+
effort: high
|
|
7
8
|
disallowedTools: Write, Edit, Bash
|
|
8
9
|
---
|
|
9
10
|
|