kairos-chain 3.55.0 → 3.57.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +35 -0
- data/lib/kairos_mcp/version.rb +1 -1
- data/templates/knowledge/multi_llm_review_workflow/multi_llm_review_workflow.md +196 -11
- data/templates/skillsets/agent/test/test_agent_complexity_review.rb +66 -0
- data/templates/skillsets/agent/tools/agent_step.rb +9 -2
- data/templates/skillsets/llm_client/lib/llm_client/claude_code_adapter.rb +31 -4
- data/templates/skillsets/llm_client/skillset.json +1 -1
- data/templates/skillsets/llm_client/test/test_claude_code_adapter_parse.rb +106 -0
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/build_review_bundle.rb +30 -2
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/consensus.rb +331 -81
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/dispatcher.rb +31 -2
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/observer_set.rb +43 -3
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/pending_state.rb +128 -16
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/persona_assembly.rb +105 -32
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/prompt_builder.rb +40 -8
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/review_serializer.rb +62 -0
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/verdict_vocabulary.rb +167 -0
- data/templates/skillsets/multi_llm_review/skillset.json +8 -4
- data/templates/skillsets/multi_llm_review/test/test_multi_llm_review.rb +532 -65
- data/templates/skillsets/multi_llm_review/test/test_mutation_survivors.rb +1124 -0
- data/templates/skillsets/multi_llm_review/test/test_observer_set.rb +147 -9
- data/templates/skillsets/multi_llm_review/test/test_observer_set_seams.rb +775 -78
- data/templates/skillsets/multi_llm_review/test/test_pending_state_v3.rb +122 -2
- data/templates/skillsets/multi_llm_review/test/test_tool_wiring.rb +429 -8
- data/templates/skillsets/multi_llm_review/tools/multi_llm_review.rb +422 -128
- data/templates/skillsets/multi_llm_review/tools/multi_llm_review_collect.rb +108 -32
- data/templates/skillsets/synoptis/lib/synoptis/attestation_engine.rb +35 -2
- data/templates/skillsets/synoptis/lib/synoptis/proof_envelope.rb +55 -3
- data/templates/skillsets/synoptis/lib/synoptis/tool_helpers.rb +9 -2
- data/templates/skillsets/synoptis/lib/synoptis/verifier.rb +32 -5
- data/templates/skillsets/synoptis/tools/attestation_verify.rb +1 -1
- metadata +4 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 8439763123a17e68d238ebc46369dcc36e6379172752809dd6736fa18f3c5ec4
|
|
4
|
+
data.tar.gz: 76fe818c7261e464411f54bee9a114bc7eb0ecc025042ef264f15bd591f2ca36
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 2385335588b313d3910519459a79a1325f921fa1f6d1a2b215f0b75144c383b14b90a95f1d9edf97999df6cdc58d95155e653f13007217299f35df097891f38a
|
|
7
|
+
data.tar.gz: 160f831a399a60089d217e2fe198dfe60692797649b55c2d8b4fbd0d0649512c6f601a75dfc41de713b9bf946be83485f01c289252018c87d33dbbb06221c1ec
|
data/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,41 @@ All notable changes to the `kairos-chain` gem will be documented in this file.
|
|
|
4
4
|
|
|
5
5
|
This project follows [Semantic Versioning](https://semver.org/).
|
|
6
6
|
|
|
7
|
+
## [3.56.0] - 2026-07-31
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
- **`multi_llm_review` SkillSet 0.7.0 — verdict reading hardened across both
|
|
12
|
+
observer paths** (review rounds R10–R15, frozen by operator declaration
|
|
13
|
+
2026-07-31 after three consecutive rounds with zero deployment-grounded
|
|
14
|
+
findings).
|
|
15
|
+
- The verdict vocabulary is one table (`VerdictVocabulary::WORDS`); the
|
|
16
|
+
prose-search and whole-value patterns are built from it, differing only
|
|
17
|
+
in anchoring and separator strictness. The rebuilt regexes are
|
|
18
|
+
source-identical to the previous literals.
|
|
19
|
+
- Verdict *reading* is `stated` (whole-value, anchored) on every path.
|
|
20
|
+
`VerdictVocabulary.classify`, `PersonaAssembly.normalize_verdict`, and
|
|
21
|
+
the dead constants `VERDICT_PATTERNS` / `ALLOWED_VERDICTS` /
|
|
22
|
+
`*_ALIASES` are deleted; their absence is pinned by tests.
|
|
23
|
+
- A persona verdict field that is not a verdict is **refused at
|
|
24
|
+
submission** (`PersonaAssembly.validate!` raises ArgumentError) instead
|
|
25
|
+
of being word-searched (≤ R12) or defaulted to REVISE (R13). The collect
|
|
26
|
+
tool validates before consuming, so a refused submission leaves the
|
|
27
|
+
pending token collectable and the corrected submission carries the vote.
|
|
28
|
+
- Declared verdicts are carried in canonical case (`extract_verdict`
|
|
29
|
+
merges the admitted form), so a lower-case declaration can no longer sit
|
|
30
|
+
in the denominator without counting.
|
|
31
|
+
- Test suite grew 401 → 478 runs / 1676 assertions, including a
|
|
32
|
+
mutation-survivor file distilled from three generative sweeps.
|
|
33
|
+
- **`llm_client` SkillSet 0.2.0 — CLI envelope diagnostics preserved.**
|
|
34
|
+
`claude_code_adapter` now picks `model_observed` by output tokens rather
|
|
35
|
+
than hash order, and passes `model_usage`, `api_error_status`,
|
|
36
|
+
`fast_mode_state` and `terminal_reason` through to the response. Four
|
|
37
|
+
model-divergence incidents (a slot requesting opus-4-6 answered by
|
|
38
|
+
haiku-4-5) were undiagnosable because these fields were discarded; the
|
|
39
|
+
divergence itself was confirmed real (the divergent reply cited
|
|
40
|
+
identifiers that exist nowhere in the reviewed code).
|
|
41
|
+
|
|
7
42
|
## [3.55.0] - 2026-07-27
|
|
8
43
|
|
|
9
44
|
### Added
|
data/lib/kairos_mcp/version.rb
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: multi_llm_review_workflow
|
|
3
3
|
description: "Multi-LLM review methodology and execution — workflow pattern, CLI tooling, consensus analysis, Persona Assembly. Applicable to design, implementation, documentation, or any artifact."
|
|
4
|
-
version: "3.
|
|
4
|
+
version: "3.8.0"
|
|
5
5
|
tags:
|
|
6
6
|
- workflow
|
|
7
7
|
- review
|
|
@@ -29,6 +29,51 @@ This skill covers:
|
|
|
29
29
|
For **WHO** (which LLM is good at what), see: `multi_llm_reviewer_evaluation`
|
|
30
30
|
For **development lifecycle** (design → implement → verify), see: `design_to_implementation_workflow`
|
|
31
31
|
|
|
32
|
+
## Step -1 — Pre-declared review spec (loop hygiene)
|
|
33
|
+
|
|
34
|
+
> Validation scope: these rules were derived from one non-converging loop
|
|
35
|
+
> (multi_llm_review R10–R15, 2026-07) where they took the P0 count from 10 to
|
|
36
|
+
> 2 in two rounds and closed the loop in three. They are instance practice
|
|
37
|
+
> until reproduced on a second, non-self-referential subject; treat the
|
|
38
|
+
> numbers below as one loop's evidence, not a law.
|
|
39
|
+
|
|
40
|
+
Before dispatching round 1 — and again whenever the review TARGET changes —
|
|
41
|
+
write a review spec and declare it frozen for the round:
|
|
42
|
+
|
|
43
|
+
1. **Pre-declare the pass condition** (per `loop_validation`: spec before
|
|
44
|
+
judgement, fail-closed). State what APPROVE requires. A loop whose target
|
|
45
|
+
drifted (e.g. from an implementation to the instrument that measures it)
|
|
46
|
+
without a re-declared spec is structurally non-converging: a 300-claim
|
|
47
|
+
artifact at any realistic per-claim error rate yields double-digit
|
|
48
|
+
findings every round regardless of quality.
|
|
49
|
+
2. **Split target from appendix.** Only the shipping deliverable is
|
|
50
|
+
P0-eligible. Instruments, sweep logs, classification tables are appendix:
|
|
51
|
+
findings against them are advisory and go to the queue. This is what
|
|
52
|
+
collapsed the claim surface from ~350 to ~30.
|
|
53
|
+
3. **Cap fixes per round (≤5)** and write one line per fix: *what this fix
|
|
54
|
+
newly claims* (values pinned, ranges narrowed, failure visibility
|
|
55
|
+
changed). A fix that cannot state its new claims is doing more than the
|
|
56
|
+
finding asked.
|
|
57
|
+
4. **Pre-flight falsifier.** Before dispatch, one agent whose only job is to
|
|
58
|
+
refute every factual claim in the spec and artifact — especially numbers
|
|
59
|
+
and "X does not exist" claims. In this loop it caught real errors before
|
|
60
|
+
every single dispatch (3 + 1 + 0 refuted across R13–R15); rounds without
|
|
61
|
+
it had returned the same errors as P0s.
|
|
62
|
+
5. **Check threshold reachability before dispatching.** Compute the maximum
|
|
63
|
+
achievable approve count from live slots; if the threshold is
|
|
64
|
+
unreachable, declare the exhaustion path up front (the frozen design's
|
|
65
|
+
own closing: findings exhausted → operator freeze declaration) instead of
|
|
66
|
+
discovering it at collect.
|
|
67
|
+
6. **Reference originals by path + sha256; do not transcribe.** Reviewers
|
|
68
|
+
read the repository; the artifact carries the manifest. Transcription
|
|
69
|
+
errors are undetectable and 100KB+ pastes rot.
|
|
70
|
+
|
|
71
|
+
Corollaries observed in the same loop: fix the *class*, and fix every copy —
|
|
72
|
+
a corrected lib comment whose refuted twin survives in a test file costs a
|
|
73
|
+
full round. When an author writes history into comments, the falsifier must
|
|
74
|
+
check the cited records; two of the loop's P0s were numbers copied from the
|
|
75
|
+
wrong document.
|
|
76
|
+
|
|
32
77
|
## Step 0 — Load reviewer characteristics (mandatory)
|
|
33
78
|
|
|
34
79
|
**Before invoking any reviewer**, fetch `multi_llm_reviewer_evaluation` via
|
|
@@ -878,26 +923,102 @@ after five consecutive non-substantive returns.
|
|
|
878
923
|
|
|
879
924
|
### Substance and the denominator
|
|
880
925
|
|
|
881
|
-
A reply
|
|
882
|
-
|
|
883
|
-
|
|
884
|
-
|
|
926
|
+
A reply counts toward the denominator when it **carries a verdict** and **says
|
|
927
|
+
something beyond it**. Both conditions, separately checked; a reply failing
|
|
928
|
+
either leaves the denominator rather than being counted, because otherwise it
|
|
929
|
+
raises the agreement everyone else must reach while contributing none of its own.
|
|
930
|
+
|
|
931
|
+
The two failures look nothing alike, and conflating them is why this rule was
|
|
932
|
+
rebuilt four times:
|
|
933
|
+
|
|
934
|
+
| Reply | Carries a verdict? | Says something? | Outcome |
|
|
935
|
+
|---|---|---|---|
|
|
936
|
+
| `APPROVE` | yes | no | `insubstantial` |
|
|
937
|
+
| `APPROVE. P0` | yes | no (a severity names no defect) | `insubstantial` |
|
|
938
|
+
| `I'll review this and give my assessment.` | **no** | yes | `no_verdict` |
|
|
939
|
+
| `競合状態あり` | no (no verdict word) | yes | `no_verdict` |
|
|
940
|
+
| `REJECT: P0 key logged in plaintext` | yes | yes | counts |
|
|
941
|
+
| `NO-GO` + findings | yes (see the vocabulary below) | yes | counts as REJECT |
|
|
942
|
+
|
|
943
|
+
The second row is the shape that retired Fable 5 — long enough to pass any
|
|
944
|
+
substance rule, and stating no judgement. Three rebuilds of the substance rule
|
|
945
|
+
missed it because they were all looking at the wrong half. A reply that states
|
|
946
|
+
no verdict used to be counted as a conservative `REVISE`, which blocked
|
|
947
|
+
convergence on a judgement its author never gave.
|
|
885
948
|
|
|
886
949
|
- The test is mechanical and has no setting: **does anything remain once the
|
|
887
950
|
verdict word is removed?** No model is asked to judge whether a review is good.
|
|
888
951
|
- It applies to every observer equally, the persona team included. A persona
|
|
889
952
|
submission whose findings are blank leaves the denominator exactly as a
|
|
890
|
-
subprocess reply would.
|
|
953
|
+
subprocess reply would — and so does a structured reply from any other slot.
|
|
954
|
+
A JSON reply is judged by **the words it carries, whatever keys they live
|
|
955
|
+
under**, not by the text it renders to and not by keys named in advance.
|
|
956
|
+
Judging it by rendered text counted its own key names as content: the residue
|
|
957
|
+
of `{"overall_verdict": "APPROVE", "findings": [], "reasoning": ""}` is three
|
|
958
|
+
words of schema and no review. Naming the keys instead threw away real
|
|
959
|
+
reviews: a REJECT whose findings used `description` was recorded as an empty
|
|
960
|
+
submission. **Answer in whatever shape you like** — nothing here requires a
|
|
961
|
+
particular key.
|
|
962
|
+
- **Open your reply with your verdict.** A structured submission states its
|
|
963
|
+
verdict as a field; a free-text reply states it in an `**Overall Verdict**`
|
|
964
|
+
header **on the first line, before anything else**. A header anywhere else is
|
|
965
|
+
not read as yours — it may be a sample you are discussing — so quoting verdict
|
|
966
|
+
headers further down is safe and needs no escaping.
|
|
967
|
+
|
|
968
|
+
Three rounds tried to tell a stated header from a displayed one by reading the
|
|
969
|
+
text more cleverly, and each attempt opened a hole in the direction that
|
|
970
|
+
passes: round 4 anchored the search to the start of a line, and a line inside
|
|
971
|
+
a fence starts a line; round 6 excluded fenced and blockquoted regions, and a
|
|
972
|
+
four-backtick fence closes on the inner three-backtick line, taking the rest
|
|
973
|
+
of the reply — including the real verdict — with it. A quotation and a
|
|
974
|
+
statement are the same characters, so no amount of reading decides between
|
|
975
|
+
them. Position does, because a quotation cannot be first unless the reviewer
|
|
976
|
+
chose to open with one.
|
|
977
|
+
|
|
978
|
+
A reply that opens with a preamble falls to a last-resort word scan over the
|
|
979
|
+
whole text. That is a guess, and it reads quotations along with everything
|
|
980
|
+
else — but it checks REJECT first, so it cannot turn a stated rejection into
|
|
981
|
+
a recorded approval.
|
|
982
|
+
- **The words that count as a verdict are one list, the same for every
|
|
983
|
+
observer.** `APPROVE / PASS / ACCEPT / LGTM / SHIP IT`,
|
|
984
|
+
`REJECT / FAIL / BLOCK / NO-GO / NACK / DENY / VETO`,
|
|
985
|
+
`REVISE / CHANGES REQUIRED / NEEDS WORK / NEEDS REVISION / REWORK`. Where a
|
|
986
|
+
reply carries more than one, the blocking reading wins (REJECT, then REVISE,
|
|
987
|
+
then APPROVE) — a reply saying two of them has qualified one judgement, not
|
|
988
|
+
stated two. Japanese verdict words are deliberately **not** in this list
|
|
989
|
+
(author decision, 2026-07-28): a Japanese review counts on its findings, but
|
|
990
|
+
its judgement must be stated in one of the words above.
|
|
891
991
|
- It is **not** a quality bar, and must not be tuned into one. `競合状態あり` is a
|
|
892
992
|
review. So is `REJECT: P0 private key logged in plaintext`. Earlier versions of
|
|
893
993
|
this rule measured length and, at one setting, discarded Japanese reviews
|
|
894
994
|
entirely — reviewers answer in the artifact's language, and this project's
|
|
895
995
|
artifacts are largely Japanese.
|
|
896
|
-
- What leaves the denominator is recorded with the reason
|
|
897
|
-
|
|
898
|
-
|
|
899
|
-
|
|
900
|
-
|
|
996
|
+
- What leaves the denominator is recorded with the reason, per reviewer and
|
|
997
|
+
again in `denominator_composition`. A slot that answered emptily and a slot
|
|
998
|
+
that never answered are both absent from the numerator and must never be
|
|
999
|
+
confused in the record, so `skip_reason` distinguishes: `insubstantial` (a
|
|
1000
|
+
reply that said nothing beyond its verdict), `no_verdict` (a reply that stated
|
|
1001
|
+
no judgement), `transport` (a call that was attempted and failed),
|
|
1002
|
+
`not_dispatched` or the dispatcher's own wording such as `dispatch_timeout`
|
|
1003
|
+
(a slot this system declined to run), and `declined` (a submission that stated
|
|
1004
|
+
SKIP for itself — reachable only by a caller that declares it, which nothing
|
|
1005
|
+
in this repository currently does). Only a token-shaped reason is carried
|
|
1006
|
+
through; anything else becomes `not_dispatched`, so a traceback or a sentence
|
|
1007
|
+
cannot land in a field documented as a small vocabulary. Nothing is defaulted:
|
|
1008
|
+
a row with no reason omits the field, because a default here states a cause
|
|
1009
|
+
the record does not know.
|
|
1010
|
+
|
|
1011
|
+
- **The escalation record is filled in only where it is wholly absent.** Pending
|
|
1012
|
+
state written before escalation existed carries no such field, and that one
|
|
1013
|
+
case is answered truthfully — a version with no escalation to offer cannot
|
|
1014
|
+
have been asked for it. A record that exists is carried as recorded, gaps and
|
|
1015
|
+
all: completing a partial one wrote `requested: false` beside
|
|
1016
|
+
`escalated: true`, a pair the producing code cannot emit. Silence about a key
|
|
1017
|
+
is legible to a reader; a value asserted where none was recorded is not.
|
|
1018
|
+
- The count beside these is `observers_reporting` — the observers that answered,
|
|
1019
|
+
which is not the number of slots the configuration named. That question is
|
|
1020
|
+
answered by `denominator_composition`, which lists every observer including
|
|
1021
|
+
the ones that never ran and why.
|
|
901
1022
|
|
|
902
1023
|
**What this rule cannot do.** It cannot tell a terse honest approval from a lazy
|
|
903
1024
|
one. `**Overall Verdict**: APPROVE / No findings` is the exact form the reviewer
|
|
@@ -1247,5 +1368,69 @@ Compression ratio: parallel agent raw → Assembly ≈ 2:1
|
|
|
1247
1368
|
the human, is the primary close — a reached ratio is neither necessary nor
|
|
1248
1369
|
sufficient on its own
|
|
1249
1370
|
|
|
1371
|
+
- Verdict determination hardened after a live failure (v3.7.1, 2026-07-27):
|
|
1372
|
+
round 4 of the escalation/persona implementation review recorded the persona
|
|
1373
|
+
team's REVISE as an APPROVE. The team's verdict was being re-derived by
|
|
1374
|
+
searching the assembled text, and a persona had quoted
|
|
1375
|
+
`{"overall_verdict": "APPROVE", ...}` inside a finding as an example of a
|
|
1376
|
+
defect; the search found the quotation before the `**Overall Verdict**:
|
|
1377
|
+
REVISE` on line 1. Three fixes: a structured submission now states its verdict
|
|
1378
|
+
as a field and nothing in its prose overrules it; the header is recognised
|
|
1379
|
+
only at the start of a line; and a JSON verdict is read by parsing a reply
|
|
1380
|
+
that *is* a JSON document rather than by scanning prose for an object. The
|
|
1381
|
+
residual limitation is recorded in § Substance and the denominator — the
|
|
1382
|
+
last-resort word heuristic still sees verdict words inside quotations, so
|
|
1383
|
+
state your verdict in the header when discussing reply shapes. Separately,
|
|
1384
|
+
INV-E2 was found to have been implemented by half: it asks that a counted
|
|
1385
|
+
reply *carry a verdict* and *have substance*, and only substance was checked,
|
|
1386
|
+
so an opening sentence with no judgement in it entered the denominator as a
|
|
1387
|
+
conservative REVISE — the exact shape that retired Fable 5, blocking
|
|
1388
|
+
convergence on nobody's verdict
|
|
1389
|
+
|
|
1390
|
+
- Three of round 4's own fixes reopened what they closed (v3.7.2, 2026-07-28):
|
|
1391
|
+
round 5 of the same implementation review found that each of the three
|
|
1392
|
+
verdict fixes recorded above had left a hole, and all three in the direction
|
|
1393
|
+
that passes. Start-of-line header matching does not exclude a quotation,
|
|
1394
|
+
because a line inside a fenced block starts a line — the header is now read
|
|
1395
|
+
from what the reviewer said, with fences and blockquotes excluded, and a
|
|
1396
|
+
reply that is entirely fenced is read whole because it is quoting nothing.
|
|
1397
|
+
Judging a structured reply by named keys (`reasoning`, `issue`) threw away a
|
|
1398
|
+
REJECT whose findings used `description` — the rule now asks whether anything
|
|
1399
|
+
was said in words, under any key. And the words that count as a verdict were
|
|
1400
|
+
two different lists, so `NO-GO` from a persona was a REJECT while `NO-GO`
|
|
1401
|
+
from an external slot was no judgement at all; there is now one list, with
|
|
1402
|
+
one precedence. Alongside these: `skip_reason` distinguishes five outcomes
|
|
1403
|
+
where it previously flattened three into `transport`; `total_configured` was
|
|
1404
|
+
renamed `observers_reporting` because it never counted what it claimed; the
|
|
1405
|
+
eased convergence rule after an exclusion now asks whether the denominator
|
|
1406
|
+
actually shrank rather than which reason fired; and a synchronous delegation
|
|
1407
|
+
writes the same storage layout as the parallel one, so the lock serialising
|
|
1408
|
+
two concurrent collects is no longer silently skipped
|
|
1409
|
+
|
|
1410
|
+
- Position replaces cleverness in verdict reading (v3.7.3, 2026-07-28): round 6
|
|
1411
|
+
was reviewed and rejected by five of six observers, all of them naming the
|
|
1412
|
+
same shape — the fixes of round 6 had reopened what they closed, in the
|
|
1413
|
+
direction that passes. Two were shipped defects: the header capture was
|
|
1414
|
+
written with `\s`, which includes the newline in Ruby, so it ran past the
|
|
1415
|
+
header into the following prose and recorded a stated APPROVE followed by
|
|
1416
|
+
"No blocking issues" as a REJECT; and the fence regex tracked neither fence
|
|
1417
|
+
length nor delimiter, so a four-backtick block quoting a three-backtick
|
|
1418
|
+
sample gave the quoted verdict the reviewer's vote. Rather than a fourth
|
|
1419
|
+
attempt at reading free text more carefully, the rule is now positional: the
|
|
1420
|
+
reviewer's verdict is the header the reply opens with, and the fence and
|
|
1421
|
+
blockquote machinery is deleted. The prompt asks for it there, and this test
|
|
1422
|
+
now pins that it does. Alongside: a partial escalation record is carried as
|
|
1423
|
+
recorded rather than completed with values the producing code cannot emit;
|
|
1424
|
+
only a token-shaped skip reason is carried into the record; and a row with no
|
|
1425
|
+
reason omits the field in the per-reviewer list as it already did in the
|
|
1426
|
+
composition.
|
|
1427
|
+
|
|
1428
|
+
The round's other lesson was about testing, and it is recorded here because
|
|
1429
|
+
it generalises past this SkillSet: **a regex pinned by one literal example is
|
|
1430
|
+
free everywhere else**. Round 6's own mutation pass replaced whole functions
|
|
1431
|
+
and killed 21 of 21; the review's mutation pass went after regex internals —
|
|
1432
|
+
fence markers, character classes, digit ranges, word boundaries — and 18 of
|
|
1433
|
+
27 survived. Mutate the inside of a pattern, not only the pattern.
|
|
1434
|
+
|
|
1250
1435
|
**Key insight**: Design reviews and implementation reviews find
|
|
1251
1436
|
**categorically different bugs**. Both phases are necessary.
|
|
@@ -689,6 +689,65 @@ assert "missing feedback_text_schema_version → reject" do
|
|
|
689
689
|
msg && msg.include?('feedback_text_schema_version missing')
|
|
690
690
|
end
|
|
691
691
|
|
|
692
|
+
# v0.7 record schema: the consumer reads reference_verdict from v2 records
|
|
693
|
+
# and still reads verdict from v1 records. Held here because reverting the
|
|
694
|
+
# read (parsed['verdict'] only) would leave every v2 review verdict-less
|
|
695
|
+
# with the rest of the suite green.
|
|
696
|
+
section "v0.7 reference_verdict consumption"
|
|
697
|
+
|
|
698
|
+
def drive_review_with(step, record)
|
|
699
|
+
fake_session = Object.new
|
|
700
|
+
def fake_session.cycle_number = 1
|
|
701
|
+
def fake_session.session_id = 'testsess'
|
|
702
|
+
ctx = Object.new
|
|
703
|
+
def ctx.derive(**_k) = nil
|
|
704
|
+
fake_session.define_singleton_method(:invocation_context) { ctx }
|
|
705
|
+
step.define_singleton_method(:invoke_tool) do |*_a, **_k|
|
|
706
|
+
[{ text: JSON.generate(record) }]
|
|
707
|
+
end
|
|
708
|
+
step.send(:run_multi_llm_review, fake_session,
|
|
709
|
+
{ 'summary' => 's', 'task_json' => { 'steps' => [] } },
|
|
710
|
+
{ level: 'high', signals: [] }, {})
|
|
711
|
+
end
|
|
712
|
+
|
|
713
|
+
assert "v2 record (reference_verdict, no top-level verdict) → verdict read" do
|
|
714
|
+
out = drive_review_with(step.dup, {
|
|
715
|
+
'status' => 'ok', 'verdict_schema_version' => 2,
|
|
716
|
+
'feedback_text_schema_version' => 1,
|
|
717
|
+
'reference_verdict' => 'REVISE', 'convergence' => {},
|
|
718
|
+
'aggregated_findings' => [], 'llm_calls' => 3, 'reviews' => [],
|
|
719
|
+
'feedback_text' => 'x'
|
|
720
|
+
})
|
|
721
|
+
out[:verdict] == 'REVISE'
|
|
722
|
+
end
|
|
723
|
+
|
|
724
|
+
assert "v1 record (verdict) → still read" do
|
|
725
|
+
out = drive_review_with(step.dup, {
|
|
726
|
+
'status' => 'ok', 'verdict_schema_version' => 1,
|
|
727
|
+
'feedback_text_schema_version' => 1,
|
|
728
|
+
'verdict' => 'APPROVE', 'convergence' => {},
|
|
729
|
+
'aggregated_findings' => [], 'llm_calls' => 3, 'reviews' => [],
|
|
730
|
+
'feedback_text' => nil
|
|
731
|
+
})
|
|
732
|
+
out[:verdict] == 'APPROVE'
|
|
733
|
+
end
|
|
734
|
+
|
|
735
|
+
# v0.7 record schema (2026-08-01): the boundary is max and max+1, not max
|
|
736
|
+
# and 99 — a probe of 99 cannot tell SUPPORTED=1 from SUPPORTED=2, which is
|
|
737
|
+
# exactly the revert these tests exist to catch.
|
|
738
|
+
assert "v0.7 schema (v=2) → accepted at the boundary" do
|
|
739
|
+
step.send(:schema_version_check, {
|
|
740
|
+
'verdict_schema_version' => 2, 'feedback_text_schema_version' => 1
|
|
741
|
+
}).nil?
|
|
742
|
+
end
|
|
743
|
+
|
|
744
|
+
assert "v=3 (max+1) → reject (fail-closed at the boundary)" do
|
|
745
|
+
msg = step.send(:schema_version_check, {
|
|
746
|
+
'verdict_schema_version' => 3, 'feedback_text_schema_version' => 1
|
|
747
|
+
})
|
|
748
|
+
msg && msg.include?('newer than supported')
|
|
749
|
+
end
|
|
750
|
+
|
|
692
751
|
assert "newer verdict_schema_version → reject (fail-closed)" do
|
|
693
752
|
msg = step.send(:schema_version_check, {
|
|
694
753
|
'verdict_schema_version' => 99, 'feedback_text_schema_version' => 1
|
|
@@ -696,6 +755,13 @@ assert "newer verdict_schema_version → reject (fail-closed)" do
|
|
|
696
755
|
msg && msg.include?('newer than supported')
|
|
697
756
|
end
|
|
698
757
|
|
|
758
|
+
assert "feedback_text v=2 (max+1) → reject (boundary, not v99)" do
|
|
759
|
+
msg = step.send(:schema_version_check, {
|
|
760
|
+
'verdict_schema_version' => 2, 'feedback_text_schema_version' => 2
|
|
761
|
+
})
|
|
762
|
+
msg && msg.include?('newer than supported')
|
|
763
|
+
end
|
|
764
|
+
|
|
699
765
|
assert "newer feedback_text_schema_version → reject" do
|
|
700
766
|
msg = step.send(:schema_version_check, {
|
|
701
767
|
'verdict_schema_version' => 1, 'feedback_text_schema_version' => 99
|
|
@@ -1696,7 +1696,12 @@ module KairosMcp
|
|
|
1696
1696
|
aggregated_findings: [], feedback_text: nil }
|
|
1697
1697
|
else
|
|
1698
1698
|
{
|
|
1699
|
-
|
|
1699
|
+
# v0.7 record schema (verdict_schema_version 2) renamed the
|
|
1700
|
+
# top-level conclusion column to reference_verdict (INV-R2: a
|
|
1701
|
+
# recorded reference value, not the run's conclusion). v1
|
|
1702
|
+
# records still carry 'verdict'; read whichever the record
|
|
1703
|
+
# speaks.
|
|
1704
|
+
verdict: parsed['reference_verdict'] || parsed['verdict'],
|
|
1700
1705
|
convergence: parsed['convergence'],
|
|
1701
1706
|
aggregated_findings: (parsed['aggregated_findings'] || []).map { |f|
|
|
1702
1707
|
f.transform_keys(&:to_sym)
|
|
@@ -1718,7 +1723,9 @@ module KairosMcp
|
|
|
1718
1723
|
|
|
1719
1724
|
# Phase 12 §3.10 fail-closed schema versioning.
|
|
1720
1725
|
# Returns nil if response schema is acceptable, else a string reason.
|
|
1721
|
-
|
|
1726
|
+
# 2 = v0.7 record schema (2026-08-01): top-level verdict became
|
|
1727
|
+
# reference_verdict; the read above handles both generations.
|
|
1728
|
+
SUPPORTED_VERDICT_SCHEMA_VERSION = 2
|
|
1722
1729
|
SUPPORTED_FEEDBACK_TEXT_SCHEMA_VERSION = 1
|
|
1723
1730
|
|
|
1724
1731
|
def schema_version_check(parsed)
|
|
@@ -183,18 +183,45 @@ module KairosMcp
|
|
|
183
183
|
result_text = data['result'] || ''
|
|
184
184
|
tool_use = extract_tool_use(result_text)
|
|
185
185
|
usage = data['usage'] || {}
|
|
186
|
+
model_usage = data['modelUsage'] || {}
|
|
187
|
+
|
|
188
|
+
# The answering model is the entry that produced the output tokens,
|
|
189
|
+
# not the first hash key. The CLI places a small internal call
|
|
190
|
+
# (claude-haiku, ~18 output tokens) beside the main call unless
|
|
191
|
+
# CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC suppresses it, so under
|
|
192
|
+
# the worker's stripped environment the envelope carries two keys,
|
|
193
|
+
# and `keys.first` attributed the reply to whichever entry the CLI
|
|
194
|
+
# inserted first. Reproduced 2/2 with byte-identical prompts
|
|
195
|
+
# (2026-07-31); this misread was the root cause of every model
|
|
196
|
+
# divergence this transport had recorded. See L2
|
|
197
|
+
# mlr_v07_design_inputs_and_haiku_root_cause_20260731.
|
|
198
|
+
observed = model_usage.max_by { |_m, u| (u || {})['outputTokens'].to_i }&.first
|
|
186
199
|
|
|
187
200
|
{
|
|
188
201
|
'content' => tool_use ? nil : result_text,
|
|
189
202
|
'tool_use' => tool_use,
|
|
190
203
|
'stop_reason' => tool_use ? 'tool_use' : map_stop_reason(data['stop_reason']),
|
|
191
|
-
'model' => requested_model ||
|
|
204
|
+
'model' => requested_model || observed || 'claude_code',
|
|
192
205
|
# What the CLI reports as having answered, when it reports it.
|
|
193
206
|
# Kept separate from 'model' (which echoes the request) so callers
|
|
194
207
|
# can tell a request from an observation and notice when the two
|
|
195
|
-
# disagree
|
|
196
|
-
#
|
|
197
|
-
|
|
208
|
+
# disagree. The four divergences recorded before the fix above
|
|
209
|
+
# (R6/R8/R10/R13, 2026-07) were all keys.first misreads — the main
|
|
210
|
+
# call was claude-opus-4-6 every time, and the R13 reply's odd
|
|
211
|
+
# content (identifiers that exist nowhere in the reviewed code) is
|
|
212
|
+
# explained by the sandboxed slot reading no repository, not by a
|
|
213
|
+
# different model answering. A divergence observed after the
|
|
214
|
+
# 2026-07-31 fix has no known benign explanation and is worth
|
|
215
|
+
# investigating.
|
|
216
|
+
'model_observed' => observed,
|
|
217
|
+
# Diagnostic envelope, previously discarded. All four divergence
|
|
218
|
+
# incidents above were undiagnosable from the record because the
|
|
219
|
+
# fields that say what happened did not survive this method.
|
|
220
|
+
# Consumers that persist reviews should carry these through.
|
|
221
|
+
'model_usage' => model_usage.empty? ? nil : model_usage,
|
|
222
|
+
'api_error_status' => data['api_error_status'],
|
|
223
|
+
'fast_mode_state' => data['fast_mode_state'],
|
|
224
|
+
'terminal_reason' => data['terminal_reason'],
|
|
198
225
|
'input_tokens' => usage['input_tokens'],
|
|
199
226
|
'output_tokens' => usage['output_tokens']
|
|
200
227
|
}
|
|
@@ -0,0 +1,106 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
# What parse_response reads from the CLI envelope, pinned after the
|
|
4
|
+
# claude_cli_opus4.6 divergence investigation (2026-07-31).
|
|
5
|
+
#
|
|
6
|
+
# Four times in the multi_llm_review loop (R6/R8/R10/R13) the record showed a
|
|
7
|
+
# slot requesting claude-opus-4-6 answered by claude-haiku-4-5. The 2026-07-31
|
|
8
|
+
# investigation reproduced the cause 2/2: the CLI places a small internal call
|
|
9
|
+
# beside the main one, the envelope carries two keys, and the old `keys.first`
|
|
10
|
+
# read attributed the reply to the internal call — the main call was
|
|
11
|
+
# claude-opus-4-6 every time. The record could not say so, because this method
|
|
12
|
+
# also discarded every diagnostic field the envelope carries (modelUsage
|
|
13
|
+
# breakdown, api_error_status, fast_mode_state, terminal_reason). These tests
|
|
14
|
+
# hold the two fixes: the answering model is chosen by output tokens rather
|
|
15
|
+
# than hash order, and the diagnostic fields survive into the response.
|
|
16
|
+
|
|
17
|
+
require 'minitest/autorun'
|
|
18
|
+
require_relative '../lib/llm_client/adapter'
|
|
19
|
+
require_relative '../lib/llm_client/claude_code_adapter'
|
|
20
|
+
|
|
21
|
+
module KairosMcp
|
|
22
|
+
module SkillSets
|
|
23
|
+
module LlmClient
|
|
24
|
+
class TestClaudeCodeAdapterParse < Minitest::Test
|
|
25
|
+
def setup
|
|
26
|
+
@adapter = ClaudeCodeAdapter.new({ 'timeout_seconds' => 30 })
|
|
27
|
+
end
|
|
28
|
+
|
|
29
|
+
def parse(payload, requested_model: 'claude-opus-4-6')
|
|
30
|
+
@adapter.send(:parse_response, JSON.generate(payload), requested_model: requested_model)
|
|
31
|
+
end
|
|
32
|
+
|
|
33
|
+
def envelope(model_usage:, **extra)
|
|
34
|
+
{
|
|
35
|
+
'type' => 'result', 'is_error' => false, 'result' => 'ok',
|
|
36
|
+
'stop_reason' => 'end_turn',
|
|
37
|
+
'usage' => { 'input_tokens' => 10, 'output_tokens' => 20 },
|
|
38
|
+
'modelUsage' => model_usage
|
|
39
|
+
}.merge(extra)
|
|
40
|
+
end
|
|
41
|
+
|
|
42
|
+
def test_a_single_model_envelope_is_that_model
|
|
43
|
+
out = parse(envelope(model_usage: {
|
|
44
|
+
'claude-opus-4-6' => { 'outputTokens' => 116 }
|
|
45
|
+
}))
|
|
46
|
+
|
|
47
|
+
assert_equal 'claude-opus-4-6', out['model_observed']
|
|
48
|
+
end
|
|
49
|
+
|
|
50
|
+
# The regression the old `keys.first` invited: when the envelope
|
|
51
|
+
# carries a second model beside the main call (reproduced 2/2 under
|
|
52
|
+
# the worker's environment), insertion order must not decide which
|
|
53
|
+
# one "answered". The output tokens do.
|
|
54
|
+
def test_the_answering_model_is_the_one_that_wrote_the_output
|
|
55
|
+
out = parse(envelope(model_usage: {
|
|
56
|
+
'claude-haiku-4-5-20251001' => { 'outputTokens' => 12 },
|
|
57
|
+
'claude-opus-4-6' => { 'outputTokens' => 2048 }
|
|
58
|
+
}))
|
|
59
|
+
|
|
60
|
+
assert_equal 'claude-opus-4-6', out['model_observed']
|
|
61
|
+
end
|
|
62
|
+
|
|
63
|
+
def test_a_divergent_answer_is_observed_as_itself
|
|
64
|
+
out = parse(envelope(model_usage: {
|
|
65
|
+
'claude-haiku-4-5-20251001' => { 'outputTokens' => 900 }
|
|
66
|
+
}))
|
|
67
|
+
|
|
68
|
+
assert_equal 'claude-haiku-4-5-20251001', out['model_observed']
|
|
69
|
+
assert_equal 'claude-opus-4-6', out['model']
|
|
70
|
+
refute_equal out['model'], out['model_observed']
|
|
71
|
+
end
|
|
72
|
+
|
|
73
|
+
def test_an_envelope_without_model_usage_observes_nothing
|
|
74
|
+
out = parse(envelope(model_usage: nil))
|
|
75
|
+
|
|
76
|
+
assert_nil out['model_observed']
|
|
77
|
+
assert_nil out['model_usage']
|
|
78
|
+
assert_equal 'claude-opus-4-6', out['model']
|
|
79
|
+
end
|
|
80
|
+
|
|
81
|
+
def test_the_diagnostic_fields_survive_into_the_response
|
|
82
|
+
out = parse(envelope(
|
|
83
|
+
model_usage: { 'claude-haiku-4-5-20251001' => { 'outputTokens' => 900 } },
|
|
84
|
+
'api_error_status' => 429,
|
|
85
|
+
'fast_mode_state' => 'off',
|
|
86
|
+
'terminal_reason' => 'completed'
|
|
87
|
+
))
|
|
88
|
+
|
|
89
|
+
assert_equal({ 'claude-haiku-4-5-20251001' => { 'outputTokens' => 900 } }, out['model_usage'])
|
|
90
|
+
assert_equal 429, out['api_error_status']
|
|
91
|
+
assert_equal 'off', out['fast_mode_state']
|
|
92
|
+
assert_equal 'completed', out['terminal_reason']
|
|
93
|
+
end
|
|
94
|
+
|
|
95
|
+
def test_a_missing_output_tokens_entry_does_not_crash_the_choice
|
|
96
|
+
out = parse(envelope(model_usage: {
|
|
97
|
+
'claude-opus-4-6' => nil,
|
|
98
|
+
'claude-haiku-4-5-20251001' => { 'outputTokens' => 5 }
|
|
99
|
+
}))
|
|
100
|
+
|
|
101
|
+
assert_equal 'claude-haiku-4-5-20251001', out['model_observed']
|
|
102
|
+
end
|
|
103
|
+
end
|
|
104
|
+
end
|
|
105
|
+
end
|
|
106
|
+
end
|
|
@@ -23,7 +23,14 @@ module KairosMcp
|
|
|
23
23
|
# the verdict shape itself (separate from feedback_text_schema_version,
|
|
24
24
|
# which lives in FeedbackFormatter). Bumped independently when verdict
|
|
25
25
|
# JSON contract changes (e.g., new field, semantic redefinition).
|
|
26
|
-
|
|
26
|
+
#
|
|
27
|
+
# 2 = v0.7 record schema (design frozen 2026-08-01): the top-level
|
|
28
|
+
# conclusion column is gone — `verdict` became `reference_verdict`
|
|
29
|
+
# (INV-R2), the composition rows carry seat marks, refusals and
|
|
30
|
+
# transport diagnostics travel with the record. Readers tell old
|
|
31
|
+
# records from new by this number; old records are not rewritten
|
|
32
|
+
# (design §4, no retroactivity).
|
|
33
|
+
VERDICT_SCHEMA_VERSION = 2
|
|
27
34
|
|
|
28
35
|
# @param artifact_content [String] sanitized at boundary; raw passthrough is caller responsibility
|
|
29
36
|
# @param artifact_name [String]
|
|
@@ -161,11 +168,32 @@ module KairosMcp
|
|
|
161
168
|
}
|
|
162
169
|
end
|
|
163
170
|
|
|
171
|
+
# INV-R7 by_reference delivery: the same canonical framing with a
|
|
172
|
+
# reference manifest where the artifact body would be. The system
|
|
173
|
+
# prompt is identical to the inline one — only the user message
|
|
174
|
+
# differs, and only in the artifact block. sha256 is computed by the
|
|
175
|
+
# caller over the raw submitted content, so a seat that reads the file
|
|
176
|
+
# can check it is reviewing what was submitted.
|
|
177
|
+
def self.build_reference_prompts(artifact_path:, artifact_sha256:,
|
|
178
|
+
artifact_name:, review_type:,
|
|
179
|
+
review_context: 'independent',
|
|
180
|
+
review_round: 1, prior_findings: nil)
|
|
181
|
+
system_prompt = PromptBuilder.build_system_prompt(review_type, review_context: review_context)
|
|
182
|
+
messages = PromptBuilder.build_messages(
|
|
183
|
+
artifact_name: artifact_name,
|
|
184
|
+
review_type: review_type,
|
|
185
|
+
review_round: review_round,
|
|
186
|
+
prior_findings: prior_findings,
|
|
187
|
+
artifact_reference: { path: artifact_path, sha256: artifact_sha256 }
|
|
188
|
+
)
|
|
189
|
+
{ system_prompt: system_prompt, messages: messages }
|
|
190
|
+
end
|
|
191
|
+
|
|
164
192
|
def self.aggregation_instructions(review_type, review_round)
|
|
165
193
|
<<~INST.strip
|
|
166
194
|
After collecting all reviewer responses, aggregate as follows:
|
|
167
195
|
1. Parse each response for verdict {APPROVE, REVISE, REJECT}.
|
|
168
|
-
2.
|
|
196
|
+
2. Compute the reference tally (e.g., 3/N APPROVE → reference APPROVE; any REJECT → REVISE; below quorum → INSUFFICIENT). It is a recorded reference value, not the run's conclusion.
|
|
169
197
|
3. Merge findings, sorted by severity (P0 first), de-dup by issue text.
|
|
170
198
|
4. For round #{review_round} #{review_type} reviews, prior findings should be verified as CLOSED/NEEDS_MORE_WORK/REOPENED.
|
|
171
199
|
INST
|