kairos-chain 3.55.0 → 3.57.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (33) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +35 -0
  3. data/lib/kairos_mcp/version.rb +1 -1
  4. data/templates/knowledge/multi_llm_review_workflow/multi_llm_review_workflow.md +196 -11
  5. data/templates/skillsets/agent/test/test_agent_complexity_review.rb +66 -0
  6. data/templates/skillsets/agent/tools/agent_step.rb +9 -2
  7. data/templates/skillsets/llm_client/lib/llm_client/claude_code_adapter.rb +31 -4
  8. data/templates/skillsets/llm_client/skillset.json +1 -1
  9. data/templates/skillsets/llm_client/test/test_claude_code_adapter_parse.rb +106 -0
  10. data/templates/skillsets/multi_llm_review/lib/multi_llm_review/build_review_bundle.rb +30 -2
  11. data/templates/skillsets/multi_llm_review/lib/multi_llm_review/consensus.rb +331 -81
  12. data/templates/skillsets/multi_llm_review/lib/multi_llm_review/dispatcher.rb +31 -2
  13. data/templates/skillsets/multi_llm_review/lib/multi_llm_review/observer_set.rb +43 -3
  14. data/templates/skillsets/multi_llm_review/lib/multi_llm_review/pending_state.rb +128 -16
  15. data/templates/skillsets/multi_llm_review/lib/multi_llm_review/persona_assembly.rb +105 -32
  16. data/templates/skillsets/multi_llm_review/lib/multi_llm_review/prompt_builder.rb +40 -8
  17. data/templates/skillsets/multi_llm_review/lib/multi_llm_review/review_serializer.rb +62 -0
  18. data/templates/skillsets/multi_llm_review/lib/multi_llm_review/verdict_vocabulary.rb +167 -0
  19. data/templates/skillsets/multi_llm_review/skillset.json +8 -4
  20. data/templates/skillsets/multi_llm_review/test/test_multi_llm_review.rb +532 -65
  21. data/templates/skillsets/multi_llm_review/test/test_mutation_survivors.rb +1124 -0
  22. data/templates/skillsets/multi_llm_review/test/test_observer_set.rb +147 -9
  23. data/templates/skillsets/multi_llm_review/test/test_observer_set_seams.rb +775 -78
  24. data/templates/skillsets/multi_llm_review/test/test_pending_state_v3.rb +122 -2
  25. data/templates/skillsets/multi_llm_review/test/test_tool_wiring.rb +429 -8
  26. data/templates/skillsets/multi_llm_review/tools/multi_llm_review.rb +422 -128
  27. data/templates/skillsets/multi_llm_review/tools/multi_llm_review_collect.rb +108 -32
  28. data/templates/skillsets/synoptis/lib/synoptis/attestation_engine.rb +35 -2
  29. data/templates/skillsets/synoptis/lib/synoptis/proof_envelope.rb +55 -3
  30. data/templates/skillsets/synoptis/lib/synoptis/tool_helpers.rb +9 -2
  31. data/templates/skillsets/synoptis/lib/synoptis/verifier.rb +32 -5
  32. data/templates/skillsets/synoptis/tools/attestation_verify.rb +1 -1
  33. metadata +4 -1
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 2291c491a077a0cfdce3d77bdd3709e8e6278fc9eb656a02f3de5508e4f66580
4
- data.tar.gz: 96fa06b0195df3167b3113c54125a093f31f4bed10955c044a2ef218fe6d9278
3
+ metadata.gz: 8439763123a17e68d238ebc46369dcc36e6379172752809dd6736fa18f3c5ec4
4
+ data.tar.gz: 76fe818c7261e464411f54bee9a114bc7eb0ecc025042ef264f15bd591f2ca36
5
5
  SHA512:
6
- metadata.gz: 72064bd0e1bde0516e8dd2b90adcbb9a85f3d079f71507b9f77210985e2f995fc9a1e8af066eee19835dda664a707f09ae8f6420ecfb6135dec62aa92449b756
7
- data.tar.gz: 2553da41e8590b33c49f8e0d3526d37ac80bc8f742bab79af45cb5798c7a425ad978e66099a47454f9ca1ade3cbce310a798b7a9d1400eedac7a781d3cfdd745
6
+ metadata.gz: 2385335588b313d3910519459a79a1325f921fa1f6d1a2b215f0b75144c383b14b90a95f1d9edf97999df6cdc58d95155e653f13007217299f35df097891f38a
7
+ data.tar.gz: 160f831a399a60089d217e2fe198dfe60692797649b55c2d8b4fbd0d0649512c6f601a75dfc41de713b9bf946be83485f01c289252018c87d33dbbb06221c1ec
data/CHANGELOG.md CHANGED
@@ -4,6 +4,41 @@ All notable changes to the `kairos-chain` gem will be documented in this file.
4
4
 
5
5
  This project follows [Semantic Versioning](https://semver.org/).
6
6
 
7
+ ## [3.56.0] - 2026-07-31
8
+
9
+ ### Changed
10
+
11
+ - **`multi_llm_review` SkillSet 0.7.0 — verdict reading hardened across both
12
+ observer paths** (review rounds R10–R15, frozen by operator declaration
13
+ 2026-07-31 after three consecutive rounds with zero deployment-grounded
14
+ findings).
15
+ - The verdict vocabulary is one table (`VerdictVocabulary::WORDS`); the
16
+ prose-search and whole-value patterns are built from it, differing only
17
+ in anchoring and separator strictness. The rebuilt regexes are
18
+ source-identical to the previous literals.
19
+ - Verdict *reading* is `stated` (whole-value, anchored) on every path.
20
+ `VerdictVocabulary.classify`, `PersonaAssembly.normalize_verdict`, and
21
+ the dead constants `VERDICT_PATTERNS` / `ALLOWED_VERDICTS` /
22
+ `*_ALIASES` are deleted; their absence is pinned by tests.
23
+ - A persona verdict field that is not a verdict is **refused at
24
+ submission** (`PersonaAssembly.validate!` raises ArgumentError) instead
25
+ of being word-searched (≤ R12) or defaulted to REVISE (R13). The collect
26
+ tool validates before consuming, so a refused submission leaves the
27
+ pending token collectable and the corrected submission carries the vote.
28
+ - Declared verdicts are carried in canonical case (`extract_verdict`
29
+ merges the admitted form), so a lower-case declaration can no longer sit
30
+ in the denominator without counting.
31
+ - Test suite grew 401 → 478 runs / 1676 assertions, including a
32
+ mutation-survivor file distilled from three generative sweeps.
33
+ - **`llm_client` SkillSet 0.2.0 — CLI envelope diagnostics preserved.**
34
+ `claude_code_adapter` now picks `model_observed` by output tokens rather
35
+ than hash order, and passes `model_usage`, `api_error_status`,
36
+ `fast_mode_state` and `terminal_reason` through to the response. Four
37
+ model-divergence incidents (a slot requesting opus-4-6 answered by
38
+ haiku-4-5) were undiagnosable because these fields were discarded; the
39
+ divergence itself was confirmed real (the divergent reply cited
40
+ identifiers that exist nowhere in the reviewed code).
41
+
7
42
  ## [3.55.0] - 2026-07-27
8
43
 
9
44
  ### Added
@@ -1,4 +1,4 @@
1
1
  module KairosMcp
2
- VERSION = "3.55.0"
2
+ VERSION = "3.57.0"
3
3
  CHANGELOG_URL = "https://github.com/masaomi/KairosChain_2026/blob/main/CHANGELOG.md"
4
4
  end
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: multi_llm_review_workflow
3
3
  description: "Multi-LLM review methodology and execution — workflow pattern, CLI tooling, consensus analysis, Persona Assembly. Applicable to design, implementation, documentation, or any artifact."
4
- version: "3.7.0"
4
+ version: "3.8.0"
5
5
  tags:
6
6
  - workflow
7
7
  - review
@@ -29,6 +29,51 @@ This skill covers:
29
29
  For **WHO** (which LLM is good at what), see: `multi_llm_reviewer_evaluation`
30
30
  For **development lifecycle** (design → implement → verify), see: `design_to_implementation_workflow`
31
31
 
32
+ ## Step -1 — Pre-declared review spec (loop hygiene)
33
+
34
+ > Validation scope: these rules were derived from one non-converging loop
35
+ > (multi_llm_review R10–R15, 2026-07) where they took the P0 count from 10 to
36
+ > 2 in two rounds and closed the loop in three. They are instance practice
37
+ > until reproduced on a second, non-self-referential subject; treat the
38
+ > numbers below as one loop's evidence, not a law.
39
+
40
+ Before dispatching round 1 — and again whenever the review TARGET changes —
41
+ write a review spec and declare it frozen for the round:
42
+
43
+ 1. **Pre-declare the pass condition** (per `loop_validation`: spec before
44
+ judgement, fail-closed). State what APPROVE requires. A loop whose target
45
+ drifted (e.g. from an implementation to the instrument that measures it)
46
+ without a re-declared spec is structurally non-converging: a 300-claim
47
+ artifact at any realistic per-claim error rate yields double-digit
48
+ findings every round regardless of quality.
49
+ 2. **Split target from appendix.** Only the shipping deliverable is
50
+ P0-eligible. Instruments, sweep logs, classification tables are appendix:
51
+ findings against them are advisory and go to the queue. This is what
52
+ collapsed the claim surface from ~350 to ~30.
53
+ 3. **Cap fixes per round (≤5)** and write one line per fix: *what this fix
54
+ newly claims* (values pinned, ranges narrowed, failure visibility
55
+ changed). A fix that cannot state its new claims is doing more than the
56
+ finding asked.
57
+ 4. **Pre-flight falsifier.** Before dispatch, one agent whose only job is to
58
+ refute every factual claim in the spec and artifact — especially numbers
59
+ and "X does not exist" claims. In this loop it caught real errors before
60
+ every single dispatch (3 + 1 + 0 refuted across R13–R15); rounds without
61
+ it had returned the same errors as P0s.
62
+ 5. **Check threshold reachability before dispatching.** Compute the maximum
63
+ achievable approve count from live slots; if the threshold is
64
+ unreachable, declare the exhaustion path up front (the frozen design's
65
+ own closing: findings exhausted → operator freeze declaration) instead of
66
+ discovering it at collect.
67
+ 6. **Reference originals by path + sha256; do not transcribe.** Reviewers
68
+ read the repository; the artifact carries the manifest. Transcription
69
+ errors are undetectable and 100KB+ pastes rot.
70
+
71
+ Corollaries observed in the same loop: fix the *class*, and fix every copy —
72
+ a corrected lib comment whose refuted twin survives in a test file costs a
73
+ full round. When an author writes history into comments, the falsifier must
74
+ check the cited records; two of the loop's P0s were numbers copied from the
75
+ wrong document.
76
+
32
77
  ## Step 0 — Load reviewer characteristics (mandatory)
33
78
 
34
79
  **Before invoking any reviewer**, fetch `multi_llm_reviewer_evaluation` via
@@ -878,26 +923,102 @@ after five consecutive non-substantive returns.
878
923
 
879
924
  ### Substance and the denominator
880
925
 
881
- A reply that carries a verdict and nothing else does not count as a review. It
882
- leaves the denominator instead of being counted as agreement — otherwise it
883
- raises the agreement everyone else must reach while contributing none of its own,
884
- which is the failure mode that retired Fable 5 from the roster.
926
+ A reply counts toward the denominator when it **carries a verdict** and **says
927
+ something beyond it**. Both conditions, separately checked; a reply failing
928
+ either leaves the denominator rather than being counted, because otherwise it
929
+ raises the agreement everyone else must reach while contributing none of its own.
930
+
931
+ The two failures look nothing alike, and conflating them is why this rule was
932
+ rebuilt four times:
933
+
934
+ | Reply | Carries a verdict? | Says something? | Outcome |
935
+ |---|---|---|---|
936
+ | `APPROVE` | yes | no | `insubstantial` |
937
+ | `APPROVE. P0` | yes | no (a severity names no defect) | `insubstantial` |
938
+ | `I'll review this and give my assessment.` | **no** | yes | `no_verdict` |
939
+ | `競合状態あり` | no (no verdict word) | yes | `no_verdict` |
940
+ | `REJECT: P0 key logged in plaintext` | yes | yes | counts |
941
+ | `NO-GO` + findings | yes (see the vocabulary below) | yes | counts as REJECT |
942
+
943
+ The second row is the shape that retired Fable 5 — long enough to pass any
944
+ substance rule, and stating no judgement. Three rebuilds of the substance rule
945
+ missed it because they were all looking at the wrong half. A reply that states
946
+ no verdict used to be counted as a conservative `REVISE`, which blocked
947
+ convergence on a judgement its author never gave.
885
948
 
886
949
  - The test is mechanical and has no setting: **does anything remain once the
887
950
  verdict word is removed?** No model is asked to judge whether a review is good.
888
951
  - It applies to every observer equally, the persona team included. A persona
889
952
  submission whose findings are blank leaves the denominator exactly as a
890
- subprocess reply would.
953
+ subprocess reply would — and so does a structured reply from any other slot.
954
+ A JSON reply is judged by **the words it carries, whatever keys they live
955
+ under**, not by the text it renders to and not by keys named in advance.
956
+ Judging it by rendered text counted its own key names as content: the residue
957
+ of `{"overall_verdict": "APPROVE", "findings": [], "reasoning": ""}` is three
958
+ words of schema and no review. Naming the keys instead threw away real
959
+ reviews: a REJECT whose findings used `description` was recorded as an empty
960
+ submission. **Answer in whatever shape you like** — nothing here requires a
961
+ particular key.
962
+ - **Open your reply with your verdict.** A structured submission states its
963
+ verdict as a field; a free-text reply states it in an `**Overall Verdict**`
964
+ header **on the first line, before anything else**. A header anywhere else is
965
+ not read as yours — it may be a sample you are discussing — so quoting verdict
966
+ headers further down is safe and needs no escaping.
967
+
968
+ Three rounds tried to tell a stated header from a displayed one by reading the
969
+ text more cleverly, and each attempt opened a hole in the direction that
970
+ passes: round 4 anchored the search to the start of a line, and a line inside
971
+ a fence starts a line; round 6 excluded fenced and blockquoted regions, and a
972
+ four-backtick fence closes on the inner three-backtick line, taking the rest
973
+ of the reply — including the real verdict — with it. A quotation and a
974
+ statement are the same characters, so no amount of reading decides between
975
+ them. Position does, because a quotation cannot be first unless the reviewer
976
+ chose to open with one.
977
+
978
+ A reply that opens with a preamble falls to a last-resort word scan over the
979
+ whole text. That is a guess, and it reads quotations along with everything
980
+ else — but it checks REJECT first, so it cannot turn a stated rejection into
981
+ a recorded approval.
982
+ - **The words that count as a verdict are one list, the same for every
983
+ observer.** `APPROVE / PASS / ACCEPT / LGTM / SHIP IT`,
984
+ `REJECT / FAIL / BLOCK / NO-GO / NACK / DENY / VETO`,
985
+ `REVISE / CHANGES REQUIRED / NEEDS WORK / NEEDS REVISION / REWORK`. Where a
986
+ reply carries more than one, the blocking reading wins (REJECT, then REVISE,
987
+ then APPROVE) — a reply saying two of them has qualified one judgement, not
988
+ stated two. Japanese verdict words are deliberately **not** in this list
989
+ (author decision, 2026-07-28): a Japanese review counts on its findings, but
990
+ its judgement must be stated in one of the words above.
891
991
  - It is **not** a quality bar, and must not be tuned into one. `競合状態あり` is a
892
992
  review. So is `REJECT: P0 private key logged in plaintext`. Earlier versions of
893
993
  this rule measured length and, at one setting, discarded Japanese reviews
894
994
  entirely — reviewers answer in the artifact's language, and this project's
895
995
  artifacts are largely Japanese.
896
- - What leaves the denominator is recorded with the reason: `skip_reason` reads
897
- `insubstantial` or `transport`, per reviewer and again in
898
- `denominator_composition`. A slot that answered emptily and a slot that never
899
- answered are both absent from the numerator and must never be confused in the
900
- record.
996
+ - What leaves the denominator is recorded with the reason, per reviewer and
997
+ again in `denominator_composition`. A slot that answered emptily and a slot
998
+ that never answered are both absent from the numerator and must never be
999
+ confused in the record, so `skip_reason` distinguishes: `insubstantial` (a
1000
+ reply that said nothing beyond its verdict), `no_verdict` (a reply that stated
1001
+ no judgement), `transport` (a call that was attempted and failed),
1002
+ `not_dispatched` or the dispatcher's own wording such as `dispatch_timeout`
1003
+ (a slot this system declined to run), and `declined` (a submission that stated
1004
+ SKIP for itself — reachable only by a caller that declares it, which nothing
1005
+ in this repository currently does). Only a token-shaped reason is carried
1006
+ through; anything else becomes `not_dispatched`, so a traceback or a sentence
1007
+ cannot land in a field documented as a small vocabulary. Nothing is defaulted:
1008
+ a row with no reason omits the field, because a default here states a cause
1009
+ the record does not know.
1010
+
1011
+ - **The escalation record is filled in only where it is wholly absent.** Pending
1012
+ state written before escalation existed carries no such field, and that one
1013
+ case is answered truthfully — a version with no escalation to offer cannot
1014
+ have been asked for it. A record that exists is carried as recorded, gaps and
1015
+ all: completing a partial one wrote `requested: false` beside
1016
+ `escalated: true`, a pair the producing code cannot emit. Silence about a key
1017
+ is legible to a reader; a value asserted where none was recorded is not.
1018
+ - The count beside these is `observers_reporting` — the observers that answered,
1019
+ which is not the number of slots the configuration named. That question is
1020
+ answered by `denominator_composition`, which lists every observer including
1021
+ the ones that never ran and why.
901
1022
 
902
1023
  **What this rule cannot do.** It cannot tell a terse honest approval from a lazy
903
1024
  one. `**Overall Verdict**: APPROVE / No findings` is the exact form the reviewer
@@ -1247,5 +1368,69 @@ Compression ratio: parallel agent raw → Assembly ≈ 2:1
1247
1368
  the human, is the primary close — a reached ratio is neither necessary nor
1248
1369
  sufficient on its own
1249
1370
 
1371
+ - Verdict determination hardened after a live failure (v3.7.1, 2026-07-27):
1372
+ round 4 of the escalation/persona implementation review recorded the persona
1373
+ team's REVISE as an APPROVE. The team's verdict was being re-derived by
1374
+ searching the assembled text, and a persona had quoted
1375
+ `{"overall_verdict": "APPROVE", ...}` inside a finding as an example of a
1376
+ defect; the search found the quotation before the `**Overall Verdict**:
1377
+ REVISE` on line 1. Three fixes: a structured submission now states its verdict
1378
+ as a field and nothing in its prose overrules it; the header is recognised
1379
+ only at the start of a line; and a JSON verdict is read by parsing a reply
1380
+ that *is* a JSON document rather than by scanning prose for an object. The
1381
+ residual limitation is recorded in § Substance and the denominator — the
1382
+ last-resort word heuristic still sees verdict words inside quotations, so
1383
+ state your verdict in the header when discussing reply shapes. Separately,
1384
+ INV-E2 was found to have been implemented by half: it asks that a counted
1385
+ reply *carry a verdict* and *have substance*, and only substance was checked,
1386
+ so an opening sentence with no judgement in it entered the denominator as a
1387
+ conservative REVISE — the exact shape that retired Fable 5, blocking
1388
+ convergence on nobody's verdict
1389
+
1390
+ - Three of round 4's own fixes reopened what they closed (v3.7.2, 2026-07-28):
1391
+ round 5 of the same implementation review found that each of the three
1392
+ verdict fixes recorded above had left a hole, and all three in the direction
1393
+ that passes. Start-of-line header matching does not exclude a quotation,
1394
+ because a line inside a fenced block starts a line — the header is now read
1395
+ from what the reviewer said, with fences and blockquotes excluded, and a
1396
+ reply that is entirely fenced is read whole because it is quoting nothing.
1397
+ Judging a structured reply by named keys (`reasoning`, `issue`) threw away a
1398
+ REJECT whose findings used `description` — the rule now asks whether anything
1399
+ was said in words, under any key. And the words that count as a verdict were
1400
+ two different lists, so `NO-GO` from a persona was a REJECT while `NO-GO`
1401
+ from an external slot was no judgement at all; there is now one list, with
1402
+ one precedence. Alongside these: `skip_reason` distinguishes five outcomes
1403
+ where it previously flattened three into `transport`; `total_configured` was
1404
+ renamed `observers_reporting` because it never counted what it claimed; the
1405
+ eased convergence rule after an exclusion now asks whether the denominator
1406
+ actually shrank rather than which reason fired; and a synchronous delegation
1407
+ writes the same storage layout as the parallel one, so the lock serialising
1408
+ two concurrent collects is no longer silently skipped
1409
+
1410
+ - Position replaces cleverness in verdict reading (v3.7.3, 2026-07-28): round 6
1411
+ was reviewed and rejected by five of six observers, all of them naming the
1412
+ same shape — the fixes of round 6 had reopened what they closed, in the
1413
+ direction that passes. Two were shipped defects: the header capture was
1414
+ written with `\s`, which includes the newline in Ruby, so it ran past the
1415
+ header into the following prose and recorded a stated APPROVE followed by
1416
+ "No blocking issues" as a REJECT; and the fence regex tracked neither fence
1417
+ length nor delimiter, so a four-backtick block quoting a three-backtick
1418
+ sample gave the quoted verdict the reviewer's vote. Rather than a fourth
1419
+ attempt at reading free text more carefully, the rule is now positional: the
1420
+ reviewer's verdict is the header the reply opens with, and the fence and
1421
+ blockquote machinery is deleted. The prompt asks for it there, and this test
1422
+ now pins that it does. Alongside: a partial escalation record is carried as
1423
+ recorded rather than completed with values the producing code cannot emit;
1424
+ only a token-shaped skip reason is carried into the record; and a row with no
1425
+ reason omits the field in the per-reviewer list as it already did in the
1426
+ composition.
1427
+
1428
+ The round's other lesson was about testing, and it is recorded here because
1429
+ it generalises past this SkillSet: **a regex pinned by one literal example is
1430
+ free everywhere else**. Round 6's own mutation pass replaced whole functions
1431
+ and killed 21 of 21; the review's mutation pass went after regex internals —
1432
+ fence markers, character classes, digit ranges, word boundaries — and 18 of
1433
+ 27 survived. Mutate the inside of a pattern, not only the pattern.
1434
+
1250
1435
  **Key insight**: Design reviews and implementation reviews find
1251
1436
  **categorically different bugs**. Both phases are necessary.
@@ -689,6 +689,65 @@ assert "missing feedback_text_schema_version → reject" do
689
689
  msg && msg.include?('feedback_text_schema_version missing')
690
690
  end
691
691
 
692
+ # v0.7 record schema: the consumer reads reference_verdict from v2 records
693
+ # and still reads verdict from v1 records. Held here because reverting the
694
+ # read (parsed['verdict'] only) would leave every v2 review verdict-less
695
+ # with the rest of the suite green.
696
+ section "v0.7 reference_verdict consumption"
697
+
698
+ def drive_review_with(step, record)
699
+ fake_session = Object.new
700
+ def fake_session.cycle_number = 1
701
+ def fake_session.session_id = 'testsess'
702
+ ctx = Object.new
703
+ def ctx.derive(**_k) = nil
704
+ fake_session.define_singleton_method(:invocation_context) { ctx }
705
+ step.define_singleton_method(:invoke_tool) do |*_a, **_k|
706
+ [{ text: JSON.generate(record) }]
707
+ end
708
+ step.send(:run_multi_llm_review, fake_session,
709
+ { 'summary' => 's', 'task_json' => { 'steps' => [] } },
710
+ { level: 'high', signals: [] }, {})
711
+ end
712
+
713
+ assert "v2 record (reference_verdict, no top-level verdict) → verdict read" do
714
+ out = drive_review_with(step.dup, {
715
+ 'status' => 'ok', 'verdict_schema_version' => 2,
716
+ 'feedback_text_schema_version' => 1,
717
+ 'reference_verdict' => 'REVISE', 'convergence' => {},
718
+ 'aggregated_findings' => [], 'llm_calls' => 3, 'reviews' => [],
719
+ 'feedback_text' => 'x'
720
+ })
721
+ out[:verdict] == 'REVISE'
722
+ end
723
+
724
+ assert "v1 record (verdict) → still read" do
725
+ out = drive_review_with(step.dup, {
726
+ 'status' => 'ok', 'verdict_schema_version' => 1,
727
+ 'feedback_text_schema_version' => 1,
728
+ 'verdict' => 'APPROVE', 'convergence' => {},
729
+ 'aggregated_findings' => [], 'llm_calls' => 3, 'reviews' => [],
730
+ 'feedback_text' => nil
731
+ })
732
+ out[:verdict] == 'APPROVE'
733
+ end
734
+
735
+ # v0.7 record schema (2026-08-01): the boundary is max and max+1, not max
736
+ # and 99 — a probe of 99 cannot tell SUPPORTED=1 from SUPPORTED=2, which is
737
+ # exactly the revert these tests exist to catch.
738
+ assert "v0.7 schema (v=2) → accepted at the boundary" do
739
+ step.send(:schema_version_check, {
740
+ 'verdict_schema_version' => 2, 'feedback_text_schema_version' => 1
741
+ }).nil?
742
+ end
743
+
744
+ assert "v=3 (max+1) → reject (fail-closed at the boundary)" do
745
+ msg = step.send(:schema_version_check, {
746
+ 'verdict_schema_version' => 3, 'feedback_text_schema_version' => 1
747
+ })
748
+ msg && msg.include?('newer than supported')
749
+ end
750
+
692
751
  assert "newer verdict_schema_version → reject (fail-closed)" do
693
752
  msg = step.send(:schema_version_check, {
694
753
  'verdict_schema_version' => 99, 'feedback_text_schema_version' => 1
@@ -696,6 +755,13 @@ assert "newer verdict_schema_version → reject (fail-closed)" do
696
755
  msg && msg.include?('newer than supported')
697
756
  end
698
757
 
758
+ assert "feedback_text v=2 (max+1) → reject (boundary, not v99)" do
759
+ msg = step.send(:schema_version_check, {
760
+ 'verdict_schema_version' => 2, 'feedback_text_schema_version' => 2
761
+ })
762
+ msg && msg.include?('newer than supported')
763
+ end
764
+
699
765
  assert "newer feedback_text_schema_version → reject" do
700
766
  msg = step.send(:schema_version_check, {
701
767
  'verdict_schema_version' => 1, 'feedback_text_schema_version' => 99
@@ -1696,7 +1696,12 @@ module KairosMcp
1696
1696
  aggregated_findings: [], feedback_text: nil }
1697
1697
  else
1698
1698
  {
1699
- verdict: parsed['verdict'],
1699
+ # v0.7 record schema (verdict_schema_version 2) renamed the
1700
+ # top-level conclusion column to reference_verdict (INV-R2: a
1701
+ # recorded reference value, not the run's conclusion). v1
1702
+ # records still carry 'verdict'; read whichever the record
1703
+ # speaks.
1704
+ verdict: parsed['reference_verdict'] || parsed['verdict'],
1700
1705
  convergence: parsed['convergence'],
1701
1706
  aggregated_findings: (parsed['aggregated_findings'] || []).map { |f|
1702
1707
  f.transform_keys(&:to_sym)
@@ -1718,7 +1723,9 @@ module KairosMcp
1718
1723
 
1719
1724
  # Phase 12 §3.10 fail-closed schema versioning.
1720
1725
  # Returns nil if response schema is acceptable, else a string reason.
1721
- SUPPORTED_VERDICT_SCHEMA_VERSION = 1
1726
+ # 2 = v0.7 record schema (2026-08-01): top-level verdict became
1727
+ # reference_verdict; the read above handles both generations.
1728
+ SUPPORTED_VERDICT_SCHEMA_VERSION = 2
1722
1729
  SUPPORTED_FEEDBACK_TEXT_SCHEMA_VERSION = 1
1723
1730
 
1724
1731
  def schema_version_check(parsed)
@@ -183,18 +183,45 @@ module KairosMcp
183
183
  result_text = data['result'] || ''
184
184
  tool_use = extract_tool_use(result_text)
185
185
  usage = data['usage'] || {}
186
+ model_usage = data['modelUsage'] || {}
187
+
188
+ # The answering model is the entry that produced the output tokens,
189
+ # not the first hash key. The CLI places a small internal call
190
+ # (claude-haiku, ~18 output tokens) beside the main call unless
191
+ # CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC suppresses it, so under
192
+ # the worker's stripped environment the envelope carries two keys,
193
+ # and `keys.first` attributed the reply to whichever entry the CLI
194
+ # inserted first. Reproduced 2/2 with byte-identical prompts
195
+ # (2026-07-31); this misread was the root cause of every model
196
+ # divergence this transport had recorded. See L2
197
+ # mlr_v07_design_inputs_and_haiku_root_cause_20260731.
198
+ observed = model_usage.max_by { |_m, u| (u || {})['outputTokens'].to_i }&.first
186
199
 
187
200
  {
188
201
  'content' => tool_use ? nil : result_text,
189
202
  'tool_use' => tool_use,
190
203
  'stop_reason' => tool_use ? 'tool_use' : map_stop_reason(data['stop_reason']),
191
- 'model' => requested_model || data.dig('modelUsage')&.keys&.first || 'claude_code',
204
+ 'model' => requested_model || observed || 'claude_code',
192
205
  # What the CLI reports as having answered, when it reports it.
193
206
  # Kept separate from 'model' (which echoes the request) so callers
194
207
  # can tell a request from an observation and notice when the two
195
- # disagree — the CLI may serve a different model than the one
196
- # asked for.
197
- 'model_observed' => data.dig('modelUsage')&.keys&.first,
208
+ # disagree. The four divergences recorded before the fix above
209
+ # (R6/R8/R10/R13, 2026-07) were all keys.first misreads — the main
210
+ # call was claude-opus-4-6 every time, and the R13 reply's odd
211
+ # content (identifiers that exist nowhere in the reviewed code) is
212
+ # explained by the sandboxed slot reading no repository, not by a
213
+ # different model answering. A divergence observed after the
214
+ # 2026-07-31 fix has no known benign explanation and is worth
215
+ # investigating.
216
+ 'model_observed' => observed,
217
+ # Diagnostic envelope, previously discarded. All four divergence
218
+ # incidents above were undiagnosable from the record because the
219
+ # fields that say what happened did not survive this method.
220
+ # Consumers that persist reviews should carry these through.
221
+ 'model_usage' => model_usage.empty? ? nil : model_usage,
222
+ 'api_error_status' => data['api_error_status'],
223
+ 'fast_mode_state' => data['fast_mode_state'],
224
+ 'terminal_reason' => data['terminal_reason'],
198
225
  'input_tokens' => usage['input_tokens'],
199
226
  'output_tokens' => usage['output_tokens']
200
227
  }
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "llm_client",
3
- "version": "0.1.0",
3
+ "version": "0.2.0",
4
4
  "description": "Pure LLM provider abstraction. One API call, returns response. No loop, no retry, no fallback.",
5
5
  "author": "Masaomi Hatakeyama",
6
6
  "layer": "L1",
@@ -0,0 +1,106 @@
1
+ # frozen_string_literal: true
2
+
3
+ # What parse_response reads from the CLI envelope, pinned after the
4
+ # claude_cli_opus4.6 divergence investigation (2026-07-31).
5
+ #
6
+ # Four times in the multi_llm_review loop (R6/R8/R10/R13) the record showed a
7
+ # slot requesting claude-opus-4-6 answered by claude-haiku-4-5. The 2026-07-31
8
+ # investigation reproduced the cause 2/2: the CLI places a small internal call
9
+ # beside the main one, the envelope carries two keys, and the old `keys.first`
10
+ # read attributed the reply to the internal call — the main call was
11
+ # claude-opus-4-6 every time. The record could not say so, because this method
12
+ # also discarded every diagnostic field the envelope carries (modelUsage
13
+ # breakdown, api_error_status, fast_mode_state, terminal_reason). These tests
14
+ # hold the two fixes: the answering model is chosen by output tokens rather
15
+ # than hash order, and the diagnostic fields survive into the response.
16
+
17
+ require 'minitest/autorun'
18
+ require_relative '../lib/llm_client/adapter'
19
+ require_relative '../lib/llm_client/claude_code_adapter'
20
+
21
+ module KairosMcp
22
+ module SkillSets
23
+ module LlmClient
24
+ class TestClaudeCodeAdapterParse < Minitest::Test
25
+ def setup
26
+ @adapter = ClaudeCodeAdapter.new({ 'timeout_seconds' => 30 })
27
+ end
28
+
29
+ def parse(payload, requested_model: 'claude-opus-4-6')
30
+ @adapter.send(:parse_response, JSON.generate(payload), requested_model: requested_model)
31
+ end
32
+
33
+ def envelope(model_usage:, **extra)
34
+ {
35
+ 'type' => 'result', 'is_error' => false, 'result' => 'ok',
36
+ 'stop_reason' => 'end_turn',
37
+ 'usage' => { 'input_tokens' => 10, 'output_tokens' => 20 },
38
+ 'modelUsage' => model_usage
39
+ }.merge(extra)
40
+ end
41
+
42
+ def test_a_single_model_envelope_is_that_model
43
+ out = parse(envelope(model_usage: {
44
+ 'claude-opus-4-6' => { 'outputTokens' => 116 }
45
+ }))
46
+
47
+ assert_equal 'claude-opus-4-6', out['model_observed']
48
+ end
49
+
50
+ # The regression the old `keys.first` invited: when the envelope
51
+ # carries a second model beside the main call (reproduced 2/2 under
52
+ # the worker's environment), insertion order must not decide which
53
+ # one "answered". The output tokens do.
54
+ def test_the_answering_model_is_the_one_that_wrote_the_output
55
+ out = parse(envelope(model_usage: {
56
+ 'claude-haiku-4-5-20251001' => { 'outputTokens' => 12 },
57
+ 'claude-opus-4-6' => { 'outputTokens' => 2048 }
58
+ }))
59
+
60
+ assert_equal 'claude-opus-4-6', out['model_observed']
61
+ end
62
+
63
+ def test_a_divergent_answer_is_observed_as_itself
64
+ out = parse(envelope(model_usage: {
65
+ 'claude-haiku-4-5-20251001' => { 'outputTokens' => 900 }
66
+ }))
67
+
68
+ assert_equal 'claude-haiku-4-5-20251001', out['model_observed']
69
+ assert_equal 'claude-opus-4-6', out['model']
70
+ refute_equal out['model'], out['model_observed']
71
+ end
72
+
73
+ def test_an_envelope_without_model_usage_observes_nothing
74
+ out = parse(envelope(model_usage: nil))
75
+
76
+ assert_nil out['model_observed']
77
+ assert_nil out['model_usage']
78
+ assert_equal 'claude-opus-4-6', out['model']
79
+ end
80
+
81
+ def test_the_diagnostic_fields_survive_into_the_response
82
+ out = parse(envelope(
83
+ model_usage: { 'claude-haiku-4-5-20251001' => { 'outputTokens' => 900 } },
84
+ 'api_error_status' => 429,
85
+ 'fast_mode_state' => 'off',
86
+ 'terminal_reason' => 'completed'
87
+ ))
88
+
89
+ assert_equal({ 'claude-haiku-4-5-20251001' => { 'outputTokens' => 900 } }, out['model_usage'])
90
+ assert_equal 429, out['api_error_status']
91
+ assert_equal 'off', out['fast_mode_state']
92
+ assert_equal 'completed', out['terminal_reason']
93
+ end
94
+
95
+ def test_a_missing_output_tokens_entry_does_not_crash_the_choice
96
+ out = parse(envelope(model_usage: {
97
+ 'claude-opus-4-6' => nil,
98
+ 'claude-haiku-4-5-20251001' => { 'outputTokens' => 5 }
99
+ }))
100
+
101
+ assert_equal 'claude-haiku-4-5-20251001', out['model_observed']
102
+ end
103
+ end
104
+ end
105
+ end
106
+ end
@@ -23,7 +23,14 @@ module KairosMcp
23
23
  # the verdict shape itself (separate from feedback_text_schema_version,
24
24
  # which lives in FeedbackFormatter). Bumped independently when verdict
25
25
  # JSON contract changes (e.g., new field, semantic redefinition).
26
- VERDICT_SCHEMA_VERSION = 1
26
+ #
27
+ # 2 = v0.7 record schema (design frozen 2026-08-01): the top-level
28
+ # conclusion column is gone — `verdict` became `reference_verdict`
29
+ # (INV-R2), the composition rows carry seat marks, refusals and
30
+ # transport diagnostics travel with the record. Readers tell old
31
+ # records from new by this number; old records are not rewritten
32
+ # (design §4, no retroactivity).
33
+ VERDICT_SCHEMA_VERSION = 2
27
34
 
28
35
  # @param artifact_content [String] sanitized at boundary; raw passthrough is caller responsibility
29
36
  # @param artifact_name [String]
@@ -161,11 +168,32 @@ module KairosMcp
161
168
  }
162
169
  end
163
170
 
171
+ # INV-R7 by_reference delivery: the same canonical framing with a
172
+ # reference manifest where the artifact body would be. The system
173
+ # prompt is identical to the inline one — only the user message
174
+ # differs, and only in the artifact block. sha256 is computed by the
175
+ # caller over the raw submitted content, so a seat that reads the file
176
+ # can check it is reviewing what was submitted.
177
+ def self.build_reference_prompts(artifact_path:, artifact_sha256:,
178
+ artifact_name:, review_type:,
179
+ review_context: 'independent',
180
+ review_round: 1, prior_findings: nil)
181
+ system_prompt = PromptBuilder.build_system_prompt(review_type, review_context: review_context)
182
+ messages = PromptBuilder.build_messages(
183
+ artifact_name: artifact_name,
184
+ review_type: review_type,
185
+ review_round: review_round,
186
+ prior_findings: prior_findings,
187
+ artifact_reference: { path: artifact_path, sha256: artifact_sha256 }
188
+ )
189
+ { system_prompt: system_prompt, messages: messages }
190
+ end
191
+
164
192
  def self.aggregation_instructions(review_type, review_round)
165
193
  <<~INST.strip
166
194
  After collecting all reviewer responses, aggregate as follows:
167
195
  1. Parse each response for verdict {APPROVE, REVISE, REJECT}.
168
- 2. Apply convergence rule (e.g., 3/N APPROVE → APPROVE; otherwise REVISE).
196
+ 2. Compute the reference tally (e.g., 3/N APPROVE → reference APPROVE; any REJECT → REVISE; below quorum → INSUFFICIENT). It is a recorded reference value, not the run's conclusion.
169
197
  3. Merge findings, sorted by severity (P0 first), de-dup by issue text.
170
198
  4. For round #{review_round} #{review_type} reviews, prior findings should be verified as CLOSED/NEEDS_MORE_WORK/REOPENED.
171
199
  INST