kairos-chain 3.61.1 → 3.63.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +79 -0
- data/lib/kairos_mcp/version.rb +1 -1
- data/templates/knowledge/multi_llm_review_workflow/multi_llm_review_workflow.md +144 -2
- data/templates/knowledge/project_orientation_report/project_orientation_report.md +39 -17
- data/templates/knowledge/project_orientation_report/references/worked_example.md +15 -6
- data/templates/knowledge/project_orientation_report/scripts/check_report.py +449 -86
- data/templates/knowledge/project_orientation_report/test/fixtures/bad_class_hidden_headings.html +1 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/bad_comment_hidden_headings.html +1 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/bad_declared_flood.html +33 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/bad_fullwidth_tokens.html +31 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/bad_svg_title_placeholder.html +31 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/bad_transform_order.html +31 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/bad_wide_punctuation_overflow.html +31 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_adjacent_cells.html +31 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_css_font_size.html +32 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_dotted_tokens.html +33 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_group_anchor.html +31 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_highlight_rect.html +31 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_indented_source.html +113 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_media_query_hidden.html +32 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_nested_summary.html +32 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_scoped_stylesheet.html +36 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_short_labels.html +31 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_style_font_size.html +31 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_style_over_attr.html +31 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_thirteen_declared.html +32 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_three_exempt.html +30 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_translated_no_x.html +31 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_wide_punctuation.html +31 -0
- data/templates/knowledge/project_orientation_report/test/fixtures/good_worked_example.html +31 -0
- data/templates/knowledge/project_orientation_report/test/test_check_report.py +35 -3
- data/templates/skillsets/multi_llm_review/bin/dispatch_worker.rb +27 -1
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/consensus.rb +92 -18
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/dispatcher.rb +26 -2
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/pending_state.rb +8 -0
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/prompt_builder.rb +19 -5
- data/templates/skillsets/multi_llm_review/lib/multi_llm_review/sanitizer.rb +11 -1
- data/templates/skillsets/multi_llm_review/skillset.json +2 -2
- data/templates/skillsets/multi_llm_review/test/test_dispatcher_usage.rb +66 -0
- data/templates/skillsets/multi_llm_review/test/test_evidence_fidelity.rb +14 -7
- data/templates/skillsets/multi_llm_review/test/test_multi_llm_review.rb +338 -3
- data/templates/skillsets/multi_llm_review/test/test_mutation_survivors.rb +5 -2
- data/templates/skillsets/multi_llm_review/tools/multi_llm_review_collect.rb +141 -13
- metadata +25 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 96a6f46091b47614f44d089e6459dde7e8435cf7caa893c9b335b7359ad25394
|
|
4
|
+
data.tar.gz: 965451d952fb8fa61d3bfcc581241242de0f25eb492dddc56ebe913e4d07683b
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 1c3812249637a923526a17b29cbbf34549f3f3039f51ae37941fa990ceec99611d68e88de280b39ff673aaf5f64e55be3133b079578b9f2887df8a6a6c81f102
|
|
7
|
+
data.tar.gz: a3688710826e2de6bd2d05ee24b0f9c6194e32e6d9dd8587c776b5e3bfa1b614549af903eee8a9bac295b5e0efc36c5e50537d5a54b8352acd40e05f2ef291c1
|
data/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,85 @@ All notable changes to the `kairos-chain` gem will be documented in this file.
|
|
|
4
4
|
|
|
5
5
|
This project follows [Semantic Versioning](https://semver.org/).
|
|
6
6
|
|
|
7
|
+
## [3.63.0] - 2026-08-06
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
- **`multi_llm_review` findings gain a weight axis, and a worker death no
|
|
12
|
+
longer discards completed seats** (SkillSet 0.9.1 → 0.10.0, frozen after one
|
|
13
|
+
review round; every deployment-grounded finding fixed in-round).
|
|
14
|
+
|
|
15
|
+
The severity of a finding said what kind of defect it was, never what it
|
|
16
|
+
costs: measured on a three-persona panel, 3 of one round's 7 P0s were
|
|
17
|
+
factually correct findings that cost nobody anything, landing at the same
|
|
18
|
+
severity as a defect that silently corrupts published output. The reviewer
|
|
19
|
+
prompt contract now requires a `[consequence: who is harmed, and how]`
|
|
20
|
+
clause on every P0, and aggregation records an unclaused P0 at P2 with the
|
|
21
|
+
stated severity and the demotion reason kept beside it. Only presence is
|
|
22
|
+
checked, mechanically; whether a stated consequence is real or trivial
|
|
23
|
+
stays the orchestrator's call. The clause is read before the byte bound
|
|
24
|
+
cuts the tail, the first non-empty clause counts, deduplication compares
|
|
25
|
+
issues without their clauses, and the demotion mark survives merging in
|
|
26
|
+
any member order. Reviewer prompts also stop naming the round number, and
|
|
27
|
+
the L1 workflow knowledge (v3.10.0) adds the matching orchestrator-side
|
|
28
|
+
rule: never tell a reviewer its finding count, the round number, or prior
|
|
29
|
+
verdicts — counts a reviewer performs select for finding-production over
|
|
30
|
+
finding-weight.
|
|
31
|
+
|
|
32
|
+
Separately, one stuck seat used to cost a whole round: a stale heartbeat
|
|
33
|
+
killed the detached worker and every completed seat's reply died with it,
|
|
34
|
+
because nothing left the worker's memory until all seats were done. The
|
|
35
|
+
worker now persists each seat's outcome the moment it is decided
|
|
36
|
+
(`partial_results.json`, atomic, single writer — including the
|
|
37
|
+
dispatcher's own deadline skips, so recovery cannot relabel a reached
|
|
38
|
+
seat as lost). The collect tool's crash and timeout branches recover the
|
|
39
|
+
completed seats, but only after the reaper CONFIRMS the worker's death: a
|
|
40
|
+
stale heartbeat is not a death certificate, and a record sealed over a
|
|
41
|
+
live worker would be permanently false once idempotent replay pins it.
|
|
42
|
+
Unconfirmed death falls back to the previous retryable total-loss report.
|
|
43
|
+
Seats the worker never reached enter the denominator as skip rows named
|
|
44
|
+
`worker_crashed_seat_lost`, and the payload carries `worker_failure`
|
|
45
|
+
naming the death and the recovered and lost seat labels.
|
|
46
|
+
|
|
47
|
+
## [3.62.0] - 2026-08-06
|
|
48
|
+
|
|
49
|
+
### Changed
|
|
50
|
+
|
|
51
|
+
- **`project_orientation_report`'s checker stops rejecting reports that are fine.**
|
|
52
|
+
Three review rounds on a fixed panel, counting findings rather than approvals:
|
|
53
|
+
16, then 6, then 7. Nearly all of the first round's findings were reports a
|
|
54
|
+
reader would accept that the tool rejected, which is the failure mode that
|
|
55
|
+
matters for a lint — a missed defect costs an ugly figure, a false rejection
|
|
56
|
+
costs the author the report.
|
|
57
|
+
|
|
58
|
+
Adjacent text nodes were concatenated with nothing between them, so two table
|
|
59
|
+
cells holding `Ruby` and `3.4.10` produced a work-internal token `Ruby3` that
|
|
60
|
+
appeared nowhere in the document. Markup indentation counted toward the length
|
|
61
|
+
budget, so a report a browser renders as 3,556 characters was rejected at
|
|
62
|
+
6,226. `text-anchor` was not inherited from an enclosing group, so centred
|
|
63
|
+
labels were reported as overflowing by an amount that did not exist. A label
|
|
64
|
+
with no `x` inside a translated group was called unmeasurable although `x`
|
|
65
|
+
defaults to zero. Ambiguous-width characters had lost their width, so arrow and
|
|
66
|
+
box-drawing labels measured 45% short in the CJK font the template mandates.
|
|
67
|
+
Two-character Japanese labels did not count as a figure at all.
|
|
68
|
+
|
|
69
|
+
Reports now declare a worked prose example with `data-example="values"`,
|
|
70
|
+
satisfying invariant 8 the way the skill always promised. Body text is
|
|
71
|
+
NFKC-normalised before the token scan, so a full-width spelling — what a
|
|
72
|
+
Japanese input method emits by default — no longer slips past. Stylesheets are
|
|
73
|
+
parsed once, with comments stripped, for hiding, preformatting and absolute
|
|
74
|
+
font sizes alike.
|
|
75
|
+
|
|
76
|
+
Left open and stated as such: a rect is treated as a label's box when it
|
|
77
|
+
contains the label's whole width, which promotes a figure border to "the box"
|
|
78
|
+
and so disables overflow detection inside it. The opposite rule rejects
|
|
79
|
+
decorative highlight bands. Both are wrong in opposite directions; this ships
|
|
80
|
+
the one that errs toward passing. Choosing between them is a design question,
|
|
81
|
+
not a patch.
|
|
82
|
+
|
|
83
|
+
46 fixtures, 24 of which had never been added to git — a clone ran the suite
|
|
84
|
+
with them missing and no way to notice.
|
|
85
|
+
|
|
7
86
|
## [3.61.1] - 2026-08-06
|
|
8
87
|
|
|
9
88
|
### Fixed
|
data/lib/kairos_mcp/version.rb
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: multi_llm_review_workflow
|
|
3
3
|
description: "Multi-LLM review methodology and execution — workflow pattern, CLI tooling, consensus analysis, Persona Assembly. Applicable to design, implementation, documentation, or any artifact."
|
|
4
|
-
version: "3.
|
|
4
|
+
version: "3.10.0"
|
|
5
5
|
tags:
|
|
6
6
|
- workflow
|
|
7
7
|
- review
|
|
@@ -498,9 +498,58 @@ The user always has the final say.
|
|
|
498
498
|
`open-question` or `defect` in the round's L2 record
|
|
499
499
|
If no (a)/(b) blocking findings → proceed to next phase
|
|
500
500
|
If any (a)/(b) finding → repeat from [2] with revised artifact
|
|
501
|
+
(the revision obeys § Revision Discipline below)
|
|
501
502
|
(c) findings are recorded as advisory; non-blocking
|
|
502
503
|
```
|
|
503
504
|
|
|
505
|
+
## Revision Discipline (between rounds)
|
|
506
|
+
|
|
507
|
+
> Evidence base: 35 recorded runs across 8 threads, 2026-08-03 → 08-06
|
|
508
|
+
> (tokens in `.kairos/multi_llm_review/pending/`), two of which ran to
|
|
509
|
+
> convergence. Same validation caveat as Step -1: a strong regularity in one
|
|
510
|
+
> instance's logs, not yet reproduced elsewhere. Analysis record: L2
|
|
511
|
+
> `mlr_p0_inflation_analysis_and_opus46_no_verdict_diagnosis_20260806`.
|
|
512
|
+
|
|
513
|
+
The strongest predictor of round N+1's raw P0 count in those logs is whether
|
|
514
|
+
the round-N revision **added mechanism** to the artifact — not the artifact's
|
|
515
|
+
size, not reviewer strictness:
|
|
516
|
+
|
|
517
|
+
- chain_history_erasure v0.5 added two invariants and a recount section
|
|
518
|
+
(draft 9.3k → 19.1k chars): raw P0 went 13 → 38, and 17 of the 38 targeted
|
|
519
|
+
the added or rewritten sections. v0.8, similar in size (18.6k) but authored
|
|
520
|
+
under an explicit "no new mechanism" rule, closed at 13 with one external
|
|
521
|
+
slot finding zero P0s.
|
|
522
|
+
- mlr_evidence_fix R1's fix added four unrequested defensive mechanisms; R2
|
|
523
|
+
returned 41 findings, nearly all of them defects inside the additions. All
|
|
524
|
+
four mechanisms were later removed, each for a measured reason.
|
|
525
|
+
- Deletions never generated findings: chain erasure v0.7 deleted three
|
|
526
|
+
mechanisms — zero new P0s against the deletions, confirmed in writing by
|
|
527
|
+
three slots.
|
|
528
|
+
- Each thread converged within 1–2 rounds of switching to subtractive
|
|
529
|
+
revisions; neither converged while revisions were additive.
|
|
530
|
+
|
|
531
|
+
Rules:
|
|
532
|
+
|
|
533
|
+
1. **A revision closes findings by deletion, correction, or naming — never by
|
|
534
|
+
default-adding.** New mechanisms, invariants, sections, or defensive
|
|
535
|
+
layers do not enter a revision unprompted. If a finding appears to require
|
|
536
|
+
new mechanism, put the question to the operator ("finding X seems to need
|
|
537
|
+
mechanism Y — add, defer to backlog, or drop?") before drafting it in.
|
|
538
|
+
Explanatory prose is a lighter form of the same risk: a sentence added
|
|
539
|
+
only to justify a retreat became the sole blocking finding of the round
|
|
540
|
+
that followed it (mlr_evidence R6 — "adding an explanation creates a new
|
|
541
|
+
claim").
|
|
542
|
+
2. **Prefer a revision author that is not the model whose additions are under
|
|
543
|
+
review.** This extends the existing separation principle — the deciding
|
|
544
|
+
context never authors what judges it — from verification to revision.
|
|
545
|
+
Opus 5 has a measured additive propensity (scope-broadening; the v0.5
|
|
546
|
+
explosion above), but the model is the pressure, not the cause: Fable 5
|
|
547
|
+
also added-and-broke (v0.6's fsync/realpath additions, v0.7's predicate 4
|
|
548
|
+
— a fatal genesis-rejecting defect) until the subtractive rule was
|
|
549
|
+
imposed, and Opus 5 converged mlr_evidence once its revisions became
|
|
550
|
+
subtractive (retreat + removal). Combine both levers: the subtractive
|
|
551
|
+
rule always, a different-model reviser when available.
|
|
552
|
+
|
|
504
553
|
## Review Types
|
|
505
554
|
|
|
506
555
|
| Type | Focus | Reviewers See | Typical Use |
|
|
@@ -548,6 +597,19 @@ and philosophy-aligned finding has been answered is converged whether or not the
|
|
|
548
597
|
numerator moved. Do not treat a reached ratio as sufficient on its own either:
|
|
549
598
|
check what the approving replies actually said before counting them.
|
|
550
599
|
|
|
600
|
+
**Count carryover and new (a)/(b) P0s separately; the machine-side signal of
|
|
601
|
+
convergence is "new P0 = 0", not the APPROVE ratio.** Require each persona to
|
|
602
|
+
state a closure verdict on its own prior-round P0s — closed / open /
|
|
603
|
+
half-closed, with grounds. This format is validated live (chain erasure
|
|
604
|
+
R6–R8) and is what makes the carryover/new split computable. A round whose
|
|
605
|
+
(a)+(b) findings are all carryover with closure verdicts, and whose revision
|
|
606
|
+
drew zero new P0s (observed without exception when the revision was
|
|
607
|
+
subtractive — see § Revision Discipline), is a freeze candidate for the
|
|
608
|
+
operator regardless of the numerator. Neither of the two 2026-08 threads
|
|
609
|
+
ever reached its APPROVE ratio; both closed by (a)+(b) exhaustion + operator
|
|
610
|
+
freeze declaration — the intended close described above, now with a
|
|
611
|
+
measurable trigger.
|
|
612
|
+
|
|
551
613
|
**Escalating raises the bar.** The rule is a ratio over the observers that
|
|
552
614
|
counted, so adding reserve observers with `escalate: true` raises the number of
|
|
553
615
|
agreements required. That is the intended cost of a wider panel, not a defect —
|
|
@@ -1027,6 +1089,26 @@ exclude it would be to exclude every honest terse approval with it. When a round
|
|
|
1027
1089
|
reaches its ratio, read what the approving replies actually said before treating
|
|
1028
1090
|
the ratio as convergence. That judgement is the human's and no rule replaces it.
|
|
1029
1091
|
|
|
1092
|
+
#### A no_verdict streak is a seat-environment signal, not a dead reviewer
|
|
1093
|
+
|
|
1094
|
+
Before treating a slot as dead, read its `raw_text_excerpt` / `stated_text`
|
|
1095
|
+
in the pending record. Diagnosed live (2026-08-06, `claude_cli_opus4.6`,
|
|
1096
|
+
four consecutive no_verdict exclusions across one implementation-review
|
|
1097
|
+
thread): the CLI itself was healthy throughout. The subprocess seat runs
|
|
1098
|
+
sandboxed — no tools, empty working directory — so on **implementation**
|
|
1099
|
+
artifacts that cite file paths, the model attempted to read code before
|
|
1100
|
+
judging: two rounds opened with pseudo-tool-call markup, one opened with
|
|
1101
|
+
"the repository is not accessible" and stated its verdict header only
|
|
1102
|
+
further down, where the positional rule correctly refuses it. The same
|
|
1103
|
+
seat, in the same period, complied on **design** artifacts (verdict on
|
|
1104
|
+
line 1, counted every round). Remedies, in order: state in the subprocess
|
|
1105
|
+
prompt that the seat has no file access and must review the artifact text
|
|
1106
|
+
alone, marking unverifiable claims `[INFERRED]` (the grounding_rules block
|
|
1107
|
+
already licenses this); keep prompt rule #6 (full artifact inline) honest
|
|
1108
|
+
for implementation reviews; only then consider `--add-dir` with read-only
|
|
1109
|
+
tools, accepting CLAUDE.md contamination. A streak that survives those
|
|
1110
|
+
remedies is a real outage.
|
|
1111
|
+
|
|
1030
1112
|
#### Async/Parallel Collect Timing — Iron Rule
|
|
1031
1113
|
|
|
1032
1114
|
When `delegation.parallel.default: true` (the v3.x default), Call 1 returns
|
|
@@ -1107,13 +1189,35 @@ Every review prompt MUST include these 7 items:
|
|
|
1107
1189
|
1. **Output filename table** — so each reviewer knows where to save
|
|
1108
1190
|
2. **Auto-execution commands** — ready-to-run CLI per reviewer
|
|
1109
1191
|
3. **Review instructions** — what to focus on, what NOT to re-review
|
|
1110
|
-
4. **
|
|
1192
|
+
4. **Prior findings to verify** (R2+) — the findings the revision addresses,
|
|
1193
|
+
so the reviewer can judge each closed / open / half-closed. Findings only:
|
|
1194
|
+
no per-reviewer verdict history, no finding counts, no round tallies
|
|
1195
|
+
(see § Reviewer incentive rule)
|
|
1111
1196
|
5. **Context** — architecture summary for reviewers unfamiliar with codebase
|
|
1112
1197
|
6. **Full artifact content inline** — reviewers may not have file access
|
|
1113
1198
|
7. **Severity ratings + output format** — structured template for review output
|
|
1114
1199
|
|
|
1115
1200
|
All prompt content MUST be in **English** for consistent parsing across LLM tools.
|
|
1116
1201
|
|
|
1202
|
+
### Reviewer incentive rule
|
|
1203
|
+
|
|
1204
|
+
**Never tell a reviewer — subprocess or persona — that its finding count is
|
|
1205
|
+
compared across rounds, which round this is, or what verdicts were given
|
|
1206
|
+
before.** A reviewer told its count is watched treats the count as the
|
|
1207
|
+
deliverable, and that selects for finding-*production* over finding-*weight*
|
|
1208
|
+
(observed 2026-08-06: an orchestrator wrote "your finding count is compared
|
|
1209
|
+
across rounds" into persona prompts during the project_orientation_report
|
|
1210
|
+
loop; of the round's 7 P0s, 3 were factually correct findings that cost
|
|
1211
|
+
nobody anything). Convergence — carryover vs new, (a)+(b) exhaustion, the
|
|
1212
|
+
ratio — is measured by the orchestrator from the record, after the replies
|
|
1213
|
+
are in. The reviewer receives the artifact, the review criteria, and the
|
|
1214
|
+
prior findings it must verify. Nothing else about the loop's state.
|
|
1215
|
+
|
|
1216
|
+
What this rule does NOT forbid: passing prior findings for closure
|
|
1217
|
+
verification (rule #4 — that is content, not score-keeping), and the
|
|
1218
|
+
carryover/new split in § Convergence Rules (that is orchestrator-side
|
|
1219
|
+
bookkeeping the reviewer never sees).
|
|
1220
|
+
|
|
1117
1221
|
### XML Block Structure for Review Prompts
|
|
1118
1222
|
|
|
1119
1223
|
Review prompts SHOULD use XML blocks to give LLMs explicit structural contracts.
|
|
@@ -1432,5 +1536,43 @@ Compression ratio: parallel agent raw → Assembly ≈ 2:1
|
|
|
1432
1536
|
fence markers, character classes, digit ranges, word boundaries — and 18 of
|
|
1433
1537
|
27 survived. Mutate the inside of a pattern, not only the pattern.
|
|
1434
1538
|
|
|
1539
|
+
- Revision discipline, new-P0 convergence signal, and no_verdict seat
|
|
1540
|
+
diagnosis (v3.9.0, 2026-08-06): cross-thread analysis of 35 recorded runs
|
|
1541
|
+
(8 threads, 2026-08-03 → 08-06) established that raw P0 growth tracks
|
|
1542
|
+
additive revisions, not reviewer severity — every mechanism a revision
|
|
1543
|
+
added became the next round's battleground, deletions drew zero new P0s
|
|
1544
|
+
in every measured case, and both threads that converged did so within
|
|
1545
|
+
1–2 rounds of switching to subtractive revisions (one under a Fable 5
|
|
1546
|
+
reviser, one under the same Opus 5 orchestrator that had produced the
|
|
1547
|
+
additive explosion). New § Revision Discipline encodes the subtractive
|
|
1548
|
+
rule and the reviser-separation preference. § Convergence Rules gains the
|
|
1549
|
+
carryover/new P0 split, with "new (a)+(b) P0 = 0" as the machine-side
|
|
1550
|
+
freeze-candidate signal — neither thread ever reached its APPROVE ratio;
|
|
1551
|
+
both closed by (a)+(b) exhaustion + operator freeze. § Substance and the
|
|
1552
|
+
denominator gains the seat-environment diagnosis: `claude_cli_opus4.6`'s
|
|
1553
|
+
four-round no_verdict streak was the sandboxed seat colliding with
|
|
1554
|
+
implementation artifacts (pseudo-tool-calls, "repository not accessible"
|
|
1555
|
+
preamble), not a dead reviewer — the same seat counted every round on
|
|
1556
|
+
design artifacts in the same period. Analysis record: L2
|
|
1557
|
+
`mlr_p0_inflation_analysis_and_opus46_no_verdict_diagnosis_20260806`
|
|
1558
|
+
|
|
1559
|
+
- Reviewer incentive rule and the finding weight axis (v3.10.0, 2026-08-06):
|
|
1560
|
+
§ Prompt Generation Rules gains the Reviewer incentive rule — reviewer
|
|
1561
|
+
prompts never mention finding counts, round numbers, or prior verdicts;
|
|
1562
|
+
convergence is measured orchestrator-side from the record, and prior
|
|
1563
|
+
findings are passed for closure verification only (rule #4 reworded
|
|
1564
|
+
accordingly, from "review history table" to "prior findings to verify").
|
|
1565
|
+
Motivating observation: an orchestrator told personas their counts were
|
|
1566
|
+
compared across rounds, and 3 of the round's 7 P0s were factually correct
|
|
1567
|
+
findings that cost nobody anything. In the same change the SkillSet
|
|
1568
|
+
(0.10.0) adds the weight axis mechanically: the prompt contract requires a
|
|
1569
|
+
`[consequence: who is harmed, and how]` clause on every P0, and
|
|
1570
|
+
aggregation records a P0 without one at P2, keeping the stated severity
|
|
1571
|
+
and the demotion reason beside it (`severity_stated` /
|
|
1572
|
+
`severity_demoted: consequence_missing`). Presence is checked
|
|
1573
|
+
mechanically; whether a stated consequence is real or trivial stays the
|
|
1574
|
+
orchestrator's call, per the (a)/(b)/(c) discipline. Handoff record: L2
|
|
1575
|
+
`handoff_mlr_finding_weight_axis_and_reviewer_incentive_20260806`
|
|
1576
|
+
|
|
1435
1577
|
**Key insight**: Design reviews and implementation reviews find
|
|
1436
1578
|
**categorically different bugs**. Both phases are necessary.
|
|
@@ -153,13 +153,19 @@ SVG の制約。守らないと壊れる。
|
|
|
153
153
|
6. **人が読む。** 通らなかった節が分かったら、**レポートではなくこの skill を直す**。
|
|
154
154
|
条件 7・8・9・10 はすべて、この段の判定から出た。
|
|
155
155
|
|
|
156
|
-
## レポート側が守る
|
|
156
|
+
## レポート側が守る 3 つの約束ごと
|
|
157
157
|
|
|
158
|
-
機械が読むための印を
|
|
158
|
+
機械が読むための印を 3 つ、レポートの HTML に入れる。
|
|
159
159
|
|
|
160
160
|
**図表の無い節は、見出しで宣言する。** `<h2 data-visual="none">決まっていないこと</h2>`。
|
|
161
161
|
宣言の無い節は図か表を要求される。位置で免除しない——節が 1 つ増えただけで免除が
|
|
162
|
-
|
|
162
|
+
隣の節へずれるため。宣言できるのは 3 節まで(表紙、「決まっていないこと」、
|
|
163
|
+
食い違いが無いときの「記録の食い違い」で使い切る想定)。
|
|
164
|
+
|
|
165
|
+
**図表の代わりに実際の値の例で満たす節は、その要素に印を付ける。**
|
|
166
|
+
`<p data-example="values">入力 X を送ると出力は Y だった。</p>`。
|
|
167
|
+
条件 8 は「図・表・実際の値による具体例のいずれか」を認めており、この印が
|
|
168
|
+
「値の入った例」を機械に伝える。印の無いただの散文は数えない。
|
|
163
169
|
|
|
164
170
|
> この宣言は v0.2 で入った。**それ以前に書かれたレポートは、表紙と「決まっていないこと」の
|
|
165
171
|
> 見出しに属性を 1 つずつ足すまで落ちる。** 読み手に見えるものは変わらないので、
|
|
@@ -170,7 +176,15 @@ SVG の制約。守らないと壊れる。
|
|
|
170
176
|
`<meta name="orientation-allowed-tokens" content="A4, Q3, gpt-5">`。
|
|
171
177
|
検出はわざと広くしてあるので、紙の大きさやモデル名のような無害な語も引っかかる。
|
|
172
178
|
それを消すには宣言を書く。**宣言を書くこと自体が規律**で、スクリプトの中の許可リストを
|
|
173
|
-
太らせるより安全である。組み込みの免除は層の名前 3
|
|
179
|
+
太らせるより安全である。組み込みの免除は層の名前 3 語だけ。宣言できるのは 16 語まで——
|
|
180
|
+
節の免除と同じで、宣言で要求を丸ごと消す形は認めない。この上限は、条件 7 を生んだ
|
|
181
|
+
「符牒 13 種で読めなかったレポート」への余白として選んだ数ではない。あの 13 種は
|
|
182
|
+
**宣言されずに**本文へ出た符牒で、上限はそこでは発火しない(自分の符牒を全部宣言してしまう形は、
|
|
183
|
+
適用範囲の表にあるとおり、意図しなければ起きない容認済みの抜け道である)。一方で、実在する
|
|
184
|
+
技術の段落が SHA-256 や Ed25519 のような公開語彙を 13 語、正当に宣言する必要があった。
|
|
185
|
+
上限が止めるのは宣言で要求を丸ごと消す形だけなので、正当な宣言の実測最大 13 語を
|
|
186
|
+
落とさない粗い天井として 16 にしてある。`v0.9` や `TLS-1.3` のような
|
|
187
|
+
語は書いたとおりに宣言すればよい(部分文字列に割って宣言し直す必要は無い)。
|
|
174
188
|
|
|
175
189
|
## 検収
|
|
176
190
|
|
|
@@ -179,13 +193,13 @@ SVG の制約。守らないと壊れる。
|
|
|
179
193
|
|
|
180
194
|
| 検査 | 合格条件 |
|
|
181
195
|
|---|---|
|
|
182
|
-
| 未記入 | 《》で囲まれた placeholder が 0
|
|
183
|
-
| 符牒 | 見えている文に、宣言されていない符牒が 0
|
|
184
|
-
| 整形済みの塊 | `<pre>` と `white-space: pre` が本体に 0
|
|
185
|
-
| 分量 | 見えている文が 6000 文字以内(`--max-body-chars`
|
|
196
|
+
| 未記入 | 《》で囲まれた placeholder が 0 個。未記入のテンプレートは通らない。SVG の `<title>` も読み上げられるので対象 |
|
|
197
|
+
| 符牒 | 見えている文に、宣言されていない符牒が 0 個。全角で書いた符牒(F1)も NFKC 正規化で同じ符牒として数える。宣言は 16 語まで |
|
|
198
|
+
| 整形済みの塊 | `<pre>` と `white-space: pre` が本体に 0 個。付録や印刷だけに効く stylesheet の規則は数えない |
|
|
199
|
+
| 分量 | 見えている文が 6000 文字以内(`--max-body-chars` で変更可)。markup の字下げは数えない |
|
|
186
200
|
| 節 | 8 節以上 |
|
|
187
|
-
| 視覚要素 |
|
|
188
|
-
| 図の収まり | すべての行が、囲んでいる `<rect>`
|
|
201
|
+
| 視覚要素 | 宣言で免除されていない全節に、**中身のある**図か表、または `data-example="values"` を付けた値の例がある(空の表は不可)。図として数える label は 2 文字以上。免除宣言は 3 節まで |
|
|
202
|
+
| 図の収まり | すべての行が、囲んでいる `<rect>` に収まる。箱とみなすのは行の全幅を含む最小の `<rect>` で、全幅を含む `<rect>` が無いときだけ開始点を含む最小の `<rect>` と照合する(強調の帯を箱と誤認しないため)。`<rect>` が無ければ図の枠に収まる |
|
|
189
203
|
|
|
190
204
|
「見えている内容」の定義がこの検収の要である。折りたたまれていない `<details open>` の中も、
|
|
191
205
|
`<svg>` の中の `<text>` も、読み手には見えているので本文に数える。`<script>` `<style>`
|
|
@@ -200,7 +214,7 @@ SVG の制約。守らないと壊れる。
|
|
|
200
214
|
|---|---|
|
|
201
215
|
| ASCII 図、未記入の欄、箱から溢れる label | 別紙の CSS が `content:` で描く文字 |
|
|
202
216
|
| 測り方が拾えない書き方の `translate` | `<iframe srcdoc>` の中身 |
|
|
203
|
-
| `em` のような相対単位の font-size |
|
|
217
|
+
| `em` のような相対単位の font-size、全角で書いた符牒 | 幅ゼロ文字を混ぜた符牒 |
|
|
204
218
|
| 全部の節を免除にしてしまう | 禁止すべき符牒を自分で許可宣言する |
|
|
205
219
|
| 見出しを `display:none` で隠す | 入れ子の `<svg>` を親の外へ置く |
|
|
206
220
|
| 折りたたんだ付録の `<summary>` に本文を書く | `textLength` で描画幅を上書きする |
|
|
@@ -212,8 +226,9 @@ SVG の制約。守らないと壊れる。
|
|
|
212
226
|
|
|
213
227
|
### 検収が測っていないもの
|
|
214
228
|
|
|
215
|
-
不変条件は 10 個あり、**機械が見ているのは 3
|
|
216
|
-
|
|
229
|
+
不変条件は 10 個あり、**機械が見ているのは 3 個だけ**(7 符牒、8 視覚要素、9 分量)。
|
|
230
|
+
図の収まりも機械が測るが、あれは「図の作り方」の規約の検査であって、10 個の不変条件の
|
|
231
|
+
うちの 1 つではない。残る 7 個は人が守るしかない。
|
|
217
232
|
|
|
218
233
|
| 未検査の条件 | なぜ機械で測れないか |
|
|
219
234
|
|---|---|
|
|
@@ -230,10 +245,14 @@ SVG の制約。守らないと壊れる。
|
|
|
230
245
|
|
|
231
246
|
## テスト
|
|
232
247
|
|
|
233
|
-
`test/test_check_report.py` が `test/fixtures/` の
|
|
248
|
+
`test/test_check_report.py` が `test/fixtures/` の 46 件を走らせる。標準ライブラリだけで動く。
|
|
234
249
|
|
|
235
|
-
fixture は **2 方向**ある。
|
|
236
|
-
|
|
250
|
+
fixture は **2 方向**ある。25 件は「読めないのに通ってしまった形」、21 件は「読めるのに
|
|
251
|
+
落とされた形」。落ちる側のうち 8 件は外部の評価者が実際に実行して示した抜け道で
|
|
252
|
+
(最初の 10 件のうち `bad_missing_section.html` と `bad_comment_details.html` の 2 件は
|
|
253
|
+
初版でも「節」で落ちていたので抜け道ではない)、
|
|
254
|
+
残りはその後のレビューで見つかった、書き手が普通に書いていて踏む事故の形と、
|
|
255
|
+
修正の後退を防ぐ検査である。
|
|
237
256
|
**塞ぐだけなら全部不合格にすれば達成できる**ので、両方向を同時に測る。
|
|
238
257
|
どの検査が捕まえるべきかも fixture ごとに宣言してあり、別の理由で落ちた場合はテストが赤くなる。
|
|
239
258
|
|
|
@@ -253,5 +272,8 @@ KairosChain の機能ではない。KairosChain 単体で動かす場合は、
|
|
|
253
272
|
|
|
254
273
|
**実装は 1 度 multi-LLM review に落ちている。** 4 者中 0 者が承認、深刻な指摘 14 件。
|
|
255
274
|
初版の検収スクリプトはタグを見ていたので、内容を除外領域へ移すだけで全部すり抜けられた。
|
|
256
|
-
評価者が実行して示した抜け道が、いま `test/fixtures/` の `bad_*` 10
|
|
275
|
+
評価者が実行して示した抜け道が、いま `test/fixtures/` の `bad_*` のうち最初の 10 件中
|
|
276
|
+
8 件になっている(`bad_missing_section.html` と `bad_comment_details.html` は初版でも
|
|
277
|
+
「節」で落ちていたので抜け道ではない。残りの `bad_*` も後のレビューで加わった事故の形と
|
|
278
|
+
後退防止で、評価者由来ではない)。
|
|
257
279
|
経緯は `references/worked_example.md`。
|
|
@@ -55,9 +55,17 @@
|
|
|
55
55
|
|
|
56
56
|
## fixture — 破られた形がそのまま検査になっている
|
|
57
57
|
|
|
58
|
-
`test/fixtures/` の
|
|
59
|
-
|
|
60
|
-
|
|
58
|
+
`test/fixtures/` の 46 件。`test/test_check_report.py` が走らせる。
|
|
59
|
+
内訳は四世代ある。初版で破られた形(下の 2 表)、レビュー 2 周目で入った
|
|
60
|
+
「書き手が普通に書いていて踏む事故の形」8 件、3 周目のレビュー指摘の
|
|
61
|
+
修正に付けた後退防止の検査、そして 4 周目(普通に書いたレポートに誤った答えを
|
|
62
|
+
返す形の修正)で加わった検査。どの fixture をどの検査が捕まえるべきかの宣言は
|
|
63
|
+
`test_check_report.py` にある。
|
|
64
|
+
|
|
65
|
+
**落ちる側の fixture のうち、初版由来の 10 件。うち抜け道は 8 件**(評価者が実行して
|
|
66
|
+
示したのは 8 件。※印の 2 件——`bad_missing_section.html` と `bad_comment_details.html`——は
|
|
67
|
+
初版(9c3eb0a)でもどちらも「節」の検査で落ちており、抜け道ではなく最初からある検査の
|
|
68
|
+
確認である。それ以降に追加された落ちる側の fixture も評価者由来ではない)
|
|
61
69
|
|
|
62
70
|
| fixture | 何を再現しているか | 捕まえる検査 |
|
|
63
71
|
|---|---|---|
|
|
@@ -68,11 +76,12 @@
|
|
|
68
76
|
| `bad_svg_overflow_rect.html` | 図の枠には収まるが箱から溢れる label | 図の収まり |
|
|
69
77
|
| `bad_svg_transform.html` | `translate` で画面外へ出る label | 図の収まり |
|
|
70
78
|
| `bad_token_forms.html` | `f1` `c-1` `step-3` `#12` のような形 | 符牒 |
|
|
71
|
-
| `bad_comment_details.html` | コメントの中の `<details` で本文を切る | 整形済みの塊 |
|
|
79
|
+
| `bad_comment_details.html` ※抜け道ではない | コメントの中の `<details` で本文を切る | 整形済みの塊 |
|
|
72
80
|
| `bad_whitespace_pre.html` | `<pre>` を使わずに ASCII 図を出す | 整形済みの塊 |
|
|
73
|
-
| `bad_missing_section.html` | 節が 1 つ足りない | 節 |
|
|
81
|
+
| `bad_missing_section.html` ※抜け道ではない | 節が 1 つ足りない | 節 |
|
|
74
82
|
|
|
75
|
-
**落としてはいけない 4
|
|
83
|
+
**落としてはいけない fixture のうち、初版からの 4 件**(その後、レビュー指摘で判明した
|
|
84
|
+
「正しく書いたのに落とされる形」を通す fixture が 17 件加わり、通る側は計 21 件)
|
|
76
85
|
|
|
77
86
|
| fixture | なぜ通らねばならないか |
|
|
78
87
|
|---|---|
|