cyber-sdd 0.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +17 -0
- package/.codex-plugin/plugin.json +17 -0
- package/.plugin/plugin.json +17 -0
- package/README.md +159 -0
- package/agents/sdd-automaton.md +97 -0
- package/agents/sdd-impl-judge.md +214 -0
- package/agents/sdd-scanner.md +120 -0
- package/agents/sdd-spec-judge.md +224 -0
- package/agents/sdd-warden.md +101 -0
- package/package.json +24 -0
- package/skills/align-spec/README.md +20 -0
- package/skills/align-spec/SKILL.md +111 -0
- package/skills/align-spec/scripts/align-spec.mts +187 -0
- package/skills/architect-impl-governance/README.md +46 -0
- package/skills/architect-impl-governance/SKILL.md +45 -0
- package/skills/architect-spec-governance/README.md +48 -0
- package/skills/architect-spec-governance/SKILL.md +59 -0
- package/skills/blast-estimate/README.md +47 -0
- package/skills/blast-estimate/SKILL.md +133 -0
- package/skills/blast-estimate/scripts/blast-estimate.mts +583 -0
- package/skills/builder-impl-governance/README.md +47 -0
- package/skills/builder-impl-governance/SKILL.md +47 -0
- package/skills/builder-spec-governance/README.md +49 -0
- package/skills/builder-spec-governance/SKILL.md +36 -0
- package/skills/check-partition-quality/README.md +22 -0
- package/skills/check-partition-quality/SKILL.md +51 -0
- package/skills/check-partition-quality/scripts/check-partition-quality.mts +336 -0
- package/skills/check-plan-safety/README.md +17 -0
- package/skills/check-plan-safety/SKILL.md +60 -0
- package/skills/check-plan-safety/scripts/check-plan-safety.mts +145 -0
- package/skills/check-project-specs/README.md +19 -0
- package/skills/check-project-specs/SKILL.md +69 -0
- package/skills/check-project-specs/scripts/check-project-specs.mts +217 -0
- package/skills/check-scenario-overlap/README.md +19 -0
- package/skills/check-scenario-overlap/SKILL.md +74 -0
- package/skills/check-scenario-overlap/scripts/check-scenario-overlap.mts +249 -0
- package/skills/check-spec-structure/README.md +17 -0
- package/skills/check-spec-structure/SKILL.md +66 -0
- package/skills/check-spec-structure/scripts/check-spec-structure.mts +346 -0
- package/skills/collision-ladder/README.md +18 -0
- package/skills/collision-ladder/SKILL.md +83 -0
- package/skills/collision-ladder/scripts/collision-ladder.mts +657 -0
- package/skills/combat-log-governance/README.md +13 -0
- package/skills/combat-log-governance/SKILL.md +257 -0
- package/skills/concept-index/README.md +13 -0
- package/skills/concept-index/SKILL.md +38 -0
- package/skills/concept-index/scripts/concept-index.mts +245 -0
- package/skills/discover-plans/README.md +16 -0
- package/skills/discover-plans/SKILL.md +74 -0
- package/skills/discover-plans/scripts/discover-plans.mts +212 -0
- package/skills/discover-specs/README.md +15 -0
- package/skills/discover-specs/SKILL.md +76 -0
- package/skills/discover-specs/scripts/discover-specs.mts +396 -0
- package/skills/doctrine-loop/README.md +15 -0
- package/skills/doctrine-loop/SKILL.md +97 -0
- package/skills/formation-loop/README.md +17 -0
- package/skills/formation-loop/SKILL.md +140 -0
- package/skills/gate-validation-governance/README.md +12 -0
- package/skills/gate-validation-governance/SKILL.md +87 -0
- package/skills/impl-producer-governance/README.md +48 -0
- package/skills/impl-producer-governance/SKILL.md +85 -0
- package/skills/init/README.md +27 -0
- package/skills/init/SKILL.md +68 -0
- package/skills/init/scripts/wire-statusline.mts +276 -0
- package/skills/lifecycle-governance/README.md +11 -0
- package/skills/lifecycle-governance/SKILL.md +168 -0
- package/skills/manage/README.md +9 -0
- package/skills/manage/SKILL.md +62 -0
- package/skills/manage-ignore/README.md +19 -0
- package/skills/manage-ignore/SKILL.md +52 -0
- package/skills/manage-ignore/scripts/manage-ignore.mts +294 -0
- package/skills/manage-scenario-bridge/README.md +20 -0
- package/skills/manage-scenario-bridge/SKILL.md +60 -0
- package/skills/manage-scenario-bridge/scripts/manage-scenario-bridge.mts +156 -0
- package/skills/manage-spec-anchors/README.md +18 -0
- package/skills/manage-spec-anchors/SKILL.md +56 -0
- package/skills/manage-spec-anchors/scripts/manage-spec-anchors.mts +328 -0
- package/skills/mission-graph/README.md +15 -0
- package/skills/mission-graph/SKILL.md +67 -0
- package/skills/mission-graph/scripts/mission-graph.mts +844 -0
- package/skills/oracle-spec-governance/README.md +45 -0
- package/skills/oracle-spec-governance/SKILL.md +45 -0
- package/skills/ownership-governance/README.md +65 -0
- package/skills/ownership-governance/SKILL.md +104 -0
- package/skills/pause-mission/README.md +18 -0
- package/skills/pause-mission/SKILL.md +112 -0
- package/skills/place-node/README.md +12 -0
- package/skills/place-node/SKILL.md +47 -0
- package/skills/place-node/scripts/place-node.mts +157 -0
- package/skills/plan-retirement/README.md +32 -0
- package/skills/plan-retirement/SKILL.md +90 -0
- package/skills/plan-retirement/scripts/retire-plans.mts +196 -0
- package/skills/plugin-contract-governance/README.md +12 -0
- package/skills/plugin-contract-governance/SKILL.md +112 -0
- package/skills/remediation-governance/README.md +46 -0
- package/skills/remediation-governance/SKILL.md +78 -0
- package/skills/resolve-governances/README.md +18 -0
- package/skills/resolve-governances/SKILL.md +50 -0
- package/skills/resolve-governances/scripts/resolve-governances.mts +515 -0
- package/skills/resolve-tracking/SKILL.md +64 -0
- package/skills/resolve-tracking/scripts/resolve-tracking.mts +213 -0
- package/skills/resume-mission/README.md +12 -0
- package/skills/resume-mission/SKILL.md +53 -0
- package/skills/scaffold-project-spec/README.md +7 -0
- package/skills/scaffold-project-spec/SKILL.md +192 -0
- package/skills/sdd/README.md +7 -0
- package/skills/sdd/SKILL.md +92 -0
- package/skills/solution-producer-governance/README.md +9 -0
- package/skills/solution-producer-governance/SKILL.md +44 -0
- package/skills/spec-format-governance/README.md +73 -0
- package/skills/spec-format-governance/SKILL.md +114 -0
- package/skills/spec-gate/README.md +26 -0
- package/skills/spec-gate/SKILL.md +201 -0
- package/skills/spec-gate/scripts/check-spec-state.mts +601 -0
- package/skills/spec-gate/scripts/check-suite.mts +501 -0
- package/skills/spec-gate/scripts/classify-edit-class.mts +411 -0
- package/skills/spec-producer-governance/README.md +7 -0
- package/skills/spec-producer-governance/SKILL.md +86 -0
- package/skills/spec-structure-governance/README.md +40 -0
- package/skills/spec-structure-governance/SKILL.md +169 -0
- package/skills/ssa-lowering/README.md +26 -0
- package/skills/ssa-lowering/SKILL.md +181 -0
- package/skills/start-mission/README.md +7 -0
- package/skills/start-mission/SKILL.md +115 -0
- package/skills/suite-format-governance/README.md +75 -0
- package/skills/suite-format-governance/SKILL.md +299 -0
- package/skills/suite-format-governance/references/rubric.md +313 -0
- package/skills/touch-set-correction/README.md +16 -0
- package/skills/touch-set-correction/SKILL.md +67 -0
- package/skills/touch-set-correction/scripts/touch-set-correction.mts +418 -0
- package/skills/verify-scenarios/README.md +17 -0
- package/skills/verify-scenarios/SKILL.md +109 -0
- package/skills/verify-scenarios/scripts/verify-scenarios.mts +386 -0
|
@@ -0,0 +1,313 @@
|
|
|
1
|
+
# Rubric form and judgment
|
|
2
|
+
|
|
3
|
+
For scenarios where behavior is a **gradient judgment** — "good enough across several dimensions" — that cannot be faithfully encoded in a single flat boolean assertion, use the rubric form. This document covers the full rubric-form lifecycle: how to author `@rubric` scenarios, decide which criteria belong in the sum, set thresholds, correct standing rubrics, and judge whether a rubric discriminates.
|
|
4
|
+
|
|
5
|
+
## Form 2 — rubric Gherkin (`@rubric`, judged by hand)
|
|
6
|
+
|
|
7
|
+
A gradient judgment ("good enough across several dimensions") cannot be faithfully encoded in a
|
|
8
|
+
single flat boolean. The rubric form admits scoring criteria into the scenario and collapses them
|
|
9
|
+
to one boolean, preserving the gate contract. It is **purely additive** — it never changes how
|
|
10
|
+
untagged scenarios work. Convention:
|
|
11
|
+
|
|
12
|
+
1. **Tag** the scenario `@rubric`.
|
|
13
|
+
2. **Embed the rubric** in a `Then` step as a docstring (`"""..."""`) — named dimensions, each with
|
|
14
|
+
a `max:` value, plus exactly one `threshold:` line.
|
|
15
|
+
3. **Close** with a boolean-collapsing `Then`: `And the rubric score is at least the threshold`.
|
|
16
|
+
|
|
17
|
+
```gherkin
|
|
18
|
+
@rubric
|
|
19
|
+
Scenario: <name>
|
|
20
|
+
Given ...
|
|
21
|
+
When ...
|
|
22
|
+
Then the judge evaluates the scenario against the rubric
|
|
23
|
+
"""
|
|
24
|
+
dimensions:
|
|
25
|
+
- name: correctness
|
|
26
|
+
max: 3
|
|
27
|
+
- name: completeness
|
|
28
|
+
max: 2
|
|
29
|
+
threshold: 4
|
|
30
|
+
"""
|
|
31
|
+
And the rubric score is at least the threshold
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
The final `Then` yields exactly one boolean — the gate sees pass/fail, not a score. The rubric is
|
|
35
|
+
internal evaluation detail, judged **by hand**.
|
|
36
|
+
|
|
37
|
+
## One criterion per dimension — split before you select
|
|
38
|
+
|
|
39
|
+
A dimension naming two criteria joined by *and* — `harness_agnostic_and_mcp_free` — is
|
|
40
|
+
**double-barreled**, and it has no honest score: a subject that is harness-agnostic but ships an MCP
|
|
41
|
+
dependency satisfies one half and fails the other, so every number you could award it reports
|
|
42
|
+
something false.
|
|
43
|
+
|
|
44
|
+
This is a **structural** defect, not a selection one, and it is caught at the **structure** check —
|
|
45
|
+
before selection, because a double-barreled criterion cannot be selected either. Its two halves
|
|
46
|
+
rarely trade the same way, so asking whether *it* is substitutable has no answer. **Split it first**,
|
|
47
|
+
then run each half through the substitutability test below on its own — the halves routinely land in
|
|
48
|
+
different forms (*that a migration is reversible* is a rule; *how thoroughly it is documented* is a
|
|
49
|
+
gradient).
|
|
50
|
+
|
|
51
|
+
## Choosing the form — the substitutability test
|
|
52
|
+
|
|
53
|
+
The two forms (Form 1 and Form 2) are not *simple* vs *complex*. They are **non-substitutable** vs **substitutable**,
|
|
54
|
+
and this is the rule that decides which form a criterion goes in.
|
|
55
|
+
|
|
56
|
+
A rubric that sums N dimensions against one threshold is a **compensatory** model: a high score on
|
|
57
|
+
one dimension **compensates** for a low score on another. That is the form's purpose and its one
|
|
58
|
+
documented cost — a subject that wholly fails one dimension can still pass by banking points
|
|
59
|
+
elsewhere. The term is **compensation**; the bad pass it produces is a **false positive
|
|
60
|
+
classification error**. It attaches to the **aggregation** — to the act of summing — and no
|
|
61
|
+
judgment method, judge quality, or threshold arithmetic touches it.
|
|
62
|
+
|
|
63
|
+
Compensation is legitimate only where the dimensions are **substitutable**: where you genuinely
|
|
64
|
+
accept that strength on one may pay for weakness on another.
|
|
65
|
+
|
|
66
|
+
> Knowing how to read English really well should not compensate for the lack of ability to speak
|
|
67
|
+
> English.
|
|
68
|
+
|
|
69
|
+
**The selection rule: a criterion belongs in a `@rubric` only if you accept that strength elsewhere
|
|
70
|
+
may pay for weakness here.** If you do not accept that trade, the criterion is not in the sum at
|
|
71
|
+
all — it is a boolean `Then` (Form 1, documented in the main governance file).
|
|
72
|
+
|
|
73
|
+
**Select before you author.** This is a decision about where a criterion **goes**, made while you are
|
|
74
|
+
writing it — not a pass that strips dimensions out of a rubric afterwards. A non-substitutable
|
|
75
|
+
criterion never becomes a dimension in the first place.
|
|
76
|
+
|
|
77
|
+
Where an **existing** suite already sums one, correcting it is never a silent strip and never the
|
|
78
|
+
simple deletion it looks like: it **adds** a boolean `Then` and it **forces the cut to be
|
|
79
|
+
re-derived**. That correction has its own procedure — *Correcting a standing rubric*, below.
|
|
80
|
+
|
|
81
|
+
Apply it per criterion, in the subject's own domain, by **saying the trade out loud**:
|
|
82
|
+
|
|
83
|
+
- *"Great scope makes up for shipping an npx dependency"* — nobody accepts this. `no_npx_dependency`
|
|
84
|
+
is a **rule**: a boolean `Then`.
|
|
85
|
+
- *"Stronger error handling makes up for thinner edge-case coverage"* — a trade a reviewer would
|
|
86
|
+
genuinely make. Both belong in the rubric.
|
|
87
|
+
|
|
88
|
+
| The criterion is | Form |
|
|
89
|
+
|---|---|
|
|
90
|
+
| **non-substitutable** — no strength elsewhere pays for failing it | **Form 1** — a boolean `Then` |
|
|
91
|
+
| **substitutable** — a reviewer would genuinely make the trade | **Form 2** — a `@rubric` dimension |
|
|
92
|
+
|
|
93
|
+
A **rule** graded as a rubric dimension becomes **tradeable**, which is the one thing a rule must
|
|
94
|
+
never be. That is a category error in the *selection*, and no threshold repairs it: a
|
|
95
|
+
non-substitutable criterion at any `max` and any `threshold` is still purchasable with points
|
|
96
|
+
earned elsewhere. One scenario routinely carries both forms — its rules as boolean `Then`s, its one
|
|
97
|
+
genuine gradient as a `@rubric`. Keeping a criterion out of the sum does not weaken the scenario: the
|
|
98
|
+
rule fails the subject outright instead of being priced.
|
|
99
|
+
|
|
100
|
+
**"Partly substitutable" means you are holding two criteria, not one.** The common case: you would
|
|
101
|
+
accept a thinner test suite paid for by strength elsewhere, but you would never accept *no tests at
|
|
102
|
+
all*. That is not one criterion with a floor under it — it is a **rule** (*there is at least one
|
|
103
|
+
test*) and a **gradient** (*how thorough it is*) that got named once. Split them and select each on
|
|
104
|
+
its own: the rule becomes a boolean `Then`, the gradient stays a dimension. Reaching instead for a
|
|
105
|
+
minimum on the dimension is the conjunctive move below, and it is the wrong one.
|
|
106
|
+
|
|
107
|
+
**Write the trade down.** Said out loud and left there, a trade leaves no trace — and a trade nobody
|
|
108
|
+
wrote is one nobody can disagree with: an **unowned selection**, exactly as a `threshold:` with no
|
|
109
|
+
recorded reason is an unowned policy. Put it in the **same record that carries the cut's reason**
|
|
110
|
+
(below, which is also where that record may live), naming the trade **and what pays for it** —
|
|
111
|
+
*"thinner edge-case coverage is paid for by stronger error handling."*
|
|
112
|
+
|
|
113
|
+
**The record is for the owner, not the judge.** Its reader is whoever reviews this suite later and
|
|
114
|
+
may disagree with it — the same audience and the same purpose as the cut's recorded reason. The
|
|
115
|
+
spec-judge does **not** grade it and never fails a dimension over it: selection is re-derived from
|
|
116
|
+
the **dimensions themselves**, exactly as where no record exists. A producer's own account of its
|
|
117
|
+
trade is not evidence, and putting it in the judge's path would buy that account a vote it has not
|
|
118
|
+
earned. What the record buys is that the selection stops being **unfalsifiable** — not detection of
|
|
119
|
+
a contested trade (*The soft spot*, below), and not any judgment the judge would otherwise skip.
|
|
120
|
+
|
|
121
|
+
The duty binds **when you author or revise a dimension**, and it is **the producer's alone**: no
|
|
122
|
+
judge reports a missing record, whoever authored it, so nothing catches you skipping it. That is the
|
|
123
|
+
honest state of it. Correcting the standing corpus is its own work — and **recording** a trade forces
|
|
124
|
+
no threshold re-derivation, changing no `max` and no cut; that is what **stripping** a
|
|
125
|
+
non-substitutable dimension forces (above).
|
|
126
|
+
|
|
127
|
+
**This is not conjunctive scoring.** Conjunctive scoring keeps every criterion in the rubric and
|
|
128
|
+
adds a per-dimension minimum each must clear. That is a worse instrument, not a safer one — the
|
|
129
|
+
**least-reliable subscore controls the outcome**, and it buys its fewer false passes with more
|
|
130
|
+
**false negative classification errors**. Per-dimension hurdles are not a safe default; do not reach
|
|
131
|
+
for them. Substitutability makes the opposite move: a non-substitutable criterion does not become a
|
|
132
|
+
graded subscore with a floor under it, it **never enters the rubric at all** and is asserted as a
|
|
133
|
+
boolean instead. Nothing is graded, so no subscore's reliability controls anything.
|
|
134
|
+
|
|
135
|
+
## The threshold is a policy call, not a derived one
|
|
136
|
+
|
|
137
|
+
Where the cut sits is decided by the **relative cost of a false positive (a wrong subject that
|
|
138
|
+
passes) against a false negative (a good subject that fails)** — a values decision the spec owner
|
|
139
|
+
makes on purpose. **It is not computable from the data.** Putting the cut where two subjects' scores
|
|
140
|
+
meet minimizes total error *only* when a false pass and a false fail cost the same and the two kinds
|
|
141
|
+
of subject are about equally likely — a claim about your domain, never a default to inherit by
|
|
142
|
+
arithmetic.
|
|
143
|
+
|
|
144
|
+
**Record why.** A `threshold:` line with no recorded reason is an unowned policy. Next to the rubric
|
|
145
|
+
or in the node's spec body, name which error the cut buys down and what it costs — *"a false pass
|
|
146
|
+
ships a broken contract and a false fail costs one more authoring round, so the cut sits high."*
|
|
147
|
+
|
|
148
|
+
**One record carries both** the cut's reason and each dimension's trade (above) — they are decided
|
|
149
|
+
together and reviewed together, and two parallel records are two things to drift.
|
|
150
|
+
|
|
151
|
+
## Correcting a standing rubric — a queue of decisions, not a sweep
|
|
152
|
+
|
|
153
|
+
Selection (above) governs authoring. A rubric that **already** sums a non-substitutable criterion is
|
|
154
|
+
a different act: **correcting** it. The correction looks mechanical and is not. Removing a dimension
|
|
155
|
+
does two things at once — it **adds** a boolean `Then` asserting the criterion (strengthening the
|
|
156
|
+
contract), and it **removes a scored dimension, which invalidates the cut**.
|
|
157
|
+
|
|
158
|
+
**The second half is the trap, because it is silent.** Strip a `max: 3` dimension out of
|
|
159
|
+
`[3, 3, 2] threshold: 6` and the survivors reach 5 against a cut of 6 — **nothing passes at all.**
|
|
160
|
+
Strip a `max: 2` out of `[2, 3, 2] threshold: 5` and the survivors total exactly 5 — only a
|
|
161
|
+
**flawless** subject passes. Neither reports anything. A rubric no subject can pass reads like a
|
|
162
|
+
**strict bar**, not a broken one.
|
|
163
|
+
|
|
164
|
+
### The re-derivation duty
|
|
165
|
+
|
|
166
|
+
**Removing a dimension changes the attainable maximum, so the `threshold:` left behind is an
|
|
167
|
+
un-re-derived cut — whether or not its number still needs to change.** A correction that removes a
|
|
168
|
+
dimension must, **in the same edit**, re-derive the cut as a fresh **policy call** (above) and record
|
|
169
|
+
its reason, and that reason names the **new attainable maximum** it was set against.
|
|
170
|
+
|
|
171
|
+
This is why the correction **cannot be computed**. The cut is set from the relative cost of a false
|
|
172
|
+
pass against a false fail (above) — not derivable from the data, so not derivable by a script
|
|
173
|
+
sweeping the corpus. Each rubric's cut is **its own decision**, and a sweep has no one to make it.
|
|
174
|
+
|
|
175
|
+
### One rubric, one decision, one ratification
|
|
176
|
+
|
|
177
|
+
A correction to a `@frozen` suite **routes to Clearance** (one of the four C's hard floors, defined
|
|
178
|
+
in `plugins/sdd/README.md`) and takes the owner's ratification. **One ratification per rubric.** A blanket approval over a batch is not a
|
|
179
|
+
faster path to the same place: it collapses N individual policy calls into **one unexamined one**,
|
|
180
|
+
which is the very thing the duty exists to prevent. Correct the corpus as a **queue** — each item
|
|
181
|
+
carrying its own re-derived cut and that cut's recorded reason — never as one sweep.
|
|
182
|
+
|
|
183
|
+
Two consequences follow, and both cut against batching:
|
|
184
|
+
|
|
185
|
+
- **The queue's membership is a judgment, not a match.** Which rubrics sum a non-substitutable
|
|
186
|
+
criterion is settled by the substitutability test per criterion, in the subject's own domain — not
|
|
187
|
+
by a name pattern. A conjunctive **name** (`..._and_...`) is a strong hint and an incomplete one:
|
|
188
|
+
a criterion can join two concerns under a name that reads singular, and a name can read conjunctive
|
|
189
|
+
while the `_or_` is one disjunctive subject rather than two criteria. A regex sweep therefore both
|
|
190
|
+
over- and under-fires, and its under-fire is the dangerous half — it **looks complete**.
|
|
191
|
+
- **Split before you correct.** A double-barreled dimension is split first (above), and its halves
|
|
192
|
+
routinely land in different forms. The correction proceeds per **half**, not per name.
|
|
193
|
+
|
|
194
|
+
### Why the routing does not depend on the edit class
|
|
195
|
+
|
|
196
|
+
State the routing from what is true of **every** such correction, not from the shape one happens to
|
|
197
|
+
take:
|
|
198
|
+
|
|
199
|
+
> **Removing a dimension always modifies the baseline rubric scenario.** So the correction is never
|
|
200
|
+
> `additive` and never `no-content-change`. It lands as **`narrowing`** (the boolean `Then` joins the
|
|
201
|
+
> same scenario) or **`mixed`** (the boolean `Then` becomes a separate `Scenario`), and **both route
|
|
202
|
+
> to Clearance.** Which one it is never changes the answer.
|
|
203
|
+
|
|
204
|
+
Do **not** encode the rule as "a mixed edit routes to Clearance." It is true and it is the wrong
|
|
205
|
+
handle: the natural in-scenario shape is `narrowing`, so a producer checking whether its edit is
|
|
206
|
+
`mixed` reads *no* and may conclude it self-clears. **It does not.** Route on the **removal**, which
|
|
207
|
+
is present in both shapes, never on the class the diff happens to report.
|
|
208
|
+
|
|
209
|
+
**The freeze sees the rubric only because the differ's pin says it does.** A `@rubric` lives wholly
|
|
210
|
+
inside a DocString. The structural differ is pinned at `gherkin-cli@0.0.2`, which hashes what a step
|
|
211
|
+
argument **says**; before that pin its scenario identity covered step text alone, and a rubric could
|
|
212
|
+
be gutted while its scenario still reported `unchanged`. **The pin is load-bearing here** — moved
|
|
213
|
+
backwards, every correction in this queue self-clears silently and Clearance never fires.
|
|
214
|
+
|
|
215
|
+
### The one mechanical check — and the bound it does not cover
|
|
216
|
+
|
|
217
|
+
**`sum(max) < threshold` is a dead rubric.** No subject, not even a perfect one, reaches the cut;
|
|
218
|
+
the scenario is `Then false` wearing a rubric's clothes. No policy could intend that of a rubric
|
|
219
|
+
meant to grade, so this is safe to **lint, fail-closed** — and it is **not a slack constant**: it
|
|
220
|
+
decrees no distance between anything, only that the passing set is **non-empty**.
|
|
221
|
+
|
|
222
|
+
**A green lint is not evidence the cut was re-derived, and the verdict says so.** The lint's bound is
|
|
223
|
+
exactly the vacuous case. `[2, 3, 2] threshold: 5` stripped to `[3, 2] threshold: 5` passes the lint
|
|
224
|
+
and is still an un-re-derived cut nobody chose — the duty above is what catches it, and only the
|
|
225
|
+
owner's ratification discharges it.
|
|
226
|
+
|
|
227
|
+
**Reach for no other arithmetic here.** A check that a surviving rubric keeps *enough* room under its
|
|
228
|
+
cut is a **slack constant** by another name, and the ban on those (*The margin is measured, not
|
|
229
|
+
decreed*, below) is not suspended because a correction prompted the question. Zero slack is not a
|
|
230
|
+
defect this bar can detect: only-a-flawless-subject-passes is a legitimate policy where the owner
|
|
231
|
+
**chose** it. The defect is that a strip nobody re-derived leaves that policy **unchosen**, and that
|
|
232
|
+
is a fact about the **edit**, not about the number.
|
|
233
|
+
|
|
234
|
+
## The soft spot — selection rests on one judgment, and that is accepted
|
|
235
|
+
|
|
236
|
+
**Do not look for a disambiguator on the contested case. There is none, and each search for one
|
|
237
|
+
repeats a settled mistake.**
|
|
238
|
+
|
|
239
|
+
The substitutability test is decisive on clear-cut criteria. On a **contested** one it gives no
|
|
240
|
+
traction, and a producer who wants a criterion to be a dimension can construct a trade that reads
|
|
241
|
+
plausibly. Nothing in this bar catches that. Each apparent guard is checked and does not reach:
|
|
242
|
+
|
|
243
|
+
- **Escalation does not reach it.** The escalate trigger keys on a judge's **inability** to
|
|
244
|
+
say whether it accepts the trade. A genuinely contested criterion produces **confident
|
|
245
|
+
disagreement** — competent judges landing on opposite sides, each sure — not inability.
|
|
246
|
+
*Contestable in fact* and *uncertain to this judge* are different properties; escalation fires on
|
|
247
|
+
the second only.
|
|
248
|
+
- **The recorded trade does not reach it.** A motivated producer writes the trade down and it reads
|
|
249
|
+
fine — and nothing grades it in any case (above).
|
|
250
|
+
- **Discrimination cannot back it up.** The subject that would expose a criterion smuggled into the
|
|
251
|
+
sum is one right about everything *except* that criterion — a **blemished good subject**, the
|
|
252
|
+
strawman the miss test bars. And selection runs **before** discrimination, so nothing downstream
|
|
253
|
+
re-asks the question.
|
|
254
|
+
|
|
255
|
+
So selection rests on **one confident judgment, with no second reader inside this bar.** That is
|
|
256
|
+
accepted deliberately, not overlooked. The backstops sit **outside** it: the cold judge's
|
|
257
|
+
independence, and an owner who can disagree with the recorded trade after the fact — which is what
|
|
258
|
+
the record is for.
|
|
259
|
+
|
|
260
|
+
The two moves that look like a fix here are already ruled out: **arithmetic** over `max` and
|
|
261
|
+
`threshold` and **per-dimension hurdles** (above). A reader who reaches for either is solving a different problem.
|
|
262
|
+
|
|
263
|
+
## Rubric-specific discrimination: loseable in arithmetic ≠ loseable in practice
|
|
264
|
+
|
|
265
|
+
Per named wrong subject, **sum what it scores; the sum sits strictly under the threshold**.
|
|
266
|
+
Strictly: the collapsing `Then` passes a score *at least* the threshold, so a wrong subject that
|
|
267
|
+
**ties** the threshold **passes**. A rubric whose free dimensions alone carry a wrong subject to the
|
|
268
|
+
threshold leaves its discriminating dimensions decorative.
|
|
269
|
+
|
|
270
|
+
> `[live 3] [restatement 3] [presence 2]`, `threshold: 6`. A memorizer banks `3+2=5` on
|
|
271
|
+
> `restatement` and `presence` without engaging `live` at all. Those two are free — a memorizer maxes
|
|
272
|
+
> both — so each fails the miss test on its own and is rewritten. The defect is the two decorative
|
|
273
|
+
> dimensions, not the distance from 5 to 6.
|
|
274
|
+
|
|
275
|
+
**Score each wrong subject at what it banks — never zero a dimension to make a point.** A wrong
|
|
276
|
+
subject scores each dimension at what that dimension's own rules award it, which for a memorizer on
|
|
277
|
+
a live dimension is rarely 0. Zeroing one dimension in turn and asking whether the rubric still
|
|
278
|
+
passes is **not** this test: it posits a subject right about everything except one thing — a
|
|
279
|
+
blemished **good** subject, precisely the strawman the miss test bars. A score profile does not
|
|
280
|
+
identify a subject: `[3, 0, 3]` is the memorizer's profile *and* a good-but-flawed subject's, and
|
|
281
|
+
they are not the same subject.
|
|
282
|
+
|
|
283
|
+
## The margin is measured, not decreed
|
|
284
|
+
|
|
285
|
+
**How far under the threshold a wrong subject must land is not a constant this bar can give you.**
|
|
286
|
+
That distance is only meaningful against the **noise of your own judge** — score the same subject
|
|
287
|
+
twice and the scores differ. The named quantity is the **conditional standard error of measurement
|
|
288
|
+
(cSEM)**: the judge's precision **at the cut score**, a **measured property of the instrument**, not
|
|
289
|
+
a number doctrine decrees. A single global reliability figure (Cronbach's alpha) is the wrong tool
|
|
290
|
+
for a pass/fail decision — it averages precision across the whole scale, and the only precision that
|
|
291
|
+
decides a pass is the precision at the cut.
|
|
292
|
+
|
|
293
|
+
**So measure it, per suite.** Score your named wrong subjects **more than once** and record whether
|
|
294
|
+
the scores reproduce. The one node in this corpus whose rubrics conform records exactly this in its
|
|
295
|
+
README — *"2.33/3 mean, measured twice, did not reproduce"* — and honestly flags that its one-point
|
|
296
|
+
slack is not evidence-backed.
|
|
297
|
+
|
|
298
|
+
**Copy that practice — the measuring and the honest record — and nothing else from it.** That node's
|
|
299
|
+
own threshold convention is its **local policy call**, not doctrine, and its design is not a template
|
|
300
|
+
this bar blesses. A slack measured against one suite's judge says nothing about yours. **The main
|
|
301
|
+
governance file says: where a node's design and this bar disagree, this bar wins.**
|
|
302
|
+
|
|
303
|
+
Any constant offered in place of that measurement — *"not by a single point"*, `gap ≥ 2`,
|
|
304
|
+
`max ≥ b + 2` — is a guess at an instrument property nobody measured. Every such constant this bar
|
|
305
|
+
has carried was wrong, and each one's repair produced the next.
|
|
306
|
+
|
|
307
|
+
**Naming subjects orders them; it never separates them.** Scoring one wrong subject and one good
|
|
308
|
+
subject and putting the cut between them is a crude instance of **contrasting groups** — a real
|
|
309
|
+
standard-setting method that scores a known-fail **group** and a known-pass **group** and reads the
|
|
310
|
+
cut off where the **distributions** separate. At one subject per group there is no distribution, no
|
|
311
|
+
variance, no intersection: two points establish an **ordering**, never a separation. The miss test is
|
|
312
|
+
a **sanity check** on a threshold you set as policy — it catches a cut that is plainly too low. It
|
|
313
|
+
never derives one.
|
|
@@ -0,0 +1,16 @@
|
|
|
1
|
+
# touch-set-correction
|
|
2
|
+
|
|
3
|
+
The concrete engine for **touch-set-correction** — a read-only, post-hoc reconciliation of a
|
|
4
|
+
Mission's declared touch-set against what its `git diff` actually changed, composing `git diff`,
|
|
5
|
+
[`resolve-governances`](../resolve-governances/SKILL.md), and `gherkin-cli diff` into the corrected
|
|
6
|
+
touch-set the mission-graph's single writer records at retirement. Built for the Op2 deferral of the
|
|
7
|
+
cyberfleet-batch change request; see
|
|
8
|
+
[`.agents/specs/sdd/touch-set-correction/README.md`](../../../../.agents/specs/sdd/touch-set-correction/README.md)
|
|
9
|
+
for the authoritative behavior description and
|
|
10
|
+
[`touch-set-correction.feature`](../../../../.agents/specs/sdd/touch-set-correction/touch-set-correction.feature)
|
|
11
|
+
for the frozen 21-scenario contract.
|
|
12
|
+
|
|
13
|
+
- **Skill contract:** [`SKILL.md`](./SKILL.md)
|
|
14
|
+
- **Script:** [`scripts/touch-set-correction.mts`](./scripts/touch-set-correction.mts)
|
|
15
|
+
- **Tests:** [`scripts/touch-set-correction.test.mts`](./scripts/touch-set-correction.test.mts)
|
|
16
|
+
(`node:test`) — one test per frozen scenario, titled `scenario: <verbatim frozen scenario name>`.
|
|
@@ -0,0 +1,67 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: touch-set-correction
|
|
3
|
+
description: "Internal skill: reconciles a Mission's declared touch-set against what its git diff actually changed — read-only, post-hoc; used to correct the mission-graph's hazard set from prediction to reality, not triggered by users directly."
|
|
4
|
+
user-invocable: false
|
|
5
|
+
metadata:
|
|
6
|
+
internal: true
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
# Touch-Set Correction
|
|
10
|
+
|
|
11
|
+
The concrete engine for **touch-set-correction** — the Op2 deferral of the cyberfleet-batch
|
|
12
|
+
self-hosting-kernel (issue #189, first bullet). A Mission's **declared touch-set** is a pre-work
|
|
13
|
+
**guess** of the work areas it will change, used by [`mission-graph`](../mission-graph/README.md)
|
|
14
|
+
to keep clashing Missions apart. This engine checks that guess against reality: given a Mission's
|
|
15
|
+
declared touch-set and its `git diff base..head`, it recovers the work areas **actually** touched
|
|
16
|
+
and lines them up against the guess:
|
|
17
|
+
|
|
18
|
+
- **confirmed** — declared ∩ actual (the guess was right)
|
|
19
|
+
- **missed** — actual − declared (touched but never guarded — the dangerous case)
|
|
20
|
+
- **overDeclared** — declared − actual (guessed but never touched — harmless)
|
|
21
|
+
- **corrected** — the actual touched set (the ground truth; what the graph's single writer records
|
|
22
|
+
at Mission retirement)
|
|
23
|
+
|
|
24
|
+
It **composes three tools**, never reimplementing any of them: `git diff --name-status` (the
|
|
25
|
+
changed files), [`resolve-governances`](../resolve-governances/SKILL.md) (each file's artifact-type,
|
|
26
|
+
best-effort — `unknown` when it doesn't resolve), and the pinned `gherkin-cli@0.0.2 diff` (a touched
|
|
27
|
+
`.feature`'s changed scenario names — the same tool `classify-edit-class` uses).
|
|
28
|
+
|
|
29
|
+
Work-area recovery is **capability-first**: a changed file under a declared project root maps to
|
|
30
|
+
`project/capability`, where the capability is the first path segment after the matched root (the
|
|
31
|
+
LONGEST matching root wins, so a nested root like `plugins/sdd/skills` beats a shallower
|
|
32
|
+
`plugins/sdd`). A spec file and its impl file sharing a capability segment collapse to one node. A
|
|
33
|
+
file under no known root is **unmapped** — surfaced, never silently dropped, never counted as
|
|
34
|
+
touched.
|
|
35
|
+
|
|
36
|
+
Pure derivations (`isFeature`, `fileToNode`, `reconcile`, `assembleCorrection`) take and return
|
|
37
|
+
plain data — no fs/network access — kept apart from a thin IO seam (`readChangedFiles`,
|
|
38
|
+
`resolveArtifactType`, `changedScenarios`, `collectChangedFiles`, `discoverLayouts`) that shells out
|
|
39
|
+
to `git`, `resolve-governances.mts`, and `npx gherkin-cli`.
|
|
40
|
+
|
|
41
|
+
## Run it
|
|
42
|
+
|
|
43
|
+
```bash
|
|
44
|
+
node "<skill>/scripts/touch-set-correction.mts" --base <ref> --declared a,b,c \
|
|
45
|
+
[--head HEAD] [--root .] \
|
|
46
|
+
[--layout 'sdd:.agents/specs/sdd,plugins/sdd/skills']... \
|
|
47
|
+
[--format toon|json]
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
- `--base` is required (the merge-base or comparison ref); `--head` defaults to `HEAD`; `--root`
|
|
51
|
+
defaults to `.`.
|
|
52
|
+
- `--layout` is repeatable (`<project>:<root1>,<root2>`). Omit it to auto-discover project layouts
|
|
53
|
+
via `discover-specs` (each project's spec-path plus an impl-root convention: `plugins/<p>` also
|
|
54
|
+
gets `plugins/<p>/skills`; `packages/<p>` also gets `packages/<p>/src`; anything else falls back
|
|
55
|
+
to the project-path itself).
|
|
56
|
+
- Default output is **TOON**; `--format json` emits the full `Correction` record: `corrected`,
|
|
57
|
+
the `confirmed`/`missed`/`overDeclared` split, `unmapped`, and per-node `files` (each `{ path,
|
|
58
|
+
artifactType }`) + `changedScenarios`.
|
|
59
|
+
|
|
60
|
+
## Boundaries
|
|
61
|
+
|
|
62
|
+
**Read-only w.r.t. the mission graph** — it never writes to it; the graph's single writer appends
|
|
63
|
+
the corrected touch-set at Mission retirement (the lifecycle loop, deferred to F3). It does **not**
|
|
64
|
+
decide whether two Missions' touch-sets collide hard or soft, run the finer-than-node ladder (file →
|
|
65
|
+
region → semantic downgrade), descend to a region/hunk tier, do SSA lowering or infer symbol-level
|
|
66
|
+
produce/consume dependencies (all later parts of issue #189), or **predict** a touch-set before the
|
|
67
|
+
work — it only corrects one, after.
|