@orkestrel/scaffold 0.0.59 → 0.0.61
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +13 -10
- package/dist/bin/main.js +632 -320
- package/dist/bin/main.js.map +1 -1
- package/dist/host/CLAUDE.md +5 -1
- package/dist/host/agents/orchestration.md +44 -19
- package/dist/host/agents/skills/enterprise-bootstrap/SKILL.md +91 -82
- package/dist/host/agents/skills/enterprise-bootstrap/references/bootstrap-reference.md +5 -5
- package/dist/host/agents/skills/enterprise-bootstrap/references/components.md +15 -15
- package/dist/host/agents/skills/enterprise-bootstrap/references/inputs.md +501 -0
- package/dist/host/agents/skills/enterprise-bootstrap/references/inspection.md +167 -0
- package/dist/host/agents/skills/enterprise-bootstrap/references/utilities.md +2 -2
- package/dist/host/agents/skills/orkestrel-debrief/SKILL.md +2 -2
- package/dist/host/agents/skills/orkestrel-debrief/references/instruction-audit.md +10 -9
- package/dist/host/agents/skills/orkestrel-falsify/SKILL.md +21 -11
- package/dist/host/agents/skills/orkestrel-falsify/references/brief.md +20 -2
- package/dist/host/agents/skills/orkestrel-harden-package/SKILL.md +1 -1
- package/dist/host/agents/skills/orkestrel-harden-package/references/hardening.md +1 -1
- package/dist/host/agents/skills/orkestrel-harden-package/references/research.md +1 -1
- package/dist/host/agents/skills/orkestrel-polish-surface/SKILL.md +4 -1
- package/dist/host/agents/skills/orkestrel-polish-surface/references/capture-harness.md +71 -50
- package/dist/host/agents/skills/orkestrel-prove-journey/SKILL.md +93 -29
- package/dist/host/agents/skills/orkestrel-prove-journey/agents/openai.yaml +1 -1
- package/dist/host/agents/skills/orkestrel-prove-journey/references/captures.md +62 -38
- package/dist/host/agents/skills/orkestrel-prove-journey/references/decide.md +68 -0
- package/dist/host/agents/skills/orkestrel-prove-journey/references/layer.md +107 -79
- package/dist/host/agents/skills/orkestrel-prove-journey/references/statechart.md +84 -0
- package/dist/host/agents/skills/orkestrel-prove-journey/references/styles.md +87 -0
- package/dist/host/agents/templates/brief.md +16 -7
- package/dist/host/agents/transports/claude.md +4 -2
- package/dist/host/agents/transports/codex.md +4 -1
- package/dist/host/claude/agents/analyst.md +3 -1
- package/dist/host/claude/agents/application.md +1 -1
- package/dist/host/claude/agents/builder.md +3 -3
- package/dist/host/claude/agents/checker.md +5 -0
- package/dist/host/claude/agents/grok.md +15 -5
- package/dist/host/claude/agents/implementer.md +1 -1
- package/dist/host/claude/agents/orkestrel.md +12 -11
- package/dist/host/claude/agents/planner.md +10 -0
- package/dist/host/claude/agents/reviewer.md +14 -8
- package/dist/host/claude/agents/sol.md +3 -1
- package/dist/host/claude/agents/verifier.md +2 -4
- package/dist/host/claude/rules/architecture.md +7 -5
- package/dist/host/claude/rules/documentation.md +1 -0
- package/dist/host/claude/rules/names.md +23 -5
- package/dist/host/claude/rules/patterns.md +1 -0
- package/dist/host/claude/rules/quality.md +1 -1
- package/dist/host/claude/rules/tests.md +3 -3
- package/dist/host/claude/rules/typescript.md +4 -1
- package/dist/host/claude/rules/writing.md +2 -2
- package/dist/host/claude/skills/orkestrel-prove-journey/SKILL.md +1 -1
- package/dist/host/codex/agents/builder.toml +6 -6
- package/dist/host/codex/agents/checker.toml +2 -1
- package/dist/host/codex/agents/grok.toml +12 -5
- package/dist/host/codex/agents/implementer.toml +2 -2
- package/dist/host/codex/agents/opus.toml +6 -1
- package/dist/host/codex/agents/planner.toml +11 -6
- package/dist/host/codex/agents/reviewer.toml +8 -6
- package/dist/host/guides/scaffold.md +39 -14
- package/dist/host/manifest.json +81 -51
- package/dist/host/scripts/codex.sh +0 -0
- package/dist/host/scripts/cursor.sh +0 -0
- package/dist/host/scripts/deps.sh +0 -0
- package/dist/host/scripts/ollama.sh +0 -0
- package/dist/src/core/index.cjs +424 -282
- package/dist/src/core/index.cjs.map +1 -1
- package/dist/src/core/index.d.cts +361 -220
- package/dist/src/core/index.d.ts +361 -220
- package/dist/src/core/index.js +421 -283
- package/dist/src/core/index.js.map +1 -1
- package/dist/src/server/index.cjs +208 -170
- package/dist/src/server/index.cjs.map +1 -1
- package/dist/src/server/index.d.cts +276 -152
- package/dist/src/server/index.d.ts +276 -152
- package/dist/src/server/index.js +200 -172
- package/dist/src/server/index.js.map +1 -1
- package/package.json +8 -7
|
@@ -0,0 +1,167 @@
|
|
|
1
|
+
# Instruments
|
|
2
|
+
|
|
3
|
+
Reach for an instrument here where a capture cannot settle the claim. Run every one with its negative
|
|
4
|
+
control, in the same conditions, on the real compiled cascade the page loads. Treat an instrument
|
|
5
|
+
whose negative control passes as broken and refuse its reading as evidence;
|
|
6
|
+
`.claude/rules/quality.md` owns that law where it is present.
|
|
7
|
+
|
|
8
|
+
Take each entry's property, population, reading, negative control, and coverage as written, and hold
|
|
9
|
+
every entry to these rules.
|
|
10
|
+
|
|
11
|
+
- **Report the population.** A reading carries the population it walked. An empty population fails
|
|
12
|
+
the run, because an extractor that quietly matched nothing satisfies every other assertion.
|
|
13
|
+
- **Draw every negative control from outside the population.** Name the membership rule first, then
|
|
14
|
+
pick a negative control the rule excludes. Reject a negative control that rule admits.
|
|
15
|
+
- **Enter a negative control through the same door the surface enters.** A negative control handed
|
|
16
|
+
straight to the reading tests the reading alone, so pair it with one appended to the tree, the stylesheet,
|
|
17
|
+
or the registry the instrument walks wherever the instrument has an extraction step.
|
|
18
|
+
|
|
19
|
+
## Contents
|
|
20
|
+
|
|
21
|
+
- [Authored class in the shipped cascade](#authored-class-in-the-shipped-cascade)
|
|
22
|
+
- [Declared class combinations](#declared-class-combinations)
|
|
23
|
+
- [Style escapes](#style-escapes)
|
|
24
|
+
- [Token discipline](#token-discipline)
|
|
25
|
+
- [Custom rule doing a utility's job](#custom-rule-doing-a-utilitys-job)
|
|
26
|
+
- [Composited contrast in both themes](#composited-contrast-in-both-themes)
|
|
27
|
+
- [One glyph, one meaning](#one-glyph-one-meaning)
|
|
28
|
+
- [When an authored rule is already earned](#when-an-authored-rule-is-already-earned)
|
|
29
|
+
|
|
30
|
+
## Authored class in the shipped cascade
|
|
31
|
+
|
|
32
|
+
- **Property.** Every class token the surface's own templates and components author has a rule in
|
|
33
|
+
the compiled CSS the page loads.
|
|
34
|
+
- **Population.** The class tokens the authored markup carries, read against every stylesheet the
|
|
35
|
+
page loads: the vendor build, each skin, and the project's own.
|
|
36
|
+
- **Reading.** Subtract the tokens the loaded stylesheets define from the tokens the markup carries.
|
|
37
|
+
A remainder fails the run and names each token with the file that authored it. Report the token
|
|
38
|
+
population walked, and fail a run that walked none.
|
|
39
|
+
- **Negative control.** A fed control and an appended control, each of which the reading must report.
|
|
40
|
+
Feed the first — a token no stylesheet defines — straight to the reading. Append the second through
|
|
41
|
+
the extraction door: an element built in the harness carrying that undefined token on an SVG
|
|
42
|
+
`class` attribute, added to the tree the reading walks. Each sits outside the population, which is
|
|
43
|
+
authored tokens the cascade resolves.
|
|
44
|
+
- **Coverage.** The fed negative control covers the subtraction. The appended negative control covers the extractor,
|
|
45
|
+
and it is what fails a reading that never leaves the root or that drops SVG tokens by reaching for
|
|
46
|
+
`className`, where the value is an `SVGAnimatedString` rather than a string. Together they prove
|
|
47
|
+
authored tokens are a subset of the cascade. The instrument says nothing about a cascade rule
|
|
48
|
+
nobody authored, a token a build step or a script adds after the read, or whether a resolved rule
|
|
49
|
+
paints what the author intended.
|
|
50
|
+
|
|
51
|
+
## Declared class combinations
|
|
52
|
+
|
|
53
|
+
- **Property.** Every multi-utility chrome string the surface reuses is declared once by name with
|
|
54
|
+
the invariant it holds, and the markup carries no undeclared combination.
|
|
55
|
+
- **Population.** The declared combinations, each with its name and its invariant, and every
|
|
56
|
+
multi-utility string the authored markup carries.
|
|
57
|
+
- **Reading.** Match each string in the markup against the declared set. An undeclared combination
|
|
58
|
+
fails and names the element that carries it.
|
|
59
|
+
- **Negative control.** A fed control and an appended control, each of which the reading must refuse.
|
|
60
|
+
Feed the first — a string one utility away from a declared combination — straight to the reading.
|
|
61
|
+
Append the second through the extraction door: an element built in the harness carrying that same
|
|
62
|
+
undeclared string on an SVG `class` attribute, added to the tree the reading walks. Each sits
|
|
63
|
+
outside the declared set.
|
|
64
|
+
- **Coverage.** The fed negative control covers the match against the declared set. The appended negative
|
|
65
|
+
control covers the extractor, and it is what fails a reading that never leaves the root or that drops SVG
|
|
66
|
+
tokens by reaching for `className`. Together they prove reused chrome is declared. The instrument
|
|
67
|
+
does not prove a declared invariant is true, and it does not read a single utility used alone.
|
|
68
|
+
- **Never substitute a cancellation heuristic.** A rule that flags a string for its utility count, or
|
|
69
|
+
for mixing categories, refuses the legitimate transparent read chrome that keeps a read view and an
|
|
70
|
+
edit view from reflowing.
|
|
71
|
+
|
|
72
|
+
## Style escapes
|
|
73
|
+
|
|
74
|
+
- **Property.** The surface's own markup carries no `style` attribute and no `<style>` element.
|
|
75
|
+
- **Population.** The elements of a freshly mounted, undriven tree — the surface as authored, before
|
|
76
|
+
any interaction drives it.
|
|
77
|
+
- **Reading.** Collect every element carrying an inline declaration or an embedded style element and
|
|
78
|
+
report it with its markup. Any hit fails.
|
|
79
|
+
- **Exemptions, declared by name.** Exempt the framework's own runtime styles and name each exemption
|
|
80
|
+
in the instrument: a Bootstrap Modal, Offcanvas, Collapse, or Dropdown writes inline styles as it
|
|
81
|
+
runs, and a conditional-visibility directive such as `v-show` emits `style="display: none"` at
|
|
82
|
+
mount. Run the reading on an undriven tree, because a reading taken after a journey drives the
|
|
83
|
+
surface reports on the framework rather than on the author.
|
|
84
|
+
- **Negative control.** An element carrying an inline declaration, built in the harness rather
|
|
85
|
+
than taken from the surface, fed to the reading. It sits outside the surface's own markup and the
|
|
86
|
+
reading must report it.
|
|
87
|
+
- **Coverage.** The instrument reads authored markup at mount. It does not see a style a component
|
|
88
|
+
writes after the person interacts, a rule authored in a stylesheet, or an escape inside a
|
|
89
|
+
third-party component's own markup.
|
|
90
|
+
|
|
91
|
+
## Token discipline
|
|
92
|
+
|
|
93
|
+
- **Property.** No authored rule carries a literal color, and every custom paint resolves through a
|
|
94
|
+
token in each color mode the product ships.
|
|
95
|
+
- **Population.** The project's own authored stylesheet rules, and the resolved value of each
|
|
96
|
+
custom-painted property in each color mode.
|
|
97
|
+
- **Reading.** A literal color in an authored declaration fails. For each custom paint, read the
|
|
98
|
+
resolved value once per mode; a mode that leaves it unresolved fails, and so does a pair of modes
|
|
99
|
+
that resolve it identically where the design says the modes differ.
|
|
100
|
+
- **Negative control.** A rule carrying a literal color, and a paint whose token the cascade does
|
|
101
|
+
not define, both fed to the reading rather than authored into the surface. Each sits outside the
|
|
102
|
+
population of authored rules that pass, and the reading must report both.
|
|
103
|
+
- **Coverage.** The instrument covers authored rules and the paints it was given. It does not judge
|
|
104
|
+
whether the chosen token is the right one, and it reads no vendor rule and no inline declaration —
|
|
105
|
+
[Style escapes](#style-escapes) covers those.
|
|
106
|
+
|
|
107
|
+
## Custom rule doing a utility's job
|
|
108
|
+
|
|
109
|
+
- **Property.** Every authored selector expresses something no shipped utility expresses, or records
|
|
110
|
+
the reason the utility does not fit.
|
|
111
|
+
- **Population.** The selectors in the project's own stylesheets.
|
|
112
|
+
- **Reading.** For each selector, name the utility that would carry the same declarations. A selector
|
|
113
|
+
a shipped utility already expresses fails unless it carries the recorded reason.
|
|
114
|
+
- **Negative control.** A rule restating a shipped utility exactly — a padding declaration matching
|
|
115
|
+
a spacing step — fed to the reading rather than authored into the stylesheet. It sits outside the
|
|
116
|
+
set of authored selectors that pass, and the reading must report it.
|
|
117
|
+
- **Coverage.** The instrument reads declarations, not intent. A rule that does a utility's job
|
|
118
|
+
alongside something else passes it, so a person still reads the authored stylesheet.
|
|
119
|
+
|
|
120
|
+
## Composited contrast in both themes
|
|
121
|
+
|
|
122
|
+
- **Property.** Every pairing the surface paints meets its bar in every theme: 4.5:1 for anything
|
|
123
|
+
information-bearing, 3:1 for textless marks and the chrome that carries state.
|
|
124
|
+
- **Population.** The pairings the surface renders, read per theme on the compiled cascade, with
|
|
125
|
+
every translucent layer composited. Exempt disabled controls, per [SKILL.md](../SKILL.md) →
|
|
126
|
+
Surfaces, color, contrast.
|
|
127
|
+
- **Reading.** Composite the painted layers, read the ratio, and fail anything under its bar with the
|
|
128
|
+
pairing named. Take the mechanics from [bootstrap-reference.md](bootstrap-reference.md) → Measuring
|
|
129
|
+
the bars.
|
|
130
|
+
- **Negative control.** An opaque pairing and a translucent stack, each of which the reading must
|
|
131
|
+
fail. Compose the opaque pairing in the harness at a ratio known to sit under the bar. Compose the
|
|
132
|
+
stack with a translucent layer over a floor, so its composited ratio sits under the bar while its
|
|
133
|
+
top layer read alone sits above it. Each is composed in the harness rather than taken from the
|
|
134
|
+
surface, so each sits outside the rendered population.
|
|
135
|
+
- **Coverage.** The opaque pairing covers the ratio arithmetic. The translucent stack covers the
|
|
136
|
+
compositing step, and it is what fails a reader that takes the top layer's declared color and skips
|
|
137
|
+
the layers under it. Together they measure what rendered, in the themes and viewports the run
|
|
138
|
+
entered. A pairing that appears only in a state the run never reached is unmeasured, so name the
|
|
139
|
+
states the run covered beside the result.
|
|
140
|
+
|
|
141
|
+
## One glyph, one meaning
|
|
142
|
+
|
|
143
|
+
- **Property.** Each status meaning takes one glyph, each glyph serves one meaning, and every
|
|
144
|
+
registered glyph resolves in the icon set the product actually ships.
|
|
145
|
+
- **Population.** The registry of meanings and glyphs the surface uses, and the shipped icon set.
|
|
146
|
+
- **Reading.** A meaning registered twice, a glyph registered against two meanings, or a glyph the
|
|
147
|
+
shipped set does not resolve fails, each named.
|
|
148
|
+
- **Negative control.** A registry entry binding a second meaning to a glyph already registered,
|
|
149
|
+
plus a glyph name the shipped set lacks. Both sit outside the registered set, and the reading must
|
|
150
|
+
report both.
|
|
151
|
+
- **Coverage.** The instrument proves the registry is consistent and resolvable. It does not prove
|
|
152
|
+
the markup draws the registered glyph for the meaning it carries, so pair it with a capture of the
|
|
153
|
+
states that use marks.
|
|
154
|
+
|
|
155
|
+
## When an authored rule is already earned
|
|
156
|
+
|
|
157
|
+
Leave rung 4 to the developer, per [SKILL.md](../SKILL.md) → When custom CSS is justified. Write an
|
|
158
|
+
authored rule without asking only when every one of these holds:
|
|
159
|
+
|
|
160
|
+
- an instrument here reports the vendor cascade failing a stated bar — the focus ring under 3:1, the
|
|
161
|
+
status text under 4.5:1, the shipped component with no class for the state the surface must
|
|
162
|
+
show;
|
|
163
|
+
- the rule cites that reading beside it, naming the instrument, the bar, and the value read;
|
|
164
|
+
- the rule restores the bar and does nothing else;
|
|
165
|
+
- the rule is written over `--bs-*` tokens, so both color modes move with the theme.
|
|
166
|
+
|
|
167
|
+
Treat anything wider as a proposal: name what the rule would buy, and stop.
|
|
@@ -39,7 +39,7 @@ Prefer `bg-body-*` and `*-subtle` over `bg-white`/`bg-light` — they track `dat
|
|
|
39
39
|
.rounded-0, .rounded-1, .rounded-2, .rounded-3, .rounded-4, .rounded-5
|
|
40
40
|
```
|
|
41
41
|
|
|
42
|
-
For borders that must stay visible in both color modes, prefer `border-*-subtle`
|
|
42
|
+
For borders that must stay visible in both color modes, prefer the `border-*-subtle` classes (theme-adaptive) over raw color borders.
|
|
43
43
|
|
|
44
44
|
### Colors (Text)
|
|
45
45
|
|
|
@@ -246,7 +246,7 @@ The composition traps in this group:
|
|
|
246
246
|
### Z-index
|
|
247
247
|
|
|
248
248
|
```css
|
|
249
|
-
.z-n1, .z-0, .z-1, .z-2, .z-3 /* NOT responsive — no breakpoint
|
|
249
|
+
.z-n1, .z-0, .z-1, .z-2, .z-3 /* NOT responsive — no breakpoint classes exist */
|
|
250
250
|
```
|
|
251
251
|
|
|
252
252
|
## Spacing Scale
|
|
@@ -71,8 +71,8 @@ a practice that worked so it repeats.
|
|
|
71
71
|
4. **Process retrospective.** Walk the campaign record for both failure and success:
|
|
72
72
|
dispatches that deviated and why; recoveries that worked (codify the mechanism that
|
|
73
73
|
saved them); estimates versus observed durations; audit rounds that caught real
|
|
74
|
-
defects versus rounds that churned; anything the orchestrator absorbed that
|
|
75
|
-
|
|
74
|
+
defects versus rounds that churned; anything the orchestrator absorbed that a dispatch
|
|
75
|
+
owned, or dispatched that the orchestrator owned.
|
|
76
76
|
5. **Instruction-set audit.** Audit the agents, rules, skills, and orchestration
|
|
77
77
|
contract themselves against the campaign record, using the adversarial method in
|
|
78
78
|
[instruction-audit.md](references/instruction-audit.md). What confused an executor is
|
|
@@ -7,11 +7,12 @@ evidence-first treatment as any surface.
|
|
|
7
7
|
## Blind passes, one brief
|
|
8
8
|
|
|
9
9
|
Run the subjective lane and the objective lane on the SAME brief, in parallel, neither
|
|
10
|
-
seeing the other's answer before both return. `
|
|
11
|
-
|
|
10
|
+
seeing the other's answer before both return. `.agents/orchestration.md` § Engine
|
|
11
|
+
assignment decides which engine and which role holds each lane, including under a dark
|
|
12
|
+
bench; read the assignment there and name it in the dispatch.
|
|
12
13
|
|
|
13
14
|
Each lane returns numbered findings, most severe first, and exactly one terminal line:
|
|
14
|
-
`INSTRAUDIT <LANE>: <
|
|
15
|
+
`INSTRAUDIT <LANE>: <finding ids, or none>`. Each charter defaults to the `orkestrel-falsify`
|
|
15
16
|
verdict shape and takes a different shape the dispatch names, so name `orkestrel-debrief`
|
|
16
17
|
in the dispatch and this shape binds.
|
|
17
18
|
|
|
@@ -25,9 +26,9 @@ record.
|
|
|
25
26
|
|
|
26
27
|
## The subjective lens list
|
|
27
28
|
|
|
28
|
-
Held by
|
|
29
|
-
role's job is one job, and whether the skill family reads as one system. This section
|
|
30
|
-
the lens list's only normative home, and the lane states its coverage against it.
|
|
29
|
+
Held by the subjective lane. It judges coherence of the role model, charter voice, whether
|
|
30
|
+
each role's job is one job, and whether the skill family reads as one system. This section
|
|
31
|
+
is the lens list's only normative home, and the lane states its coverage against it.
|
|
31
32
|
|
|
32
33
|
- **Role-job singularity.** Is each charter's work cohesive? A charter describing bundled
|
|
33
34
|
jobs is either a role to split or a bundle no dispatch sends whole.
|
|
@@ -48,9 +49,9 @@ the lens list's only normative home, and the lane states its coverage against it
|
|
|
48
49
|
|
|
49
50
|
## The objective lens list
|
|
50
51
|
|
|
51
|
-
Held by
|
|
52
|
-
record. This section is the lens list's only normative home, and the lane states
|
|
53
|
-
against it.
|
|
52
|
+
Held by the objective lane. It runs evidence-only sweeps of the actual files and the
|
|
53
|
+
campaign record. This section is the lens list's only normative home, and the lane states
|
|
54
|
+
its coverage against it.
|
|
54
55
|
|
|
55
56
|
- **Duplication diff.** Whole-line and obligation-level comparison across charters, rules,
|
|
56
57
|
and skills. A charter that restates a rule drifts from it; a rule restated elsewhere has
|
|
@@ -72,13 +72,15 @@ nobody claimed.
|
|
|
72
72
|
brief that points at a rule file, section, or guide the executor cannot find delivers nothing while
|
|
73
73
|
looking like authority — and it fails silently, because an auditor does not report a heading it
|
|
74
74
|
never saw. Check before dispatch; propagate the missing file rather than restating its contents in
|
|
75
|
-
the brief. This is the reason restatement felt necessary, and it is the wrong cure.
|
|
75
|
+
the brief. This is the reason restatement felt necessary, and it is the wrong cure. Where the
|
|
76
|
+
executor's tree carries a superseded vendored copy of an authority, take the stale-authority
|
|
77
|
+
branch in `references/brief.md` § "What not to put in a brief".
|
|
76
78
|
- Run the **adversarial pass** on one identical brief: a subjective lane and an objective
|
|
77
79
|
lane, each a fresh subagent with a clean context, blind to each other. Reconcile them yourself.
|
|
78
80
|
`.agents/orchestration.md` owns lane definitions, engine assignment, and what happens when an
|
|
79
81
|
engine is dark; do not restate them here.
|
|
80
|
-
-
|
|
81
|
-
|
|
82
|
+
- `.agents/orchestration.md` § Execution loop owns which lanes a round runs and the deviation a
|
|
83
|
+
short round records; § Engine assignment owns the substitution.
|
|
82
84
|
- **Pair every finder with an independent refuter when the round fans out past the subjective and
|
|
83
85
|
objective lanes.** The
|
|
84
86
|
refuter receives one slice's findings, never that finder's work, and is briefed to BREAK them
|
|
@@ -135,27 +137,35 @@ comparable; a round that invents its own cannot be read against the last one.
|
|
|
135
137
|
| --------------- | -------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
136
138
|
| `CONFIRMED` | attacked and it held | as the Falsification law requires |
|
|
137
139
|
| `BROKEN` | falsified | as the Falsification law requires — note it says _input, **state, or interleaving**_, so a concurrency claim is falsified by an interleaving, not by an input |
|
|
138
|
-
| `UNRESOLVED` | cannot be decided from the evidence available | what would settle it
|
|
140
|
+
| `UNRESOLVED` | cannot be decided from the evidence available | what would settle it; a claim whose only evidence is the writer's own report takes this value |
|
|
139
141
|
| `NOT-EVIDENCED` | a claim about a rendered or externally driven surface the supplied capture cannot show | which capture is missing |
|
|
140
142
|
|
|
141
143
|
`CONFIRMED` and `BROKEN` defer; only `UNRESOLVED` and `NOT-EVIDENCED` are this skill's, because
|
|
142
144
|
the law does not name them. `BROKEN` and `UNRESOLVED` are **separate**: a claim nobody could
|
|
143
145
|
decide has not been falsified, and it cannot supply the fields falsification requires.
|
|
144
|
-
`NOT-EVIDENCED` is the
|
|
145
|
-
not re-invented.
|
|
146
|
+
`NOT-EVIDENCED` is the auditing lanes' own token; it is kept, not re-invented.
|
|
146
147
|
|
|
147
148
|
2. **Findings fitting no claim**, if any, each substantiated to the same standard as `BROKEN`.
|
|
148
149
|
|
|
149
|
-
3. **
|
|
150
|
+
3. **Attacked and held**: the claims this round attacked and could not break, each with the attack
|
|
151
|
+
that failed, and the adjacent behaviour that looks like the defect and is correct. This is where
|
|
152
|
+
the next round reads what has already been tried, so a claim listed here without its attack is
|
|
153
|
+
worth nothing to it. A claim's own `CONFIRMED` line already carries the evidence that convinced
|
|
154
|
+
the auditor, so list here only the attacks no verdict line carries and the adjacent behaviour.
|
|
155
|
+
|
|
156
|
+
4. **One terminal line**, and only one:
|
|
150
157
|
|
|
151
158
|
```text
|
|
152
|
-
VERDICT: PASS
|
|
153
|
-
VERDICT: FAIL
|
|
159
|
+
VERDICT: PASS
|
|
160
|
+
VERDICT: FAIL <claim numbers>; outside the claims: <finding ids>
|
|
154
161
|
```
|
|
155
162
|
|
|
156
|
-
|
|
163
|
+
The claim numbers are every claim that is not `CONFIRMED`. Write `none` in a slot the round
|
|
164
|
+
leaves empty, so a `FAIL` driven only by a finding outside the claims still names both.
|
|
165
|
+
|
|
166
|
+
**`PASS` requires all of this**: every claim `CONFIRMED`, nothing `UNRESOLVED`, nothing
|
|
157
167
|
`NOT-EVIDENCED`, and no substantiated finding outside the claims. A single substantiated finding
|
|
158
|
-
forces `FAIL`
|
|
168
|
+
forces `FAIL` however the numbered claims landed — otherwise a round can report a real
|
|
159
169
|
defect and still emit the word that authorises the release.
|
|
160
170
|
|
|
161
171
|
No process diary. No summary of what was read.
|
|
@@ -36,6 +36,21 @@ an executor inventing an answer and building on it silently.
|
|
|
36
36
|
**The threshold.** State that a finding is worth more than a clean pass, and why: the alternative is
|
|
37
37
|
a consumer finding it after publication, when the version is already spent.
|
|
38
38
|
|
|
39
|
+
## The read-only audit lane's brief
|
|
40
|
+
|
|
41
|
+
An audit lane writes nothing and runs nothing, so its brief carries fewer rows than a writing
|
|
42
|
+
unit's. Give a lane every row § Anatomy names — the subject with the evidence
|
|
43
|
+
§ "Evidence, by subject type" requires of each row it occupies, what the round decides, already
|
|
44
|
+
established, the numbered falsifiable claims, the unknowns, and the threshold — plus its own
|
|
45
|
+
**Role and lane** row (the role, its engine, and which lane it holds) and **Output** row (the
|
|
46
|
+
verdict shape and its single terminal line). Review evidence folds into the subject.
|
|
47
|
+
|
|
48
|
+
Omit the rows a writer needs and a lane cannot use: owned, shared, and off-limits files; the
|
|
49
|
+
Execution line's writer form, keeping the sentence that the lane performs the assignment directly
|
|
50
|
+
and spawns nothing; and acceptance criteria stated as gate commands. A lane that holds no shell
|
|
51
|
+
cannot close a gate criterion, so a brief handing it one is asking for a ruling on the writer's
|
|
52
|
+
report.
|
|
53
|
+
|
|
39
54
|
## The successor rule
|
|
40
55
|
|
|
41
56
|
A re-run **amends**; it never restates. Rewriting a brief from scratch loses the shape of what has
|
|
@@ -71,7 +86,7 @@ repair to carry the next defect; a round that finds them is converging, not fail
|
|
|
71
86
|
accumulated damage no single diff shows.
|
|
72
87
|
- **"The self-declared sound-and-unchanged verdicts are sound."** A writer's table saying a site
|
|
73
88
|
needed no change is a claim like any other, made by the party least able to test it. Require the
|
|
74
|
-
auditor to pick the ones it considers most likely wrong and actually attack them
|
|
89
|
+
auditor to pick the ones it considers most likely wrong and actually attack them.
|
|
75
90
|
|
|
76
91
|
## Instructions that change auditor behaviour
|
|
77
92
|
|
|
@@ -92,7 +107,10 @@ the brief, by the orchestrator.
|
|
|
92
107
|
## What not to put in a brief
|
|
93
108
|
|
|
94
109
|
- Laws already binding from `AGENTS.md` and the rule files. Reference them; restating invites drift
|
|
95
|
-
between the copy and the original.
|
|
110
|
+
between the copy and the original. Where the executor's tree carries a vendored copy of an
|
|
111
|
+
authority the canon has since superseded, quote the landed text with its canonical path and mark
|
|
112
|
+
the quotation as superseding the vendored copy, because a bare reference resolves to the stale
|
|
113
|
+
copy the executor holds.
|
|
96
114
|
- Any hint of what the other auditor is finding, or has found.
|
|
97
115
|
- Your own hypothesis about where the defect is, beyond what the claims state. An auditor handed a
|
|
98
116
|
suspect investigates the suspect and stops.
|
|
@@ -47,7 +47,7 @@ Load [hardening.md](references/hardening.md) for the hardening lane and for any
|
|
|
47
47
|
6. **Prove each defect before repairing it.** A repair begins with a test that fails for that defect: record the exact command and its failing count before the fix and the same command's passing count after. A repair with no red-then-green record is unproven.
|
|
48
48
|
7. **Consolidate.** Run the complete centralization and wrapper sweep. Update all call sites to the real symbol rather than leaving aliases or 1:1 delegates.
|
|
49
49
|
8. **Challenge seams.** Add deterministic tests for invariants, boundaries, failures, lifecycle, cleanup, cancellation, concurrency, hostile input, and resource pressure as applicable, under the test rules' real-implementation law.
|
|
50
|
-
9. **Use live services deliberately.** Put real external services/models in their dedicated project, require readiness, and make each request minimally sufficient,
|
|
50
|
+
9. **Use live services deliberately.** Put real external services/models in their dedicated project, require readiness, and make each request minimally sufficient, stable across the service's nondeterminism, and behaviorally meaningful. When the claim is that a foreign client can use this package, drive one representative real client end to end.
|
|
51
51
|
10. **Document the final behavior.** Update the governing guide, examples, method tables, limitations, and parity coverage. Document architectural limits honestly.
|
|
52
52
|
11. **Audit completion.** Inspect test discovery, `.todo`/`.skip`/conditional skip use, source/test helper duplication, exports, environment isolation, unexpected text corruption, and the entire diff.
|
|
53
53
|
12. **Verify.** Run the repository-prescribed gates in order and inspect the generated outputs relevant to the request.
|
|
@@ -50,7 +50,7 @@ Keep live tests in a dedicated project with explicit readiness, setup, timeout,
|
|
|
50
50
|
- For model tests, constrain temperature/seed/options when the real API supports it, but do not claim determinism the provider does not promise.
|
|
51
51
|
- Increase context or workload incrementally only when the scenario requires it.
|
|
52
52
|
- Test instruction precedence, long-context behavior, summarization, tool calls, scopes, and state transitions through observable outcomes.
|
|
53
|
-
- Avoid redundant expensive calls; one request
|
|
53
|
+
- Avoid redundant expensive calls; one request proves one primary claim.
|
|
54
54
|
|
|
55
55
|
Keep live projects outside the fast default suite when repository policy requires it, while making their explicit command authoritative for the campaign.
|
|
56
56
|
|
|
@@ -35,7 +35,7 @@ Include:
|
|
|
35
35
|
|
|
36
36
|
- public and internal behavior needed by real consumers;
|
|
37
37
|
- official upstream capabilities that fit the requested scope;
|
|
38
|
-
- architectural limitations that cannot or
|
|
38
|
+
- architectural limitations that cannot or must not be copied;
|
|
39
39
|
- legacy features worth salvaging;
|
|
40
40
|
- every `TODO`, deferred branch, placeholder, or documented omission in scope.
|
|
41
41
|
|
|
@@ -26,7 +26,10 @@ contract and this skill as the workflow. Preserve dirty and user-owned work.
|
|
|
26
26
|
A claim about a rendered surface is proven by capture, never by reading the code that was
|
|
27
27
|
supposed to produce it. Source-reading review passes a component that renders nothing.
|
|
28
28
|
|
|
29
|
-
-
|
|
29
|
+
- Generate the portfolio from the journey suite's capture family wherever a Vitest
|
|
30
|
+
browser project can drive the surface, and from a spawned harness only where none
|
|
31
|
+
can ([capture-harness.md](references/capture-harness.md)).
|
|
32
|
+
- The portfolio IS the review input: captures at every viewport and every theme the surface declares, an
|
|
30
33
|
accessibility snapshot, and an interaction log.
|
|
31
34
|
- Source is corroboration for a mechanism, never the proof that the surface shows it.
|
|
32
35
|
- A claim the portfolio cannot show is unproven, not passed. Say so.
|
|
@@ -1,82 +1,103 @@
|
|
|
1
1
|
# The capture harness
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
3
|
+
Take the portfolio from the journey suite's capture family wherever a Vitest browser project can
|
|
4
|
+
mount and drive the surface. Build the spawned script this file describes only for a surface no such
|
|
5
|
+
project can host — a served page, a foreign client, a process the runner cannot start inside a test.
|
|
6
|
+
Choose one source per surface, and never judge a round against a portfolio that is part
|
|
7
|
+
journey-generated and part spawned.
|
|
8
|
+
|
|
9
|
+
- Read the `orkestrel-prove-journey` skill for how the journey run generates a portfolio. This file
|
|
10
|
+
adds only what the review requires of a portfolio and how a spawned harness produces one.
|
|
11
|
+
- Own the spawned harness as the campaign owner. Never let a verdict lane write or edit it.
|
|
12
|
+
- Treat the spawned harness as a throwaway instrument: written for this surface, rebuilt or deleted
|
|
13
|
+
when the surface changes.
|
|
6
14
|
|
|
7
15
|
## One call, one lifecycle
|
|
8
16
|
|
|
9
|
-
|
|
10
|
-
verdict round spent on a half-dead harness is a wasted round.
|
|
17
|
+
Apply this section to a spawned harness only.
|
|
11
18
|
|
|
12
|
-
- Write the harness as one self-contained script that spawns its own children, waits for
|
|
13
|
-
|
|
14
|
-
- Never leave a child running across calls or expect one to survive its parent.
|
|
15
|
-
|
|
19
|
+
- Write the harness as one self-contained script that spawns its own children, waits for readiness,
|
|
20
|
+
does the capture, and kills them before returning.
|
|
21
|
+
- Never leave a child running across calls or expect one to survive its parent. A process started
|
|
22
|
+
inside one tool call dies with that call's process group.
|
|
23
|
+
- Give every child a pinned working directory. A process that resolves assets, configuration, or
|
|
16
24
|
fixtures relative to the current directory dies silently when launched from elsewhere.
|
|
17
|
-
- Pipe child standard error somewhere readable and print it on failure.
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
leaves no orphaned server, browser, or port.
|
|
25
|
+
- Pipe child standard error somewhere readable and print it on failure.
|
|
26
|
+
- Wait on an observable readiness signal — a served response, a printed line, a health probe — never
|
|
27
|
+
on a fixed sleep.
|
|
28
|
+
- Tear down on every exit path, including assertion and setup failure, so a failed capture leaves no
|
|
29
|
+
orphaned server, browser, or port.
|
|
23
30
|
|
|
24
31
|
## Validate the seed before capturing
|
|
25
32
|
|
|
26
|
-
|
|
27
|
-
|
|
33
|
+
Apply this section to a spawned harness only. In the journey run the acceptance journey is the seed,
|
|
34
|
+
and `orkestrel-prove-journey` fixes how it enters and what it may reach past.
|
|
28
35
|
|
|
29
|
-
- Build seed payloads from the surface's own published contract,
|
|
30
|
-
near-miss field name
|
|
31
|
-
- Assert the seeded state is present before shooting: the row exists, the prompt is parked,
|
|
32
|
-
|
|
33
|
-
- Drive the surface through its real entry path so the captured state is one a
|
|
34
|
-
|
|
35
|
-
- Reset to a known state between scenarios; a capture that inherits the previous scenario's
|
|
36
|
+
- Build seed payloads from the surface's own published contract, never from memory of it; a
|
|
37
|
+
near-miss field name renders an empty screen that reads as a product defect.
|
|
38
|
+
- Assert the seeded state is present before shooting: the row exists, the prompt is parked, the list
|
|
39
|
+
is non-empty.
|
|
40
|
+
- Drive the surface through its real entry path, so the captured state is one a person can reach.
|
|
41
|
+
- Reset to a known state between scenarios. A capture that inherits the previous scenario's
|
|
36
42
|
selection, focus, or scroll proves nothing about either.
|
|
37
43
|
|
|
38
44
|
## Capture the full portfolio
|
|
39
45
|
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
|
45
|
-
|
|
|
46
|
-
|
|
|
47
|
-
|
|
|
48
|
-
|
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
-
|
|
56
|
-
|
|
46
|
+
Produce every artifact in the following table each round, for every scenario in scope. The Source
|
|
47
|
+
column names what produces the artifact in the journey run; a spawned harness produces each one
|
|
48
|
+
itself.
|
|
49
|
+
|
|
50
|
+
| Artifact | Source in the journey run | What the artifact must show |
|
|
51
|
+
| ---------------------- | ------------------------------------ | ------------------------------------------------------------------------------------- |
|
|
52
|
+
| Viewport captures | `place` across the declared variants | Every breakpoint the surface declares, never one convenient size |
|
|
53
|
+
| Theme captures | `place` across the declared variants | Every theme the surface ships, at every declared viewport |
|
|
54
|
+
| Accessibility snapshot | `describeTree` and `describeFocus` | The rendered roles, names, and states, and the focus order the walk took |
|
|
55
|
+
| Interaction log | The journal's `steps` | Each interaction, its trigger, and the result observed after it |
|
|
56
|
+
| Console and error log | The journal's `output` | Everything the page emitted, including an uncaught error |
|
|
57
|
+
| Statechart outcome | The harness after a play-all run | The terminal status, the passed, failed, and total tallies, and each row's own result |
|
|
58
|
+
|
|
59
|
+
- Require the statechart outcome of every surface that declares the statechart family, and omit the
|
|
60
|
+
row only where the surface declares none.
|
|
61
|
+
- Read the journey run's per-variant written artifact for the accessibility snapshot and the logs,
|
|
62
|
+
and cite the statechart harness by its deep link where a lane must watch the widget move rather
|
|
63
|
+
than read a still of it. `orkestrel-prove-journey` fixes what each holds and what each is named
|
|
64
|
+
for.
|
|
65
|
+
- Shoot the whole surface before selecting or focusing anything inside it. A capture taken after a
|
|
66
|
+
selection reports a duplicate or highlighted artifact that does not exist.
|
|
67
|
+
- Start a keyboard walk from a neutral state, never from an already-focused control, or the log will
|
|
68
|
+
report a broken order no person meets.
|
|
69
|
+
- Name a spawned harness's artifacts so a verdict cites one exactly: scenario, viewport, theme,
|
|
70
|
+
step. The journey run's filename law already does this.
|
|
71
|
+
- Keep the artifacts of each round beside its verdicts. A round judged against the previous round's
|
|
72
|
+
captures is not a round.
|
|
57
73
|
|
|
58
74
|
## Preflight before spending a round
|
|
59
75
|
|
|
60
|
-
|
|
76
|
+
Open every artifact yourself before dispatching a verdict lane, whichever source produced it.
|
|
77
|
+
Confirm each of the following:
|
|
61
78
|
|
|
62
79
|
- each capture shows the scenario it claims, in the theme and viewport it claims;
|
|
63
80
|
- the seeded state is visible;
|
|
64
81
|
- the accessibility snapshot is non-empty and matches the captured screen;
|
|
65
82
|
- the interaction log records the interactions the brief asked for;
|
|
83
|
+
- the statechart outcome reads a terminal status, with a zero failed tally and a passed tally equal
|
|
84
|
+
to the total;
|
|
66
85
|
- nothing in the console log indicates the harness, rather than the surface, failed.
|
|
67
86
|
|
|
68
|
-
|
|
69
|
-
|
|
87
|
+
Repair a portfolio that fails preflight before dispatching it. Never spend a verdict round
|
|
88
|
+
discovering a harness defect.
|
|
70
89
|
|
|
71
90
|
## Triage missing evidence to the harness first
|
|
72
91
|
|
|
73
|
-
|
|
74
|
-
|
|
92
|
+
Treat a not-evidenced verdict item as a harness fault before treating it as a product finding. Work
|
|
93
|
+
these in order:
|
|
75
94
|
|
|
76
|
-
1. Confirm the artifact that
|
|
95
|
+
1. Confirm the artifact that decides the item exists and is named as the brief said.
|
|
77
96
|
2. Confirm the scenario reached the state the item is about.
|
|
78
97
|
3. Confirm the seed and the entry path match the surface's real contract.
|
|
79
|
-
4. Only then
|
|
98
|
+
4. Only then record it as a product finding.
|
|
80
99
|
|
|
81
|
-
|
|
82
|
-
|
|
100
|
+
- Repair every harness gap a round exposes before the recapture, and record the repair with the
|
|
101
|
+
round.
|
|
102
|
+
- Route a journey-run gap to the journey suite that owns it, and recapture from the repaired
|
|
103
|
+
journeys.
|