@orkestrel/scaffold 0.0.59 → 0.0.61

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (76) hide show
  1. package/README.md +13 -10
  2. package/dist/bin/main.js +632 -320
  3. package/dist/bin/main.js.map +1 -1
  4. package/dist/host/CLAUDE.md +5 -1
  5. package/dist/host/agents/orchestration.md +44 -19
  6. package/dist/host/agents/skills/enterprise-bootstrap/SKILL.md +91 -82
  7. package/dist/host/agents/skills/enterprise-bootstrap/references/bootstrap-reference.md +5 -5
  8. package/dist/host/agents/skills/enterprise-bootstrap/references/components.md +15 -15
  9. package/dist/host/agents/skills/enterprise-bootstrap/references/inputs.md +501 -0
  10. package/dist/host/agents/skills/enterprise-bootstrap/references/inspection.md +167 -0
  11. package/dist/host/agents/skills/enterprise-bootstrap/references/utilities.md +2 -2
  12. package/dist/host/agents/skills/orkestrel-debrief/SKILL.md +2 -2
  13. package/dist/host/agents/skills/orkestrel-debrief/references/instruction-audit.md +10 -9
  14. package/dist/host/agents/skills/orkestrel-falsify/SKILL.md +21 -11
  15. package/dist/host/agents/skills/orkestrel-falsify/references/brief.md +20 -2
  16. package/dist/host/agents/skills/orkestrel-harden-package/SKILL.md +1 -1
  17. package/dist/host/agents/skills/orkestrel-harden-package/references/hardening.md +1 -1
  18. package/dist/host/agents/skills/orkestrel-harden-package/references/research.md +1 -1
  19. package/dist/host/agents/skills/orkestrel-polish-surface/SKILL.md +4 -1
  20. package/dist/host/agents/skills/orkestrel-polish-surface/references/capture-harness.md +71 -50
  21. package/dist/host/agents/skills/orkestrel-prove-journey/SKILL.md +93 -29
  22. package/dist/host/agents/skills/orkestrel-prove-journey/agents/openai.yaml +1 -1
  23. package/dist/host/agents/skills/orkestrel-prove-journey/references/captures.md +62 -38
  24. package/dist/host/agents/skills/orkestrel-prove-journey/references/decide.md +68 -0
  25. package/dist/host/agents/skills/orkestrel-prove-journey/references/layer.md +107 -79
  26. package/dist/host/agents/skills/orkestrel-prove-journey/references/statechart.md +84 -0
  27. package/dist/host/agents/skills/orkestrel-prove-journey/references/styles.md +87 -0
  28. package/dist/host/agents/templates/brief.md +16 -7
  29. package/dist/host/agents/transports/claude.md +4 -2
  30. package/dist/host/agents/transports/codex.md +4 -1
  31. package/dist/host/claude/agents/analyst.md +3 -1
  32. package/dist/host/claude/agents/application.md +1 -1
  33. package/dist/host/claude/agents/builder.md +3 -3
  34. package/dist/host/claude/agents/checker.md +5 -0
  35. package/dist/host/claude/agents/grok.md +15 -5
  36. package/dist/host/claude/agents/implementer.md +1 -1
  37. package/dist/host/claude/agents/orkestrel.md +12 -11
  38. package/dist/host/claude/agents/planner.md +10 -0
  39. package/dist/host/claude/agents/reviewer.md +14 -8
  40. package/dist/host/claude/agents/sol.md +3 -1
  41. package/dist/host/claude/agents/verifier.md +2 -4
  42. package/dist/host/claude/rules/architecture.md +7 -5
  43. package/dist/host/claude/rules/documentation.md +1 -0
  44. package/dist/host/claude/rules/names.md +23 -5
  45. package/dist/host/claude/rules/patterns.md +1 -0
  46. package/dist/host/claude/rules/quality.md +1 -1
  47. package/dist/host/claude/rules/tests.md +3 -3
  48. package/dist/host/claude/rules/typescript.md +4 -1
  49. package/dist/host/claude/rules/writing.md +2 -2
  50. package/dist/host/claude/skills/orkestrel-prove-journey/SKILL.md +1 -1
  51. package/dist/host/codex/agents/builder.toml +6 -6
  52. package/dist/host/codex/agents/checker.toml +2 -1
  53. package/dist/host/codex/agents/grok.toml +12 -5
  54. package/dist/host/codex/agents/implementer.toml +2 -2
  55. package/dist/host/codex/agents/opus.toml +6 -1
  56. package/dist/host/codex/agents/planner.toml +11 -6
  57. package/dist/host/codex/agents/reviewer.toml +8 -6
  58. package/dist/host/guides/scaffold.md +39 -14
  59. package/dist/host/manifest.json +81 -51
  60. package/dist/host/scripts/codex.sh +0 -0
  61. package/dist/host/scripts/cursor.sh +0 -0
  62. package/dist/host/scripts/deps.sh +0 -0
  63. package/dist/host/scripts/ollama.sh +0 -0
  64. package/dist/src/core/index.cjs +424 -282
  65. package/dist/src/core/index.cjs.map +1 -1
  66. package/dist/src/core/index.d.cts +361 -220
  67. package/dist/src/core/index.d.ts +361 -220
  68. package/dist/src/core/index.js +421 -283
  69. package/dist/src/core/index.js.map +1 -1
  70. package/dist/src/server/index.cjs +208 -170
  71. package/dist/src/server/index.cjs.map +1 -1
  72. package/dist/src/server/index.d.cts +276 -152
  73. package/dist/src/server/index.d.ts +276 -152
  74. package/dist/src/server/index.js +200 -172
  75. package/dist/src/server/index.js.map +1 -1
  76. package/package.json +8 -7
@@ -1,85 +1,112 @@
1
1
  # The journey layer
2
2
 
3
- Build every capability here before writing the first journey, and route every journey step through
4
- it. Treat a journey that works around a missing capability as a layer defect.
3
+ Route every journey step through the published layer. Treat a journey that works around a missing
4
+ capability by reaching for a selector as a layer defect.
5
5
 
6
- Implement the signatures below as a contract in the workspace's browser test setup module; never
7
- copy them as source. Name each helper for the human act it performs.
6
+ ## Import, never implement
7
+
8
+ Import every journey helper from `@orkestrel/test/browser`. Write one of your own only where that
9
+ package publishes none for the act.
10
+
11
+ - Place a helper you write in the workspace's browser test setup module, export it from there, and
12
+ name it for the human act it performs.
13
+ - Read the package's own exports before writing anything. A helper that renames a published one is a
14
+ defect under `AGENTS.md`, and a second implementation of one drifts from the first.
15
+ - Code every journey against the vocabulary in this file, which is the published one. Diagnose a
16
+ target that stops resolving here, and fix it in the application.
8
17
 
9
18
  ## What it drives
10
19
 
11
- - Drive the real browser through the installed Vitest browser provider. Import its `page` locators
12
- and `userEvent` from `vitest/browser`; the `@vitest/browser/context` specifier is deprecated and
13
- is not the import a new layer uses.
14
- - Use the provider verbs for input: `click`, `keyboard`, `tab`, `type`, `clear`, `fill`, and
15
- `selectOptions`. Use `page.viewport` and `page.screenshot` for captures, and the runner's file
16
- command to read a written capture back.
17
- - Never dispatch a constructed event from the layer or from a journey. The centralized event
18
- factories `.claude/rules/tests.md` prescribes serve unit tests whose subject is the handler; a
19
- journey drives input through the provider verbs only.
20
- - Never let a helper take an element, a component instance, or a selector from the caller. Every
21
- helper resolves its own target from role and accessible name.
20
+ - Drive the real browser through the installed Vitest browser provider. The published verbs import
21
+ `page` and `userEvent` from `vitest/browser`; the `@vitest/browser/context` specifier is
22
+ deprecated and is not the import a workspace helper uses.
23
+ - Never dispatch a constructed event from a journey. The published `createPointerEvent`,
24
+ `createDragEvent`, `typeInput`, and `commitInput` serve a unit test whose subject is the handler; a
25
+ journey drives input through `clickAccessible`, `clickAccessibleWithin`, `clickDisclosure`,
26
+ `typeAccessible`, `fillAccessible`, `pressKeys`, and `traverseAccessible` only.
27
+ - Yield with `waitForFrame` where a step needs the browser to paint before the next reading. Never
28
+ guard a fact with a fixed delay.
29
+
30
+ ## Which helpers take an element
31
+
32
+ A journey verb resolves its own target from role and accessible name, and refuses an element, a
33
+ component instance, or a selector from the caller. A reader, a fixture builder, and a capture each
34
+ take one, because their subject is a node the caller already holds.
35
+
36
+ | Population | Takes an element |
37
+ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------- |
38
+ | `resolveAccessible`, `resolveRendered`, `clickAccessible`, `clickAccessibleWithin`, `clickDisclosure`, `typeAccessible`, `fillAccessible`, `pressKeys`, `traverseAccessible`, `readPerception`, `readPage`, `readFocus`, `readValue` | No |
39
+ | `readText`, `readRole`, `readName`, `readStates`, `describeTree`, `describeFocus`, `isReachable`, `isRendered` | Yes |
40
+ | `readContrast`, `readRing`, `readLayers`, `readBackdrop`, `readStyle`, `readToken`, `readPixels`, `readClasses`, `extractStyles`, `extractOrphans`, `readRows` | Yes |
41
+ | `mount`, `typeInput`, `commitInput`, `captureFrame`, and a portfolio's `place` | Yes |
42
+
43
+ Never pass an element to a verb from a journey step. Read a step that would pass one as a missing
44
+ verb, and add the verb instead.
22
45
 
23
46
  ## The resolver
24
47
 
25
- ```ts
26
- resolveAccessible(name: string): HTMLElement
27
- resolveAccessible(role: string, name: string): HTMLElement
28
- ```
48
+ `resolveAccessible(name)` resolves across the published interactive roles, and
49
+ `resolveAccessible(role, name)` resolves within exactly the named role. `resolveRendered` sits
50
+ beneath it and applies the same match without the scroll, so an act does not fail on a target the
51
+ act itself scrolls into view.
29
52
 
30
53
  - Match the accessible name exactly, never a substring.
31
- - Search a fixed set of interactive roles for the bare-name form, and exactly the named role for
32
- the two-argument form.
33
- - Count a match as reachable only when every condition holds: it is connected; it passes a
34
- visibility check that honours opacity and CSS; its box has non-zero width and height; its
35
- `tabIndex` is at least zero; it matches neither `:disabled` nor `[aria-disabled="true"]`; and it
36
- has no `[inert]` ancestor.
37
- - Scroll a wholly off-viewport target into view once, then measure reachability again. Count a
38
- control a person can scroll to as reachable, and one that stays outside the viewport as
39
- unreachable.
40
- - Give the layer a rendered-only resolver beneath the public one, and use it from the acting verbs
41
- so a click does not fail on a target the act itself scrolls into view.
54
+ - Read `ACCESSIBLE_ROLES` for the bare-name search set and `FOCUSABLE_SELECTOR` for what the layer
55
+ counts as focusable, rather than restating either list in a workspace.
56
+ - Read reachability with `isReachable` and rendering with `isRendered` where a test needs the
57
+ condition rather than the throw. A reachable match is connected, passes a visibility check that
58
+ honours opacity and CSS, has a box with non-zero width and height, carries a `tabIndex` of at
59
+ least zero, matches neither `:disabled` nor `[aria-disabled="true"]`, and has no `[inert]`
60
+ ancestor.
61
+ - Count a control a person can scroll to as reachable, and one that stays outside the viewport after
62
+ the scroll as unreachable.
42
63
 
43
64
  ### The failure voices
44
65
 
45
- Keep these distinct, and never merge any of them into one message.
46
-
47
- | Condition | The voice it must throw |
48
- | ------------------------------------------ | ---------------------------------------------------------------- |
49
- | No element carries the name | `No interactive element has the accessible name "<name>"` |
50
- | Every match fails a reachability condition | `Interactive target "<name>" is not visible and focus-reachable` |
51
- | Several matches are reachable | `Interactive target "<name>" is ambiguous across <n> elements` |
52
- | Still off-viewport after being scrolled to | `Interactive target "<name>" is unreachable after scrolling` |
66
+ Assert the one voice the case means. Never write an assertion that accepts more than one.
67
+
68
+ | Condition | The voice thrown | Thrown by |
69
+ | ------------------------------------------------ | ------------------------------------------------------------------------------------- | ----------------------- |
70
+ | No element carries the name | `No interactive element has the accessible name "<name>"` | `resolveRendered` |
71
+ | Every match fails a reachability condition | `Interactive target "<name>" is not visible and focus-reachable` | `resolveRendered` |
72
+ | Several matches are reachable | `Interactive target "<name>" is ambiguous across <n> elements` | `resolveRendered` |
73
+ | Still off-viewport after being scrolled to | `Interactive target "<name>" is unreachable after scrolling` | `resolveAccessible` |
74
+ | The region holds no reachable match | `Interactive target "<name>" is not reachable inside "<region>"` | `clickAccessibleWithin` |
75
+ | The region holds several reachable matches | `Interactive target "<name>" is ambiguous across <n> elements inside "<region>"` | `clickAccessibleWithin` |
76
+ | No native disclosure is reachable under the name | `Native disclosure "<name>" is not visible and focus-reachable` | `clickDisclosure` |
77
+ | Several native disclosures carry the name | `Native disclosure "<name>" is ambiguous across <n> elements` | `clickDisclosure` |
78
+ | Forward Tab never lands on the target | `Interactive target "<name>" is not reachable through forward Tab traversal: <trail>` | `traverseAccessible` |
79
+ | The named region is hidden | `Named region "<name>" is not visible` | `readPerception` |
80
+ | Several named regions carry the name | `Named region "<name>" is ambiguous across <n> elements` | `readPerception` |
81
+ | The resolved control renders no value | `Interactive target "<name>" does not carry a value` | `readValue` |
53
82
 
54
83
  - Report an absent control and a present-but-unreachable one as different findings: absence names
55
84
  a missing control, and unreachability names the interface gating one that exists.
56
- - Report ambiguity as a finding about the surface. Name the match count in the message, and
85
+ - Report ambiguity as a finding about the surface. Quote the match count from the message, and
57
86
  re-target the journey by role or region.
87
+ - Read the package's own voice table before asserting a message this file does not list. A workspace
88
+ that transcribes a voice by memory asserts a sentence the package does not throw.
58
89
 
59
90
  ## Role vocabulary
60
91
 
61
- Never infer a role from markup. Confirm the computed role in the browser whenever a target stops
62
- resolving.
92
+ Never infer a role from markup. Confirm the computed role in the browser with `readRole` whenever a
93
+ target stops resolving, and read the exposed name with `readName` and the exposed state with
94
+ `readStates` beside it.
63
95
 
64
96
  - A `list`-bearing input computes `combobox`, not `textbox`. Attaching native suggestion machinery
65
97
  to a field is a role change: re-target every journey that names that field, and read a resolver
66
98
  miss immediately after such a change as this before treating the element as missing.
67
99
  - Always target a tab by its role. A tab and its panel collide on a bare name by construction,
68
100
  because the panel is labelled by its tab.
69
- - `<summary>` is exposed as a native disclosure rather than through a role the provider's role
70
- locators accept. Give the layer a separate disclosure verb keyed to the summary's rendered text,
71
- applying the same reachability conditions and its own failure voices.
101
+ - Drive a `<summary>` with `clickDisclosure`, which applies the same reachability conditions and
102
+ throws its own voices. The provider's role locators do not resolve it, because it is exposed as a
103
+ native disclosure rather than through a role.
72
104
 
73
105
  ## Region-scoped resolution
74
106
 
75
- ```ts
76
- clickAccessibleWithin(region: string, role: string, name: string): Promise<void>
77
- ```
78
-
79
- - Provide this form for repeated short verbs such as `Add`, and for a control whose accessible name
80
- is completed by a status the row renders.
81
- - Apply the same reachability conditions inside the region, and throw voices that name the region
82
- as well as the target.
107
+ Reach for `clickAccessibleWithin(region, role, name)` where a short verb such as `Add` repeats
108
+ across a page, and where a rendered status completes a control's accessible name. It applies the
109
+ same reachability conditions inside the region and names the region in every voice it throws.
83
110
 
84
111
  ## Input and traversal
85
112
 
@@ -88,44 +115,45 @@ clickAccessibleWithin(region: string, role: string, name: string): Promise<void>
88
115
  | `typeAccessible(name, text)` | Focus the field, select all, delete, then send real keystrokes. Escape the provider's key syntax in the text. |
89
116
  | `fillAccessible(name, text)` | Replace the value in one operation for text too long to type. The real element still publishes real input. |
90
117
  | `pressKeys(keys)` | Send a provider keyboard sequence for Enter, arrows, modifiers, and combinations. |
91
- | `traverseAccessible(name)` | Move focus by forward Tab from wherever focus is, and return the target once focus lands on it. |
118
+ | `traverseAccessible(name)` | Move focus by forward Tab from wherever focus is, and return the target after focus lands on it. |
92
119
 
93
- - Count a traversal step only when focus actually lands on an element, and never charge the step
94
- bound for a Tab that moved nothing.
95
- - End the traversal when focus revisits an element, which is one complete cycle of the tab order,
96
- and throw the traversal's own voice, carrying the trail of what focus did reach. Keep a hard cap
97
- above the cycle so a page with no tab order fails instead of hanging.
98
- - Re-resolve the target by role and name on every step, and never hold a node reference across
99
- steps. A framework may replace the node between resolution and focus arrival.
100
- - Never call the browser's focus method to place focus.
120
+ - Reach for `typeAccessible` where the keystrokes are part of what the journey claims, and
121
+ `fillAccessible` where the text is only a payload the person pastes.
122
+ - Let `traverseAccessible` end the walk. It counts a step only where focus lands, stops at one
123
+ complete cycle of the tab order, carries the trail of what focus reached in its voice, and holds a
124
+ cap above the cycle so a page with no tab order fails instead of hanging.
125
+ - Never call the browser's focus method to place focus, and never hold a node reference across
126
+ traversal steps. A framework may replace the node between resolution and focus arrival.
101
127
 
102
128
  ## Perception
103
129
 
104
- ```ts
105
- readPerception(name: string): string
106
- ```
130
+ `readPerception(name)` returns the normalized `innerText` of exactly one visible named region,
131
+ dialog, table, tab panel, alert, or status. Quote that text in assertions.
107
132
 
108
- - Return the normalized `innerText` of exactly one visible named region, dialog, table, tab panel,
109
- alert, or status. Collapse whitespace runs to single spaces and trim.
110
133
  - Read `innerText`, never `textContent`: `innerText` applies CSS transforms and leaves out content
111
- the layout hides. Quote that text in assertions.
112
- - Include descendant visually-hidden text, which a screen reader perceives and which a clip-based
113
- hiding technique leaves in `innerText`.
114
- - Throw when the named region is absent, hidden, or ambiguous.
115
- - Give the layer a whole-page perception reader for cross-region sentences and the vocabulary
116
- sweep, a focus reader that returns the active element's rendered text, and a value reader that
117
- returns a resolved control's value. A perception assertion may quote that value: it is a rendered
118
- fact rather than internal state.
134
+ the layout hides. Descendant visually-hidden text stays in, which is what a screen reader
135
+ perceives.
136
+ - Reach for `readPage` for a cross-region sentence and the vocabulary sweep, `readFocus` for the
137
+ active element's rendered text, and `readValue` for a resolved control's rendered value. A
138
+ perception assertion may quote that value: it is a rendered fact rather than internal state.
139
+ - Reach for `describeTree` and `describeFocus` where the subject is the accessible tree itself —
140
+ roles, names, states, and focus order. They are the accessibility snapshot a written artifact
141
+ composes ([decide.md](decide.md) → The rendered artifact).
119
142
 
120
143
  ## Mounting and cleanup
121
144
 
122
145
  - Mount the shipped root component with its real provisions and return an idempotent cleanup that
123
- unmounts the app and removes its container.
146
+ unmounts the app and removes its container. Reach for the published `mount`, `render`, and `build`
147
+ where a fixture needs a node rather than the application.
124
148
  - Undo everything a journey changed after each test: unmount, destroy the session, reset the theme,
125
- clear the keys the application persisted, and return the route to its entry. Never let a journey
126
- inherit the previous journey's state.
149
+ return the route to its entry, and clear what the application persisted with `clearStorage` and
150
+ `removeDatabase`. Never let a journey inherit the previous journey's state.
151
+ - Record what a journey did with `createJournal`, started inside the journey and stopped in a
152
+ `finally`. Its `steps` and `output` are the evidence a failing journey hands back, and the input
153
+ the run's written artifact composes.
127
154
 
128
155
  ## The capture hook
129
156
 
130
- Give the layer exactly one capture helper. [captures.md](captures.md) owns the registry, the
131
- variant matrix, and the proofs that read what it wrote.
157
+ `createPortfolio` is the capture door and `place(state, element?)` is the hook.
158
+ [captures.md](captures.md) owns the registry, the variant matrix, and the proofs that read what it
159
+ wrote.
@@ -0,0 +1,84 @@
1
+ # The statechart family
2
+
3
+ Declare one transition table, and give it to the automated run that asserts it and to the
4
+ harness a person watches. Never write a second table for the harness.
5
+
6
+ ## Declare the table
7
+
8
+ - Declare each transition as a `StateTransition` carrying its `name`, its `from` state, the `event`,
9
+ and its `to` state. Type it on the entity's own state and event unions, so a row naming a state the
10
+ entity does not have fails to typecheck.
11
+ - Write one `StateScenario` per transition, carrying that `transition` plus `arrange`, `act`, and
12
+ `assert`. Each phase receives the context and the part of the transition it owns.
13
+ - Put the scenarios in the workspace's browser test setup module. Put the table beside them until a
14
+ harness page ships; from then on declare it in the application's own constants module, typed on the
15
+ entity's unions, and import it from the page and from the setup alike, because a page cannot
16
+ import from `tests/` and a second table is what this reference forbids.
17
+ - Declare a transition for every event the surface accepts in every state it accepts it, including
18
+ the event that leaves the state unchanged. A table that lists only the moves the happy path takes
19
+ proves the happy path.
20
+
21
+ ## Run the table
22
+
23
+ - Run the table with `executeScenarios(scenarios, build)` from `@orkestrel/test`. It drives the rows
24
+ in the order they are written, builds a context per row, and stops at the first row that throws.
25
+ - Run one row with `executeScenario(scenario, context)` where a single transition is the subject.
26
+ - Build the context in `build`, which receives the row it is building for. That is what lets one
27
+ table mix fixtures.
28
+ - Let the runner name the failure. It prepends the transition's `name` to whatever the row threw and
29
+ carries the original as the `cause`, so a bare assertion message still says which row failed.
30
+ - Never assert the entity's internal state in `assert` where the transition is one a person drives.
31
+ Assert what the interface renders, through `readPerception`, `readValue`, or `readStates`.
32
+
33
+ ## Drive the act the way the transition happens
34
+
35
+ - Drive `act` through the journey verbs — `clickAccessible`, `clickDisclosure`, `typeAccessible`,
36
+ `pressKeys`, `traverseAccessible` — for every transition a person can cause.
37
+ - Drive `act` through the entity's own API only where the transition is the entity's rather than the
38
+ person's: a lifecycle event, a transport reply, a timer the surface owns.
39
+ - Say which door each row used, in the row's `name`. A table that mixes the doors silently reads as
40
+ a set of user transitions and proves something else.
41
+
42
+ ## Build the harness a person watches
43
+
44
+ Mount the harness on the same table. It stays in the repository that owns the surface: only
45
+ `StateTransition`, `StateScenario`, `executeScenario`, `executeScenarios`, `STATECHART_ATTRIBUTES`,
46
+ and `STATECHART_STATUSES` are published, and a workspace that spells a `data-statechart-*` string of
47
+ its own has left the contract.
48
+
49
+ - Render one play control per transition and one play-all control, each disabled while a run is in
50
+ flight.
51
+ - Render a state badge carrying the entity's current state, an event log of what the entity emitted,
52
+ and a `role="status"` announcer that narrates each step in a sentence, so a screen reader and a
53
+ vision model both read the run without visual chrome.
54
+ - Publish `STATECHART_ATTRIBUTES.status` on the harness root, and cycle its value through
55
+ `STATECHART_STATUSES`: `pending` before a run has a result for every row, `idle` standing ready,
56
+ `running` in flight, and `passed` or `failed` as the terminal reading. Publish
57
+ `STATECHART_ATTRIBUTES.passed`, `.failed`, and `.total` on the same element, so a gate finds the
58
+ harness and reads the tally from one node.
59
+ - Publish `STATECHART_ATTRIBUTES.scenario` and `STATECHART_ATTRIBUTES.result` on each row, and
60
+ `STATECHART_ATTRIBUTES.state` on the element rendering the entity's current state.
61
+ - Write every attribute from the map rather than from a literal. A harness that sets an attribute
62
+ the gate does not read fails silently as a run that never completes.
63
+ - Deep-link one transition and the play-all run from the route, so a decision round names the exact
64
+ link it wants looked at ([decide.md](decide.md) → The rendered artifact).
65
+ - Give the harness a demo step that leaves the widget in its most legible state after the run, for a
66
+ person or a vision model deciding on a look.
67
+ - Pace the harness for a person to watch. The gate inherits that wall time, so budget the gate from
68
+ the row count and the pause rather than from a fixed timeout.
69
+
70
+ ## Gate the harness
71
+
72
+ Prove the harness from the browser project, through the interface:
73
+
74
+ - Mount the harness page and clear the route's query first, so a leftover deep link cannot start the
75
+ walk before the gate does.
76
+ - Press the play-all control through `clickAccessible`, never through a constructed event and never
77
+ by setting the deep link.
78
+ - Poll `STATECHART_ATTRIBUTES.status` until it reads `passed` or `failed`. Never assert it from one
79
+ read after the click.
80
+ - Assert the status reads `passed`, the failed tally reads zero, and the passed tally equals the
81
+ total. Name the failing rows from `STATECHART_ATTRIBUTES.scenario` in the failure message, so a red
82
+ gate says which transition broke.
83
+ - Assert the harness inventory before the rows: a page that mounted no transition passes every tally
84
+ assertion.
@@ -0,0 +1,87 @@
1
+ # Proving what the browser resolved
2
+
3
+ Prove a style from what the browser resolved on the mounted surface. The `enterprise-bootstrap`
4
+ skill's [instruments reference](../../enterprise-bootstrap/references/inspection.md) names each
5
+ instrument's property, its population, and the negative controls that must fail. Take those from
6
+ there and the reading from here.
7
+
8
+ ## Assert the resolved value
9
+
10
+ - Read one property with `readStyle(element, property)` and a length with `readPixels(element, property)`. A
11
+ class present in the markup and absent from the cascade resolves to nothing, and an assertion on
12
+ the class list passes on it.
13
+ - Never substitute `findRule` for a resolved read. It proves a declaration exists, and another rule
14
+ may still win; reach for it where the stylesheet itself is the subject.
15
+ - Compare a color through `matchesColor` or `parseCSSColor` rather than by string. A browser normalizes a color
16
+ expression, so a literal comparison fails on a value that resolved correctly.
17
+
18
+ ## Run per variant
19
+
20
+ The run axis is fixed in [SKILL.md](../SKILL.md) → Read the variant once. The matrix family's own
21
+ readings follow.
22
+
23
+ - Apply each variant's `apply` and its `width` and `height` before the readings, and take every
24
+ reading for that variant before moving to the next.
25
+ - Name the attribute the surface actually reads in `apply`; a Bootstrap surface switches on
26
+ `data-bs-theme`. An `apply` that sets another attribute leaves the run in the default theme, where
27
+ every reading passes.
28
+ - Assert that the run read every declared variant. A matrix that silently walked one variant reports
29
+ a pass for the theme nobody exercised.
30
+ - Report which variants a result covers beside it. A pairing that appears only in a state the run
31
+ never entered is unmeasured.
32
+
33
+ ## Contrast and focus chrome
34
+
35
+ - Read a text pairing with `readContrast(element)`, which composites the painted ancestors to the first
36
+ opaque layer, so a translucent tint reads as a tint over what shows through it.
37
+ - Pass `floor` only where the surface the stack really sits on is known, and pass `CANVAS_COLOR`
38
+ where the browser paints onto its own canvas. Omitting `floor` refuses a stack whose painted layers
39
+ are all translucent; take that refusal as the reading, because an assumed white canvas turns "this
40
+ surface declares no background" into a number that reads like a measurement.
41
+ - Read focus chrome with `readRing(control)`, after focus arrived through `traverseAccessible`,
42
+ `pressKeys`, or a real click. Pass `worn` where the chrome is painted onto a second element such
43
+ as a label. It reports `undefined` for a control not matching `:focus-visible`, for the browser's
44
+ own automatic ring, and for a focus style that only repaints the fill — treat each as a finding
45
+ about the surface rather than as a pass.
46
+ - Carry the negative controls the composited-contrast instrument names, in the same run and composed
47
+ in the harness rather than taken from the surface. An instrument whose negative control passes is
48
+ broken, and its readings are not evidence.
49
+ - Reach for `measureContrast`, `measureLuminance`, `blendColor`, `readLayers`, and `readBackdrop`
50
+ only where the composite itself is the subject. Never re-derive `readContrast` from them.
51
+
52
+ ## The authored-class census
53
+
54
+ Take the property, the population, and the negative controls from the instruments reference →
55
+ Authored class in the shipped cascade. This is the reading.
56
+
57
+ - Read the census as the difference between `readClasses(root)` of the mounted surface and
58
+ `readCascade()`. A token in the difference is a class the markup uses and no loaded stylesheet
59
+ declares.
60
+ - Take `root` from the mounted surface, so the census covers what rendered rather than what a
61
+ template file spells.
62
+ - Report what `readClasses` walked as the population, and fail a run that walked none.
63
+ - Append the extraction-door negative control to that same `root`, so it reaches the difference
64
+ through `readClasses` rather than beside it.
65
+
66
+ ## Style escapes
67
+
68
+ Take the property, the population, the named exemptions, and the negative control from the
69
+ instruments reference → Style escapes. This is the reading.
70
+
71
+ - Read escapes with `extractStyles(root)`, which returns the markup of every hit it found.
72
+ - Take the reading before any journey drives the surface, because the population is the undriven
73
+ tree.
74
+ - Append the harness-built negative control element to that same `root`, so it reaches the reading
75
+ through `extractStyles`.
76
+ - Reach for `extractOrphans` where the finding is a child element rendered outside its required
77
+ parent, and `readRows` where the subject is a repeated row's rendered text.
78
+
79
+ ## Tokens
80
+
81
+ - Read a token with `readToken(element, name)` where inheritance matters and `readRootToken(name)` where the
82
+ document declares it. The leading dashes are optional in each.
83
+ - Compare values, never presence. An absent token and one declared empty both read as `''`, and a
84
+ `var()` naming an undeclared property resolves to the inherited color rather than refusing, so an
85
+ assertion on presence passes on a token nobody declared.
86
+ - Read each token once per variant and assert the values differ where the design says the variants
87
+ differ. A pair of variants that resolves a token identically is a theme that did not switch.
@@ -10,6 +10,10 @@ copy `# Unit UNIT_ID — SHORT_SUBJECT`. Delete each italic reminder as you fill
10
10
  under, and leave no row blank: fill a row you cannot close with a named unknown label, and
11
11
  describe that label under § Unknowns with how the unit reports back on it.
12
12
 
13
+ For a read-only audit lane, fill the rows
14
+ `.agents/skills/orkestrel-falsify/references/brief.md` § "The read-only audit lane's brief" names
15
+ and delete the rest of this template.
16
+
13
17
  ## Role and engine
14
18
 
15
19
  ROLE_NAME on ENGINE_NAME, reached as TRANSPORT.
@@ -73,7 +77,7 @@ answer around._
73
77
  _Grant a behaviour with the tests that pin it, a constant with every fixture and expectation derived
74
78
  from it, a template with the materialized copy the package generates from it, and a mechanism with
75
79
  the prose describing it: the comment beside the code it edits and the guide passage stating the
76
- behaviour it moves._
80
+ behaviour it moves. A unit that moves a published symbol owns the package `README.md`._
77
81
 
78
82
  **Shared (report-only).** SHARED_FILES
79
83
 
@@ -84,21 +88,25 @@ edits nothing in this row._
84
88
 
85
89
  _Name each file the unit must not touch, and read every acceptance criterion against this row line by
86
90
  line. Grant the file a criterion needs, or strike that criterion. A file the change breaks that
87
- appears in no row of this section is unscoped._
91
+ appears in no row of this section is unscoped. Never list `tmp/probe/` off-limits: it is the unit's
92
+ probe home, and `.claude/rules/tests.md` puts every runtime probe there._
88
93
 
89
94
  **What asserts the state this change ends.** FILES_THE_RESULT_MAKES_FALSE
90
95
 
91
96
  _List every file the result makes false rather than every file that declares the thing changing: the
92
97
  test asserting the reversed behaviour, the fixture carrying the raised value, the golden digest over
93
98
  generated output, the consumer script naming the removed union member. Derive the list by running the
94
- suite; where you cannot run it, name the search's bound so the unit re-derives the list. End each
95
- entry in Owned, in Shared, or with a named carrier dispatched before this change ships._
99
+ suite; where you cannot run it, name the search's bound so the unit re-derives the list. A rename's
100
+ search bound is a word-boundary sweep over the old name followed by a case-insensitive sweep over
101
+ its `-s`, `-ed`, and `-ing` inflections. End each entry in Owned, in Shared, or with a named carrier
102
+ dispatched before this change ships._
96
103
 
97
104
  **Tools and limits.** ALLOWED_TOOLS, PERMISSION_LIMITS
98
105
 
99
106
  _Check the § Output mechanism and every acceptance criterion's verification method against this
100
107
  allowlist. A read-only lane writes no report file and runs no probe, so hand it the rendered evidence
101
- instead._
108
+ instead. A rename moves the file with the shell's `mv`, never `git mv`; `git add -N` is permitted
109
+ only to render diff evidence._
102
110
 
103
111
  ## Execution
104
112
 
@@ -138,8 +146,9 @@ criterion behind it. Where the change edits a file the repository vendors or dig
138
146
  regeneration step ahead of every gate that reads the generated artifact. Ask what the change does
139
147
  to every fact you measured, and fix each criterion to the state the unit finishes in. Close each
140
148
  criterion with owned files alone, and name the property the unit must change; record a consequence
141
- you expect to follow as an observation, never as a criterion. A scoped run over the unit's own
142
- owned files stays a legitimate criterion._
149
+ you expect to follow as an observation, never as a criterion. A criterion that closes on an
150
+ instrument's reading names that instrument's negative control and the class the control proves the
151
+ instrument can see. A scoped run over the unit's own owned files stays a legitimate criterion._
143
152
 
144
153
  **Observations, not criteria.** TIMING_SENSITIVE_OR_WHOLE_SUITE_GATES
145
154
 
@@ -21,8 +21,10 @@ with the permission mode the route pins. Never substitute a fixed Claude model i
21
21
 
22
22
  Verify that the `claude` CLI resolves and is authenticated before first use. On either
23
23
  failure return it immediately with the fallback named, so the Sol main session records
24
- Opus unavailable for the round. Never install, authenticate, or substitute an API key,
25
- access token, or copied auth file.
24
+ Opus unavailable for the round. Where the binary is absent, that report names the install
25
+ command for the `claude` CLI, so the Orchestrator can put it to the user in the same turn
26
+ it records the bench dark and re-probe when the user answers. Never install, authenticate,
27
+ or substitute an API key, access token, or copied auth file.
26
28
 
27
29
  Journal every run: redirect --output-format stream-json to tmp/claude/<unit>.jsonl,
28
30
  which is gitignored, and record the session id. A bench unit with no journal ran on its
@@ -119,7 +119,7 @@ skill owns the value set and the terminal line. Point the brief at both; restate
119
119
 
120
120
  ## Implementer route
121
121
 
122
- Sandbox `workspace-write`, main checkout, sole serial writer from a clean committed
122
+ Sandbox `workspace-write`, the checkout the route writes in, its sole serial writer from a clean committed
123
123
  baseline, with owned files, off-limits files, and a deviation contract. The brief forbids
124
124
  dependency installation, commits, pushes, publishing, credentials, destructive commands,
125
125
  shared-file edits, and tree-wide mutating gates.
@@ -147,6 +147,9 @@ negative tests is unaffected.
147
147
 
148
148
  - Verify `codex --version` before first use. On Windows `codex` resolves in Bash through
149
149
  the extensionless npm shim; if it does not, invoke `codex.cmd`.
150
+ - Binary absent: report it so the Orchestrator can record the bench dark and, in the same turn,
151
+ name to the user the install command for `@openai/codex` and the bench it unblocks. The
152
+ Orchestrator re-probes when the user answers. Never install it yourself.
150
153
  - Binary present but authentication unavailable: report it so the Orchestrator can start
151
154
  device-auth recovery in the same turn. It backgrounds `codex login --device-auth` with
152
155
  output captured to `tmp/codex/login.log`, surfaces the verification URL and one-time code
@@ -34,7 +34,9 @@ Everything `.agents/orchestration.md`'s dispatch contract requires, plus:
34
34
  - **Every authority the brief references must exist in the tree the exec is rooted in.** Check
35
35
  before dispatch. A brief citing a rule file or section the executor cannot find delivers nothing
36
36
  while looking like authority, and it fails silently — an auditor does not report a heading it
37
- never saw. Propagate the missing file; do not restate its contents in the brief.
37
+ never saw. Propagate the missing file rather than restating its contents in the brief, and take
38
+ the stale-authority branch in `.agents/skills/orkestrel-falsify/references/brief.md` § "What not
39
+ to put in a brief" where the executor's tree carries a superseded vendored copy.
38
40
  - For an audit: the subject as numbered falsifiable claims, and the skill that fixes the verdict
39
41
  shape. The Falsification section of `.claude/rules/quality.md` owns the method and the evidence
40
42
  each verdict carries. The verdict shape defaults to `orkestrel-falsify`; a dispatch may name a
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: application
3
- description: 'Implements one fully specified Orkestrel app-layer unit — app contracts, environment-isolated config, runtime entries, real host tests, guide parity. Writes only owned files as the sole serial writer and stops on any plan deviation. Nontrivial app design belongs to GPT-5.6 Sol or Opus 5.'
3
+ description: 'Implements one fully specified Orkestrel app-layer unit — app contracts, environment-isolated config, runtime entries, real host tests, guide parity. Writes only owned files in the checkout the unit writes as the sole serial writer and stops on any plan deviation. Nontrivial app design belongs to GPT-5.6 Sol or Opus 5.'
4
4
  tools: Read, Grep, Glob, Edit, Write, Bash
5
5
  model: sonnet
6
6
  effort: low
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: builder
3
- description: 'Implements one small, fully specified, taste-free unit exactly as dispatched. Writes only owned files in the main checkout as the sole serial writer, validates narrowly, and stops on any plan deviation. Nontrivial implementation belongs to GPT-5.6 Sol or Opus 5.'
3
+ description: 'Implements one small, fully specified, taste-free unit exactly as dispatched. Writes only owned files in the checkout the unit writes as the sole serial writer, validates narrowly, and stops on any plan deviation. Nontrivial implementation belongs to GPT-5.6 Sol or Opus 5.'
4
4
  tools: Read, Grep, Glob, Edit, Write, Bash
5
5
  model: sonnet
6
6
  effort: low
@@ -19,8 +19,8 @@ dispatch contract.
19
19
 
20
20
  - Before writing, read **AGENTS.md**, every applicable `.claude/rules/*.md`, the
21
21
  dispatch-named skill and required references, and the governing guide/spec. All bind you
22
- as written there, and this charter restates none of them. An app-layer unit additionally
23
- binds `.claude/rules/application.md` and `.claude/rules/workspace.md`.
22
+ as written there, and this charter restates none of them.
23
+ - An app-layer unit belongs to `application`: stop and say so.
24
24
  - Write ONLY the owned files named in your dispatch. Shared or off-limits files are
25
25
  report-only: if one needs a change, RETURN the exact patch — never edit it.
26
26
  - NO tree-wide or mutating commands: never `format`, lint `--fix`, or `build`.
@@ -37,6 +37,11 @@ No judgment calls: a question that needs one becomes a **referral** — specific
37
37
  evidenced, addressed to the subjective lane when it is running and to the
38
38
  Orchestrator when it is not — never a guess and never a verdict of yours.
39
39
 
40
+ Rule a claim whose only evidence is the writer's report `UNRESOLVED`, never
41
+ `CONFIRMED`, whatever the brief says. A quoted command and exit code inside a
42
+ report is the writer quoting itself, so it evidences nothing until a lane that ran
43
+ the command supplies the reading.
44
+
40
45
  ## Output contract
41
46
 
42
47
  Return the shape fixed by the dispatch.
@@ -40,7 +40,8 @@ shim that answered once does not clear it, and when it does fire it leaves only
40
40
  `SetConsoleWindowTitle` trace — which reads as a bench that returned nothing rather than as a launch
41
41
  that never happened. The versioned entry has no console dependency and no such failure mode.
42
42
 
43
- Read an empty shim run as a launch failure until its log is checked for that trace.
43
+ Read an empty shim run as a launch failure until its `.err` journal is checked for that
44
+ trace.
44
45
 
45
46
  If nothing responds the bench is dark. Stop with a deviation naming the fallback from the root
46
47
  tedious-work ladder — Luna, then Sonnet. Never hand the reading to the Orchestrator, `planner`, or
@@ -48,14 +49,21 @@ tedious-work ladder — Luna, then Sonnet. Never hand the reading to the Orchest
48
49
 
49
50
  Create `tmp/cursor/` first. Write any brief longer than a couple of sentences to
50
51
  `tmp/cursor/<unit>-brief.md` and make the prompt a pointer to it; briefs never travel as
51
- fragile shell arguments. Every run journals its output, so the user can tail progress live
52
- and an interrupted run leaves its partial distillate on disk:
52
+ fragile shell arguments. Every run journals its event stream, so the user can tail progress
53
+ live and an interrupted run leaves its partial distillate on disk:
53
54
 
54
- `<resolved-entry> -p --trust --mode=ask --model "$CURSOR_GROK_MODEL" "<brief or pointer>" | tee tmp/cursor/<unit>.log`
55
+ `<resolved-entry> -p --trust --mode=ask --model "$CURSOR_GROK_MODEL" --output-format stream-json "<pointer>" > tmp/cursor/<unit>.jsonl 2> tmp/cursor/<unit>.err`
55
56
 
56
57
  Write that chain to `tmp/cursor/run.sh` and run the file, so the resolution, the model, and the
57
58
  journalling are one artifact the next run reuses.
58
59
 
60
+ The journal's first event is the `init` event, and its `session_id` is the run's recovery
61
+ handle. The journal's `result` event carries the final answer. Return the journal path and
62
+ that session id with the result, so the Orchestrator can confirm the bench ran. Read the
63
+ `.err` file before calling a run empty; a launch that never reached the model leaves its
64
+ trace only there. Resume an interrupted run through the CLI's `--resume` option, probed
65
+ before its first use.
66
+
59
67
  Run that yourself only for a short bounded ask finishing in about two minutes. For anything
60
68
  longer your job ends at drafting: return the brief path, the exact resolved command, and the
61
69
  journal path, and let the Orchestrator launch it as a harness-tracked background command under
@@ -79,7 +87,9 @@ Return only:
79
87
  - `Question`: one line.
80
88
  - `Evidence`: concise facts with `file:line` or primary-source pointers.
81
89
  - `Distillate`: the smallest context the next engine needs.
82
- - `Unknowns`: unresolved facts, not recommendations.
90
+ - `Unknowns`: unresolved facts, not recommendations, naming every input row the
91
+ distillate did not reach.
92
+ - `Journal`: the journal path and the session id from its `init` event.
83
93
  - `Deviation`: unavailable CLI, model, or auth; command failure; dirty containment.
84
94
 
85
95
  Grok's output is evidence, never a decision or a verdict.