@orkestrel/scaffold 0.0.59 → 0.0.61
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +13 -10
- package/dist/bin/main.js +632 -320
- package/dist/bin/main.js.map +1 -1
- package/dist/host/CLAUDE.md +5 -1
- package/dist/host/agents/orchestration.md +44 -19
- package/dist/host/agents/skills/enterprise-bootstrap/SKILL.md +91 -82
- package/dist/host/agents/skills/enterprise-bootstrap/references/bootstrap-reference.md +5 -5
- package/dist/host/agents/skills/enterprise-bootstrap/references/components.md +15 -15
- package/dist/host/agents/skills/enterprise-bootstrap/references/inputs.md +501 -0
- package/dist/host/agents/skills/enterprise-bootstrap/references/inspection.md +167 -0
- package/dist/host/agents/skills/enterprise-bootstrap/references/utilities.md +2 -2
- package/dist/host/agents/skills/orkestrel-debrief/SKILL.md +2 -2
- package/dist/host/agents/skills/orkestrel-debrief/references/instruction-audit.md +10 -9
- package/dist/host/agents/skills/orkestrel-falsify/SKILL.md +21 -11
- package/dist/host/agents/skills/orkestrel-falsify/references/brief.md +20 -2
- package/dist/host/agents/skills/orkestrel-harden-package/SKILL.md +1 -1
- package/dist/host/agents/skills/orkestrel-harden-package/references/hardening.md +1 -1
- package/dist/host/agents/skills/orkestrel-harden-package/references/research.md +1 -1
- package/dist/host/agents/skills/orkestrel-polish-surface/SKILL.md +4 -1
- package/dist/host/agents/skills/orkestrel-polish-surface/references/capture-harness.md +71 -50
- package/dist/host/agents/skills/orkestrel-prove-journey/SKILL.md +93 -29
- package/dist/host/agents/skills/orkestrel-prove-journey/agents/openai.yaml +1 -1
- package/dist/host/agents/skills/orkestrel-prove-journey/references/captures.md +62 -38
- package/dist/host/agents/skills/orkestrel-prove-journey/references/decide.md +68 -0
- package/dist/host/agents/skills/orkestrel-prove-journey/references/layer.md +107 -79
- package/dist/host/agents/skills/orkestrel-prove-journey/references/statechart.md +84 -0
- package/dist/host/agents/skills/orkestrel-prove-journey/references/styles.md +87 -0
- package/dist/host/agents/templates/brief.md +16 -7
- package/dist/host/agents/transports/claude.md +4 -2
- package/dist/host/agents/transports/codex.md +4 -1
- package/dist/host/claude/agents/analyst.md +3 -1
- package/dist/host/claude/agents/application.md +1 -1
- package/dist/host/claude/agents/builder.md +3 -3
- package/dist/host/claude/agents/checker.md +5 -0
- package/dist/host/claude/agents/grok.md +15 -5
- package/dist/host/claude/agents/implementer.md +1 -1
- package/dist/host/claude/agents/orkestrel.md +12 -11
- package/dist/host/claude/agents/planner.md +10 -0
- package/dist/host/claude/agents/reviewer.md +14 -8
- package/dist/host/claude/agents/sol.md +3 -1
- package/dist/host/claude/agents/verifier.md +2 -4
- package/dist/host/claude/rules/architecture.md +7 -5
- package/dist/host/claude/rules/documentation.md +1 -0
- package/dist/host/claude/rules/names.md +23 -5
- package/dist/host/claude/rules/patterns.md +1 -0
- package/dist/host/claude/rules/quality.md +1 -1
- package/dist/host/claude/rules/tests.md +3 -3
- package/dist/host/claude/rules/typescript.md +4 -1
- package/dist/host/claude/rules/writing.md +2 -2
- package/dist/host/claude/skills/orkestrel-prove-journey/SKILL.md +1 -1
- package/dist/host/codex/agents/builder.toml +6 -6
- package/dist/host/codex/agents/checker.toml +2 -1
- package/dist/host/codex/agents/grok.toml +12 -5
- package/dist/host/codex/agents/implementer.toml +2 -2
- package/dist/host/codex/agents/opus.toml +6 -1
- package/dist/host/codex/agents/planner.toml +11 -6
- package/dist/host/codex/agents/reviewer.toml +8 -6
- package/dist/host/guides/scaffold.md +39 -14
- package/dist/host/manifest.json +81 -51
- package/dist/host/scripts/codex.sh +0 -0
- package/dist/host/scripts/cursor.sh +0 -0
- package/dist/host/scripts/deps.sh +0 -0
- package/dist/host/scripts/ollama.sh +0 -0
- package/dist/src/core/index.cjs +424 -282
- package/dist/src/core/index.cjs.map +1 -1
- package/dist/src/core/index.d.cts +361 -220
- package/dist/src/core/index.d.ts +361 -220
- package/dist/src/core/index.js +421 -283
- package/dist/src/core/index.js.map +1 -1
- package/dist/src/server/index.cjs +208 -170
- package/dist/src/server/index.cjs.map +1 -1
- package/dist/src/server/index.d.cts +276 -152
- package/dist/src/server/index.d.ts +276 -152
- package/dist/src/server/index.js +200 -172
- package/dist/src/server/index.js.map +1 -1
- package/package.json +8 -7
|
@@ -1,85 +1,112 @@
|
|
|
1
1
|
# The journey layer
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
|
|
3
|
+
Route every journey step through the published layer. Treat a journey that works around a missing
|
|
4
|
+
capability by reaching for a selector as a layer defect.
|
|
5
5
|
|
|
6
|
-
|
|
7
|
-
|
|
6
|
+
## Import, never implement
|
|
7
|
+
|
|
8
|
+
Import every journey helper from `@orkestrel/test/browser`. Write one of your own only where that
|
|
9
|
+
package publishes none for the act.
|
|
10
|
+
|
|
11
|
+
- Place a helper you write in the workspace's browser test setup module, export it from there, and
|
|
12
|
+
name it for the human act it performs.
|
|
13
|
+
- Read the package's own exports before writing anything. A helper that renames a published one is a
|
|
14
|
+
defect under `AGENTS.md`, and a second implementation of one drifts from the first.
|
|
15
|
+
- Code every journey against the vocabulary in this file, which is the published one. Diagnose a
|
|
16
|
+
target that stops resolving here, and fix it in the application.
|
|
8
17
|
|
|
9
18
|
## What it drives
|
|
10
19
|
|
|
11
|
-
- Drive the real browser through the installed Vitest browser provider.
|
|
12
|
-
and `userEvent` from `vitest/browser`; the `@vitest/browser/context` specifier is
|
|
13
|
-
is not the import a
|
|
14
|
-
-
|
|
15
|
-
`
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
20
|
+
- Drive the real browser through the installed Vitest browser provider. The published verbs import
|
|
21
|
+
`page` and `userEvent` from `vitest/browser`; the `@vitest/browser/context` specifier is
|
|
22
|
+
deprecated and is not the import a workspace helper uses.
|
|
23
|
+
- Never dispatch a constructed event from a journey. The published `createPointerEvent`,
|
|
24
|
+
`createDragEvent`, `typeInput`, and `commitInput` serve a unit test whose subject is the handler; a
|
|
25
|
+
journey drives input through `clickAccessible`, `clickAccessibleWithin`, `clickDisclosure`,
|
|
26
|
+
`typeAccessible`, `fillAccessible`, `pressKeys`, and `traverseAccessible` only.
|
|
27
|
+
- Yield with `waitForFrame` where a step needs the browser to paint before the next reading. Never
|
|
28
|
+
guard a fact with a fixed delay.
|
|
29
|
+
|
|
30
|
+
## Which helpers take an element
|
|
31
|
+
|
|
32
|
+
A journey verb resolves its own target from role and accessible name, and refuses an element, a
|
|
33
|
+
component instance, or a selector from the caller. A reader, a fixture builder, and a capture each
|
|
34
|
+
take one, because their subject is a node the caller already holds.
|
|
35
|
+
|
|
36
|
+
| Population | Takes an element |
|
|
37
|
+
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------- |
|
|
38
|
+
| `resolveAccessible`, `resolveRendered`, `clickAccessible`, `clickAccessibleWithin`, `clickDisclosure`, `typeAccessible`, `fillAccessible`, `pressKeys`, `traverseAccessible`, `readPerception`, `readPage`, `readFocus`, `readValue` | No |
|
|
39
|
+
| `readText`, `readRole`, `readName`, `readStates`, `describeTree`, `describeFocus`, `isReachable`, `isRendered` | Yes |
|
|
40
|
+
| `readContrast`, `readRing`, `readLayers`, `readBackdrop`, `readStyle`, `readToken`, `readPixels`, `readClasses`, `extractStyles`, `extractOrphans`, `readRows` | Yes |
|
|
41
|
+
| `mount`, `typeInput`, `commitInput`, `captureFrame`, and a portfolio's `place` | Yes |
|
|
42
|
+
|
|
43
|
+
Never pass an element to a verb from a journey step. Read a step that would pass one as a missing
|
|
44
|
+
verb, and add the verb instead.
|
|
22
45
|
|
|
23
46
|
## The resolver
|
|
24
47
|
|
|
25
|
-
|
|
26
|
-
resolveAccessible(name
|
|
27
|
-
|
|
28
|
-
|
|
48
|
+
`resolveAccessible(name)` resolves across the published interactive roles, and
|
|
49
|
+
`resolveAccessible(role, name)` resolves within exactly the named role. `resolveRendered` sits
|
|
50
|
+
beneath it and applies the same match without the scroll, so an act does not fail on a target the
|
|
51
|
+
act itself scrolls into view.
|
|
29
52
|
|
|
30
53
|
- Match the accessible name exactly, never a substring.
|
|
31
|
-
-
|
|
32
|
-
|
|
33
|
-
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
has no `[inert]`
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
unreachable.
|
|
40
|
-
- Give the layer a rendered-only resolver beneath the public one, and use it from the acting verbs
|
|
41
|
-
so a click does not fail on a target the act itself scrolls into view.
|
|
54
|
+
- Read `ACCESSIBLE_ROLES` for the bare-name search set and `FOCUSABLE_SELECTOR` for what the layer
|
|
55
|
+
counts as focusable, rather than restating either list in a workspace.
|
|
56
|
+
- Read reachability with `isReachable` and rendering with `isRendered` where a test needs the
|
|
57
|
+
condition rather than the throw. A reachable match is connected, passes a visibility check that
|
|
58
|
+
honours opacity and CSS, has a box with non-zero width and height, carries a `tabIndex` of at
|
|
59
|
+
least zero, matches neither `:disabled` nor `[aria-disabled="true"]`, and has no `[inert]`
|
|
60
|
+
ancestor.
|
|
61
|
+
- Count a control a person can scroll to as reachable, and one that stays outside the viewport after
|
|
62
|
+
the scroll as unreachable.
|
|
42
63
|
|
|
43
64
|
### The failure voices
|
|
44
65
|
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
| Condition
|
|
48
|
-
|
|
|
49
|
-
| No element carries the name
|
|
50
|
-
| Every match fails a reachability condition
|
|
51
|
-
| Several matches are reachable
|
|
52
|
-
| Still off-viewport after being scrolled to
|
|
66
|
+
Assert the one voice the case means. Never write an assertion that accepts more than one.
|
|
67
|
+
|
|
68
|
+
| Condition | The voice thrown | Thrown by |
|
|
69
|
+
| ------------------------------------------------ | ------------------------------------------------------------------------------------- | ----------------------- |
|
|
70
|
+
| No element carries the name | `No interactive element has the accessible name "<name>"` | `resolveRendered` |
|
|
71
|
+
| Every match fails a reachability condition | `Interactive target "<name>" is not visible and focus-reachable` | `resolveRendered` |
|
|
72
|
+
| Several matches are reachable | `Interactive target "<name>" is ambiguous across <n> elements` | `resolveRendered` |
|
|
73
|
+
| Still off-viewport after being scrolled to | `Interactive target "<name>" is unreachable after scrolling` | `resolveAccessible` |
|
|
74
|
+
| The region holds no reachable match | `Interactive target "<name>" is not reachable inside "<region>"` | `clickAccessibleWithin` |
|
|
75
|
+
| The region holds several reachable matches | `Interactive target "<name>" is ambiguous across <n> elements inside "<region>"` | `clickAccessibleWithin` |
|
|
76
|
+
| No native disclosure is reachable under the name | `Native disclosure "<name>" is not visible and focus-reachable` | `clickDisclosure` |
|
|
77
|
+
| Several native disclosures carry the name | `Native disclosure "<name>" is ambiguous across <n> elements` | `clickDisclosure` |
|
|
78
|
+
| Forward Tab never lands on the target | `Interactive target "<name>" is not reachable through forward Tab traversal: <trail>` | `traverseAccessible` |
|
|
79
|
+
| The named region is hidden | `Named region "<name>" is not visible` | `readPerception` |
|
|
80
|
+
| Several named regions carry the name | `Named region "<name>" is ambiguous across <n> elements` | `readPerception` |
|
|
81
|
+
| The resolved control renders no value | `Interactive target "<name>" does not carry a value` | `readValue` |
|
|
53
82
|
|
|
54
83
|
- Report an absent control and a present-but-unreachable one as different findings: absence names
|
|
55
84
|
a missing control, and unreachability names the interface gating one that exists.
|
|
56
|
-
- Report ambiguity as a finding about the surface.
|
|
85
|
+
- Report ambiguity as a finding about the surface. Quote the match count from the message, and
|
|
57
86
|
re-target the journey by role or region.
|
|
87
|
+
- Read the package's own voice table before asserting a message this file does not list. A workspace
|
|
88
|
+
that transcribes a voice by memory asserts a sentence the package does not throw.
|
|
58
89
|
|
|
59
90
|
## Role vocabulary
|
|
60
91
|
|
|
61
|
-
Never infer a role from markup. Confirm the computed role in the browser whenever a
|
|
62
|
-
resolving
|
|
92
|
+
Never infer a role from markup. Confirm the computed role in the browser with `readRole` whenever a
|
|
93
|
+
target stops resolving, and read the exposed name with `readName` and the exposed state with
|
|
94
|
+
`readStates` beside it.
|
|
63
95
|
|
|
64
96
|
- A `list`-bearing input computes `combobox`, not `textbox`. Attaching native suggestion machinery
|
|
65
97
|
to a field is a role change: re-target every journey that names that field, and read a resolver
|
|
66
98
|
miss immediately after such a change as this before treating the element as missing.
|
|
67
99
|
- Always target a tab by its role. A tab and its panel collide on a bare name by construction,
|
|
68
100
|
because the panel is labelled by its tab.
|
|
69
|
-
- `<summary>`
|
|
70
|
-
|
|
71
|
-
|
|
101
|
+
- Drive a `<summary>` with `clickDisclosure`, which applies the same reachability conditions and
|
|
102
|
+
throws its own voices. The provider's role locators do not resolve it, because it is exposed as a
|
|
103
|
+
native disclosure rather than through a role.
|
|
72
104
|
|
|
73
105
|
## Region-scoped resolution
|
|
74
106
|
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
- Provide this form for repeated short verbs such as `Add`, and for a control whose accessible name
|
|
80
|
-
is completed by a status the row renders.
|
|
81
|
-
- Apply the same reachability conditions inside the region, and throw voices that name the region
|
|
82
|
-
as well as the target.
|
|
107
|
+
Reach for `clickAccessibleWithin(region, role, name)` where a short verb such as `Add` repeats
|
|
108
|
+
across a page, and where a rendered status completes a control's accessible name. It applies the
|
|
109
|
+
same reachability conditions inside the region and names the region in every voice it throws.
|
|
83
110
|
|
|
84
111
|
## Input and traversal
|
|
85
112
|
|
|
@@ -88,44 +115,45 @@ clickAccessibleWithin(region: string, role: string, name: string): Promise<void>
|
|
|
88
115
|
| `typeAccessible(name, text)` | Focus the field, select all, delete, then send real keystrokes. Escape the provider's key syntax in the text. |
|
|
89
116
|
| `fillAccessible(name, text)` | Replace the value in one operation for text too long to type. The real element still publishes real input. |
|
|
90
117
|
| `pressKeys(keys)` | Send a provider keyboard sequence for Enter, arrows, modifiers, and combinations. |
|
|
91
|
-
| `traverseAccessible(name)` | Move focus by forward Tab from wherever focus is, and return the target
|
|
118
|
+
| `traverseAccessible(name)` | Move focus by forward Tab from wherever focus is, and return the target after focus lands on it. |
|
|
92
119
|
|
|
93
|
-
-
|
|
94
|
-
|
|
95
|
-
-
|
|
96
|
-
|
|
97
|
-
above the cycle so a page with no tab order fails instead of hanging.
|
|
98
|
-
-
|
|
99
|
-
steps. A framework may replace the node between resolution and focus arrival.
|
|
100
|
-
- Never call the browser's focus method to place focus.
|
|
120
|
+
- Reach for `typeAccessible` where the keystrokes are part of what the journey claims, and
|
|
121
|
+
`fillAccessible` where the text is only a payload the person pastes.
|
|
122
|
+
- Let `traverseAccessible` end the walk. It counts a step only where focus lands, stops at one
|
|
123
|
+
complete cycle of the tab order, carries the trail of what focus reached in its voice, and holds a
|
|
124
|
+
cap above the cycle so a page with no tab order fails instead of hanging.
|
|
125
|
+
- Never call the browser's focus method to place focus, and never hold a node reference across
|
|
126
|
+
traversal steps. A framework may replace the node between resolution and focus arrival.
|
|
101
127
|
|
|
102
128
|
## Perception
|
|
103
129
|
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
```
|
|
130
|
+
`readPerception(name)` returns the normalized `innerText` of exactly one visible named region,
|
|
131
|
+
dialog, table, tab panel, alert, or status. Quote that text in assertions.
|
|
107
132
|
|
|
108
|
-
- Return the normalized `innerText` of exactly one visible named region, dialog, table, tab panel,
|
|
109
|
-
alert, or status. Collapse whitespace runs to single spaces and trim.
|
|
110
133
|
- Read `innerText`, never `textContent`: `innerText` applies CSS transforms and leaves out content
|
|
111
|
-
the layout hides.
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
134
|
+
the layout hides. Descendant visually-hidden text stays in, which is what a screen reader
|
|
135
|
+
perceives.
|
|
136
|
+
- Reach for `readPage` for a cross-region sentence and the vocabulary sweep, `readFocus` for the
|
|
137
|
+
active element's rendered text, and `readValue` for a resolved control's rendered value. A
|
|
138
|
+
perception assertion may quote that value: it is a rendered fact rather than internal state.
|
|
139
|
+
- Reach for `describeTree` and `describeFocus` where the subject is the accessible tree itself —
|
|
140
|
+
roles, names, states, and focus order. They are the accessibility snapshot a written artifact
|
|
141
|
+
composes ([decide.md](decide.md) → The rendered artifact).
|
|
119
142
|
|
|
120
143
|
## Mounting and cleanup
|
|
121
144
|
|
|
122
145
|
- Mount the shipped root component with its real provisions and return an idempotent cleanup that
|
|
123
|
-
unmounts the app and removes its container.
|
|
146
|
+
unmounts the app and removes its container. Reach for the published `mount`, `render`, and `build`
|
|
147
|
+
where a fixture needs a node rather than the application.
|
|
124
148
|
- Undo everything a journey changed after each test: unmount, destroy the session, reset the theme,
|
|
125
|
-
|
|
126
|
-
inherit the previous journey's state.
|
|
149
|
+
return the route to its entry, and clear what the application persisted with `clearStorage` and
|
|
150
|
+
`removeDatabase`. Never let a journey inherit the previous journey's state.
|
|
151
|
+
- Record what a journey did with `createJournal`, started inside the journey and stopped in a
|
|
152
|
+
`finally`. Its `steps` and `output` are the evidence a failing journey hands back, and the input
|
|
153
|
+
the run's written artifact composes.
|
|
127
154
|
|
|
128
155
|
## The capture hook
|
|
129
156
|
|
|
130
|
-
|
|
131
|
-
variant matrix, and the proofs that read what it
|
|
157
|
+
`createPortfolio` is the capture door and `place(state, element?)` is the hook.
|
|
158
|
+
[captures.md](captures.md) owns the registry, the variant matrix, and the proofs that read what it
|
|
159
|
+
wrote.
|
|
@@ -0,0 +1,84 @@
|
|
|
1
|
+
# The statechart family
|
|
2
|
+
|
|
3
|
+
Declare one transition table, and give it to the automated run that asserts it and to the
|
|
4
|
+
harness a person watches. Never write a second table for the harness.
|
|
5
|
+
|
|
6
|
+
## Declare the table
|
|
7
|
+
|
|
8
|
+
- Declare each transition as a `StateTransition` carrying its `name`, its `from` state, the `event`,
|
|
9
|
+
and its `to` state. Type it on the entity's own state and event unions, so a row naming a state the
|
|
10
|
+
entity does not have fails to typecheck.
|
|
11
|
+
- Write one `StateScenario` per transition, carrying that `transition` plus `arrange`, `act`, and
|
|
12
|
+
`assert`. Each phase receives the context and the part of the transition it owns.
|
|
13
|
+
- Put the scenarios in the workspace's browser test setup module. Put the table beside them until a
|
|
14
|
+
harness page ships; from then on declare it in the application's own constants module, typed on the
|
|
15
|
+
entity's unions, and import it from the page and from the setup alike, because a page cannot
|
|
16
|
+
import from `tests/` and a second table is what this reference forbids.
|
|
17
|
+
- Declare a transition for every event the surface accepts in every state it accepts it, including
|
|
18
|
+
the event that leaves the state unchanged. A table that lists only the moves the happy path takes
|
|
19
|
+
proves the happy path.
|
|
20
|
+
|
|
21
|
+
## Run the table
|
|
22
|
+
|
|
23
|
+
- Run the table with `executeScenarios(scenarios, build)` from `@orkestrel/test`. It drives the rows
|
|
24
|
+
in the order they are written, builds a context per row, and stops at the first row that throws.
|
|
25
|
+
- Run one row with `executeScenario(scenario, context)` where a single transition is the subject.
|
|
26
|
+
- Build the context in `build`, which receives the row it is building for. That is what lets one
|
|
27
|
+
table mix fixtures.
|
|
28
|
+
- Let the runner name the failure. It prepends the transition's `name` to whatever the row threw and
|
|
29
|
+
carries the original as the `cause`, so a bare assertion message still says which row failed.
|
|
30
|
+
- Never assert the entity's internal state in `assert` where the transition is one a person drives.
|
|
31
|
+
Assert what the interface renders, through `readPerception`, `readValue`, or `readStates`.
|
|
32
|
+
|
|
33
|
+
## Drive the act the way the transition happens
|
|
34
|
+
|
|
35
|
+
- Drive `act` through the journey verbs — `clickAccessible`, `clickDisclosure`, `typeAccessible`,
|
|
36
|
+
`pressKeys`, `traverseAccessible` — for every transition a person can cause.
|
|
37
|
+
- Drive `act` through the entity's own API only where the transition is the entity's rather than the
|
|
38
|
+
person's: a lifecycle event, a transport reply, a timer the surface owns.
|
|
39
|
+
- Say which door each row used, in the row's `name`. A table that mixes the doors silently reads as
|
|
40
|
+
a set of user transitions and proves something else.
|
|
41
|
+
|
|
42
|
+
## Build the harness a person watches
|
|
43
|
+
|
|
44
|
+
Mount the harness on the same table. It stays in the repository that owns the surface: only
|
|
45
|
+
`StateTransition`, `StateScenario`, `executeScenario`, `executeScenarios`, `STATECHART_ATTRIBUTES`,
|
|
46
|
+
and `STATECHART_STATUSES` are published, and a workspace that spells a `data-statechart-*` string of
|
|
47
|
+
its own has left the contract.
|
|
48
|
+
|
|
49
|
+
- Render one play control per transition and one play-all control, each disabled while a run is in
|
|
50
|
+
flight.
|
|
51
|
+
- Render a state badge carrying the entity's current state, an event log of what the entity emitted,
|
|
52
|
+
and a `role="status"` announcer that narrates each step in a sentence, so a screen reader and a
|
|
53
|
+
vision model both read the run without visual chrome.
|
|
54
|
+
- Publish `STATECHART_ATTRIBUTES.status` on the harness root, and cycle its value through
|
|
55
|
+
`STATECHART_STATUSES`: `pending` before a run has a result for every row, `idle` standing ready,
|
|
56
|
+
`running` in flight, and `passed` or `failed` as the terminal reading. Publish
|
|
57
|
+
`STATECHART_ATTRIBUTES.passed`, `.failed`, and `.total` on the same element, so a gate finds the
|
|
58
|
+
harness and reads the tally from one node.
|
|
59
|
+
- Publish `STATECHART_ATTRIBUTES.scenario` and `STATECHART_ATTRIBUTES.result` on each row, and
|
|
60
|
+
`STATECHART_ATTRIBUTES.state` on the element rendering the entity's current state.
|
|
61
|
+
- Write every attribute from the map rather than from a literal. A harness that sets an attribute
|
|
62
|
+
the gate does not read fails silently as a run that never completes.
|
|
63
|
+
- Deep-link one transition and the play-all run from the route, so a decision round names the exact
|
|
64
|
+
link it wants looked at ([decide.md](decide.md) → The rendered artifact).
|
|
65
|
+
- Give the harness a demo step that leaves the widget in its most legible state after the run, for a
|
|
66
|
+
person or a vision model deciding on a look.
|
|
67
|
+
- Pace the harness for a person to watch. The gate inherits that wall time, so budget the gate from
|
|
68
|
+
the row count and the pause rather than from a fixed timeout.
|
|
69
|
+
|
|
70
|
+
## Gate the harness
|
|
71
|
+
|
|
72
|
+
Prove the harness from the browser project, through the interface:
|
|
73
|
+
|
|
74
|
+
- Mount the harness page and clear the route's query first, so a leftover deep link cannot start the
|
|
75
|
+
walk before the gate does.
|
|
76
|
+
- Press the play-all control through `clickAccessible`, never through a constructed event and never
|
|
77
|
+
by setting the deep link.
|
|
78
|
+
- Poll `STATECHART_ATTRIBUTES.status` until it reads `passed` or `failed`. Never assert it from one
|
|
79
|
+
read after the click.
|
|
80
|
+
- Assert the status reads `passed`, the failed tally reads zero, and the passed tally equals the
|
|
81
|
+
total. Name the failing rows from `STATECHART_ATTRIBUTES.scenario` in the failure message, so a red
|
|
82
|
+
gate says which transition broke.
|
|
83
|
+
- Assert the harness inventory before the rows: a page that mounted no transition passes every tally
|
|
84
|
+
assertion.
|
|
@@ -0,0 +1,87 @@
|
|
|
1
|
+
# Proving what the browser resolved
|
|
2
|
+
|
|
3
|
+
Prove a style from what the browser resolved on the mounted surface. The `enterprise-bootstrap`
|
|
4
|
+
skill's [instruments reference](../../enterprise-bootstrap/references/inspection.md) names each
|
|
5
|
+
instrument's property, its population, and the negative controls that must fail. Take those from
|
|
6
|
+
there and the reading from here.
|
|
7
|
+
|
|
8
|
+
## Assert the resolved value
|
|
9
|
+
|
|
10
|
+
- Read one property with `readStyle(element, property)` and a length with `readPixels(element, property)`. A
|
|
11
|
+
class present in the markup and absent from the cascade resolves to nothing, and an assertion on
|
|
12
|
+
the class list passes on it.
|
|
13
|
+
- Never substitute `findRule` for a resolved read. It proves a declaration exists, and another rule
|
|
14
|
+
may still win; reach for it where the stylesheet itself is the subject.
|
|
15
|
+
- Compare a color through `matchesColor` or `parseCSSColor` rather than by string. A browser normalizes a color
|
|
16
|
+
expression, so a literal comparison fails on a value that resolved correctly.
|
|
17
|
+
|
|
18
|
+
## Run per variant
|
|
19
|
+
|
|
20
|
+
The run axis is fixed in [SKILL.md](../SKILL.md) → Read the variant once. The matrix family's own
|
|
21
|
+
readings follow.
|
|
22
|
+
|
|
23
|
+
- Apply each variant's `apply` and its `width` and `height` before the readings, and take every
|
|
24
|
+
reading for that variant before moving to the next.
|
|
25
|
+
- Name the attribute the surface actually reads in `apply`; a Bootstrap surface switches on
|
|
26
|
+
`data-bs-theme`. An `apply` that sets another attribute leaves the run in the default theme, where
|
|
27
|
+
every reading passes.
|
|
28
|
+
- Assert that the run read every declared variant. A matrix that silently walked one variant reports
|
|
29
|
+
a pass for the theme nobody exercised.
|
|
30
|
+
- Report which variants a result covers beside it. A pairing that appears only in a state the run
|
|
31
|
+
never entered is unmeasured.
|
|
32
|
+
|
|
33
|
+
## Contrast and focus chrome
|
|
34
|
+
|
|
35
|
+
- Read a text pairing with `readContrast(element)`, which composites the painted ancestors to the first
|
|
36
|
+
opaque layer, so a translucent tint reads as a tint over what shows through it.
|
|
37
|
+
- Pass `floor` only where the surface the stack really sits on is known, and pass `CANVAS_COLOR`
|
|
38
|
+
where the browser paints onto its own canvas. Omitting `floor` refuses a stack whose painted layers
|
|
39
|
+
are all translucent; take that refusal as the reading, because an assumed white canvas turns "this
|
|
40
|
+
surface declares no background" into a number that reads like a measurement.
|
|
41
|
+
- Read focus chrome with `readRing(control)`, after focus arrived through `traverseAccessible`,
|
|
42
|
+
`pressKeys`, or a real click. Pass `worn` where the chrome is painted onto a second element such
|
|
43
|
+
as a label. It reports `undefined` for a control not matching `:focus-visible`, for the browser's
|
|
44
|
+
own automatic ring, and for a focus style that only repaints the fill — treat each as a finding
|
|
45
|
+
about the surface rather than as a pass.
|
|
46
|
+
- Carry the negative controls the composited-contrast instrument names, in the same run and composed
|
|
47
|
+
in the harness rather than taken from the surface. An instrument whose negative control passes is
|
|
48
|
+
broken, and its readings are not evidence.
|
|
49
|
+
- Reach for `measureContrast`, `measureLuminance`, `blendColor`, `readLayers`, and `readBackdrop`
|
|
50
|
+
only where the composite itself is the subject. Never re-derive `readContrast` from them.
|
|
51
|
+
|
|
52
|
+
## The authored-class census
|
|
53
|
+
|
|
54
|
+
Take the property, the population, and the negative controls from the instruments reference →
|
|
55
|
+
Authored class in the shipped cascade. This is the reading.
|
|
56
|
+
|
|
57
|
+
- Read the census as the difference between `readClasses(root)` of the mounted surface and
|
|
58
|
+
`readCascade()`. A token in the difference is a class the markup uses and no loaded stylesheet
|
|
59
|
+
declares.
|
|
60
|
+
- Take `root` from the mounted surface, so the census covers what rendered rather than what a
|
|
61
|
+
template file spells.
|
|
62
|
+
- Report what `readClasses` walked as the population, and fail a run that walked none.
|
|
63
|
+
- Append the extraction-door negative control to that same `root`, so it reaches the difference
|
|
64
|
+
through `readClasses` rather than beside it.
|
|
65
|
+
|
|
66
|
+
## Style escapes
|
|
67
|
+
|
|
68
|
+
Take the property, the population, the named exemptions, and the negative control from the
|
|
69
|
+
instruments reference → Style escapes. This is the reading.
|
|
70
|
+
|
|
71
|
+
- Read escapes with `extractStyles(root)`, which returns the markup of every hit it found.
|
|
72
|
+
- Take the reading before any journey drives the surface, because the population is the undriven
|
|
73
|
+
tree.
|
|
74
|
+
- Append the harness-built negative control element to that same `root`, so it reaches the reading
|
|
75
|
+
through `extractStyles`.
|
|
76
|
+
- Reach for `extractOrphans` where the finding is a child element rendered outside its required
|
|
77
|
+
parent, and `readRows` where the subject is a repeated row's rendered text.
|
|
78
|
+
|
|
79
|
+
## Tokens
|
|
80
|
+
|
|
81
|
+
- Read a token with `readToken(element, name)` where inheritance matters and `readRootToken(name)` where the
|
|
82
|
+
document declares it. The leading dashes are optional in each.
|
|
83
|
+
- Compare values, never presence. An absent token and one declared empty both read as `''`, and a
|
|
84
|
+
`var()` naming an undeclared property resolves to the inherited color rather than refusing, so an
|
|
85
|
+
assertion on presence passes on a token nobody declared.
|
|
86
|
+
- Read each token once per variant and assert the values differ where the design says the variants
|
|
87
|
+
differ. A pair of variants that resolves a token identically is a theme that did not switch.
|
|
@@ -10,6 +10,10 @@ copy `# Unit UNIT_ID — SHORT_SUBJECT`. Delete each italic reminder as you fill
|
|
|
10
10
|
under, and leave no row blank: fill a row you cannot close with a named unknown label, and
|
|
11
11
|
describe that label under § Unknowns with how the unit reports back on it.
|
|
12
12
|
|
|
13
|
+
For a read-only audit lane, fill the rows
|
|
14
|
+
`.agents/skills/orkestrel-falsify/references/brief.md` § "The read-only audit lane's brief" names
|
|
15
|
+
and delete the rest of this template.
|
|
16
|
+
|
|
13
17
|
## Role and engine
|
|
14
18
|
|
|
15
19
|
ROLE_NAME on ENGINE_NAME, reached as TRANSPORT.
|
|
@@ -73,7 +77,7 @@ answer around._
|
|
|
73
77
|
_Grant a behaviour with the tests that pin it, a constant with every fixture and expectation derived
|
|
74
78
|
from it, a template with the materialized copy the package generates from it, and a mechanism with
|
|
75
79
|
the prose describing it: the comment beside the code it edits and the guide passage stating the
|
|
76
|
-
behaviour it moves._
|
|
80
|
+
behaviour it moves. A unit that moves a published symbol owns the package `README.md`._
|
|
77
81
|
|
|
78
82
|
**Shared (report-only).** SHARED_FILES
|
|
79
83
|
|
|
@@ -84,21 +88,25 @@ edits nothing in this row._
|
|
|
84
88
|
|
|
85
89
|
_Name each file the unit must not touch, and read every acceptance criterion against this row line by
|
|
86
90
|
line. Grant the file a criterion needs, or strike that criterion. A file the change breaks that
|
|
87
|
-
appears in no row of this section is unscoped.
|
|
91
|
+
appears in no row of this section is unscoped. Never list `tmp/probe/` off-limits: it is the unit's
|
|
92
|
+
probe home, and `.claude/rules/tests.md` puts every runtime probe there._
|
|
88
93
|
|
|
89
94
|
**What asserts the state this change ends.** FILES_THE_RESULT_MAKES_FALSE
|
|
90
95
|
|
|
91
96
|
_List every file the result makes false rather than every file that declares the thing changing: the
|
|
92
97
|
test asserting the reversed behaviour, the fixture carrying the raised value, the golden digest over
|
|
93
98
|
generated output, the consumer script naming the removed union member. Derive the list by running the
|
|
94
|
-
suite; where you cannot run it, name the search's bound so the unit re-derives the list.
|
|
95
|
-
|
|
99
|
+
suite; where you cannot run it, name the search's bound so the unit re-derives the list. A rename's
|
|
100
|
+
search bound is a word-boundary sweep over the old name followed by a case-insensitive sweep over
|
|
101
|
+
its `-s`, `-ed`, and `-ing` inflections. End each entry in Owned, in Shared, or with a named carrier
|
|
102
|
+
dispatched before this change ships._
|
|
96
103
|
|
|
97
104
|
**Tools and limits.** ALLOWED_TOOLS, PERMISSION_LIMITS
|
|
98
105
|
|
|
99
106
|
_Check the § Output mechanism and every acceptance criterion's verification method against this
|
|
100
107
|
allowlist. A read-only lane writes no report file and runs no probe, so hand it the rendered evidence
|
|
101
|
-
instead.
|
|
108
|
+
instead. A rename moves the file with the shell's `mv`, never `git mv`; `git add -N` is permitted
|
|
109
|
+
only to render diff evidence._
|
|
102
110
|
|
|
103
111
|
## Execution
|
|
104
112
|
|
|
@@ -138,8 +146,9 @@ criterion behind it. Where the change edits a file the repository vendors or dig
|
|
|
138
146
|
regeneration step ahead of every gate that reads the generated artifact. Ask what the change does
|
|
139
147
|
to every fact you measured, and fix each criterion to the state the unit finishes in. Close each
|
|
140
148
|
criterion with owned files alone, and name the property the unit must change; record a consequence
|
|
141
|
-
you expect to follow as an observation, never as a criterion. A
|
|
142
|
-
|
|
149
|
+
you expect to follow as an observation, never as a criterion. A criterion that closes on an
|
|
150
|
+
instrument's reading names that instrument's negative control and the class the control proves the
|
|
151
|
+
instrument can see. A scoped run over the unit's own owned files stays a legitimate criterion._
|
|
143
152
|
|
|
144
153
|
**Observations, not criteria.** TIMING_SENSITIVE_OR_WHOLE_SUITE_GATES
|
|
145
154
|
|
|
@@ -21,8 +21,10 @@ with the permission mode the route pins. Never substitute a fixed Claude model i
|
|
|
21
21
|
|
|
22
22
|
Verify that the `claude` CLI resolves and is authenticated before first use. On either
|
|
23
23
|
failure return it immediately with the fallback named, so the Sol main session records
|
|
24
|
-
Opus unavailable for the round.
|
|
25
|
-
|
|
24
|
+
Opus unavailable for the round. Where the binary is absent, that report names the install
|
|
25
|
+
command for the `claude` CLI, so the Orchestrator can put it to the user in the same turn
|
|
26
|
+
it records the bench dark and re-probe when the user answers. Never install, authenticate,
|
|
27
|
+
or substitute an API key, access token, or copied auth file.
|
|
26
28
|
|
|
27
29
|
Journal every run: redirect --output-format stream-json to tmp/claude/<unit>.jsonl,
|
|
28
30
|
which is gitignored, and record the session id. A bench unit with no journal ran on its
|
|
@@ -119,7 +119,7 @@ skill owns the value set and the terminal line. Point the brief at both; restate
|
|
|
119
119
|
|
|
120
120
|
## Implementer route
|
|
121
121
|
|
|
122
|
-
Sandbox `workspace-write`,
|
|
122
|
+
Sandbox `workspace-write`, the checkout the route writes in, its sole serial writer from a clean committed
|
|
123
123
|
baseline, with owned files, off-limits files, and a deviation contract. The brief forbids
|
|
124
124
|
dependency installation, commits, pushes, publishing, credentials, destructive commands,
|
|
125
125
|
shared-file edits, and tree-wide mutating gates.
|
|
@@ -147,6 +147,9 @@ negative tests is unaffected.
|
|
|
147
147
|
|
|
148
148
|
- Verify `codex --version` before first use. On Windows `codex` resolves in Bash through
|
|
149
149
|
the extensionless npm shim; if it does not, invoke `codex.cmd`.
|
|
150
|
+
- Binary absent: report it so the Orchestrator can record the bench dark and, in the same turn,
|
|
151
|
+
name to the user the install command for `@openai/codex` and the bench it unblocks. The
|
|
152
|
+
Orchestrator re-probes when the user answers. Never install it yourself.
|
|
150
153
|
- Binary present but authentication unavailable: report it so the Orchestrator can start
|
|
151
154
|
device-auth recovery in the same turn. It backgrounds `codex login --device-auth` with
|
|
152
155
|
output captured to `tmp/codex/login.log`, surfaces the verification URL and one-time code
|
|
@@ -34,7 +34,9 @@ Everything `.agents/orchestration.md`'s dispatch contract requires, plus:
|
|
|
34
34
|
- **Every authority the brief references must exist in the tree the exec is rooted in.** Check
|
|
35
35
|
before dispatch. A brief citing a rule file or section the executor cannot find delivers nothing
|
|
36
36
|
while looking like authority, and it fails silently — an auditor does not report a heading it
|
|
37
|
-
never saw. Propagate the missing file
|
|
37
|
+
never saw. Propagate the missing file rather than restating its contents in the brief, and take
|
|
38
|
+
the stale-authority branch in `.agents/skills/orkestrel-falsify/references/brief.md` § "What not
|
|
39
|
+
to put in a brief" where the executor's tree carries a superseded vendored copy.
|
|
38
40
|
- For an audit: the subject as numbered falsifiable claims, and the skill that fixes the verdict
|
|
39
41
|
shape. The Falsification section of `.claude/rules/quality.md` owns the method and the evidence
|
|
40
42
|
each verdict carries. The verdict shape defaults to `orkestrel-falsify`; a dispatch may name a
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: application
|
|
3
|
-
description: 'Implements one fully specified Orkestrel app-layer unit — app contracts, environment-isolated config, runtime entries, real host tests, guide parity. Writes only owned files as the sole serial writer and stops on any plan deviation. Nontrivial app design belongs to GPT-5.6 Sol or Opus 5.'
|
|
3
|
+
description: 'Implements one fully specified Orkestrel app-layer unit — app contracts, environment-isolated config, runtime entries, real host tests, guide parity. Writes only owned files in the checkout the unit writes as the sole serial writer and stops on any plan deviation. Nontrivial app design belongs to GPT-5.6 Sol or Opus 5.'
|
|
4
4
|
tools: Read, Grep, Glob, Edit, Write, Bash
|
|
5
5
|
model: sonnet
|
|
6
6
|
effort: low
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: builder
|
|
3
|
-
description: 'Implements one small, fully specified, taste-free unit exactly as dispatched. Writes only owned files in the
|
|
3
|
+
description: 'Implements one small, fully specified, taste-free unit exactly as dispatched. Writes only owned files in the checkout the unit writes as the sole serial writer, validates narrowly, and stops on any plan deviation. Nontrivial implementation belongs to GPT-5.6 Sol or Opus 5.'
|
|
4
4
|
tools: Read, Grep, Glob, Edit, Write, Bash
|
|
5
5
|
model: sonnet
|
|
6
6
|
effort: low
|
|
@@ -19,8 +19,8 @@ dispatch contract.
|
|
|
19
19
|
|
|
20
20
|
- Before writing, read **AGENTS.md**, every applicable `.claude/rules/*.md`, the
|
|
21
21
|
dispatch-named skill and required references, and the governing guide/spec. All bind you
|
|
22
|
-
as written there, and this charter restates none of them.
|
|
23
|
-
|
|
22
|
+
as written there, and this charter restates none of them.
|
|
23
|
+
- An app-layer unit belongs to `application`: stop and say so.
|
|
24
24
|
- Write ONLY the owned files named in your dispatch. Shared or off-limits files are
|
|
25
25
|
report-only: if one needs a change, RETURN the exact patch — never edit it.
|
|
26
26
|
- NO tree-wide or mutating commands: never `format`, lint `--fix`, or `build`.
|
|
@@ -37,6 +37,11 @@ No judgment calls: a question that needs one becomes a **referral** — specific
|
|
|
37
37
|
evidenced, addressed to the subjective lane when it is running and to the
|
|
38
38
|
Orchestrator when it is not — never a guess and never a verdict of yours.
|
|
39
39
|
|
|
40
|
+
Rule a claim whose only evidence is the writer's report `UNRESOLVED`, never
|
|
41
|
+
`CONFIRMED`, whatever the brief says. A quoted command and exit code inside a
|
|
42
|
+
report is the writer quoting itself, so it evidences nothing until a lane that ran
|
|
43
|
+
the command supplies the reading.
|
|
44
|
+
|
|
40
45
|
## Output contract
|
|
41
46
|
|
|
42
47
|
Return the shape fixed by the dispatch.
|
|
@@ -40,7 +40,8 @@ shim that answered once does not clear it, and when it does fire it leaves only
|
|
|
40
40
|
`SetConsoleWindowTitle` trace — which reads as a bench that returned nothing rather than as a launch
|
|
41
41
|
that never happened. The versioned entry has no console dependency and no such failure mode.
|
|
42
42
|
|
|
43
|
-
Read an empty shim run as a launch failure until its
|
|
43
|
+
Read an empty shim run as a launch failure until its `.err` journal is checked for that
|
|
44
|
+
trace.
|
|
44
45
|
|
|
45
46
|
If nothing responds the bench is dark. Stop with a deviation naming the fallback from the root
|
|
46
47
|
tedious-work ladder — Luna, then Sonnet. Never hand the reading to the Orchestrator, `planner`, or
|
|
@@ -48,14 +49,21 @@ tedious-work ladder — Luna, then Sonnet. Never hand the reading to the Orchest
|
|
|
48
49
|
|
|
49
50
|
Create `tmp/cursor/` first. Write any brief longer than a couple of sentences to
|
|
50
51
|
`tmp/cursor/<unit>-brief.md` and make the prompt a pointer to it; briefs never travel as
|
|
51
|
-
fragile shell arguments. Every run journals its
|
|
52
|
-
and an interrupted run leaves its partial distillate on disk:
|
|
52
|
+
fragile shell arguments. Every run journals its event stream, so the user can tail progress
|
|
53
|
+
live and an interrupted run leaves its partial distillate on disk:
|
|
53
54
|
|
|
54
|
-
`<resolved-entry> -p --trust --mode=ask --model "$CURSOR_GROK_MODEL" "<
|
|
55
|
+
`<resolved-entry> -p --trust --mode=ask --model "$CURSOR_GROK_MODEL" --output-format stream-json "<pointer>" > tmp/cursor/<unit>.jsonl 2> tmp/cursor/<unit>.err`
|
|
55
56
|
|
|
56
57
|
Write that chain to `tmp/cursor/run.sh` and run the file, so the resolution, the model, and the
|
|
57
58
|
journalling are one artifact the next run reuses.
|
|
58
59
|
|
|
60
|
+
The journal's first event is the `init` event, and its `session_id` is the run's recovery
|
|
61
|
+
handle. The journal's `result` event carries the final answer. Return the journal path and
|
|
62
|
+
that session id with the result, so the Orchestrator can confirm the bench ran. Read the
|
|
63
|
+
`.err` file before calling a run empty; a launch that never reached the model leaves its
|
|
64
|
+
trace only there. Resume an interrupted run through the CLI's `--resume` option, probed
|
|
65
|
+
before its first use.
|
|
66
|
+
|
|
59
67
|
Run that yourself only for a short bounded ask finishing in about two minutes. For anything
|
|
60
68
|
longer your job ends at drafting: return the brief path, the exact resolved command, and the
|
|
61
69
|
journal path, and let the Orchestrator launch it as a harness-tracked background command under
|
|
@@ -79,7 +87,9 @@ Return only:
|
|
|
79
87
|
- `Question`: one line.
|
|
80
88
|
- `Evidence`: concise facts with `file:line` or primary-source pointers.
|
|
81
89
|
- `Distillate`: the smallest context the next engine needs.
|
|
82
|
-
- `Unknowns`: unresolved facts, not recommendations
|
|
90
|
+
- `Unknowns`: unresolved facts, not recommendations, naming every input row the
|
|
91
|
+
distillate did not reach.
|
|
92
|
+
- `Journal`: the journal path and the session id from its `init` event.
|
|
83
93
|
- `Deviation`: unavailable CLI, model, or auth; command failure; dirty containment.
|
|
84
94
|
|
|
85
95
|
Grok's output is evidence, never a decision or a verdict.
|