@orkestrel/scaffold 0.0.76 → 0.0.78
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/agents/skills/orkestrel-dispatch/scripts/bench.js +204 -0
- package/dist/agents/skills/orkestrel-dispatch/scripts/brief.js +102 -0
- package/dist/agents/skills/orkestrel-dispatch/scripts/cite.js +95 -0
- package/dist/agents/skills/orkestrel-dispatch/scripts/helpers.js +207 -0
- package/dist/agents/skills/orkestrel-dispatch/scripts/launch.js +108 -0
- package/dist/agents/skills/orkestrel-dispatch/scripts/login.js +114 -0
- package/dist/agents/skills/orkestrel-dispatch/scripts/result.js +108 -0
- package/dist/agents/skills/orkestrel-dispatch/scripts/sweep.js +156 -0
- package/dist/agents/skills/orkestrel-harden/scripts/discovery.js +196 -0
- package/dist/agents/skills/orkestrel-publish/scripts/compare.js +206 -0
- package/dist/agents/skills/orkestrel-publish/scripts/pins.js +93 -0
- package/dist/agents/skills/orkestrel-publish/scripts/wave.js +458 -0
- package/dist/agents/skills/orkestrel-publish/scripts/window.js +188 -0
- package/dist/agents/skills/orkestrel-scout/scripts/map.js +300 -0
- package/dist/agents/templates/brief.md +55 -0
- package/dist/bin/main.js +4 -2
- package/dist/bin/main.js.map +1 -1
- package/dist/host/AGENTS.md +77 -135
- package/dist/host/agents/orchestration.md +147 -927
- package/dist/host/agents/skills/enterprise-bootstrap/SKILL.md +2 -2
- package/dist/host/agents/skills/enterprise-bootstrap/references/inspection.md +1 -1
- package/dist/host/agents/skills/{orkestrel-align-packages → orkestrel-align}/SKILL.md +6 -13
- package/dist/host/agents/skills/{orkestrel-align-packages → orkestrel-align}/agents/openai.yaml +1 -1
- package/dist/host/agents/skills/{orkestrel-align-packages → orkestrel-align}/references/fleet.md +5 -7
- package/dist/host/agents/skills/{orkestrel-build-application → orkestrel-build}/SKILL.md +11 -22
- package/dist/host/agents/skills/{orkestrel-build-application → orkestrel-build}/agents/openai.yaml +1 -1
- package/dist/host/agents/skills/orkestrel-debrief/SKILL.md +8 -16
- package/dist/host/agents/skills/orkestrel-debrief/references/instruction-audit.md +3 -3
- package/dist/host/agents/skills/orkestrel-debrief/references/retention.md +13 -13
- package/dist/host/agents/skills/orkestrel-dispatch/SKILL.md +61 -0
- package/dist/host/agents/skills/orkestrel-dispatch/agents/openai.yaml +4 -0
- package/dist/host/agents/skills/orkestrel-dispatch/references/bench.md +25 -0
- package/dist/host/agents/skills/orkestrel-dispatch/references/launch.md +32 -0
- package/dist/host/agents/skills/orkestrel-dispatch/scripts/bench.ts +259 -0
- package/dist/host/agents/skills/orkestrel-dispatch/scripts/brief.ts +110 -0
- package/dist/host/agents/skills/orkestrel-dispatch/scripts/cite.ts +115 -0
- package/dist/host/agents/skills/orkestrel-dispatch/scripts/helpers.ts +239 -0
- package/dist/host/agents/skills/orkestrel-dispatch/scripts/launch.ts +124 -0
- package/dist/host/agents/skills/orkestrel-dispatch/scripts/login.ts +123 -0
- package/dist/host/agents/skills/orkestrel-dispatch/scripts/result.ts +129 -0
- package/dist/host/agents/skills/orkestrel-dispatch/scripts/sweep.ts +157 -0
- package/dist/host/agents/skills/orkestrel-falsify/SKILL.md +42 -193
- package/dist/host/agents/skills/orkestrel-falsify/references/brief.md +38 -108
- package/dist/host/agents/skills/orkestrel-falsify/references/reconcile.md +35 -134
- package/dist/host/agents/skills/{orkestrel-harden-package → orkestrel-harden}/SKILL.md +10 -14
- package/dist/host/agents/skills/{orkestrel-harden-package → orkestrel-harden}/agents/openai.yaml +1 -1
- package/dist/host/agents/skills/{orkestrel-harden-package → orkestrel-harden}/references/hardening.md +3 -4
- package/dist/host/agents/skills/orkestrel-harden/scripts/discovery.ts +228 -0
- package/dist/host/agents/skills/{orkestrel-prove-journey → orkestrel-journey}/SKILL.md +15 -23
- package/dist/host/agents/skills/orkestrel-journey/agents/openai.yaml +4 -0
- package/dist/host/agents/skills/{orkestrel-prove-journey → orkestrel-journey}/references/captures.md +1 -1
- package/dist/host/agents/skills/{orkestrel-polish-surface → orkestrel-polish}/SKILL.md +25 -33
- package/dist/host/agents/skills/{orkestrel-polish-surface → orkestrel-polish}/agents/openai.yaml +1 -1
- package/dist/host/agents/skills/{orkestrel-polish-surface → orkestrel-polish}/references/capture-harness.md +3 -3
- package/dist/host/agents/skills/orkestrel-publish/SKILL.md +33 -20
- package/dist/host/agents/skills/orkestrel-publish/references/release.md +39 -0
- package/dist/host/agents/skills/orkestrel-publish/references/wave.md +22 -21
- package/dist/host/agents/skills/orkestrel-publish/references/window.md +27 -14
- package/dist/host/agents/skills/orkestrel-publish/scripts/compare.ts +220 -0
- package/dist/host/agents/skills/orkestrel-publish/scripts/pins.ts +114 -0
- package/dist/host/agents/skills/orkestrel-publish/scripts/wave.ts +629 -0
- package/dist/host/agents/skills/orkestrel-publish/scripts/window.ts +242 -0
- package/dist/host/agents/skills/orkestrel-scout/SKILL.md +28 -0
- package/dist/host/agents/skills/orkestrel-scout/agents/openai.yaml +4 -0
- package/dist/host/agents/skills/orkestrel-scout/scripts/map.ts +352 -0
- package/dist/host/agents/templates/brief.md +21 -142
- package/dist/host/agents/transports/claude-cli.md +21 -0
- package/dist/host/agents/transports/codex.md +38 -159
- package/dist/host/agents/transports/cursor.md +16 -65
- package/dist/host/claude/AGENTS.md +38 -0
- package/dist/host/claude/agents/analyst.md +14 -53
- package/dist/host/claude/agents/astra.md +26 -0
- package/dist/host/claude/agents/builder.md +14 -30
- package/dist/host/claude/agents/checker.md +13 -57
- package/dist/host/claude/agents/distiller.md +11 -26
- package/dist/host/claude/agents/grok.md +12 -35
- package/dist/host/claude/agents/opus.md +14 -30
- package/dist/host/claude/agents/planner.md +10 -44
- package/dist/host/claude/agents/researcher.md +11 -30
- package/dist/host/claude/agents/reviewer.md +11 -95
- package/dist/host/claude/agents/scout.md +9 -23
- package/dist/host/claude/agents/verifier.md +15 -33
- package/dist/host/claude/rules/documentation.md +8 -2
- package/dist/host/claude/rules/portability.md +7 -1
- package/dist/host/claude/rules/quality.md +36 -96
- package/dist/host/claude/rules/styles.md +3 -0
- package/dist/host/claude/rules/tests.md +6 -3
- package/dist/host/claude/rules/workspace.md +19 -15
- package/dist/host/claude/rules/writing.md +57 -108
- package/dist/host/claude/settings.json +5 -3
- package/dist/host/claude/skills/enterprise-bootstrap/SKILL.md +1 -1
- package/dist/host/claude/skills/{orkestrel-align-packages → orkestrel-align}/SKILL.md +2 -2
- package/dist/host/claude/skills/{orkestrel-build-application → orkestrel-build}/SKILL.md +2 -2
- package/dist/host/claude/skills/orkestrel-dispatch/SKILL.md +11 -0
- package/dist/host/claude/skills/orkestrel-falsify/SKILL.md +2 -1
- package/dist/host/claude/skills/{orkestrel-harden-package → orkestrel-harden}/SKILL.md +2 -2
- package/dist/host/claude/skills/{orkestrel-prove-journey → orkestrel-journey}/SKILL.md +2 -2
- package/dist/host/claude/skills/orkestrel-polish/SKILL.md +12 -0
- package/dist/host/claude/skills/orkestrel-scout/SKILL.md +11 -0
- package/dist/host/codex/agents/analyst.toml +15 -32
- package/dist/host/codex/agents/astra.toml +25 -0
- package/dist/host/codex/agents/builder.toml +13 -20
- package/dist/host/codex/agents/checker.toml +13 -27
- package/dist/host/codex/agents/distiller.toml +9 -22
- package/dist/host/codex/agents/grok.toml +11 -30
- package/dist/host/codex/agents/opus.toml +14 -22
- package/dist/host/codex/agents/orkestrel.toml +1 -1
- package/dist/host/codex/agents/planner.toml +11 -28
- package/dist/host/codex/agents/researcher.toml +10 -22
- package/dist/host/codex/agents/reviewer.toml +11 -27
- package/dist/host/codex/agents/scout.toml +11 -17
- package/dist/host/codex/agents/verifier.toml +16 -12
- package/dist/host/codex/config.toml +20 -23
- package/dist/host/cursor/mcp.json +0 -4
- package/dist/host/cursor/rules/orchestration.mdc +12 -20
- package/dist/host/dotfiles/mcp.json +0 -4
- package/dist/host/dotfiles/oxlintrc.json +7 -0
- package/dist/host/guides/probe.md +9 -9
- package/dist/host/guides/scaffold.md +117 -71
- package/dist/host/guides/test.md +1 -1
- package/dist/host/manifest.json +322 -185
- package/dist/host/scripts/codex.sh +2 -2
- package/dist/host/tests/config.test.ts +68 -46
- package/dist/host/tests/policy.test.ts +1 -5
- package/dist/host/tests/setupPolicy.ts +179 -4
- package/dist/src/core/index.cjs +255 -84
- package/dist/src/core/index.cjs.map +1 -1
- package/dist/src/core/index.d.cts +94 -29
- package/dist/src/core/index.d.ts +94 -29
- package/dist/src/core/index.js +253 -85
- package/dist/src/core/index.js.map +1 -1
- package/dist/src/server/index.cjs +55 -9
- package/dist/src/server/index.cjs.map +1 -1
- package/dist/src/server/index.d.cts +29 -4
- package/dist/src/server/index.d.ts +29 -4
- package/dist/src/server/index.js +56 -11
- package/dist/src/server/index.js.map +1 -1
- package/package.json +15 -11
- package/dist/host/CLAUDE.md +0 -61
- package/dist/host/agents/skills/orkestrel-prove-journey/agents/openai.yaml +0 -4
- package/dist/host/agents/transports/claude.md +0 -49
- package/dist/host/claude/agents/application.md +0 -36
- package/dist/host/claude/agents/sol.md +0 -61
- package/dist/host/claude/skills/orkestrel-polish-surface/SKILL.md +0 -12
- package/dist/host/codex/agents/application.toml +0 -25
- package/dist/host/codex/agents/sol.toml +0 -19
- /package/dist/host/agents/skills/{orkestrel-align-packages → orkestrel-align}/references/integration.md +0 -0
- /package/dist/host/agents/skills/{orkestrel-harden-package → orkestrel-harden}/references/centralization.md +0 -0
- /package/dist/host/agents/skills/{orkestrel-harden-package → orkestrel-harden}/references/contract.md +0 -0
- /package/dist/host/agents/skills/{orkestrel-harden-package → orkestrel-harden}/references/research.md +0 -0
- /package/dist/host/agents/skills/{orkestrel-prove-journey → orkestrel-journey}/references/decide.md +0 -0
- /package/dist/host/agents/skills/{orkestrel-prove-journey → orkestrel-journey}/references/layer.md +0 -0
- /package/dist/host/agents/skills/{orkestrel-prove-journey → orkestrel-journey}/references/statechart.md +0 -0
- /package/dist/host/agents/skills/{orkestrel-prove-journey → orkestrel-journey}/references/styles.md +0 -0
|
@@ -1,107 +1,23 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: reviewer
|
|
3
|
-
description: '
|
|
3
|
+
description: 'Opus 5.5 review lane over implemented work, read-only: design fit, API and vocabulary, architecture shape, guide voice by default; correctness, constraints, and test sufficiency when the dispatch assigns the objective lane. Attacks numbered claims and returns the falsify verdict. Never edits.'
|
|
4
4
|
tools: Read, Grep, Glob
|
|
5
5
|
model: opus
|
|
6
6
|
effort: high
|
|
7
7
|
permissionMode: dontAsk
|
|
8
8
|
---
|
|
9
9
|
|
|
10
|
-
You
|
|
11
|
-
role set. You are independent of the builder: their self-assessment carries no
|
|
12
|
-
weight with you. You are an Executor: do the audit yourself, spawn nothing.
|
|
10
|
+
You review by trying to break the claims. You hold no edit tool and run no command; your final message is the verdict.
|
|
13
11
|
|
|
14
|
-
|
|
15
|
-
dispatch contract.
|
|
12
|
+
## Do
|
|
16
13
|
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
14
|
+
1. Read the claims file, the actual diff and `git status --porcelain` the dispatch supplies, `AGENTS.md`, the rules whose `paths` match the changed files, the guide, and enough surrounding source to judge. Return a deviation if the diff is missing.
|
|
15
|
+
2. Hold the lane the dispatch names and say which one. Subjective: the requested shape and voice are present, names and boundaries are coherent, the work sits at the right layer, each concept earns its place, the guide matches the code. Objective: correctness under adverse orderings, what the contracts and installed declarations permit, dependency and range truth, the missing seam or the assertion that cannot fail, the letter of the rules.
|
|
16
|
+
3. Attack each claim with a concrete input, state, or interleaving. Before confirming a claim about a proof, name the mutation that would make the proof fail and whether the assertions distinguish it.
|
|
17
|
+
4. Rule a claim whose only evidence is the writer's report `UNRESOLVED`. Cite a capture for every claim about a rendered surface; mark what the portfolio cannot show `NOT-EVIDENCED`.
|
|
18
|
+
5. Refer a question outside your lane to the other lane or the Orchestrator with its evidence; never rule on it.
|
|
19
|
+
6. Treat a bench's finding as a proposal to test against the code, never as authority.
|
|
22
20
|
|
|
23
|
-
##
|
|
21
|
+
## Return
|
|
24
22
|
|
|
25
|
-
|
|
26
|
-
dispatch-named skill and required references, the governing guide/spec, the actual
|
|
27
|
-
diff and status evidence supplied by the Orchestrator, and enough surrounding
|
|
28
|
-
source to judge it. If the dispatch omits the diff, return a deviation instead of
|
|
29
|
-
reconstructing it.
|
|
30
|
-
|
|
31
|
-
While you hold the subjective lane, audit the changed work through Opus 5's
|
|
32
|
-
subjective and creative lens:
|
|
33
|
-
|
|
34
|
-
1. **Design acceptance criteria** — the requested experience, shape, and voice are
|
|
35
|
-
actually present, not merely approximated.
|
|
36
|
-
2. **API and vocabulary** — names, ergonomics, conceptual boundaries, and the
|
|
37
|
-
single shared language feel deliberate and coherent.
|
|
38
|
-
3. **Architecture fit** — the work belongs at the chosen layers and abstractions,
|
|
39
|
-
composes naturally, and does not introduce awkward conceptual machinery.
|
|
40
|
-
4. **Simplification** — the design earns each concept and wrapper and leaves the
|
|
41
|
-
package plain for a human to understand.
|
|
42
|
-
5. **Guide voice and product coherence** — documentation reads as the package's
|
|
43
|
-
current, self-contained human guide and matches the experience the code presents.
|
|
44
|
-
|
|
45
|
-
While you hold the objective lane, audit the changed work through these lenses with
|
|
46
|
-
the same weight:
|
|
47
|
-
|
|
48
|
-
1. **Correctness under adverse orderings** — the adverse conditions
|
|
49
|
-
`.claude/rules/quality.md` § Falsification names.
|
|
50
|
-
2. **Constraints** — what the declared contracts, the installed declarations, and the
|
|
51
|
-
permission floor permit.
|
|
52
|
-
3. **Dependency and range truth** — the installed `@orkestrel/*` capabilities and the
|
|
53
|
-
ranges that reach a consumer.
|
|
54
|
-
4. **Test sufficiency** — the missing seam, the assertion that cannot fail, the probe
|
|
55
|
-
with no control.
|
|
56
|
-
5. **Mechanical conformance** — the letter of `AGENTS.md` and the applicable rules.
|
|
57
|
-
|
|
58
|
-
Test a design claim by asking whether the shipped artifact still matches it — a
|
|
59
|
-
guide, charter, or name that described the work two revisions ago is drift, and
|
|
60
|
-
that question is what finds it. Anything you cannot settle within your lane becomes
|
|
61
|
-
a referral — to the other lane when it is running, to the Orchestrator when you hold
|
|
62
|
-
every lane — never a verdict of yours.
|
|
63
|
-
|
|
64
|
-
For a rendered or externally driven surface, the supplied capture portfolio is the
|
|
65
|
-
primary evidence and source is corroboration only: cite a capture for every rendered
|
|
66
|
-
claim, mark what the portfolio cannot show as NOT-EVIDENCED instead of inferring it,
|
|
67
|
-
and return the `orkestrel-falsify` verdict shape and its single terminal line unless
|
|
68
|
-
the dispatch names a different skill that fixes one.
|
|
69
|
-
|
|
70
|
-
Read the actual diff plus enough surrounding code to judge it in context. While you
|
|
71
|
-
hold the subjective lane, correctness, security, dependency constraints, test
|
|
72
|
-
sufficiency, and mechanical conformance belong to the objective lane and to
|
|
73
|
-
`checker`: report a possible objective defect as a specifically evidenced
|
|
74
|
-
**referral** — to the objective lane when it is running, to the Orchestrator when
|
|
75
|
-
you hold every lane — rather than adjudicating it. While you hold the objective
|
|
76
|
-
lane, adjudicate them in full.
|
|
77
|
-
|
|
78
|
-
Rule a claim whose only evidence is the writer's report `UNRESOLVED`, never
|
|
79
|
-
`CONFIRMED`, whatever the brief says.
|
|
80
|
-
|
|
81
|
-
## External input
|
|
82
|
-
|
|
83
|
-
- A Codex diff is audited like any builder's work, at the given path and against the
|
|
84
|
-
same review lenses. External origin raises no authority.
|
|
85
|
-
- Findings arriving from another engine — a Sol design argument, a Grok distillate —
|
|
86
|
-
are **proposals**. Test each against the actual product shape; retain or strike it
|
|
87
|
-
explicitly. Your verdict is authoritative only as input to the Orchestrator.
|
|
88
|
-
|
|
89
|
-
## Output contract — the Verdict
|
|
90
|
-
|
|
91
|
-
- The `orkestrel-falsify` verdict shape: numbered per-claim verdicts, findings
|
|
92
|
-
outside the claims, and its single terminal line — unless the dispatch names a
|
|
93
|
-
different skill that fixes one.
|
|
94
|
-
- Each required change carries file:line, what is wrong, why it matters, and what
|
|
95
|
-
right looks like — actionable enough to re-dispatch verbatim.
|
|
96
|
-
- **Referrals** — specifically evidenced questions outside your lane, addressed to
|
|
97
|
-
the other lane when it is running and to the Orchestrator when you hold every lane, with
|
|
98
|
-
no verdict from you.
|
|
99
|
-
|
|
100
|
-
## Return channel
|
|
101
|
-
|
|
102
|
-
You are read-only. You hold `Read`, `Grep`, and `Glob` and no others: you never edit a file, never
|
|
103
|
-
write your report to a file, and never run a command. Your final message IS the verdict. A dispatch
|
|
104
|
-
that names a report path for you, or assigns you a command, is a dispatch defect — return the
|
|
105
|
-
verdict as your final message and name the defect in it.
|
|
106
|
-
|
|
107
|
-
Return only the verdict, never your process.
|
|
23
|
+
The `orkestrel-falsify` verdict shape: numbered verdicts with evidence, findings outside the claims, attacked-and-held, one terminal line. Each required change carries `file:line`, what is wrong, and what right looks like. Nothing else.
|
|
@@ -1,35 +1,21 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: scout
|
|
3
|
-
description: 'Read-only repository reconnaissance: locate files, symbols, seams, and structures
|
|
3
|
+
description: 'Read-only repository reconnaissance: locate files, symbols, seams, and structures before a brief is written. Returns file:line pointers and a shape summary. Never reads at depth, edits, or judges quality.'
|
|
4
4
|
tools: Read, Grep, Glob
|
|
5
5
|
model: sonnet
|
|
6
6
|
effort: low
|
|
7
7
|
permissionMode: dontAsk
|
|
8
|
+
omitClaudeMd: true
|
|
8
9
|
---
|
|
9
10
|
|
|
10
|
-
You
|
|
11
|
-
role set. You answer "where does X live, what shape is it, what touches it" so the
|
|
12
|
-
Orchestrator can write a precise dispatch. You are an Executor: spawn nothing.
|
|
11
|
+
You locate. You do not read at depth, edit, or judge.
|
|
13
12
|
|
|
14
|
-
|
|
15
|
-
dispatch contract.
|
|
13
|
+
## Do
|
|
16
14
|
|
|
17
|
-
|
|
15
|
+
1. Take one bounded question: what to find and where to stop.
|
|
16
|
+
2. Read the map the dispatch supplies (the Orchestrator runs the `orkestrel-scout` skill's `map.ts` and names its path); then search by name, symbol, export, and call site for what the map leaves open. Open a file only far enough to confirm a match. Refuse a dispatch that names no map and asks for one.
|
|
17
|
+
3. Return every hit as `file:line` with a one-line shape note, grouped by the question's parts, and name the search patterns and roots you used so the coverage is checkable.
|
|
18
18
|
|
|
19
|
-
|
|
20
|
-
answer. This charter restates nothing they own.
|
|
21
|
-
- Locate, do not absorb: read excerpts sufficient to identify a seam, an owner,
|
|
22
|
-
or a shape. Deep reading and synthesis belong to the `grok` bench, and quality
|
|
23
|
-
judgment belongs to the review roles. If the question needs either, say so
|
|
24
|
-
instead of drifting into it.
|
|
25
|
-
- Return pointers, not prose: `file:line` for every claim, the minimal shape
|
|
26
|
-
summary the question needs, and an explicit list of places searched that came
|
|
27
|
-
up empty — an absence claim is only as good as its named search.
|
|
28
|
-
- Never speculate past the evidence.
|
|
19
|
+
## Return
|
|
29
20
|
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
You are read-only. You hold `Read`, `Grep`, and `Glob` and no others: you never edit a file, never
|
|
33
|
-
write your report to a file, and never run a command. Your final message IS the answer. A dispatch
|
|
34
|
-
that names a report path for you, or assigns you a command, is a dispatch defect — return the
|
|
35
|
-
answer as your final message and name the defect in it.
|
|
21
|
+
`Question`, `Hits` (cited), `Shape` (under ten lines), `Not found` (patterns that returned nothing). Nothing else.
|
|
@@ -1,47 +1,29 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: verifier
|
|
3
|
-
description: 'Runs the exact
|
|
3
|
+
description: 'Runs the exact gates or evidence commands the dispatch names, scoped first, and reports exit-code truth with exact failure excerpts. Independent of every writer; never fixes.'
|
|
4
4
|
tools: Read, Grep, Glob, Bash
|
|
5
5
|
model: sonnet
|
|
6
6
|
effort: low
|
|
7
7
|
permissionMode: default
|
|
8
|
+
omitClaudeMd: true
|
|
8
9
|
---
|
|
9
10
|
|
|
10
|
-
You
|
|
11
|
-
No builder's self-report counts as gate evidence. You are an Executor: run the
|
|
12
|
-
gates yourself, spawn nothing.
|
|
11
|
+
You run gates and report their true result. You never edit a file and never fix a failure.
|
|
13
12
|
|
|
14
|
-
|
|
15
|
-
dispatch contract.
|
|
13
|
+
## Do
|
|
16
14
|
|
|
17
|
-
|
|
15
|
+
1. Run exactly the commands the dispatch names, in order. Default scoped sweep: the `check:` script and `test:` script of each project the dispatch names. Default tree-wide sweep, only when the dispatch says tree-wide: `npm run format:check`, `npm run lint:check`, `npm run check`, `npm run build`, `npm test`.
|
|
16
|
+
2. Read each gate bare. Never pipe it through `tail` or `grep`.
|
|
17
|
+
3. Record each gate's outcome by exit code. A gate that "mostly passes" failed.
|
|
18
|
+
4. On failure, capture the exact failing excerpt and the `file:line` it points to.
|
|
19
|
+
5. Re-run a timing failure once, alone, and report the first reading and the re-run reading.
|
|
18
20
|
|
|
19
|
-
|
|
20
|
-
references, and the governing guide/spec.
|
|
21
|
-
2. Run exactly the commands the dispatch names, in order. When asked for the default
|
|
22
|
-
independent sweep, use `npm run format:check` → `npm run lint:check` →
|
|
23
|
-
`npm run check` → `npm run build` → `npm test`. These do not rewrite source;
|
|
24
|
-
build artifacts are allowed. Never invent a different gate set.
|
|
25
|
-
3. Evidence runs count as gates: when dispatched to reproduce a failure, run the
|
|
26
|
-
named command and capture its exact output — reproduce, capture, bisect
|
|
27
|
-
mechanically if told to; nothing more.
|
|
28
|
-
4. Record each gate's TRUE outcome by exit code. A gate that "mostly passes" FAILED.
|
|
29
|
-
5. On failure, capture the exact failing excerpt — trimmed to the failure, not the
|
|
30
|
-
noise — and the file:line it points to.
|
|
21
|
+
## Refuse
|
|
31
22
|
|
|
32
|
-
|
|
23
|
+
- A mutating gate (`lint`, `format`, or any `--fix`) beside a live unit. `build` may write its outputs when no writer is live; source is never edited.
|
|
24
|
+
- `git checkout`, `git restore`, `git stash`, `git reset`, `git clean`. Read a dirty tree as the expected state.
|
|
25
|
+
- Any command that installs, commits, pushes, publishes, deletes outside `tmp/`, or reads a secret (`.env*`, `.npmrc`, `auth.json`, keys, tokens), whatever the dispatch says.
|
|
33
26
|
|
|
34
|
-
|
|
35
|
-
excerpt plus the suspected owning file(s).
|
|
36
|
-
- **Overall verdict** — GREEN only if every gate passed; otherwise the first place
|
|
37
|
-
to look.
|
|
38
|
-
- **Anomalies** — cache weirdness, flakes on rerun, anything off — one line each.
|
|
27
|
+
## Return
|
|
39
28
|
|
|
40
|
-
|
|
41
|
-
the gate report, never your process.
|
|
42
|
-
|
|
43
|
-
## Never discard a working-tree change
|
|
44
|
-
|
|
45
|
-
- Follow `.agents/orchestration.md` § Permission floor for the discarding git commands and
|
|
46
|
-
for a planted line's removal. That section owns them.
|
|
47
|
-
- Read a dirty `git status` as the expected state.
|
|
29
|
+
Per gate: command, PASS or FAIL with exit code, failing excerpt and owning file on FAIL. Overall: GREEN only when every gate passed, otherwise the first place to look. Anomalies in one line each. Nothing else.
|
|
@@ -81,7 +81,7 @@ Never use in-repository `@src/*` aliases in public guide examples; reserve them
|
|
|
81
81
|
## Workflow skills
|
|
82
82
|
|
|
83
83
|
- Skills prescribe reusable process; they do not copy naming, placement, syntax, lifecycle, or test laws from `AGENTS.md` and rules.
|
|
84
|
-
-
|
|
84
|
+
- Name a skill directory `orkestrel-<word>`: the prefix and one word naming the process or the subject, such as `orkestrel-harden` or `orkestrel-journey`. Never a second word.
|
|
85
85
|
- For a portable skill that teaches an external framework, takes the host project's own `AGENTS.md` file as its code law, and binds none of this repository's rule files, drop the `orkestrel-` prefix and name the skill for the framework a reader searches for. Add no other exception. The `enterprise-bootstrap` skill meets that test and keeps its name: it carries Bootstrap 5.3 craft, assumes no stack, and names the `.agents/orchestration.md` and `.claude/rules/quality.md` files only where they are present. Do not rename it into the namespace.
|
|
86
86
|
- Keep `SKILL.md` concise and route conditional detail to one-level `references/`.
|
|
87
87
|
- Name every Markdown file in a skill's `references/` from its `SKILL.md`, and delete a reference nothing names.
|
|
@@ -89,7 +89,13 @@ Never use in-repository `@src/*` aliases in public guide examples; reserve them
|
|
|
89
89
|
- Set `name` to the skill's own directory name.
|
|
90
90
|
- Write `description` as a single-line scalar or a folded `>-` block, and no other shape. Include in it a sentence beginning `Use ` that names when to invoke the skill.
|
|
91
91
|
- Do not put model routing or package version catalogs in a skill.
|
|
92
|
-
- Validate every referenced resource, leave no template TODOs, and limit each skill directory to `SKILL.md`, `agents/openai.yaml`,
|
|
92
|
+
- Validate every referenced resource, leave no template TODOs, and limit each skill directory to `SKILL.md`, `agents/openai.yaml`, the `references/*.md` files its `SKILL.md` names, and the `scripts/*.ts` files its `SKILL.md` names; add no other file or directory.
|
|
93
|
+
- Put a repeatable or idempotent step of a skill in `scripts/<name>.ts`, run with `node .agents/skills/<skill>/scripts/<name>.ts` from the checkout root under Node's type stripping, and name the script, its arguments, and its exit codes from `SKILL.md` so an executor runs it instead of re-deriving the step. Name another skill's script by its full path under `.agents/skills/<skill>/scripts/`; the policy sweep attributes a bare `scripts/<name>.ts` token to the skill whose `SKILL.md` carries it and requires a qualified path to exist.
|
|
94
|
+
- A skill script obeys `AGENTS.md`, `.claude/rules/typescript.md`, `.claude/rules/names.md`, and `.claude/rules/portability.md` as written, plus the no-nested-function and wrapper laws in `.claude/rules/architecture.md`, with these adaptations: an executable script exports nothing, runs `main` at the bottom, and carries the usage line and the exit codes in its opening comment; the types, constants, and helpers it alone needs live in the file, because the kind-file placement law stops at a self-contained script; what two scripts share lives in `.agents/skills/orkestrel-dispatch/scripts/helpers.ts`, the one shared module, which exports declarations with TSDoc, runs no entry point, and carries no usage line; a script imports `node:` modules and that module and nothing else, so the `@orkestrel/process` spawn rule yields to `child_process` there (a target may not declare that package), and imports it with the `.ts` extension, because the checkout runs a script unbuilt; and it spawns `process.execPath` or an executable by name with an argument array and never a shell, except the npm shim where no `npm-cli.js` sits beside the binary. Every script has a mirrored proof at `tests/agents/skills/<skill>/scripts/<name>.test.ts`: an executable script's proof drives it as a child process against a scratch fixture, and the shared module's proof imports it and tests each export. The policy sweep refuses a script without a proof, and a proof grows a case for every defect found.
|
|
95
|
+
- Run a script in a target through its built twin, `node node_modules/@orkestrel/scaffold/dist/agents/skills/<skill>/scripts/<name>.js`, because Node refuses to strip types for a file under `node_modules`; `.claude/rules/workspace.md` names the `build:skills` script that emits the twins. When a script spawns a sibling script, take the extension from the running file (`extname(fileURLToPath(import.meta.url))`), so the twin spawns `.js` and the checkout spawns `.ts`. When a script resolves a file under `.agents/templates/`, add that file to the copy in `build:skills`.
|
|
96
|
+
- Resolve a sibling canon file (a template, another script) from the script's own location with `import.meta.url`, and the manifest, `tmp/`, and the test tree from the working directory, because a target runs the script from `node_modules/@orkestrel/scaffold/dist/agents/` against its own checkout.
|
|
97
|
+
- Read arguments as `--flag value` pairs through the shared helpers: the first occurrence of a flag wins, a value that opens with `--` is a missing value, and `--flag=value` is not read. Where flags select different work (`--visit` or `--plan`), the script names one operation mode per run and refuses two; where each flag adds a section to one report (`--tree` with `--headings`), the script combines them and its usage line says so. Read a flag that repeats (`--range`) through `readOptions` and say in the usage line that it repeats. Exit 64 on every usage refusal and say on stderr what was refused: the usage line when no mode or two modes were given, and the flag with the value it takes when a flag was malformed or given with no value.
|
|
98
|
+
- Print to the terminal by default. Offer `--json` for a reader that parses, and `--out FILE` (`.txt` for a list, `.json` for a structure) where a lane reads the result later or the output would crowd a context; a written file saves the tokens a terminal dump spends. Choose per script; state each form in its usage line.
|
|
93
99
|
- Verify each API a skill instructs an executor to call against the installed package's public entry before landing the instruction, and name the entry you read. A skill that names a symbol its package does not export teaches an executor to write a dangling import.
|
|
94
100
|
- Put every symbol a skill teaches in a named import inside a Markdown fence in `SKILL.md` or a named reference. The policy sweep reads fenced value and type imports from `@orkestrel/*` declaration entries; it does not read identifiers in prose or table cells, or validate call signatures and runtime behavior.
|
|
95
101
|
- Import only packages in `BASE_DEV_DEPENDENCIES` in those fences. The sweep refuses a package outside that set because a generated workspace need not install it.
|
|
@@ -5,6 +5,7 @@ paths:
|
|
|
5
5
|
- 'configs/**/*'
|
|
6
6
|
- 'tests/**/*'
|
|
7
7
|
- 'scripts/**/*'
|
|
8
|
+
- '.agents/skills/*/scripts/*.ts'
|
|
8
9
|
- 'guides/**/*'
|
|
9
10
|
- 'package.json'
|
|
10
11
|
- '.gitattributes'
|
|
@@ -60,6 +61,10 @@ and the form of a conditional skip.
|
|
|
60
61
|
- Resolve, spawn, and terminate a child through `@orkestrel/process` where the package declares it.
|
|
61
62
|
- Where it is not declared, spawn `process.execPath` with a JavaScript entry. Never spawn a `.bin`
|
|
62
63
|
shim, and never add `shell: true` to reach one.
|
|
64
|
+
- Merge a child environment by case-folded name: before spawning, drop every inherited variable
|
|
65
|
+
whose lower-cased name an added variable claims. A Windows environment block folds names by case,
|
|
66
|
+
so an inherited `NPM_CONFIG_CACHE` outranks an added `npm_config_cache`, and npm reads its
|
|
67
|
+
`npm_config_` variables case-insensitively on every host.
|
|
63
68
|
- Take a resolver's first match only after splitting its output into lines and trimming each one.
|
|
64
69
|
`where` prints a match per line and can name a file that is not an executable.
|
|
65
70
|
- Treat `fs.constants.X_OK` as an existence check on Windows, where `accessSync` passes on a plain
|
|
@@ -89,7 +94,8 @@ and the form of a conditional skip.
|
|
|
89
94
|
- Write every `package.json` script as a portable command: a Node invocation or an installed binary.
|
|
90
95
|
Never name a `.sh` file there.
|
|
91
96
|
- Keep `#!/usr/bin/env node` on an npm bin. npm writes the Windows shim from it.
|
|
92
|
-
- Keep
|
|
97
|
+
- Keep `scripts/` for the Claude Code Cloud session hooks alone, in bash, with `.gitattributes` `eol=lf`. Every other script is TypeScript run by Node per `AGENTS.md`; a skill's scripts live in its `scripts/` directory.
|
|
98
|
+
- Spawn a command from a script with an argument array and no shell, so quoting is the same on every host; resolve a `.cmd` shim to the JavaScript entry it wraps rather than spawning the shim.
|
|
93
99
|
|
|
94
100
|
## Claims
|
|
95
101
|
|
|
@@ -4,117 +4,57 @@ paths:
|
|
|
4
4
|
- 'app/**/*'
|
|
5
5
|
- 'tests/**/*'
|
|
6
6
|
- 'guides/**/*'
|
|
7
|
-
- 'package.json'
|
|
8
|
-
- 'vite.config.ts'
|
|
9
|
-
- 'tsconfig.json'
|
|
10
|
-
- '.agents/skills/**/*'
|
|
11
|
-
- '.claude/skills/**/*'
|
|
12
7
|
---
|
|
13
8
|
|
|
14
|
-
#
|
|
9
|
+
# Evidence, probes, and completion rules
|
|
15
10
|
|
|
16
11
|
## Evidence before change
|
|
17
12
|
|
|
18
|
-
-
|
|
19
|
-
- Use current primary sources for external capabilities, and the exact installed declarations and guides for dependencies. Separate verified fact from inference.
|
|
20
|
-
- Read authoritative types and named decision-bearing implementation files first-hand. Delegate bulk supporting context, never the owning design decision.
|
|
13
|
+
- Read the authoritative `types.ts` and the decision-bearing implementation first-hand. Delegate bulk reading, never the owning decision.
|
|
21
14
|
- Treat existing code, tests, `old/`, branches, and copied projects as evidence, not authority.
|
|
22
|
-
-
|
|
23
|
-
-
|
|
24
|
-
- Never end a row as "hardened further." Replace any evaluative phrase with the concrete condition that closes the row.
|
|
25
|
-
- Record a finding outside the matrix against the row that owns it, for the next matrix. Do not reopen this one.
|
|
15
|
+
- Research when the user asks, when comparing an upstream or legacy implementation, or when current external behavior changes the design. Use primary sources for external capabilities and the installed declarations for dependencies. Separate verified fact from inference.
|
|
16
|
+
- For a broad API or production-readiness change, build a capability/defect matrix before editing. Every row ends as implement, repair, retain, or exclude with evidence. The matrix is the definition of done; fix it when the change starts. Record a finding outside it for the next matrix.
|
|
26
17
|
|
|
27
18
|
## Probes before arguments
|
|
28
19
|
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
-
|
|
32
|
-
-
|
|
33
|
-
-
|
|
34
|
-
-
|
|
35
|
-
-
|
|
36
|
-
-
|
|
37
|
-
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
-
|
|
42
|
-
-
|
|
43
|
-
-
|
|
44
|
-
-
|
|
45
|
-
-
|
|
46
|
-
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
A review that reads a diff finds what the diff shows. A review that tries to break named claims finds what the diff hides. Code that has passed diff review several times can still carry a defect nobody has tried to trigger.
|
|
51
|
-
|
|
52
|
-
### Writing the claims
|
|
53
|
-
|
|
54
|
-
- State the subject as a numbered list of the claims the work makes, never as "review this diff." Each claim is falsifiable: a property some concrete input, state, or interleaving could show false.
|
|
55
|
-
- Instruct the auditor to attempt refutation, not confirmation. A claim it cannot break is CONFIRMED with the evidence that convinced it. A claim it breaks is BROKEN with the exact failing input, state, or interleaving, plus the smallest correct fix.
|
|
56
|
-
- Derive claims from what the change asserts under adverse conditions: cancellation, restart, concurrency, partial failure, hostile input, resource exhaustion, and the orderings a happy path never reaches.
|
|
57
|
-
- Read the installed declaration or implementation of every substrate a claim depends on. A claim about `stop()` is unfalsifiable until you know what `stop()` does when the status is not the one the caller assumed.
|
|
58
|
-
|
|
59
|
-
### Instruments
|
|
60
|
-
|
|
61
|
-
- An instrument is not evidence until it has failed. Pair every probe, comparison, or matrix with a negative control that must report failure, run under the same conditions. An identity check whose control reports "same" has measured nothing.
|
|
62
|
-
- When a question about a TypeScript edit can supply a workspace project, a case of workspace files with a test, and a negative control naming its files, its test, the stage it must fail at, and why, call the `prove` tool the `probe` MCP server registers before relying on the answer. When the question supplies no project, no case, or no control, follow `.claude/rules/tests.md` § Probes and report the fallback instrument's own control and coverage.
|
|
63
|
-
- When no `probe` server is registered in the session, register one outside the repository, in the harness's own local or user MCP scope — in Claude Code, `claude mcp add` outside project scope — naming the installed `node_modules/@orkestrel/probe/dist/bin/main.js` entry, and start it in the repository whose projects the question names, because the server fixes its workspace from its own working directory. The registration cannot live in the tree: a scaffold target holds no `.mcp.json` file, that path is instruction canon, and a copy at it reports as foreign drift on every `scaffold audit` run. Where the harness registers no server at all, treat the question as supplying no project and take the preceding rule's fallback.
|
|
64
|
-
- Quote the closing line of the `prove` answer verbatim in every report, brief, and audit verdict that rests on the claim: the `receipt probe:<digest>:…` line when the case ran clean and the control broke exactly where the claim declared it would, and the `no receipt` line otherwise. A `no receipt` line leaves the claim unproved — report it with the stage that refused.
|
|
65
|
-
- Read a receipt as evidence about its claim, never as a gate result. The gate chain still runs, and `verifier` still owns its result.
|
|
66
|
-
- Draw the negative control from outside the population the instrument covers. Name the instrument's membership rule first, then pick a control that rule excludes. A control sampled from constructs the instrument already handles proves only that it discriminates among those constructs, and says nothing about the class it silently cannot reach.
|
|
67
|
-
- State an instrument's coverage beside its result. A conclusion inherits the instrument's scope, not the question's. An unstated coverage claim is read as complete, and it never is. A search proves something about the paths it walked, so name them.
|
|
68
|
-
- Match the instrument to the question. A text search reports on text, so a claim about declarations, call sites, or structure needs the compiler or a parser instead. A pattern written for one spelling of a construct reports on that spelling alone. A path check answers relative to the directory it runs from, so resolve the inputs against their own base before reading a miss as a finding.
|
|
69
|
-
- Name the rival reading the instrument must exclude, and show it reports differently under that reading. Give independent measurements independent state: one counter shared across members reports read order and per-member read count identically, so a result consistent with both measured neither.
|
|
70
|
-
- Report a question unanswered rather than answering it with a weaker instrument. A fallback that measures something adjacent returns a confident wrong answer, and nothing downstream can tell that answer from the real one — searching commit messages for a release when the question is where a version changed will match some release, but not the one asked about. Name the substitute and what it actually measures, or say the question is open.
|
|
71
|
-
- State what the controls established and what they did not. An instrument certified only from the inside is trusted exactly where it has never been tested.
|
|
72
|
-
- Treat a gap between what an instrument says it checks and what it actually matches as a defect in the instrument, not as a documented limit. A recorded blind spot buys trust only when everything outside it is genuinely covered.
|
|
73
|
-
- Measure the product, not the harness. A recorded baseline that counts something about its own fixture is not evidence about the shipped surface, however often a guide quotes it.
|
|
74
|
-
- Baseline a published-artifact claim against the published artifact. "Did my change move the surface" and "does this release differ from the last one" are different questions, and a diff against your own starting point answers only the first. A toolchain that re-emits declarations moves the artifact without any source edit, so every writer can correctly report an unmoved surface while the package's published contract has changed. Fetch what consumers actually have — the tarball, the deployed asset — and compare against that.
|
|
75
|
-
- Prove a module cycle by loading the built artifact, not by a green suite. Tests import through the source graph and a bundler resolves it differently, so a cycle that is fatal at module-init in the shipped form can stay invisible under every test. Import each published entry point and read an export from it.
|
|
76
|
-
- Adopt an instrument that settled a claim as a test before accepting the work it settled. The probe that proved a fix, carrying the control that proved the probe, is that fix's regression guard. A verification that runs once is a rehearsal, not a gate.
|
|
77
|
-
|
|
78
|
-
### Rounds and verdicts
|
|
79
|
-
|
|
80
|
-
- Treat an all-confirmed round as a legitimate result, and put the brief on trial rather than the subject. Re-read the claims and ask whether any could have been falsified by evidence the round actually had. If none could, the claims were descriptive and the round proved nothing — sharpen them and re-run. If they could have been and were not, the pass stands.
|
|
81
|
-
- Name the claims you could not break either way, so the next round knows what has already been attacked.
|
|
82
|
-
- Never tell an auditor that a clean round means it did not try. That instructs it to manufacture a finding, and a manufactured finding costs a fix unit, an argument, and the credibility of the true findings beside it.
|
|
83
|
-
- Treat a repaired claim as a new claim, not a settled one. Re-ask it at every entry point that reaches the same rule, not only the door the defect arrived through. The engine that wrote the fix is least able to see this, because re-verifying where the fix is feels like verifying the fix.
|
|
84
|
-
- A fix that adopts the auditor's prescription verbatim may close with a mutation probe in place of a fresh audit round: disable the load-bearing line, watch the adopted pin fail, restore it, and commit the pin as the regression guard. A fix that departs from the prescription gets the cross-engine round.
|
|
85
|
-
- Let reachability bound the fix. A defect reachable through the package's own shipped code or a documented extension seam falsifies its claim and is repaired now.
|
|
86
|
-
- Document the obligation instead when a defect is reachable only through a hypothetical foreign implementation of a contract this package publishes. State it on the interface that owns it and prove the documentation. Do not build coordination machinery against a requirement nobody wrote down. Attacks are unlimited; reachable ones are not, and only the reachable set is a work list.
|
|
87
|
-
- **Three rounds at one seam is the budget, and reaching it switches the search strategy rather than stopping the work.** Repeated rounds against one seam are evidence about the design, not evidence of diligence. At the third round, name what the audit is trying to accomplish, then pick the successor strategy from what the recurrence shows.
|
|
88
|
-
- Recurrence with a direction — each fix relocates the class along one stream, or the evidence points up a dependency chain — ends the depth search. Run one breadth round that probes the stream's stations in parallel to locate the source, in the shape `.agents/orchestration.md` § Context and decomposition prescribes, then plan the downstream work from the source and size it by the sweep's map.
|
|
89
|
-
- Recurrence with no direction makes the seam itself the question. The next unit is a ruling — on the threat model, the mechanism, or the boundary — taken with the same adversarial pass a design gets, not a fourth repair.
|
|
90
|
-
- A subject that reprices itself on every edit — a count, a census, a total over prose — has no closing condition and is not a seam. Drop the claim, or recast it as the property the tally stood in for.
|
|
91
|
-
- Write the round count down in the capability/defect matrix row that owns the seam, when the seam opens, so it is a fact rather than a feeling. A seam that has consumed more rounds than the rest of the matrix combined has already answered the question.
|
|
92
|
-
- State the ruling that ends a seam as the invariant the code will obey, the constraint bounding it against over-correction, and the interface where a consumer meets the obligation. A ruling that names only the defect it replaces produces the opposite defect next round.
|
|
93
|
-
- Give every behavioural audit the means to run its attacks. An auditor that cannot execute cannot falsify a behavioural claim: it returns derivations, and a derivation reads exactly like a verdict while being a different thing — it will confirm a claim one probe would break. Treat a report with no executed evidence as a review of the source, and label it as such.
|
|
20
|
+
- Settle a question about behavior by running it: what a function returns, what a config resolves to, whether a path is reached. Run before stating the belief.
|
|
21
|
+
- When the argument about a behavior grows past a few sentences, stop and run it.
|
|
22
|
+
- Label an unverified assertion as unverified. An unverified claim in a brief becomes the premise of every downstream unit.
|
|
23
|
+
- Bound a search before starting it: the benchmark, the population, or the row that ends it.
|
|
24
|
+
- A passing negative probe proves nothing until its input is shown to reach the code under test.
|
|
25
|
+
- A clean reproduction of a reported defect means your vector was weaker than the reporter's; get theirs before ruling. A vector the compiler rejects refutes the vector, not the finding.
|
|
26
|
+
- "No tests found", an empty match, or a runner that resolved nothing reports on the harness. Confirm the probe was collected before reading a result.
|
|
27
|
+
- Prefer an observation to a derivation. When a measurement and an argument disagree, the argument is wrong until the measurement is shown broken.
|
|
28
|
+
- Read a gate bare. A `| tail` or `| grep` behind it hides the failing lines and reports the pipe's exit code.
|
|
29
|
+
|
|
30
|
+
## Instruments
|
|
31
|
+
|
|
32
|
+
- An instrument counts as evidence only after it has failed: pair every probe, comparison, or matrix with a negative control that must fail under the same conditions, drawn from outside the population the instrument covers.
|
|
33
|
+
- When a TypeScript question can name a workspace project, a case (files plus a test), and a control (files, test, the stage it must fail at, and why), call the `prove` tool the `probe` MCP server registers. Quote its closing line (`receipt probe:<digest>:…` or `no receipt`) where the claim is reported. A `no receipt` line leaves the claim unproved; report the stage that refused.
|
|
34
|
+
- When the question supplies no project, no case, or no control, write a probe per `.claude/rules/tests.md` § Probes and report the probe's control and coverage.
|
|
35
|
+
- When no `probe` server is registered, register one outside the repository in the harness's user or local MCP scope, pointing at `node_modules/@orkestrel/probe/dist/bin/main.js`, and start it in the repository whose projects the question names. A scaffold target holds no `.mcp.json`.
|
|
36
|
+
- Read a receipt as evidence about its claim, never as a gate result.
|
|
37
|
+
- State the instrument's coverage beside its result. Match the instrument to the question: a text search reports on text; a claim about declarations or call sites needs the compiler or a parser.
|
|
38
|
+
- Report a question unanswered rather than answering it with a weaker instrument.
|
|
39
|
+
- Baseline a published-artifact claim against the published artifact (the tarball, the deployed asset). Prove a module cycle by loading the built entry points.
|
|
40
|
+
- Promote an instrument that settled a claim into a test, with its control, before accepting the work it settled.
|
|
94
41
|
|
|
95
42
|
## Ecosystem reuse
|
|
96
43
|
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
-
|
|
100
|
-
- Treat downstream friction as valid evidence of a reusable upstream defect, not as automatic proof. Fix the lowest package that owns the general mechanism and keep product policy downstream.
|
|
101
|
-
- Never re-export a dependency's symbol to soften a consumer's import.
|
|
44
|
+
- Prove a semantic difference before keeping a local variant of an installed primitive.
|
|
45
|
+
- Fix a reusable defect in the lowest package that owns the mechanism; keep product policy downstream.
|
|
46
|
+
- Never re-export a dependency's symbol.
|
|
102
47
|
|
|
103
48
|
## Production hardening
|
|
104
49
|
|
|
105
|
-
- Translate "enterprise-grade" or "production-ready" into
|
|
106
|
-
- Grade
|
|
107
|
-
- Test observable invariants at each
|
|
108
|
-
-
|
|
109
|
-
- Audit test discovery, counts, skipped and todo tests, cleanup, and assertion adequacy
|
|
110
|
-
-
|
|
111
|
-
- Treat a claim that a surface works with an external client as unproven until one representative real client of that class has driven it end to end. Protocol tests prove the protocol, not the integration.
|
|
112
|
-
- Add an independent adversarial review for security, destructive paths, concurrency, protocols, or untrusted external input.
|
|
50
|
+
- Translate "enterprise-grade" or "production-ready" into a risk and seam matrix: inputs, states, failures, cleanup, cancellation, concurrency, resource ownership, hostile boundaries, environment isolation, serialization, package consumption.
|
|
51
|
+
- Grade the matrix on coverage of applicable seams, not on further interleavings against a seam already proven.
|
|
52
|
+
- Test observable invariants at each seam with real implementations; use a dedicated real-service project for external behavior.
|
|
53
|
+
- Treat "works with an external client" as unproven until one representative real client has driven it end to end.
|
|
54
|
+
- Audit test discovery, counts, skipped and todo tests, cleanup, and assertion adequacy before acceptance. Coverage is not adequacy.
|
|
55
|
+
- Add an independent review for security, destructive paths, concurrency, protocols, or untrusted input, per the size gate in `.agents/orchestration.md`.
|
|
113
56
|
|
|
114
57
|
## Completion
|
|
115
58
|
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
- Perform a final centralization, wrapper, test-helper, and text-integrity sweep after implementation and before gates.
|
|
119
|
-
- Produce local quality gates and relevant output inspection as required evidence.
|
|
120
|
-
- Stop once the enumerated scope is closed and those gates are green. Stopping there is correct, and the next scope is the deliverable. A further pass over the same surface requires a new instruction from the user, not an auditor's remaining appetite. Depth is owed to what the scope names, not to what an engine can still imagine about it.
|
|
59
|
+
- Sweep centralization, wrappers, test helpers, and text integrity after implementation and before gates.
|
|
60
|
+
- Stop when the enumerated scope is closed and the gates the size gate names are green. The next scope is the deliverable; a further pass needs a further instruction from the user.
|
|
@@ -43,6 +43,9 @@ SCSS mirrors TypeScript centralization. Concrete token prefixes are project-spec
|
|
|
43
43
|
literal color may appear in is `_tokens.scss`, where the token itself is declared.
|
|
44
44
|
- Never repeat per-color/per-variant blocks; drive shared structure with one `@each` over a shared list.
|
|
45
45
|
- If a pattern appears in at least two partials, move it to `_mixins.scss`.
|
|
46
|
+
- Treat a declaration block two partials share because each records an external value as a
|
|
47
|
+
coincidence, never a pattern: keep both copies inline, and move a block only where its callers
|
|
48
|
+
share one decision whose divergence is a defect.
|
|
46
49
|
- A one-partial pattern stays inline; do not create a mixin for one caller.
|
|
47
50
|
- Never `@extend` across partials; share through tokens/mixins.
|
|
48
51
|
- Never declare a `transition:` without `prefers-reduced-motion: reduce`. Use the project transition mixin, which emits both.
|
|
@@ -12,7 +12,7 @@ paths:
|
|
|
12
12
|
|
|
13
13
|
- Mirror module/application structure:
|
|
14
14
|
`tests/{src,app}/[environment]/[domain]/[module].test.ts`.
|
|
15
|
-
- The mirrored population is `src` and `app`
|
|
15
|
+
- The mirrored population is `src` and `app`, plus the skill scripts: `.agents/skills/<skill>/scripts/<name>.ts` is proved by `tests/agents/skills/<skill>/scripts/<name>.test.ts`, one proof per script, and the policy sweep refuses a script with no proof. `configs/` is a source directory and is
|
|
16
16
|
deliberately not a mirrored root: its leaves produce the workspace's configuration rather than ship
|
|
17
17
|
in it, and they are proved from `tests/config.test.ts` beside the configuration they produce. Do
|
|
18
18
|
not add `tests/configs/`.
|
|
@@ -54,6 +54,7 @@ its own:
|
|
|
54
54
|
| `tests/config.test.ts` | Root configuration resolves its aliases, projects, and outputs, and the `configs/` leaves behind them |
|
|
55
55
|
| `tests/guides.test.ts` | Every documented API exists, every public API is documented, every compared summary, example, and pitch equals its source, and every executable fence returns what the guide says it returns |
|
|
56
56
|
| `tests/conformance.test.ts` | Where this package drifts from the official tooling it tracks |
|
|
57
|
+
| `tests/agents/**/*.test.ts` | Each skill script does what its `SKILL.md` states, driven as a child process against a scratch fixture from its mirrored proof; the scaffold checkout alone carries them |
|
|
57
58
|
| `tests/distribution.test.ts` | The packed package installs and resolves through its public exports |
|
|
58
59
|
| `tests/integration.test.ts` | The package's features work together end to end across environments |
|
|
59
60
|
| `tests/setup*.test.ts` | Reusable behavior exported from sibling `tests/setup*.ts` modules works as the workspace's suites require |
|
|
@@ -109,7 +110,7 @@ The kinds split by which tool has to see the probe:
|
|
|
109
110
|
lives in the source tree beside what it measures. Delete it before the unit returns; a leaked one
|
|
110
111
|
fails the `policy` plugin's placement rules, because a probe filename is not a centralized kind
|
|
111
112
|
file.
|
|
112
|
-
- A **runtime probe** is collected by a Vitest project, so it lives in `tmp/
|
|
113
|
+
- A **runtime probe** is collected by a Vitest project, so it lives in `tmp/probes/` and runs through
|
|
113
114
|
the `probe` project. `tmp/` is ignored by git, so no probe enters a commit by accident, and every
|
|
114
115
|
test script names its project, so no gate runs the `probe` project.
|
|
115
116
|
- A **bench** is read by Vitest's benchmark mode, so it lives inside a test file as a block behind
|
|
@@ -321,7 +322,9 @@ Coverage rules:
|
|
|
321
322
|
|
|
322
323
|
## Discovery and adequacy audit
|
|
323
324
|
|
|
324
|
-
|
|
325
|
+
For a change the size gate in `AGENTS.md` § Work loop names medium or large, run
|
|
326
|
+
`node .agents/skills/orkestrel-harden/scripts/discovery.ts` before acceptance for the census of
|
|
327
|
+
projects, gates, collected files, and skip markers, then:
|
|
325
328
|
|
|
326
329
|
- prove every intended test file is discovered by the correct project;
|
|
327
330
|
- prove every declared project is reachable from a gate. A project registered in the root
|