@arjunkhera/atlas 0.3.7 → 0.3.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (67) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/agents/artifact-renderer.md +22 -22
  3. package/agents/verifier.md +29 -0
  4. package/door/cli.mjs +62 -12
  5. package/door/kit-releases.json +4 -0
  6. package/door/lib/design-build.mjs +409 -0
  7. package/door/lib/design.mjs +199 -118
  8. package/door/lib/markdown.mjs +160 -0
  9. package/door/lib/proof.mjs +127 -0
  10. package/door/lib/tests.mjs +190 -0
  11. package/package.json +2 -1
  12. package/skills/lead/SKILL.md +4 -1
  13. package/skills/sdlc-task/SKILL.md +33 -24
  14. package/skills/sdlc-task/design/README.md +187 -0
  15. package/skills/sdlc-task/design/parts/actors.md +26 -0
  16. package/skills/sdlc-task/design/parts/alternatives.md +24 -0
  17. package/skills/sdlc-task/design/parts/build.md +23 -0
  18. package/skills/sdlc-task/design/parts/calls.md +25 -0
  19. package/skills/sdlc-task/design/parts/change.md +27 -0
  20. package/skills/sdlc-task/design/parts/data.md +22 -0
  21. package/skills/sdlc-task/design/parts/done.md +23 -0
  22. package/skills/sdlc-task/design/parts/edges.md +24 -0
  23. package/skills/sdlc-task/design/parts/goals.md +27 -0
  24. package/skills/sdlc-task/design/parts/key.md +25 -0
  25. package/skills/sdlc-task/design/parts/migration.md +22 -0
  26. package/skills/sdlc-task/design/parts/order.md +24 -0
  27. package/skills/sdlc-task/design/parts/problem.md +22 -0
  28. package/skills/sdlc-task/design/parts/proof.md +24 -0
  29. package/skills/sdlc-task/design/parts/proposal.md +24 -0
  30. package/skills/sdlc-task/design/parts/records.md +24 -0
  31. package/skills/sdlc-task/design/parts/repos.md +25 -0
  32. package/skills/sdlc-task/design/parts/risks.md +24 -0
  33. package/skills/sdlc-task/design/parts/rollout.md +24 -0
  34. package/skills/sdlc-task/design/parts/routes.md +24 -0
  35. package/skills/sdlc-task/design/parts/scorecard.md +25 -0
  36. package/skills/sdlc-task/design/parts/security.md +22 -0
  37. package/skills/sdlc-task/design/parts/shared-decisions.md +24 -0
  38. package/skills/sdlc-task/design/parts/states.md +25 -0
  39. package/skills/sdlc-task/design/parts/stories.md +26 -0
  40. package/skills/sdlc-task/design/parts/summary.md +33 -0
  41. package/skills/sdlc-task/design/parts/why.md +22 -0
  42. package/skills/sdlc-task/design/parts/words.md +29 -0
  43. package/skills/sdlc-task/design/parts/yardstick.md +25 -0
  44. package/skills/sdlc-task/design/parts.yaml +306 -0
  45. package/skills/sdlc-task/lifecycle.yaml +2 -2
  46. package/skills/tests/SKILL.md +234 -0
  47. package/tests/contract.mjs +398 -0
  48. package/tests/drivers/function.mjs +30 -0
  49. package/tests/drivers/http.mjs +68 -0
  50. package/tests/drivers/index.mjs +65 -0
  51. package/tests/drivers/mcp-stdio.mjs +174 -0
  52. package/tests/environment.mjs +227 -0
  53. package/tests/errors.mjs +27 -0
  54. package/tests/evidence.mjs +113 -0
  55. package/tests/fresh.mjs +42 -0
  56. package/tests/guards.mjs +159 -0
  57. package/tests/index.mjs +11 -0
  58. package/tests/link-check.mjs +576 -0
  59. package/tests/procs.mjs +43 -0
  60. package/tests/redact.mjs +58 -0
  61. package/tests/scenario.mjs +325 -0
  62. package/tests/stand-in.mjs +74 -0
  63. package/tests/tests-yaml.mjs +258 -0
  64. package/tests/wait.mjs +44 -0
  65. package/tests/yaml.mjs +327 -0
  66. package/agents/artifact-format/walkthrough.html +0 -706
  67. package/skills/sdlc-task/templates/design-doc.md +0 -126
@@ -0,0 +1,22 @@
1
+ # Why one initiative
2
+
3
+ It holds: Why these designs belong together.
4
+
5
+ Needs: Decide: none, Flow: none, Contract: none, Initiative: needed.
6
+
7
+ Rules:
8
+
9
+ 1. Say why these designs belong together.
10
+ 2. Name the one goal they serve.
11
+ 3. Say what would go wrong if each went alone.
12
+
13
+ Example, in a file of its own:
14
+
15
+ ````markdown
16
+ ---
17
+ part: why
18
+ ---
19
+ # Why one initiative
20
+
21
+ Each file type needs its own reader, but all share one viewer and one queue. Alone, each would build its own queue.
22
+ ````
@@ -0,0 +1,29 @@
1
+ # Your words
2
+
3
+ It holds: The owner's words, word for word. Claude's reading stays apart.
4
+
5
+ Needs: Decide: needed, Flow: needed, Contract: needed, Initiative: needed.
6
+
7
+ Rules:
8
+
9
+ 1. Quote the owner word for word, in a block quote, with the date and the place.
10
+ 2. Put your own reading after the quote, as a short list headed "What we take from it".
11
+ 3. Never change a word inside the quote.
12
+
13
+ Example, in a file of its own:
14
+
15
+ ````markdown
16
+ ---
17
+ part: words
18
+ ---
19
+ # Your words
20
+
21
+ > "I want tests to run before the merge, not after."
22
+
23
+ The owner, 9 October, in the project thread.
24
+
25
+ What we take from it:
26
+
27
+ - A red test stops the merge.
28
+ - The owner sees the failing check, not a log.
29
+ ````
@@ -0,0 +1,25 @@
1
+ # Yardstick
2
+
3
+ It holds: The tests every option must pass, set before the options.
4
+
5
+ Needs: Decide: needed, Flow: none, Contract: none, Initiative: none.
6
+
7
+ Rules:
8
+
9
+ 1. Set the tests every option must pass before you list the options.
10
+ 2. Give each test a short name, so the scorecard can use it.
11
+ 3. Keep it to five tests or fewer.
12
+
13
+ Example, in a file of its own:
14
+
15
+ ````markdown
16
+ ---
17
+ part: yardstick
18
+ ---
19
+ # Yardstick
20
+
21
+ | Test | What it asks |
22
+ |---|---|
23
+ | Stops red | Does a red test stop the merge? |
24
+ | Fast | Does the check finish in ten minutes? |
25
+ ````
@@ -0,0 +1,306 @@
1
+ # The parts of an Atlas design, and which kind needs each one.
2
+ #
3
+ # A design is a folder. design.yaml names its kinds. Each part is one
4
+ # markdown file in the folder, with `part: <id>` in its front matter.
5
+ # The design check reads this file to find a missing or empty part.
6
+ # The page builder reads it for the order and the groups of the rail.
7
+ # The guide in design/README.md shows the same table; a test keeps the two equal.
8
+ #
9
+ # Each part says what each kind needs:
10
+ # needed the part must be in the folder and must not be empty
11
+ # optional write it when it applies
12
+ # none the kind does not use it
13
+ # shared needed when design.yaml says `shared: true`, optional when not
14
+ #
15
+ # A design keeps no state. The tracker parts at the end are read from the
16
+ # tracker at build time. No part file may hold them.
17
+
18
+ kinds:
19
+ - id: decide
20
+ name: Decide
21
+ pick: There are several ways, and one must win.
22
+ - id: flow
23
+ name: Flow
24
+ pick: The behaviour is the hard part.
25
+ - id: contract
26
+ name: Contract
27
+ pick: An interface, a schema or a command changes.
28
+ - id: initiative
29
+ name: Initiative
30
+ pick: Many designs serve one goal.
31
+
32
+ groups:
33
+ - id: start
34
+ name: Start here
35
+ - id: decide
36
+ name: Decide
37
+ - id: flow
38
+ name: Flow
39
+ - id: contract
40
+ name: Contract
41
+ - id: initiative
42
+ name: Initiative
43
+ - id: prove
44
+ name: Prove it
45
+ - id: tracker
46
+ name: From the tracker
47
+
48
+ parts:
49
+ - id: summary
50
+ name: Summary
51
+ group: start
52
+ holds: The answer in a few lines, and one picture.
53
+ decide: needed
54
+ flow: needed
55
+ contract: needed
56
+ initiative: needed
57
+ - id: words
58
+ name: Your words
59
+ group: start
60
+ holds: The owner's words, word for word. Claude's reading stays apart.
61
+ decide: needed
62
+ flow: needed
63
+ contract: needed
64
+ initiative: needed
65
+ - id: goals
66
+ name: Goals and non-goals
67
+ group: start
68
+ holds: What it is for, what it is not, and how we know it worked.
69
+ decide: needed
70
+ flow: needed
71
+ contract: needed
72
+ initiative: needed
73
+
74
+ - id: problem
75
+ name: Problem today
76
+ group: decide
77
+ holds: What happens today, with a real example.
78
+ decide: needed
79
+ flow: optional
80
+ contract: optional
81
+ initiative: none
82
+ - id: yardstick
83
+ name: Yardstick
84
+ group: decide
85
+ holds: The tests every option must pass, set before the options.
86
+ decide: needed
87
+ flow: none
88
+ contract: none
89
+ initiative: none
90
+ - id: proposal
91
+ name: Proposal
92
+ group: decide
93
+ holds: The option we pick, and how it works.
94
+ decide: needed
95
+ flow: none
96
+ contract: none
97
+ initiative: none
98
+ - id: alternatives
99
+ name: Alternatives
100
+ group: decide
101
+ holds: The other options, and why each one lost.
102
+ decide: needed
103
+ flow: optional
104
+ contract: optional
105
+ initiative: none
106
+ - id: scorecard
107
+ name: Scorecard
108
+ group: decide
109
+ holds: Each option against each yardstick test.
110
+ decide: needed
111
+ flow: none
112
+ contract: none
113
+ initiative: none
114
+ - id: build
115
+ name: Build list
116
+ group: decide
117
+ holds: The pieces to build, in order.
118
+ decide: needed
119
+ flow: optional
120
+ contract: optional
121
+ initiative: none
122
+
123
+ - id: actors
124
+ name: Who and what
125
+ group: flow
126
+ holds: The people, agents and systems in the flow.
127
+ decide: optional
128
+ flow: needed
129
+ contract: optional
130
+ initiative: none
131
+ - id: stories
132
+ name: Stories
133
+ group: flow
134
+ holds: One walkthrough for each story, step by step.
135
+ decide: optional
136
+ flow: needed
137
+ contract: optional
138
+ initiative: none
139
+ - id: edges
140
+ name: Edge cases
141
+ group: flow
142
+ holds: What happens when a step fails or comes in a strange order.
143
+ decide: optional
144
+ flow: needed
145
+ contract: optional
146
+ initiative: none
147
+ - id: records
148
+ name: Records
149
+ group: flow
150
+ holds: The records each step reads and writes.
151
+ decide: none
152
+ flow: needed
153
+ contract: optional
154
+ initiative: none
155
+ - id: states
156
+ name: States
157
+ group: flow
158
+ holds: The states a thing moves through, and what moves it.
159
+ decide: none
160
+ flow: optional
161
+ contract: none
162
+ initiative: none
163
+
164
+ - id: routes
165
+ name: Routes and interface
166
+ group: contract
167
+ holds: Every route, exported function, command, tool or schema object that is added, changed or removed.
168
+ decide: none
169
+ flow: none
170
+ contract: needed
171
+ initiative: none
172
+ - id: calls
173
+ name: Calls and answers
174
+ group: contract
175
+ holds: A real call and its real answer for each entry of the interface. A schema repo shows a real query; a plugin shows a tool call.
176
+ decide: none
177
+ flow: none
178
+ contract: needed
179
+ initiative: none
180
+ - id: data
181
+ name: Data change
182
+ group: contract
183
+ holds: The schema or data change. Write "No change" and why, when there is none, as a library or a CLI often does.
184
+ decide: none
185
+ flow: none
186
+ contract: needed
187
+ initiative: none
188
+ - id: migration
189
+ name: Migration and undo
190
+ group: contract
191
+ holds: How old data moves to the new shape, and how the data change is undone.
192
+ decide: none
193
+ flow: none
194
+ contract: needed
195
+ initiative: none
196
+ - id: security
197
+ name: Security
198
+ group: contract
199
+ holds: Who may call it, what it may reach, and what it never leaks.
200
+ decide: optional
201
+ flow: optional
202
+ contract: needed
203
+ initiative: none
204
+ - id: rollout
205
+ name: Rollout
206
+ group: contract
207
+ holds: The order of release, and how we watch it.
208
+ decide: optional
209
+ flow: optional
210
+ contract: needed
211
+ initiative: none
212
+
213
+ - id: why
214
+ name: Why one initiative
215
+ group: initiative
216
+ holds: Why these designs belong together.
217
+ decide: none
218
+ flow: none
219
+ contract: none
220
+ initiative: needed
221
+ - id: shared-decisions
222
+ name: Shared decisions
223
+ group: initiative
224
+ holds: The decisions every design in the family follows, and why.
225
+ decide: none
226
+ flow: none
227
+ contract: none
228
+ initiative: needed
229
+ - id: order
230
+ name: Order and dependencies
231
+ group: initiative
232
+ holds: Which design comes first, and what each one waits on.
233
+ decide: none
234
+ flow: none
235
+ contract: none
236
+ initiative: needed
237
+
238
+ - id: risks
239
+ name: Risks
240
+ group: prove
241
+ holds: What breaks, and what we do. "Accepted" with a reason is valid.
242
+ decide: needed
243
+ flow: needed
244
+ contract: needed
245
+ initiative: needed
246
+ - id: proof
247
+ name: Test and proof plan
248
+ group: prove
249
+ holds: How each piece proves itself.
250
+ decide: needed
251
+ flow: needed
252
+ contract: needed
253
+ initiative: none
254
+ - id: repos
255
+ name: Each kind of repo
256
+ group: prove
257
+ holds: What the design means for a service, a library, a schema, a CLI and a plugin.
258
+ decide: shared
259
+ flow: shared
260
+ contract: shared
261
+ initiative: none
262
+ - id: change
263
+ name: Change and undo
264
+ group: prove
265
+ holds: The kind of change, its impact, how to undo it, and each existing test it changes.
266
+ decide: needed
267
+ flow: needed
268
+ contract: needed
269
+ initiative: optional
270
+ - id: done
271
+ name: Done line
272
+ group: prove
273
+ holds: Lines someone else can check. The approval lives in the tracker.
274
+ decide: needed
275
+ flow: needed
276
+ contract: needed
277
+ initiative: needed
278
+ - id: key
279
+ name: Key
280
+ group: prove
281
+ holds: Every word and code the design uses.
282
+ decide: needed
283
+ flow: needed
284
+ contract: needed
285
+ initiative: needed
286
+
287
+ # Read from the tracker at build time. The page shows them read only.
288
+ tracker:
289
+ - id: questions
290
+ name: Questions
291
+ holds: Each question to the owner, open or answered, with the answer.
292
+ - id: decisions
293
+ name: Decisions
294
+ holds: Each decision, in the owner's words, with its date.
295
+ - id: approval
296
+ name: Approval and lock
297
+ holds: The owner's approval and the lock, with the words and the date.
298
+ - id: since
299
+ name: Since approval
300
+ holds: Each change after approval, the part it touched, and the owner's words.
301
+ - id: family
302
+ name: Family
303
+ holds: The designs linked to this one, with their state.
304
+ - id: rollup
305
+ name: Roll-up
306
+ holds: How far the family has come. Initiative designs only.
@@ -25,7 +25,7 @@ tripwires: [definition_of_done, security, schema_migrations, money_cost]
25
25
  tiers: [hotfix, standard, initiative]
26
26
  tier_escalators: [migrations, auth_rls_surface, new_external_dependency, money_cost]
27
27
 
28
- # Both templates travel with the package, beside the kernel.
28
+ # The design format and the resume anchor travel with the package, beside the kernel.
29
29
  templates:
30
- design_doc: templates/design-doc.md
30
+ design_doc: design/README.md
31
31
  resume_anchor: templates/resume-anchor.md
@@ -0,0 +1,234 @@
1
+ ---
2
+ name: tests
3
+ description: >
4
+ Scenario tests for an area of a repo. Trigger when the owner asks to put tests on an area,
5
+ to review the words of a scenario, to write or convert a test from a scenario, when a test
6
+ is red, or to prove that tests catch faults. It holds five procedures: onboard an area,
7
+ review words, convert, sort a red test, and plant faults.
8
+ ---
9
+
10
+ # Tests
11
+
12
+ A scenario is a Markdown file. It says what must be true. A test is code
13
+ that checks it. Atlas ships the code that joins the two: the contract
14
+ library, the drivers, the stand-in helper and the link check. This skill
15
+ holds the steps to use them.
16
+
17
+ ## The files of an area
18
+
19
+ | File | What it holds | Who changes it |
20
+ |---|---|---|
21
+ | `atlas/tests.yaml` | How to start the product, the ways in, the stand-ins, the personas, the guards | You, at onboarding. The owner approves its guard parts |
22
+ | `scenarios/<id>.md` | The words: start, steps, waits, case tables, assertions | The design. The builder copies the approved words |
23
+ | `scenarios/map.md` | How to reach each thing: calls, page parts, error forms, gaps | The builder, as it learns facts |
24
+ | `test/atlas/` | The Atlas kit, copied byte for byte | Only `atlas tests write` |
25
+ | `test/scenarios/` | One test file for each scenario | The builder |
26
+ | `test/support/` | The stand-ins, the actions, the makers, the adapter | The builder |
27
+ | `test/evidence/` | The evidence of each run | The library. Git ignores it |
28
+
29
+ In this skill, "you" is the agent that runs the skill. It is never the owner.
30
+
31
+ Never edit a file in `test/atlas/`. A hand edit stops the next
32
+ `atlas tests write`, and `atlas tests check` reports it. A person merges
33
+ every file that command writes.
34
+
35
+ ## The three commands
36
+
37
+ 1. `atlas tests write --root <repo>` puts the kit in `test/atlas/`. It
38
+ refuses to replace a file that you edited, and it refuses to write
39
+ through a link. It writes `test/evidence/.gitignore`. It adds a `.env.test`
40
+ line to the `.gitignore` at the repo root. It says what it added.
41
+ 2. `atlas tests check --root <repo>` runs the link check. It prints one line
42
+ for each finding, as `file:line code words`. Run it before you ask for a merge.
43
+ 3. `atlas tests proof --root <repo>` prints the proof table of one run.
44
+
45
+ The link check has three halves. Run the first two before the tests and the
46
+ third after them:
47
+
48
+ 1. `--halves static,lint` reads the scenarios, `tests.yaml` and the code. It
49
+ also checks each file in `test/atlas/` against `manifest.json`. The manifest
50
+ must match a kit that a release of Atlas shipped.
51
+ 2. `--halves run --run <id>` reads the evidence of that one run. With no
52
+ `--run`, it reads the newest run of each scenario, and `atlas tests proof`
53
+ does the same. It fails on an assertion that the run did not check. It fails
54
+ on evidence of an older version of a scenario file.
55
+ 3. `--guards-from <file>` takes `tests.yaml` from the main branch. A guard
56
+ part that differs is a finding. CI must run the check from the main branch.
57
+
58
+ The environment variable `ATLAS_GUARDS_FROM` names that file for both the
59
+ check and the scenario run. A scenario run whose guard parts differ from
60
+ that file ends `blocked` and starts nothing. A local run without it uses the branch file.
61
+
62
+ ## Run the tests in CI
63
+
64
+ CI runs the tests from the main branch, so a pull request cannot change its own judge.
65
+
66
+ 1. Fetch the main branch.
67
+ 2. Run the copy of `<area>/test/atlas/` that main holds. If main has none
68
+ yet, run the copy of the branch, and print a warning that says so.
69
+ 3. Set `ATLAS_GUARDS_FROM` to `atlas/tests.yaml` of the main branch.
70
+ 4. Run the scenarios with the runner command of `tests.yaml`. Set `ATLAS_RUN_ID`
71
+ or let CI make the run id.
72
+ 5. Run the link check with `--run <that run id>`.
73
+ 6. Upload `<area>/test/evidence/` for 14 days.
74
+
75
+ ## Onboard an area
76
+
77
+ Path 1 is for an area that has integration tests which call its routes,
78
+ and which people keep using. Old unit tests that call functions do not make path 1.
79
+ Use path 2: the area has no integration test that calls its routes. If it has
80
+ route tests that people keep, stop. Tell the owner that path 1 comes later.
81
+
82
+ 1. Send the `atlas:code-explorer` crew to map the area. Ask for routes,
83
+ tools, pages, storage, calls to outside services and existing tests.
84
+ 2. Run `atlas tests write --root <repo>`.
85
+ 3. Draft `atlas/tests.yaml`. Give it `local` and `ci`, each way in, a persona
86
+ for each caller, and `guards.hosts` with local addresses only.
87
+ 4. Give the product no real secret. Write `{run.secret}` for each secret
88
+ of the product. The library makes it new for each run.
89
+ 5. Write a stand-in for each outside service. Use `serveStandIn` from
90
+ `test/atlas/stand-in.mjs`. Give the product a fake key for the real service.
91
+ 6. Write the contract scenario of each stand-in (kind `contract`). It proves
92
+ that the stand-in answers like the real service.
93
+ 7. Put two personas on one secret only with `same-identity` on both.
94
+ 8. Draft `scenarios/map.md` from the code digest and the docs of the repo.
95
+ 9. Write two or three first scenarios. Start with the smallest.
96
+ 10. Review the words (next section). Run `atlas ste` on each scenario.
97
+ 11. Convert each scenario into a test.
98
+ 12. Run `atlas tests check --root <repo> --halves static,lint`, then the tests,
99
+ then `atlas tests check --root <repo> --halves run`.
100
+ 13. Ask the `atlas:verifier` crew to start each environment once and run every scenario.
101
+ 14. Plant a fault for each assertion (last section).
102
+ 15. Open one pull request. Show the owner the first scenarios, and the guard
103
+ parts of `tests.yaml` as a short list in plain words.
104
+
105
+ Worked run, on a meal planner with a web page, an MCP server and storage in
106
+ a note service:
107
+
108
+ 1. Input: the code digest names the routes, the MCP server and the page. It
109
+ finds 11 tests that call functions and none that call a route.
110
+ 2. Process: path 2. Draft `tests.yaml` with two environments, three ways in,
111
+ a stand-in for the note service and four personas. Write the stand-in
112
+ and two scenarios.
113
+ 3. Output: one pull request. The owner reads two scenarios and three lines of guards.
114
+
115
+ ## Write assertions
116
+
117
+ Use these steps when you write or change an assertion. They come from
118
+ sections 9.2 and 6.10 of the design.
119
+
120
+ 1. Write the definition of done as assertions. A feature gets full scenarios.
121
+ A small fix gets a short card with its assertions.
122
+ 2. Name each assertion that you add, change or remove by its full id in the design.
123
+ 3. Write one outcome in each assertion. Then one planted fault points at one assertion.
124
+ 4. Mark it `exact` when code checks one right answer. Mark it `judged` only
125
+ when it is about meaning, and say why in the design.
126
+ 5. Never let a judged assertion decide an access rule alone.
127
+ 6. Mark `gate` on an assertion that later steps need.
128
+ 7. Start a "when" assertion with the step that makes it true: "step 3", or a kept result.
129
+ 8. Name what it reads in backticks.
130
+ 9. Write a "not" assertion so that a blank answer fails it.
131
+ 10. Add `repeat` only when the subject varies by design, such as a model. Say why.
132
+
133
+ ## Review the words
134
+
135
+ A separate agent reads each new or changed scenario before any code exists.
136
+ It proposes words. It never edits. It asks ten questions of each assertion:
137
+
138
+ | Problem | The question |
139
+ |---|---|
140
+ | two-readings | Can two careful agents write different tests from it? |
141
+ | not-exact | Does an "exact" assertion have one right answer? |
142
+ | many-claims | Does it hold more than one outcome? |
143
+ | duplicate | Does another assertion check the same thing? |
144
+ | weak | How could the product pass it while the claim is false? |
145
+ | hidden-precondition | Does it need a state that `Start with` does not name? |
146
+ | missing-data | Does it need a value that the scenario does not give? |
147
+ | shared-data | Can another run change the data it reads? |
148
+ | empty-when | Does a "when" assertion name a step that makes it true? |
149
+ | loose-not | Does a "not" assertion pass on a blank or broken answer? |
150
+
151
+ Then run `atlas ste` on the scenario. The owner reads these pages.
152
+
153
+ Rules for the words:
154
+
155
+ 1. One outcome for each assertion. Mark `exact` or `judged`.
156
+ 2. Mark `gate` on an assertion that later steps need.
157
+ 3. A wait ends with "Give up after <time>". The link check refuses it otherwise.
158
+ 4. A "when" assertion names its step: "step 3", or a result that a step keeps.
159
+ 5. Put each name in backticks. A name is a persona, way in, data rule, setup, fixture, table or kept result.
160
+ 6. An access rule is always exact. A judge never decides it alone.
161
+
162
+ ## Convert a scenario into a test
163
+
164
+ 1. Read the scenario, the map, the actions and one worked example. Do not
165
+ read the product source if the map covers the scenario.
166
+ 2. Write one test file for each scenario, with `scenario(path, body, { adapter })`
167
+ from `test/atlas/contract.mjs`. The library runs the body once for each way in.
168
+ 3. Name each assertion by its full id, as a plain string:
169
+ `run.check('<scenario>/e1#<fingerprint>', () => assert...)`.
170
+ Use `run.gate` for a `gate` assertion. Get the fingerprints from the failure message of `atlas tests check`.
171
+ 4. Use an action, or add one. An action for a new call also adds its row to the map.
172
+ 5. Make data with `run.fresh(rule)`. Read a case table with `run.table`, `run.cells` or `run.cases`.
173
+ 6. Wait with `run.wait` and a limit. Never sleep for a fixed time.
174
+ 7. Use `run.notExercised(id, why)` for a "when" assertion that the run never made true.
175
+ 8. Mark each guess in a comment. List the guesses in the pull request.
176
+ 9. Never change the words to make a test pass. Only the design changes the words.
177
+ 10. Run the test before the change. A test for new behaviour must fail.
178
+
179
+ The adapter binds the library to the product. All its parts are optional:
180
+ `standIns`, `warmUp`, `drivers`, `functions`, `loadModule`, `makers`, `signIns`,
181
+ `data` and `actions`. The header of `test/atlas/contract.mjs` describes each one.
182
+
183
+ A result is `pass`, `fail`, `not checked`, `not exercised`, `blocked` or
184
+ `not here`. A test is red when any assertion is not `pass`. The one
185
+ exception is `not here`: the scenario does not run in that environment, it is
186
+ not counted, and the test passes.
187
+
188
+ After the product is ready, a lost connection, a server exit or a time-out is
189
+ `fail`. `blocked` is for a guard refusal, a start-up fault and the first-call check.
190
+
191
+ ## Sort a red test
192
+
193
+ Do not edit a red test before you know why it is red. There are four causes:
194
+
195
+ | Cause | Sign | What happens | What may change |
196
+ |---|---|---|---|
197
+ | The product broke | The words still hold. The product does not | Fix the code | Product code only |
198
+ | The behaviour changed on purpose | The approved design names the assertion | Change the assertion and its test | The assertions that the design names, and their tests |
199
+ | The test is out of date | A page part moved, or the form of an answer changed. The verifier's own check passes | Repair the actions, the map or the finders | Actions, map and finders. No assertion and no assert statement |
200
+ | The environment failed | The result is `blocked`: the product did not start, or a stand-in is down | Fix the setup | `tests.yaml` without its guard parts, and the support code |
201
+
202
+ Two checks follow:
203
+
204
+ 1. A repair that changes an assertion or an assert statement is not a repair. Refuse it unless the design names the assertion.
205
+ 2. A repair must keep the verifier's own check green. The verifier does not use the actions.
206
+
207
+ A `blocked` result because a guard refused, or because the stand-in saw no
208
+ first call from the product, is a finding, not noise.
209
+ Find out why the product reached for something else.
210
+
211
+ ## Plant faults
212
+
213
+ A fault proves that an assertion can go red. Plant one for each new or
214
+ changed assertion, and for each assertion at onboarding.
215
+
216
+ 1. Commit first. A clone holds committed files only, so a change that is not
217
+ committed is not in the clone. Then clone the worktree to a throwaway
218
+ folder outside it, with `git clone --no-hardlinks`.
219
+ 2. Remove every remote from the clone. Run `git remote` there. It must print nothing.
220
+ 3. Plant faults in product code only. Never touch `tests.yaml`, the support code or the kit.
221
+ 4. Use the cheap model first: Haiku. After two failed tries on one assertion, use Sonnet.
222
+ 5. Give the planter one assertion and the product code. Ask for one fault that makes only that claim false.
223
+ 6. Run the scenario tests from the real worktree. Set `ATLAS_PRODUCT_ROOT` to
224
+ the area folder in the clone, and `ATLAS_FAULT_RUN` to a label. The product,
225
+ the MCP server and the module of the `function` driver start in the clone.
226
+ The guards, the kit and `tests.yaml` load from the worktree.
227
+ 7. The test for that assertion must go red.
228
+ 8. Ask the `atlas:verifier` crew for its own check of that assertion, on the same clone. It uses a raw driver, not the actions.
229
+ 9. A fault is valid only when the verifier's own check also fails. Discard any other fault and try again.
230
+ 10. Report each assertion as caught, missed, or "no fault of its own" after Sonnet also fails.
231
+ 11. Delete the clone. Run `git status` in the worktree. It must show no change from the faults.
232
+
233
+ Evidence from a fault run carries the label. The link check ignores it, so it never counts as proof.
234
+ Never push from the clone.