@arjunkhera/atlas 0.3.7 → 0.3.9
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/agents/artifact-renderer.md +22 -22
- package/agents/verifier.md +29 -0
- package/door/cli.mjs +62 -12
- package/door/kit-releases.json +4 -0
- package/door/lib/design-build.mjs +409 -0
- package/door/lib/design.mjs +199 -118
- package/door/lib/markdown.mjs +160 -0
- package/door/lib/proof.mjs +127 -0
- package/door/lib/tests.mjs +190 -0
- package/package.json +2 -1
- package/skills/lead/SKILL.md +4 -1
- package/skills/sdlc-task/SKILL.md +33 -24
- package/skills/sdlc-task/design/README.md +187 -0
- package/skills/sdlc-task/design/parts/actors.md +26 -0
- package/skills/sdlc-task/design/parts/alternatives.md +24 -0
- package/skills/sdlc-task/design/parts/build.md +23 -0
- package/skills/sdlc-task/design/parts/calls.md +25 -0
- package/skills/sdlc-task/design/parts/change.md +27 -0
- package/skills/sdlc-task/design/parts/data.md +22 -0
- package/skills/sdlc-task/design/parts/done.md +23 -0
- package/skills/sdlc-task/design/parts/edges.md +24 -0
- package/skills/sdlc-task/design/parts/goals.md +27 -0
- package/skills/sdlc-task/design/parts/key.md +25 -0
- package/skills/sdlc-task/design/parts/migration.md +22 -0
- package/skills/sdlc-task/design/parts/order.md +24 -0
- package/skills/sdlc-task/design/parts/problem.md +22 -0
- package/skills/sdlc-task/design/parts/proof.md +24 -0
- package/skills/sdlc-task/design/parts/proposal.md +24 -0
- package/skills/sdlc-task/design/parts/records.md +24 -0
- package/skills/sdlc-task/design/parts/repos.md +25 -0
- package/skills/sdlc-task/design/parts/risks.md +24 -0
- package/skills/sdlc-task/design/parts/rollout.md +24 -0
- package/skills/sdlc-task/design/parts/routes.md +24 -0
- package/skills/sdlc-task/design/parts/scorecard.md +25 -0
- package/skills/sdlc-task/design/parts/security.md +22 -0
- package/skills/sdlc-task/design/parts/shared-decisions.md +24 -0
- package/skills/sdlc-task/design/parts/states.md +25 -0
- package/skills/sdlc-task/design/parts/stories.md +26 -0
- package/skills/sdlc-task/design/parts/summary.md +33 -0
- package/skills/sdlc-task/design/parts/why.md +22 -0
- package/skills/sdlc-task/design/parts/words.md +29 -0
- package/skills/sdlc-task/design/parts/yardstick.md +25 -0
- package/skills/sdlc-task/design/parts.yaml +306 -0
- package/skills/sdlc-task/lifecycle.yaml +2 -2
- package/skills/tests/SKILL.md +234 -0
- package/tests/contract.mjs +398 -0
- package/tests/drivers/function.mjs +30 -0
- package/tests/drivers/http.mjs +68 -0
- package/tests/drivers/index.mjs +65 -0
- package/tests/drivers/mcp-stdio.mjs +174 -0
- package/tests/environment.mjs +227 -0
- package/tests/errors.mjs +27 -0
- package/tests/evidence.mjs +113 -0
- package/tests/fresh.mjs +42 -0
- package/tests/guards.mjs +159 -0
- package/tests/index.mjs +11 -0
- package/tests/link-check.mjs +576 -0
- package/tests/procs.mjs +43 -0
- package/tests/redact.mjs +58 -0
- package/tests/scenario.mjs +325 -0
- package/tests/stand-in.mjs +74 -0
- package/tests/tests-yaml.mjs +258 -0
- package/tests/wait.mjs +44 -0
- package/tests/yaml.mjs +327 -0
- package/agents/artifact-format/walkthrough.html +0 -706
- package/skills/sdlc-task/templates/design-doc.md +0 -126
|
@@ -0,0 +1,22 @@
|
|
|
1
|
+
# Why one initiative
|
|
2
|
+
|
|
3
|
+
It holds: Why these designs belong together.
|
|
4
|
+
|
|
5
|
+
Needs: Decide: none, Flow: none, Contract: none, Initiative: needed.
|
|
6
|
+
|
|
7
|
+
Rules:
|
|
8
|
+
|
|
9
|
+
1. Say why these designs belong together.
|
|
10
|
+
2. Name the one goal they serve.
|
|
11
|
+
3. Say what would go wrong if each went alone.
|
|
12
|
+
|
|
13
|
+
Example, in a file of its own:
|
|
14
|
+
|
|
15
|
+
````markdown
|
|
16
|
+
---
|
|
17
|
+
part: why
|
|
18
|
+
---
|
|
19
|
+
# Why one initiative
|
|
20
|
+
|
|
21
|
+
Each file type needs its own reader, but all share one viewer and one queue. Alone, each would build its own queue.
|
|
22
|
+
````
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
# Your words
|
|
2
|
+
|
|
3
|
+
It holds: The owner's words, word for word. Claude's reading stays apart.
|
|
4
|
+
|
|
5
|
+
Needs: Decide: needed, Flow: needed, Contract: needed, Initiative: needed.
|
|
6
|
+
|
|
7
|
+
Rules:
|
|
8
|
+
|
|
9
|
+
1. Quote the owner word for word, in a block quote, with the date and the place.
|
|
10
|
+
2. Put your own reading after the quote, as a short list headed "What we take from it".
|
|
11
|
+
3. Never change a word inside the quote.
|
|
12
|
+
|
|
13
|
+
Example, in a file of its own:
|
|
14
|
+
|
|
15
|
+
````markdown
|
|
16
|
+
---
|
|
17
|
+
part: words
|
|
18
|
+
---
|
|
19
|
+
# Your words
|
|
20
|
+
|
|
21
|
+
> "I want tests to run before the merge, not after."
|
|
22
|
+
|
|
23
|
+
The owner, 9 October, in the project thread.
|
|
24
|
+
|
|
25
|
+
What we take from it:
|
|
26
|
+
|
|
27
|
+
- A red test stops the merge.
|
|
28
|
+
- The owner sees the failing check, not a log.
|
|
29
|
+
````
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
# Yardstick
|
|
2
|
+
|
|
3
|
+
It holds: The tests every option must pass, set before the options.
|
|
4
|
+
|
|
5
|
+
Needs: Decide: needed, Flow: none, Contract: none, Initiative: none.
|
|
6
|
+
|
|
7
|
+
Rules:
|
|
8
|
+
|
|
9
|
+
1. Set the tests every option must pass before you list the options.
|
|
10
|
+
2. Give each test a short name, so the scorecard can use it.
|
|
11
|
+
3. Keep it to five tests or fewer.
|
|
12
|
+
|
|
13
|
+
Example, in a file of its own:
|
|
14
|
+
|
|
15
|
+
````markdown
|
|
16
|
+
---
|
|
17
|
+
part: yardstick
|
|
18
|
+
---
|
|
19
|
+
# Yardstick
|
|
20
|
+
|
|
21
|
+
| Test | What it asks |
|
|
22
|
+
|---|---|
|
|
23
|
+
| Stops red | Does a red test stop the merge? |
|
|
24
|
+
| Fast | Does the check finish in ten minutes? |
|
|
25
|
+
````
|
|
@@ -0,0 +1,306 @@
|
|
|
1
|
+
# The parts of an Atlas design, and which kind needs each one.
|
|
2
|
+
#
|
|
3
|
+
# A design is a folder. design.yaml names its kinds. Each part is one
|
|
4
|
+
# markdown file in the folder, with `part: <id>` in its front matter.
|
|
5
|
+
# The design check reads this file to find a missing or empty part.
|
|
6
|
+
# The page builder reads it for the order and the groups of the rail.
|
|
7
|
+
# The guide in design/README.md shows the same table; a test keeps the two equal.
|
|
8
|
+
#
|
|
9
|
+
# Each part says what each kind needs:
|
|
10
|
+
# needed the part must be in the folder and must not be empty
|
|
11
|
+
# optional write it when it applies
|
|
12
|
+
# none the kind does not use it
|
|
13
|
+
# shared needed when design.yaml says `shared: true`, optional when not
|
|
14
|
+
#
|
|
15
|
+
# A design keeps no state. The tracker parts at the end are read from the
|
|
16
|
+
# tracker at build time. No part file may hold them.
|
|
17
|
+
|
|
18
|
+
kinds:
|
|
19
|
+
- id: decide
|
|
20
|
+
name: Decide
|
|
21
|
+
pick: There are several ways, and one must win.
|
|
22
|
+
- id: flow
|
|
23
|
+
name: Flow
|
|
24
|
+
pick: The behaviour is the hard part.
|
|
25
|
+
- id: contract
|
|
26
|
+
name: Contract
|
|
27
|
+
pick: An interface, a schema or a command changes.
|
|
28
|
+
- id: initiative
|
|
29
|
+
name: Initiative
|
|
30
|
+
pick: Many designs serve one goal.
|
|
31
|
+
|
|
32
|
+
groups:
|
|
33
|
+
- id: start
|
|
34
|
+
name: Start here
|
|
35
|
+
- id: decide
|
|
36
|
+
name: Decide
|
|
37
|
+
- id: flow
|
|
38
|
+
name: Flow
|
|
39
|
+
- id: contract
|
|
40
|
+
name: Contract
|
|
41
|
+
- id: initiative
|
|
42
|
+
name: Initiative
|
|
43
|
+
- id: prove
|
|
44
|
+
name: Prove it
|
|
45
|
+
- id: tracker
|
|
46
|
+
name: From the tracker
|
|
47
|
+
|
|
48
|
+
parts:
|
|
49
|
+
- id: summary
|
|
50
|
+
name: Summary
|
|
51
|
+
group: start
|
|
52
|
+
holds: The answer in a few lines, and one picture.
|
|
53
|
+
decide: needed
|
|
54
|
+
flow: needed
|
|
55
|
+
contract: needed
|
|
56
|
+
initiative: needed
|
|
57
|
+
- id: words
|
|
58
|
+
name: Your words
|
|
59
|
+
group: start
|
|
60
|
+
holds: The owner's words, word for word. Claude's reading stays apart.
|
|
61
|
+
decide: needed
|
|
62
|
+
flow: needed
|
|
63
|
+
contract: needed
|
|
64
|
+
initiative: needed
|
|
65
|
+
- id: goals
|
|
66
|
+
name: Goals and non-goals
|
|
67
|
+
group: start
|
|
68
|
+
holds: What it is for, what it is not, and how we know it worked.
|
|
69
|
+
decide: needed
|
|
70
|
+
flow: needed
|
|
71
|
+
contract: needed
|
|
72
|
+
initiative: needed
|
|
73
|
+
|
|
74
|
+
- id: problem
|
|
75
|
+
name: Problem today
|
|
76
|
+
group: decide
|
|
77
|
+
holds: What happens today, with a real example.
|
|
78
|
+
decide: needed
|
|
79
|
+
flow: optional
|
|
80
|
+
contract: optional
|
|
81
|
+
initiative: none
|
|
82
|
+
- id: yardstick
|
|
83
|
+
name: Yardstick
|
|
84
|
+
group: decide
|
|
85
|
+
holds: The tests every option must pass, set before the options.
|
|
86
|
+
decide: needed
|
|
87
|
+
flow: none
|
|
88
|
+
contract: none
|
|
89
|
+
initiative: none
|
|
90
|
+
- id: proposal
|
|
91
|
+
name: Proposal
|
|
92
|
+
group: decide
|
|
93
|
+
holds: The option we pick, and how it works.
|
|
94
|
+
decide: needed
|
|
95
|
+
flow: none
|
|
96
|
+
contract: none
|
|
97
|
+
initiative: none
|
|
98
|
+
- id: alternatives
|
|
99
|
+
name: Alternatives
|
|
100
|
+
group: decide
|
|
101
|
+
holds: The other options, and why each one lost.
|
|
102
|
+
decide: needed
|
|
103
|
+
flow: optional
|
|
104
|
+
contract: optional
|
|
105
|
+
initiative: none
|
|
106
|
+
- id: scorecard
|
|
107
|
+
name: Scorecard
|
|
108
|
+
group: decide
|
|
109
|
+
holds: Each option against each yardstick test.
|
|
110
|
+
decide: needed
|
|
111
|
+
flow: none
|
|
112
|
+
contract: none
|
|
113
|
+
initiative: none
|
|
114
|
+
- id: build
|
|
115
|
+
name: Build list
|
|
116
|
+
group: decide
|
|
117
|
+
holds: The pieces to build, in order.
|
|
118
|
+
decide: needed
|
|
119
|
+
flow: optional
|
|
120
|
+
contract: optional
|
|
121
|
+
initiative: none
|
|
122
|
+
|
|
123
|
+
- id: actors
|
|
124
|
+
name: Who and what
|
|
125
|
+
group: flow
|
|
126
|
+
holds: The people, agents and systems in the flow.
|
|
127
|
+
decide: optional
|
|
128
|
+
flow: needed
|
|
129
|
+
contract: optional
|
|
130
|
+
initiative: none
|
|
131
|
+
- id: stories
|
|
132
|
+
name: Stories
|
|
133
|
+
group: flow
|
|
134
|
+
holds: One walkthrough for each story, step by step.
|
|
135
|
+
decide: optional
|
|
136
|
+
flow: needed
|
|
137
|
+
contract: optional
|
|
138
|
+
initiative: none
|
|
139
|
+
- id: edges
|
|
140
|
+
name: Edge cases
|
|
141
|
+
group: flow
|
|
142
|
+
holds: What happens when a step fails or comes in a strange order.
|
|
143
|
+
decide: optional
|
|
144
|
+
flow: needed
|
|
145
|
+
contract: optional
|
|
146
|
+
initiative: none
|
|
147
|
+
- id: records
|
|
148
|
+
name: Records
|
|
149
|
+
group: flow
|
|
150
|
+
holds: The records each step reads and writes.
|
|
151
|
+
decide: none
|
|
152
|
+
flow: needed
|
|
153
|
+
contract: optional
|
|
154
|
+
initiative: none
|
|
155
|
+
- id: states
|
|
156
|
+
name: States
|
|
157
|
+
group: flow
|
|
158
|
+
holds: The states a thing moves through, and what moves it.
|
|
159
|
+
decide: none
|
|
160
|
+
flow: optional
|
|
161
|
+
contract: none
|
|
162
|
+
initiative: none
|
|
163
|
+
|
|
164
|
+
- id: routes
|
|
165
|
+
name: Routes and interface
|
|
166
|
+
group: contract
|
|
167
|
+
holds: Every route, exported function, command, tool or schema object that is added, changed or removed.
|
|
168
|
+
decide: none
|
|
169
|
+
flow: none
|
|
170
|
+
contract: needed
|
|
171
|
+
initiative: none
|
|
172
|
+
- id: calls
|
|
173
|
+
name: Calls and answers
|
|
174
|
+
group: contract
|
|
175
|
+
holds: A real call and its real answer for each entry of the interface. A schema repo shows a real query; a plugin shows a tool call.
|
|
176
|
+
decide: none
|
|
177
|
+
flow: none
|
|
178
|
+
contract: needed
|
|
179
|
+
initiative: none
|
|
180
|
+
- id: data
|
|
181
|
+
name: Data change
|
|
182
|
+
group: contract
|
|
183
|
+
holds: The schema or data change. Write "No change" and why, when there is none, as a library or a CLI often does.
|
|
184
|
+
decide: none
|
|
185
|
+
flow: none
|
|
186
|
+
contract: needed
|
|
187
|
+
initiative: none
|
|
188
|
+
- id: migration
|
|
189
|
+
name: Migration and undo
|
|
190
|
+
group: contract
|
|
191
|
+
holds: How old data moves to the new shape, and how the data change is undone.
|
|
192
|
+
decide: none
|
|
193
|
+
flow: none
|
|
194
|
+
contract: needed
|
|
195
|
+
initiative: none
|
|
196
|
+
- id: security
|
|
197
|
+
name: Security
|
|
198
|
+
group: contract
|
|
199
|
+
holds: Who may call it, what it may reach, and what it never leaks.
|
|
200
|
+
decide: optional
|
|
201
|
+
flow: optional
|
|
202
|
+
contract: needed
|
|
203
|
+
initiative: none
|
|
204
|
+
- id: rollout
|
|
205
|
+
name: Rollout
|
|
206
|
+
group: contract
|
|
207
|
+
holds: The order of release, and how we watch it.
|
|
208
|
+
decide: optional
|
|
209
|
+
flow: optional
|
|
210
|
+
contract: needed
|
|
211
|
+
initiative: none
|
|
212
|
+
|
|
213
|
+
- id: why
|
|
214
|
+
name: Why one initiative
|
|
215
|
+
group: initiative
|
|
216
|
+
holds: Why these designs belong together.
|
|
217
|
+
decide: none
|
|
218
|
+
flow: none
|
|
219
|
+
contract: none
|
|
220
|
+
initiative: needed
|
|
221
|
+
- id: shared-decisions
|
|
222
|
+
name: Shared decisions
|
|
223
|
+
group: initiative
|
|
224
|
+
holds: The decisions every design in the family follows, and why.
|
|
225
|
+
decide: none
|
|
226
|
+
flow: none
|
|
227
|
+
contract: none
|
|
228
|
+
initiative: needed
|
|
229
|
+
- id: order
|
|
230
|
+
name: Order and dependencies
|
|
231
|
+
group: initiative
|
|
232
|
+
holds: Which design comes first, and what each one waits on.
|
|
233
|
+
decide: none
|
|
234
|
+
flow: none
|
|
235
|
+
contract: none
|
|
236
|
+
initiative: needed
|
|
237
|
+
|
|
238
|
+
- id: risks
|
|
239
|
+
name: Risks
|
|
240
|
+
group: prove
|
|
241
|
+
holds: What breaks, and what we do. "Accepted" with a reason is valid.
|
|
242
|
+
decide: needed
|
|
243
|
+
flow: needed
|
|
244
|
+
contract: needed
|
|
245
|
+
initiative: needed
|
|
246
|
+
- id: proof
|
|
247
|
+
name: Test and proof plan
|
|
248
|
+
group: prove
|
|
249
|
+
holds: How each piece proves itself.
|
|
250
|
+
decide: needed
|
|
251
|
+
flow: needed
|
|
252
|
+
contract: needed
|
|
253
|
+
initiative: none
|
|
254
|
+
- id: repos
|
|
255
|
+
name: Each kind of repo
|
|
256
|
+
group: prove
|
|
257
|
+
holds: What the design means for a service, a library, a schema, a CLI and a plugin.
|
|
258
|
+
decide: shared
|
|
259
|
+
flow: shared
|
|
260
|
+
contract: shared
|
|
261
|
+
initiative: none
|
|
262
|
+
- id: change
|
|
263
|
+
name: Change and undo
|
|
264
|
+
group: prove
|
|
265
|
+
holds: The kind of change, its impact, how to undo it, and each existing test it changes.
|
|
266
|
+
decide: needed
|
|
267
|
+
flow: needed
|
|
268
|
+
contract: needed
|
|
269
|
+
initiative: optional
|
|
270
|
+
- id: done
|
|
271
|
+
name: Done line
|
|
272
|
+
group: prove
|
|
273
|
+
holds: Lines someone else can check. The approval lives in the tracker.
|
|
274
|
+
decide: needed
|
|
275
|
+
flow: needed
|
|
276
|
+
contract: needed
|
|
277
|
+
initiative: needed
|
|
278
|
+
- id: key
|
|
279
|
+
name: Key
|
|
280
|
+
group: prove
|
|
281
|
+
holds: Every word and code the design uses.
|
|
282
|
+
decide: needed
|
|
283
|
+
flow: needed
|
|
284
|
+
contract: needed
|
|
285
|
+
initiative: needed
|
|
286
|
+
|
|
287
|
+
# Read from the tracker at build time. The page shows them read only.
|
|
288
|
+
tracker:
|
|
289
|
+
- id: questions
|
|
290
|
+
name: Questions
|
|
291
|
+
holds: Each question to the owner, open or answered, with the answer.
|
|
292
|
+
- id: decisions
|
|
293
|
+
name: Decisions
|
|
294
|
+
holds: Each decision, in the owner's words, with its date.
|
|
295
|
+
- id: approval
|
|
296
|
+
name: Approval and lock
|
|
297
|
+
holds: The owner's approval and the lock, with the words and the date.
|
|
298
|
+
- id: since
|
|
299
|
+
name: Since approval
|
|
300
|
+
holds: Each change after approval, the part it touched, and the owner's words.
|
|
301
|
+
- id: family
|
|
302
|
+
name: Family
|
|
303
|
+
holds: The designs linked to this one, with their state.
|
|
304
|
+
- id: rollup
|
|
305
|
+
name: Roll-up
|
|
306
|
+
holds: How far the family has come. Initiative designs only.
|
|
@@ -25,7 +25,7 @@ tripwires: [definition_of_done, security, schema_migrations, money_cost]
|
|
|
25
25
|
tiers: [hotfix, standard, initiative]
|
|
26
26
|
tier_escalators: [migrations, auth_rls_surface, new_external_dependency, money_cost]
|
|
27
27
|
|
|
28
|
-
#
|
|
28
|
+
# The design format and the resume anchor travel with the package, beside the kernel.
|
|
29
29
|
templates:
|
|
30
|
-
design_doc:
|
|
30
|
+
design_doc: design/README.md
|
|
31
31
|
resume_anchor: templates/resume-anchor.md
|
|
@@ -0,0 +1,234 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: tests
|
|
3
|
+
description: >
|
|
4
|
+
Scenario tests for an area of a repo. Trigger when the owner asks to put tests on an area,
|
|
5
|
+
to review the words of a scenario, to write or convert a test from a scenario, when a test
|
|
6
|
+
is red, or to prove that tests catch faults. It holds five procedures: onboard an area,
|
|
7
|
+
review words, convert, sort a red test, and plant faults.
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# Tests
|
|
11
|
+
|
|
12
|
+
A scenario is a Markdown file. It says what must be true. A test is code
|
|
13
|
+
that checks it. Atlas ships the code that joins the two: the contract
|
|
14
|
+
library, the drivers, the stand-in helper and the link check. This skill
|
|
15
|
+
holds the steps to use them.
|
|
16
|
+
|
|
17
|
+
## The files of an area
|
|
18
|
+
|
|
19
|
+
| File | What it holds | Who changes it |
|
|
20
|
+
|---|---|---|
|
|
21
|
+
| `atlas/tests.yaml` | How to start the product, the ways in, the stand-ins, the personas, the guards | You, at onboarding. The owner approves its guard parts |
|
|
22
|
+
| `scenarios/<id>.md` | The words: start, steps, waits, case tables, assertions | The design. The builder copies the approved words |
|
|
23
|
+
| `scenarios/map.md` | How to reach each thing: calls, page parts, error forms, gaps | The builder, as it learns facts |
|
|
24
|
+
| `test/atlas/` | The Atlas kit, copied byte for byte | Only `atlas tests write` |
|
|
25
|
+
| `test/scenarios/` | One test file for each scenario | The builder |
|
|
26
|
+
| `test/support/` | The stand-ins, the actions, the makers, the adapter | The builder |
|
|
27
|
+
| `test/evidence/` | The evidence of each run | The library. Git ignores it |
|
|
28
|
+
|
|
29
|
+
In this skill, "you" is the agent that runs the skill. It is never the owner.
|
|
30
|
+
|
|
31
|
+
Never edit a file in `test/atlas/`. A hand edit stops the next
|
|
32
|
+
`atlas tests write`, and `atlas tests check` reports it. A person merges
|
|
33
|
+
every file that command writes.
|
|
34
|
+
|
|
35
|
+
## The three commands
|
|
36
|
+
|
|
37
|
+
1. `atlas tests write --root <repo>` puts the kit in `test/atlas/`. It
|
|
38
|
+
refuses to replace a file that you edited, and it refuses to write
|
|
39
|
+
through a link. It writes `test/evidence/.gitignore`. It adds a `.env.test`
|
|
40
|
+
line to the `.gitignore` at the repo root. It says what it added.
|
|
41
|
+
2. `atlas tests check --root <repo>` runs the link check. It prints one line
|
|
42
|
+
for each finding, as `file:line code words`. Run it before you ask for a merge.
|
|
43
|
+
3. `atlas tests proof --root <repo>` prints the proof table of one run.
|
|
44
|
+
|
|
45
|
+
The link check has three halves. Run the first two before the tests and the
|
|
46
|
+
third after them:
|
|
47
|
+
|
|
48
|
+
1. `--halves static,lint` reads the scenarios, `tests.yaml` and the code. It
|
|
49
|
+
also checks each file in `test/atlas/` against `manifest.json`. The manifest
|
|
50
|
+
must match a kit that a release of Atlas shipped.
|
|
51
|
+
2. `--halves run --run <id>` reads the evidence of that one run. With no
|
|
52
|
+
`--run`, it reads the newest run of each scenario, and `atlas tests proof`
|
|
53
|
+
does the same. It fails on an assertion that the run did not check. It fails
|
|
54
|
+
on evidence of an older version of a scenario file.
|
|
55
|
+
3. `--guards-from <file>` takes `tests.yaml` from the main branch. A guard
|
|
56
|
+
part that differs is a finding. CI must run the check from the main branch.
|
|
57
|
+
|
|
58
|
+
The environment variable `ATLAS_GUARDS_FROM` names that file for both the
|
|
59
|
+
check and the scenario run. A scenario run whose guard parts differ from
|
|
60
|
+
that file ends `blocked` and starts nothing. A local run without it uses the branch file.
|
|
61
|
+
|
|
62
|
+
## Run the tests in CI
|
|
63
|
+
|
|
64
|
+
CI runs the tests from the main branch, so a pull request cannot change its own judge.
|
|
65
|
+
|
|
66
|
+
1. Fetch the main branch.
|
|
67
|
+
2. Run the copy of `<area>/test/atlas/` that main holds. If main has none
|
|
68
|
+
yet, run the copy of the branch, and print a warning that says so.
|
|
69
|
+
3. Set `ATLAS_GUARDS_FROM` to `atlas/tests.yaml` of the main branch.
|
|
70
|
+
4. Run the scenarios with the runner command of `tests.yaml`. Set `ATLAS_RUN_ID`
|
|
71
|
+
or let CI make the run id.
|
|
72
|
+
5. Run the link check with `--run <that run id>`.
|
|
73
|
+
6. Upload `<area>/test/evidence/` for 14 days.
|
|
74
|
+
|
|
75
|
+
## Onboard an area
|
|
76
|
+
|
|
77
|
+
Path 1 is for an area that has integration tests which call its routes,
|
|
78
|
+
and which people keep using. Old unit tests that call functions do not make path 1.
|
|
79
|
+
Use path 2: the area has no integration test that calls its routes. If it has
|
|
80
|
+
route tests that people keep, stop. Tell the owner that path 1 comes later.
|
|
81
|
+
|
|
82
|
+
1. Send the `atlas:code-explorer` crew to map the area. Ask for routes,
|
|
83
|
+
tools, pages, storage, calls to outside services and existing tests.
|
|
84
|
+
2. Run `atlas tests write --root <repo>`.
|
|
85
|
+
3. Draft `atlas/tests.yaml`. Give it `local` and `ci`, each way in, a persona
|
|
86
|
+
for each caller, and `guards.hosts` with local addresses only.
|
|
87
|
+
4. Give the product no real secret. Write `{run.secret}` for each secret
|
|
88
|
+
of the product. The library makes it new for each run.
|
|
89
|
+
5. Write a stand-in for each outside service. Use `serveStandIn` from
|
|
90
|
+
`test/atlas/stand-in.mjs`. Give the product a fake key for the real service.
|
|
91
|
+
6. Write the contract scenario of each stand-in (kind `contract`). It proves
|
|
92
|
+
that the stand-in answers like the real service.
|
|
93
|
+
7. Put two personas on one secret only with `same-identity` on both.
|
|
94
|
+
8. Draft `scenarios/map.md` from the code digest and the docs of the repo.
|
|
95
|
+
9. Write two or three first scenarios. Start with the smallest.
|
|
96
|
+
10. Review the words (next section). Run `atlas ste` on each scenario.
|
|
97
|
+
11. Convert each scenario into a test.
|
|
98
|
+
12. Run `atlas tests check --root <repo> --halves static,lint`, then the tests,
|
|
99
|
+
then `atlas tests check --root <repo> --halves run`.
|
|
100
|
+
13. Ask the `atlas:verifier` crew to start each environment once and run every scenario.
|
|
101
|
+
14. Plant a fault for each assertion (last section).
|
|
102
|
+
15. Open one pull request. Show the owner the first scenarios, and the guard
|
|
103
|
+
parts of `tests.yaml` as a short list in plain words.
|
|
104
|
+
|
|
105
|
+
Worked run, on a meal planner with a web page, an MCP server and storage in
|
|
106
|
+
a note service:
|
|
107
|
+
|
|
108
|
+
1. Input: the code digest names the routes, the MCP server and the page. It
|
|
109
|
+
finds 11 tests that call functions and none that call a route.
|
|
110
|
+
2. Process: path 2. Draft `tests.yaml` with two environments, three ways in,
|
|
111
|
+
a stand-in for the note service and four personas. Write the stand-in
|
|
112
|
+
and two scenarios.
|
|
113
|
+
3. Output: one pull request. The owner reads two scenarios and three lines of guards.
|
|
114
|
+
|
|
115
|
+
## Write assertions
|
|
116
|
+
|
|
117
|
+
Use these steps when you write or change an assertion. They come from
|
|
118
|
+
sections 9.2 and 6.10 of the design.
|
|
119
|
+
|
|
120
|
+
1. Write the definition of done as assertions. A feature gets full scenarios.
|
|
121
|
+
A small fix gets a short card with its assertions.
|
|
122
|
+
2. Name each assertion that you add, change or remove by its full id in the design.
|
|
123
|
+
3. Write one outcome in each assertion. Then one planted fault points at one assertion.
|
|
124
|
+
4. Mark it `exact` when code checks one right answer. Mark it `judged` only
|
|
125
|
+
when it is about meaning, and say why in the design.
|
|
126
|
+
5. Never let a judged assertion decide an access rule alone.
|
|
127
|
+
6. Mark `gate` on an assertion that later steps need.
|
|
128
|
+
7. Start a "when" assertion with the step that makes it true: "step 3", or a kept result.
|
|
129
|
+
8. Name what it reads in backticks.
|
|
130
|
+
9. Write a "not" assertion so that a blank answer fails it.
|
|
131
|
+
10. Add `repeat` only when the subject varies by design, such as a model. Say why.
|
|
132
|
+
|
|
133
|
+
## Review the words
|
|
134
|
+
|
|
135
|
+
A separate agent reads each new or changed scenario before any code exists.
|
|
136
|
+
It proposes words. It never edits. It asks ten questions of each assertion:
|
|
137
|
+
|
|
138
|
+
| Problem | The question |
|
|
139
|
+
|---|---|
|
|
140
|
+
| two-readings | Can two careful agents write different tests from it? |
|
|
141
|
+
| not-exact | Does an "exact" assertion have one right answer? |
|
|
142
|
+
| many-claims | Does it hold more than one outcome? |
|
|
143
|
+
| duplicate | Does another assertion check the same thing? |
|
|
144
|
+
| weak | How could the product pass it while the claim is false? |
|
|
145
|
+
| hidden-precondition | Does it need a state that `Start with` does not name? |
|
|
146
|
+
| missing-data | Does it need a value that the scenario does not give? |
|
|
147
|
+
| shared-data | Can another run change the data it reads? |
|
|
148
|
+
| empty-when | Does a "when" assertion name a step that makes it true? |
|
|
149
|
+
| loose-not | Does a "not" assertion pass on a blank or broken answer? |
|
|
150
|
+
|
|
151
|
+
Then run `atlas ste` on the scenario. The owner reads these pages.
|
|
152
|
+
|
|
153
|
+
Rules for the words:
|
|
154
|
+
|
|
155
|
+
1. One outcome for each assertion. Mark `exact` or `judged`.
|
|
156
|
+
2. Mark `gate` on an assertion that later steps need.
|
|
157
|
+
3. A wait ends with "Give up after <time>". The link check refuses it otherwise.
|
|
158
|
+
4. A "when" assertion names its step: "step 3", or a result that a step keeps.
|
|
159
|
+
5. Put each name in backticks. A name is a persona, way in, data rule, setup, fixture, table or kept result.
|
|
160
|
+
6. An access rule is always exact. A judge never decides it alone.
|
|
161
|
+
|
|
162
|
+
## Convert a scenario into a test
|
|
163
|
+
|
|
164
|
+
1. Read the scenario, the map, the actions and one worked example. Do not
|
|
165
|
+
read the product source if the map covers the scenario.
|
|
166
|
+
2. Write one test file for each scenario, with `scenario(path, body, { adapter })`
|
|
167
|
+
from `test/atlas/contract.mjs`. The library runs the body once for each way in.
|
|
168
|
+
3. Name each assertion by its full id, as a plain string:
|
|
169
|
+
`run.check('<scenario>/e1#<fingerprint>', () => assert...)`.
|
|
170
|
+
Use `run.gate` for a `gate` assertion. Get the fingerprints from the failure message of `atlas tests check`.
|
|
171
|
+
4. Use an action, or add one. An action for a new call also adds its row to the map.
|
|
172
|
+
5. Make data with `run.fresh(rule)`. Read a case table with `run.table`, `run.cells` or `run.cases`.
|
|
173
|
+
6. Wait with `run.wait` and a limit. Never sleep for a fixed time.
|
|
174
|
+
7. Use `run.notExercised(id, why)` for a "when" assertion that the run never made true.
|
|
175
|
+
8. Mark each guess in a comment. List the guesses in the pull request.
|
|
176
|
+
9. Never change the words to make a test pass. Only the design changes the words.
|
|
177
|
+
10. Run the test before the change. A test for new behaviour must fail.
|
|
178
|
+
|
|
179
|
+
The adapter binds the library to the product. All its parts are optional:
|
|
180
|
+
`standIns`, `warmUp`, `drivers`, `functions`, `loadModule`, `makers`, `signIns`,
|
|
181
|
+
`data` and `actions`. The header of `test/atlas/contract.mjs` describes each one.
|
|
182
|
+
|
|
183
|
+
A result is `pass`, `fail`, `not checked`, `not exercised`, `blocked` or
|
|
184
|
+
`not here`. A test is red when any assertion is not `pass`. The one
|
|
185
|
+
exception is `not here`: the scenario does not run in that environment, it is
|
|
186
|
+
not counted, and the test passes.
|
|
187
|
+
|
|
188
|
+
After the product is ready, a lost connection, a server exit or a time-out is
|
|
189
|
+
`fail`. `blocked` is for a guard refusal, a start-up fault and the first-call check.
|
|
190
|
+
|
|
191
|
+
## Sort a red test
|
|
192
|
+
|
|
193
|
+
Do not edit a red test before you know why it is red. There are four causes:
|
|
194
|
+
|
|
195
|
+
| Cause | Sign | What happens | What may change |
|
|
196
|
+
|---|---|---|---|
|
|
197
|
+
| The product broke | The words still hold. The product does not | Fix the code | Product code only |
|
|
198
|
+
| The behaviour changed on purpose | The approved design names the assertion | Change the assertion and its test | The assertions that the design names, and their tests |
|
|
199
|
+
| The test is out of date | A page part moved, or the form of an answer changed. The verifier's own check passes | Repair the actions, the map or the finders | Actions, map and finders. No assertion and no assert statement |
|
|
200
|
+
| The environment failed | The result is `blocked`: the product did not start, or a stand-in is down | Fix the setup | `tests.yaml` without its guard parts, and the support code |
|
|
201
|
+
|
|
202
|
+
Two checks follow:
|
|
203
|
+
|
|
204
|
+
1. A repair that changes an assertion or an assert statement is not a repair. Refuse it unless the design names the assertion.
|
|
205
|
+
2. A repair must keep the verifier's own check green. The verifier does not use the actions.
|
|
206
|
+
|
|
207
|
+
A `blocked` result because a guard refused, or because the stand-in saw no
|
|
208
|
+
first call from the product, is a finding, not noise.
|
|
209
|
+
Find out why the product reached for something else.
|
|
210
|
+
|
|
211
|
+
## Plant faults
|
|
212
|
+
|
|
213
|
+
A fault proves that an assertion can go red. Plant one for each new or
|
|
214
|
+
changed assertion, and for each assertion at onboarding.
|
|
215
|
+
|
|
216
|
+
1. Commit first. A clone holds committed files only, so a change that is not
|
|
217
|
+
committed is not in the clone. Then clone the worktree to a throwaway
|
|
218
|
+
folder outside it, with `git clone --no-hardlinks`.
|
|
219
|
+
2. Remove every remote from the clone. Run `git remote` there. It must print nothing.
|
|
220
|
+
3. Plant faults in product code only. Never touch `tests.yaml`, the support code or the kit.
|
|
221
|
+
4. Use the cheap model first: Haiku. After two failed tries on one assertion, use Sonnet.
|
|
222
|
+
5. Give the planter one assertion and the product code. Ask for one fault that makes only that claim false.
|
|
223
|
+
6. Run the scenario tests from the real worktree. Set `ATLAS_PRODUCT_ROOT` to
|
|
224
|
+
the area folder in the clone, and `ATLAS_FAULT_RUN` to a label. The product,
|
|
225
|
+
the MCP server and the module of the `function` driver start in the clone.
|
|
226
|
+
The guards, the kit and `tests.yaml` load from the worktree.
|
|
227
|
+
7. The test for that assertion must go red.
|
|
228
|
+
8. Ask the `atlas:verifier` crew for its own check of that assertion, on the same clone. It uses a raw driver, not the actions.
|
|
229
|
+
9. A fault is valid only when the verifier's own check also fails. Discard any other fault and try again.
|
|
230
|
+
10. Report each assertion as caught, missed, or "no fault of its own" after Sonnet also fails.
|
|
231
|
+
11. Delete the clone. Run `git status` in the worktree. It must show no change from the faults.
|
|
232
|
+
|
|
233
|
+
Evidence from a fault run carries the label. The link check ignores it, so it never counts as proof.
|
|
234
|
+
Never push from the clone.
|