@aotter/mantle 0.0.11-alpha.63 → 0.0.11-alpha.65

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/README.md +87 -12
  2. package/dist/cli.d.ts +3 -0
  3. package/dist/cli.d.ts.map +1 -0
  4. package/dist/cli.js +52 -0
  5. package/dist/cli.js.map +1 -0
  6. package/dist/generate.d.ts +2 -0
  7. package/dist/generate.d.ts.map +1 -0
  8. package/dist/generate.js +181 -0
  9. package/dist/generate.js.map +1 -0
  10. package/dist/skills.d.ts +2 -0
  11. package/dist/skills.d.ts.map +1 -0
  12. package/dist/skills.js +80 -0
  13. package/dist/skills.js.map +1 -0
  14. package/dist/update.d.ts +2 -0
  15. package/dist/update.d.ts.map +1 -0
  16. package/dist/update.js +387 -0
  17. package/dist/update.js.map +1 -0
  18. package/docs/adr/0001-four-atom-manifest-model.md +6 -7
  19. package/docs/adr/0007-ai-as-primary-author.md +100 -138
  20. package/docs/adr/0008-structured-diagnostic-shape.md +79 -99
  21. package/docs/adr/0009-consumer-supplied-manifests.md +101 -228
  22. package/docs/adr/0012-views-as-public-rest.md +43 -15
  23. package/docs/adr/0014-auth-better-auth-and-multi-tenant-mcp.md +43 -17
  24. package/docs/adr/0018-core-starters-repository-boundary.md +155 -0
  25. package/docs/adr/README.md +8 -6
  26. package/docs/cloudflare-low-level-composition.md +94 -0
  27. package/docs/design-atoms.md +59 -57
  28. package/docs/design-references/editorial-blog-2026-05-05.md +7 -7
  29. package/docs/labels.md +1 -1
  30. package/docs/media-uploads.md +1 -1
  31. package/docs/release-process.md +156 -523
  32. package/package.json +9 -6
  33. package/skills/README.md +20 -16
  34. package/skills/develop/SKILL.md +4 -4
  35. package/skills/install/SKILL.md +16 -2
  36. package/skills/plugin/SKILL.md +1 -1
  37. package/skills/provision/SKILL.md +1 -1
  38. package/skills/theme/SKILL.md +12 -10
  39. package/skills/update/SKILL.md +31 -19
  40. package/skills/customize-design/SKILL.md +0 -215
  41. package/skills/extend/SKILL.md +0 -257
@@ -1,8 +1,9 @@
1
- # ADR-0007: AI-as-primary-author — three feedback loops, two role surfaces
1
+ # ADR-0007: AI-as-primary-author — three pre-serve loops, two role surfaces
2
2
 
3
- **Status:** Carried over from POC v0.0.x; refreshed and folded for v0.1.0 (incorporates POC ADR-0013).
3
+ **Status:** Carried over from POC v0.0.x; amended for the shipped v0.1 CLI,
4
+ testing, plugin, and MCP surfaces (incorporates POC ADR-0013).
4
5
 
5
- **Date**: 2026-05-03
6
+ **Date:** 2026-05-03; last amended 2026-08-03
6
7
 
7
8
  **Deciders**: phsu
8
9
 
@@ -74,11 +75,12 @@ client like Claude Desktop, talking to the deployed Worker over MCP).
74
75
 
75
76
  ## Decision
76
77
 
77
- ### Part A — Three feedback loops for the coder agent
78
+ ### Part A — Three pre-serve feedback loops for the coder agent
78
79
 
79
- The SDK provides **three explicit feedback loops**, ordered by
80
- "leftward shift" — catch the failure as early in the author's
81
- workflow as physically possible:
80
+ The SDK provides **three explicit pre-serve feedback loops**, ordered by
81
+ "leftward shift" — catch the failure as early in the author's workflow as
82
+ physically possible. Runtime diagnostics remain the final, serving-time
83
+ boundary rather than a fourth authoring gate.
82
84
 
83
85
  #### Loop 1 — Static validation (`mantle validate`)
84
86
 
@@ -87,8 +89,9 @@ network. Runs in the AI's terminal in milliseconds. Reads YAML files
87
89
  and handler registration source files; emits structured diagnostics;
88
90
  exits non-zero on any error.
89
91
 
90
- The CLI ships in the `@aotter/mantle-spec` package — it's the
91
- spec authority for what a v0.1 manifest must look like.
92
+ The implementation lives in `@aotter/mantle-spec`. Adopters run it through the
93
+ umbrella package's `mantle` binary; a direct spec install exposes the fallback
94
+ `mantle-spec` binary.
92
95
 
93
96
  What it catches:
94
97
  - Manifest envelope (`apiVersion`, `kind`, `metadata.name`)
@@ -101,61 +104,32 @@ What it catches:
101
104
  - `requires.auth.all` predicates are in the v0.1 vocabulary
102
105
  - `Trigger.source.path` does not collide with another Trigger
103
106
  - `metadata.name` is unique within `kind`
104
- - Each `Procedure.handler.ref` has a corresponding
105
- `sdk.registerHandler("<ref>", ...)` call somewhere in the
106
- consumer's TS source (textual grep, not runtime — boot loop is
107
- Loop 2's job)
107
+ - Each `Procedure.handler.ref` appears as a key in the consumer's TypeScript
108
+ handler map (textual grep, not runtime — boot Loop 3 verifies the actual
109
+ map)
108
110
 
109
111
  This is the cheapest loop; it must be runnable without booting
110
112
  anything. AI authors should run it after every manifest edit.
111
113
 
112
- #### Test harness (planned public package surface)
114
+ #### Loop 2 — Local tests and measured harnesses
113
115
 
114
- Supporting tooling that consumers use inside their own test suites —
115
- it is not itself a feedback loop, it's the harness those tests run
116
- against. In-memory dispatcher invocation: the consumer's test suite
117
- imports a public testing module that spins up a complete dispatcher
118
- against an in-memory D1 substitute, seeds rows, and calls Procedures
119
- or Views directly without going through HTTP.
116
+ Consumer tests exercise exported runtime use cases or the adapter application
117
+ with project-owned fakes appropriate to that test. Core does not ship a second
118
+ in-memory Worker or dispatcher abstraction.
120
119
 
121
- Target API sketch:
120
+ The public Node-only `@aotter/mantle/runtime/testing` surface instead provides
121
+ the two shared measurements that are expensive for each consumer to rebuild:
122
122
 
123
- ```ts
124
- import { createTestDispatcher } from "@aotter/mantle-runtime/testing";
125
- import { manifests } from "../manifests"; // loaded by build hook
126
- import { handlers } from "../handlers"; // map of ref → fn
123
+ - crowded real-SQLite View execution and `EXPLAIN QUERY PLAN` index coverage;
124
+ - HTTP sampling for a running Worker, including optional test-only query/row
125
+ metric headers.
127
126
 
128
- const dispatcher = await createTestDispatcher({ manifests, handlers });
129
- await dispatcher.seed({ collection: "posts", rows: [...] });
127
+ The umbrella package exposes both through `mantle-harness`. These checks use the
128
+ real manifest validator, migrations, View compiler, SQL, and HTTP surface; they
129
+ do not pretend an in-memory adapter proves Cloudflare behavior. Project tests
130
+ still own handler logic, auth combinations, and product-specific fixtures.
130
131
 
131
- const result = await dispatcher.invoke(
132
- "send-contact-message",
133
- { name: "Alex", message: "hi" },
134
- { user: { id: "user-uuid" } },
135
- );
136
- expect(result).toEqual({ ok: true });
137
-
138
- const view = await dispatcher.queryView("recent-published");
139
- expect(view.items).toHaveLength(2);
140
- ```
141
-
142
- Catches:
143
- - Handler logic bugs (validation rules, branching, error returns)
144
- - Auth predicate evaluation (handler called with wrong user/staff
145
- context returns `AUTH_DENIED`)
146
- - Input/output schema enforcement (caller-supplied `x-mantle-bind` field
147
- is rejected; handler-returned shape mismatching `output` is
148
- rejected)
149
- - View result correctness against fixture data
150
- - Cross-Trigger correctness when one Procedure has multiple bindings
151
-
152
- The test harness is **not** mocked — it should share the real
153
- dispatcher, validator, handler registry, and in-memory runtime ports.
154
- The public `testing` export has not shipped yet; current v0.1.0 code
155
- uses runtime package tests plus starter integration smokes to cover
156
- the same contract until the consumer-facing harness is promoted.
157
-
158
- #### Loop 2 — Boot-time fail-fast
132
+ #### Loop 3 — Boot-time fail-fast
159
133
 
160
134
  Dispatcher build phase, after manifests are parsed and handlers are
161
135
  registered, before the serve loop starts. Walks the entire manifest
@@ -163,38 +137,37 @@ graph and the in-memory handler registry; refuses to start serving if
163
137
  anything is missing or inconsistent.
164
138
 
165
139
  What it catches that Loop 1 cannot:
166
- - `Procedure.handler.ref` that has a `sdk.registerHandler(...)` call
167
- in source (Loop 1 saw it textually) but did not actually execute
168
- (e.g. the registration file wasn't imported, the boot function was
169
- skipped behind a feature flag)
170
- - D1 schema drift — table state in the connected D1 does not match
171
- the DDL derived from current manifests; SDK refuses to serve until
172
- a migration is applied or the drift is acknowledged
140
+ - `Procedure.handler.ref` that appeared textually in source (Loop 1 saw it) but
141
+ is absent from the actual `handlers` map passed to runtime assembly
142
+ - Current site locale state that conflicts with localized/translation Schemas
143
+ - Generated index state that cannot be reconciled with declared Schema indexes
173
144
  - Cross-manifest references that a recent deploy introduced without
174
145
  the corresponding atom (e.g. updated Trigger pointed at a Procedure
175
146
  that wasn't included in the deploy bundle)
176
147
 
177
- Failure mode is **process exit non-zero**, not "log a warning and
178
- serve anyway." A missing handler must not surface as a runtime 500 to
179
- a customer; it must surface as a boot failure to the deploy pipeline,
180
- which is visible to the AI author who just shipped.
148
+ Failure mode is a rejected `bootInit()` promise carrying structured boot
149
+ diagnostics, not "log a warning and serve anyway." The conventional adapter
150
+ does not dispatch through an unbooted runtime and resets a rejected lazy-boot
151
+ promise so a transient infrastructure failure can retry. Deployment workflows
152
+ must exercise the Worker readiness/smoke path if they need the failure before
153
+ traffic.
181
154
 
182
- #### Loop 3 — Runtime errors
155
+ #### Runtime diagnostics
183
156
 
184
157
  Runtime diagnostics (see ADR-0008 and the design-atoms reference)
185
158
  cover the cases that reach this far. The contract's job is to make
186
- Loop 3 the layer that handles **input it had no way to predict**
159
+ the serving boundary handle **input it had no way to predict**
187
160
  (request shape, auth state, transient infrastructure failures), not
188
- the layer that handles **author-side mistakes** (those are Loops 1
189
- and 2's job).
161
+ the layer that handles **author-side mistakes** (those belong in the earliest
162
+ applicable pre-serve loop).
190
163
 
191
164
  #### Phase ordering rationale
192
165
 
193
166
  The author's lifecycle is:
194
167
 
195
168
  ```
196
- edit YAML → run validate (Loop 1) → run tests (test harness)
197
- → push → deploy boots (Loop 2) → serves (Loop 3)
169
+ edit YAML → run validate (Loop 1) → run tests/harnesses (Loop 2)
170
+ → push → deploy boots (Loop 3) → serves (runtime diagnostics)
198
171
  ```
199
172
 
200
173
  The earlier a class of error is caught, the cheaper it is to fix and
@@ -225,13 +198,11 @@ Two AI agents read and write to a `@aotter/mantle-*` deployment:
225
198
  and network access in the consumer's dev environment.
226
199
 
227
200
  2. **The operator agent** — a content-editing client (Claude Desktop,
228
- claude.ai, or any MCP-speaking host) connected to the deployed
229
- Worker's `/mcp` endpoint over OAuth. Calls the Day-1 MCP tool
230
- catalog: generic tools (`list_entries`, `get_entry`,
231
- `request_publish`, `unpublish_entry`, `archive_entry`) plus
232
- per-collection authoring tools (`create_draft_<collection>`,
233
- `update_draft_<collection>`). Has *only* the MCP tools the Worker
234
- exposes; no shell, no filesystem, no git.
201
+ claude.ai, or any MCP-speaking host) connected to the deployed Worker over
202
+ OAuth. `/mcp` carries public View queries and explicitly public MCP
203
+ Triggers; `/mcp/staff` carries staff View queries, authoring/lifecycle
204
+ builtins, and explicitly staff MCP Triggers. The agent has *only* the tools
205
+ on its authorized catalog; no shell, filesystem, or git.
235
206
 
236
207
  These are different roles with different *capability surfaces*. The
237
208
  coder agent has many more capabilities than the operator agent *by
@@ -266,25 +237,24 @@ SDK TS API instead of needing a matching MCP tool).
266
237
  operator never sees the manifest layer.
267
238
  - Anything that requires **shell or filesystem access** —
268
239
  scaffolding, test running, deploy.
269
- - Anything **dev-loop shaped** — the three feedback loops above
270
- (validate / boot / runtime) plus the test harness are
271
- coder-surface concerns, not operator-surface.
240
+ - Anything **dev-loop shaped** — static validation, local
241
+ tests/measurements, and boot verification are coder-surface concerns, not
242
+ operator-surface.
272
243
  - Anything that **mutates code** in the consumer's repo — registering
273
244
  handlers, wiring imports, updating `wrangler.toml`.
274
245
  - **Skill files** carrying domain context. cc reads these; MCP clients
275
246
  do not.
276
247
 
277
248
  Concrete artifacts today:
278
- - `skills/<name>/SKILL.md` files in the
279
- [`aotter/mantle`](https://github.com/aotter/mantle)
280
- repo, discoverable by URL. Distribution as a Claude plugin is
281
- **optional** (a v0.1.x convenience), not required — the canonical
282
- location is the in-repo path, and any agent that can fetch a URL
283
- can consume them.
284
- - `mantle validate` CLI (shipped in `@aotter/mantle-spec`)
285
- - `mantle emit-openapi` CLI (likewise)
286
- - Public testing harness (planned; current coverage is SDK tests +
287
- starter integration smokes)
249
+
250
+ - Version-matched Core skills ship inside `@aotter/mantle`, through the Mantle
251
+ agent plugin, and through exact-byte repo-local projections written by
252
+ `mantle skills`.
253
+ - The umbrella `mantle` CLI exposes `validate`, `generate`, `skills`, `update`,
254
+ `introspect`, `emit-openapi`, and `emit-types`; direct spec installs expose
255
+ the authoring subset through `mantle-spec`.
256
+ - `@aotter/mantle/runtime/testing` and `mantle-harness` expose the measured
257
+ SQLite/index and live-HTTP checks described in Loop 2.
288
258
 
289
259
  #### What goes on MCP (operator surface)
290
260
 
@@ -301,20 +271,18 @@ Concrete artifacts today:
301
271
 
302
272
  #### What goes on neither (the SDK TS API only)
303
273
 
304
- A small third category exists: APIs that only the **runtime itself**
305
- consumes (the dispatcher, the boot validator, the diagnostic types).
306
- These are exported from `@aotter/mantle-spec` for the coder to
307
- import inside their handler code, but are not CLI commands and not
308
- MCP tools — they're library types.
274
+ A small third category exists: library APIs such as manifest/diagnostic types,
275
+ runtime use cases, dispatchers, and the boot validator. They are exported from
276
+ the matching `@aotter/mantle` subpath for composition or tests, but are not
277
+ separate CLI commands or MCP tools.
309
278
 
310
279
  #### What does NOT go on either surface
311
280
 
312
- - **Consumer Procedures auto-emitted as MCP tools.** This was
313
- considered as `Trigger.source.kind: mcp` and rejected. A
314
- consumer-defined Procedure carries arbitrary code; auto-emitting it
315
- to MCP would let the operator agent trigger arbitrary side effects
316
- defined by the coder agent. The role boundary collapses. Ops-role
317
- verbs land as SDK builtins, not as auto-bound consumer Procedures.
281
+ - **Unbound consumer Procedures auto-emitted as MCP tools.** A Procedure carries
282
+ arbitrary code and is transport-neutral by itself. Exposure requires an
283
+ explicit `Trigger.source.kind: mcp`, a declared `public | staff` surface, and
284
+ the target's authorization predicates/guard. Core also emits its closed
285
+ authoring/lifecycle builtins and View query tools on the matching surface.
318
286
  - **CLI commands that mutate runtime state.** The CLI is for the
319
287
  pre-deploy authoring loop. Once deployed, runtime state changes via
320
288
  MCP (operator) or via direct SDK calls inside handler code
@@ -331,11 +299,11 @@ MCP tools — they're library types.
331
299
  just executed. Boot errors → deploy pipeline log the deploy is
332
300
  blocked on. Runtime errors → customer-visible failures (the only
333
301
  place runtime errors should appear).
334
- - One diagnostic format (ADR-0008) across all loops means the coder
335
- agent writes one parser, not three.
336
- - Test harness as a public SDK export means consumer apps can ship
337
- handler tests as part of normal CI without wiring up their own
338
- dispatcher harness.
302
+ - One diagnostic format (ADR-0008) across validation, boot, and runtime
303
+ failures means the coder agent writes one parser. Measured harnesses use
304
+ separate stable report types because they are results, not failures.
305
+ - Public measured harnesses let consumers gate index access paths and live HTTP
306
+ behavior without rebuilding Core's migration/compiler instrumentation.
339
307
  - Boot fail-fast turns a class of "runtime 500 reaches customer"
340
308
  failures into "deploy blocked, fix forward" failures.
341
309
  - The framing — coder agent as primary author, operator agent as
@@ -356,11 +324,8 @@ MCP tools — they're library types.
356
324
  - v0.1 ships more than just a runtime: also the `mantle validate`
357
325
  CLI, a public testing module, and a boot validator. Roughly +30%
358
326
  of the v0.1 dispatcher work.
359
- - Test harness needs an in-memory D1 substitute (most plausible:
360
- `better-sqlite3` or an equivalent SQLite binding running the same
361
- DDL the production D1 adapter applies). Adds a dev-dependency to
362
- the SDK and a compatibility surface to keep honest against real D1
363
- quirks.
327
+ - The Node-only planner uses the platform's `node:sqlite`; it is a compatibility
328
+ surface to keep honest against real D1 and does not enter Worker bundles.
364
329
  - Diagnostic format is locked early (ADR-0008); any field-shape
365
330
  change forces a doc revise and a code change across all loops.
366
331
  - The error catalog (one named code per failure mode, per loop) must
@@ -372,9 +337,8 @@ MCP tools — they're library types.
372
337
  versioning, and contracts apply to both. Mitigated by the fact
373
338
  that the two surfaces have little overlap by design — most features
374
339
  touch only one.
375
- - No cross-surface convenience. A hypothetical "consumer defines a
376
- Procedure that's also an MCP tool with one YAML block" would be
377
- ergonomically nice but is structurally rejected by Part B.
340
+ - No implicit cross-surface convenience. A consumer Procedure reaches MCP only
341
+ through an explicit MCP Trigger and its surface/authorization declaration.
378
342
 
379
343
  ### Risks
380
344
 
@@ -383,20 +347,18 @@ MCP tools — they're library types.
383
347
  Mitigation: validator MUST err on the permissive side; ambiguous
384
348
  cases emit `severity: warning` not `error`; exit code 0 unless an
385
349
  error is found.
386
- - **Test harness diverges from runtime**: in-memory SQLite has
387
- slightly different SQL semantics (collation, type affinity, JSON1
388
- edge cases) from production D1. Tests pass here, prod fails. Same
389
- trap as mocking the database. Mitigation: SDK itself runs starter
390
- integration tests against real D1 in CI; document the
391
- known-divergence list in the testing module's README.
350
+ - **SQLite planner diverges from D1**: local SQLite can differ in collation,
351
+ type affinity, JSON, or planner choices. Mitigation: Core CI also runs the
352
+ Wrangler-local Worker benchmark; local planner output gates query shape and
353
+ access paths, not provider equivalence.
392
354
  - **Boot fail-fast aggressiveness**: if boot refuses to start on
393
355
  cosmetic issues (a Schema property has an unrecognized but harmless
394
356
  `x-` extension), the deploy pipeline becomes hostile. Mitigation:
395
357
  boot validates only **load-bearing** invariants (missing handler
396
358
  refs, dangling Trigger targets, schema drift). Cosmetic / advisory
397
- checks are warnings via `--strict` flag, not errors.
359
+ checks remain non-blocking diagnostics rather than boot errors.
398
360
  - **AI authors over-trust the contract**: "the SDK said it was fine
399
- therefore it is fine." Loop 3 (runtime) still exists for reasons
361
+ therefore it is fine." Runtime diagnostics still exist for reasons
400
362
  the contract cannot prevent. Mitigation: docs make clear that
401
363
  runtime errors are still possible and that the contract reduces —
402
364
  not eliminates — surface area.
@@ -425,21 +387,20 @@ UI for that author. Failures must land in the loop the author is
425
387
  currently in.
426
388
 
427
389
  **(b) Static validation only.**
428
- Rejected as insufficient: cannot catch handler logic bugs, cannot
429
- verify that a `registerHandler` call actually executes (only that it
430
- appears textually), cannot test data-flow against fixtures.
390
+ Rejected as insufficient: cannot catch handler logic bugs, cannot verify that
391
+ a textual handler key reaches the assembled runtime map, and cannot test
392
+ data-flow against fixtures.
431
393
 
432
394
  **(c) Test harness only.**
433
395
  Rejected: better than (a) but still lets misconfigurations land in
434
396
  deploys (handler renamed but not re-registered, manifest references
435
397
  stale schema name). Static is cheaper and finds these faster.
436
398
 
437
- **(d) Skip the test harness; rely on consumer's own integration
438
- tests.**
439
- Rejected: it is the SDK's job to make handler-level testing trivial;
440
- pushing the dispatcher-spin-up burden to every consumer is exactly
441
- the kind of boilerplate that demotivates writing tests in the first
442
- place.
399
+ **(d) Ship a second full in-memory Worker/dispatcher harness.**
400
+ Deferred. Consumer tests can call the public use cases or adapter app directly;
401
+ Core only centralizes the real-SQLite planner and live-HTTP measurements whose
402
+ instrumentation would otherwise be duplicated. Add a fuller factory only when
403
+ multiple consumers demonstrate the same unavoidable setup.
443
404
 
444
405
  **(e) Prose error messages.**
445
406
  Rejected: see ADR-0008. AI parseability outweighs human readability
@@ -479,16 +440,16 @@ When proposing a new SDK feature or behavior, declare:
479
440
  CLI/skills, operator never sees it). Real "both" cases should
480
441
  be vanishingly rare.
481
442
 
482
- 2. **Which loop owns the failure mode?** (Loop 1 static, Loop 2
483
- boot, Loop 3 runtime)
443
+ 2. **Which loop owns the failure mode?** (Loop 1 static, Loop 2 local
444
+ tests/harnesses, Loop 3 boot; unpredictable serving conditions stay runtime)
484
445
 
485
446
  3. **Could it move left?** (e.g. a runtime check that could be a
486
447
  boot check, a boot check that could be a static check)
487
448
 
488
449
  4. **Does it emit the structured diagnostic shape?** (ADR-0008)
489
450
 
490
- 5. **Can a consumer reproduce the failure offline?** (i.e. is it
491
- visible to Loop 1 or to the test harness, not only Loop 3)
451
+ 5. **Can a consumer reproduce the failure offline?** (i.e. is it visible to
452
+ Loop 1 or Loop 2, not only after boot or at runtime)
492
453
 
493
454
  6. **Does the feature mutate the consumer's repo?** → CLI (or SDK
494
455
  API the CLI uses). Never MCP — the operator agent is not in the
@@ -498,7 +459,8 @@ When proposing a new SDK feature or behavior, declare:
498
459
  Worker?** → MCP tool, with the ops-role description in the
499
460
  tool's `description` field. Or HTTP Trigger + handler if the
500
461
  feature is consumer-shaped (a Procedure the consumer wrote).
501
- Don't auto-bridge consumer Procedures to MCP.
462
+ An MCP Trigger is the explicit bridge when a consumer intentionally exposes
463
+ a Procedure to an operator surface.
502
464
 
503
465
  Reviewers should treat "this lands as a runtime 500 only" as a
504
466
  yellow flag — sometimes correct (genuine runtime conditions), often