@marver-design/marver 0.21.0 → 0.22.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/CHANGELOG.md +134 -15
  2. package/README.md +17 -4
  3. package/dist/{bake-BID6mo-N.mjs → bake-CvY8QR3L.mjs} +1 -1
  4. package/dist/board-status-CWGHIdo_.mjs +1307 -0
  5. package/dist/boards-BwKI0qWd.mjs +188 -0
  6. package/dist/{build-C7MqQ7hq.mjs → build-B1kavcpc.mjs} +74 -13
  7. package/dist/cli.mjs +69 -9
  8. package/dist/config-DJxMRVD8.mjs +373 -0
  9. package/dist/context-D4t59wDi.mjs +600 -0
  10. package/dist/{daemon-CRZFpl6K.mjs → daemon-PmNqoOAk.mjs} +1 -1
  11. package/dist/{dev-BNZF4Mup.mjs → dev-Cf4wLGe2.mjs} +15 -7
  12. package/dist/{init-C34BY3R4.mjs → init-C_K-cjBQ.mjs} +41 -45
  13. package/dist/managed-write-Bo-oPc-i.mjs +71 -0
  14. package/dist/{manifest-mMfUhPtL.mjs → manifest-CcdWx7ud.mjs} +42 -376
  15. package/dist/{plugin-omHLCn91.mjs → plugin-C5u07Yhv.mjs} +234 -105
  16. package/dist/{poster-BvxiAzy1.mjs → poster-BduzQBYz.mjs} +1 -1
  17. package/dist/{publish-bakes-BqzAAa3w.mjs → publish-bakes-b0jUwvM_.mjs} +2 -2
  18. package/dist/{shot-DswS4iRK.mjs → shot-BAR8hmU9.mjs} +2 -2
  19. package/docs/boards-and-folders.md +161 -0
  20. package/docs/context.md +117 -0
  21. package/docs/sharing.md +12 -2
  22. package/package.json +1 -1
  23. package/src/client/shell/BoardList.tsx +33 -3
  24. package/src/client/shell/ContextMenu.tsx +16 -5
  25. package/src/client/shell/StatusPicker.tsx +91 -0
  26. package/src/client/shell/board-icons.tsx +75 -0
  27. package/src/client/shell/store.ts +88 -9
  28. package/src/client/shell/styles.css +34 -3
  29. package/src/shared/board-tree.ts +30 -13
  30. package/src/shared/board-types.ts +103 -0
  31. package/src/shared/context.ts +265 -0
  32. package/src/shared/status.ts +157 -0
  33. package/templates/AGENTS-embedded.md +12 -0
  34. package/templates/AGENTS-studio.md +12 -0
  35. package/templates/context/INDEX.md +50 -0
  36. package/templates/context/map.json +6 -0
  37. package/templates/context/shipped-knowledge.md +19 -0
  38. package/templates/context/shipped.md +28 -0
  39. package/templates/instructions/boards.md +33 -0
  40. package/templates/instructions/context.md +123 -0
  41. package/templates/playbooks/publish-canvas/PLAYBOOK.md +72 -0
  42. package/templates/playbooks/reorganize-context/PLAYBOOK.md +232 -0
  43. package/templates/playbooks/reorganize-context/eval.md +93 -0
  44. package/dist/boards-BwiDAmPf.mjs +0 -337
  45. package/dist/boards-DnLewfj8.mjs +0 -71
@@ -0,0 +1,123 @@
1
+ # Context - what the product is, kept true
2
+
3
+ The project's knowledge lives in `context/`, beside `design/`. Code says what is implemented;
4
+ `context/` says what is available, how it works, why, and what is next - with evidence. The canvas
5
+ reads the same files: every feature and project board wears the status they record.
6
+
7
+ ## At the start of a session
8
+
9
+ Read `context/INDEX.md`, then at most two more files per question. The index routes; it never
10
+ explains. If there is no `context/`, see "Setting up" below before you invent answers.
11
+
12
+ ## The files
13
+
14
+ | Path | What it holds |
15
+ |---|---|
16
+ | `INDEX.md` | routing only, under 800 words; its capability table is generated from `map.json` |
17
+ | `shipped.md` | **the only place availability is written** - services by environment, capabilities by who has them; for knowledge work, what was delivered to whom |
18
+ | `product/<capability>.md` | one current contract per capability |
19
+ | `map.json` | which code each capability lives in - the pull-request rule reads it |
20
+ | `playbooks/<name>/PLAYBOOK.md` | how we do things - steps with checks, traps, a run log |
21
+ | `feedback/` | the inbox: one row per piece of feedback, with its state |
22
+ | `plans/` | proposals in flight - a plan folds into its contract when it lands |
23
+
24
+ ## Evidence has levels
25
+
26
+ Every availability claim in `shipped.md` carries one, with a citation:
27
+ - `confirmed` - observed in a system of record: the deploy job that ran for that service and
28
+ succeeded, a ci run that executed the suite. Not an environment marked "deployed" - some
29
+ pipelines mark it when the run stood down.
30
+ - `reported` - a written claim, cited: a changelog line, an ops note.
31
+ - `unknown` - neither, with the reason. An honest unknown is a correct answer.
32
+
33
+ Never upgrade a level because a claim sounds sure. Commit ancestry says a change is inside a
34
+ revision, not that the revision is running.
35
+
36
+ ## A contract
37
+
38
+ ```yaml
39
+ ---
40
+ state: current # current | proposed | historical
41
+ capability: quote-intake # the same slug as its feature board
42
+ audience: team # team | publishable | restricted
43
+ reviewed: { revision: 606ad4e, scope: "apps/quote, db schema", method: "read code and tests" }
44
+ ---
45
+ ```
46
+
47
+ Then: what it does and its limits; who can use it and when; how it works; what it touches (frontend,
48
+ backend, database, workers, integrations); invariants; code and tests; decisions; its frames
49
+ (`scene/frame` ids - say which are designed but not built); the PRs that built it; what is still
50
+ proposed; and `## Availability` pointing at its row in `shipped.md`.
51
+
52
+ - **Read the code to write it.** Cite `path:line` for every behaviour; a test file's existence is
53
+ not a run.
54
+ - **No status line, no availability** ("live", "on production", "since 30 Sep") - that is
55
+ `shipped.md`'s alone. The check fails on both.
56
+ - **One answer per rule.** When a plan amends a rule, write the amended rule - never "section 20
57
+ wins where it differs".
58
+ - **When code and an accepted document disagree,** the contract says what the code does and the
59
+ conflict goes to the human. Never decide what was intended.
60
+
61
+ ## When you change things
62
+
63
+ - **Behaviour** - update its contract in the same change. A change that leaves behaviour alone says
64
+ `no-contract-change: <capability> - <why>` in the pull request body.
65
+ - **A deploy** - write its row in `shipped.md` with the run that proves it (`playbooks/` usually
66
+ has the release recipe).
67
+ - **A decision** - in the project's decision log, numbered, with what it applies to.
68
+ - **Feedback** - one row per item: source, who, when, what they said, theme, state (`new`,
69
+ `triaged`, `proposed`, `planned`, then `shipped` or `declined`), and what resolved it. An item is
70
+ `shipped` only when `shipped.md` shows its fix `confirmed` where it was raised; a `reported`
71
+ deploy leaves it `planned`, saying what is missing. On closing, draft - never send - a note back.
72
+
73
+ ## Status on the canvas
74
+
75
+ Feature and project boards wear a status read from these files (the first that matches wins):
76
+ the board's own `"status"` - `archived`, `paused`, or `blocked` with a `"reason"`; Unknown when the
77
+ evidence cannot be read; In progress when an open plan in `plans/` names the capability (filling by
78
+ phase: `<cap>-specs`, `<cap>-lofi`, `<cap>`); Done when `shipped.md` shows it `confirmed` in an
79
+ availability clause (one per `;`) with no negation or pending word ("nowhere", "not yet", "rolled
80
+ back", "planned") and no pre-production place (staging, preview, dev, local, testing) unless it also
81
+ names production; Done, reported when only `reported`; To do when its contract is `state: current`; else Backlog.
82
+ **Done is never set by hand** - `"status": "done"` fails the check. A board finds its capability by
83
+ its own name, or `"capability": "<slug>"` in its JSON.
84
+
85
+ ## Audiences
86
+
87
+ `team` - everyone with the repository (the default); `publishable` - may appear on a published
88
+ canvas; `restricted` - never tracked by git (keep it under a gitignored `context/private/`). A
89
+ frame that renders a `team` file onto a published board fails the check. A digest of restricted
90
+ material is restricted, however it is labelled.
91
+
92
+ ## Setting up - when there is no `context/`
93
+
94
+ Context starts with people. Ask, in one message:
95
+ 1. Who is involved - the team and their roles; clients, users, partners?
96
+ 2. Where does each relationship happen - channels, email, recorded meetings, docs, in-product
97
+ feedback, canvas comments?
98
+ 3. What is private, and who may see what?
99
+ 4. How is this project named in shared spaces - keywords, tags?
100
+
101
+ Then, in order - this is the whole recipe; "set up our context" asks for all of it:
102
+ 1. `npx marver context init --kind product` - or `--kind knowledge` when the work is deliverables
103
+ for clients rather than a product people use. Record the answers in `context/people.md`.
104
+ 2. Draft from evidence: the shipped record from deploy runs and the changelog (`unknown` where
105
+ nothing proves it), a map from the code, contracts for the capabilities that matter now.
106
+ 3. Let the canvas read it: `npx marver init --kind <the same kind>` adds the typed folders (it never
107
+ moves a board); put each feature board in Features (a project in Projects), named after its
108
+ capability or saying `"capability": "<slug>"`; then `npx marver boards new home --type start
109
+ --folder start-here`.
110
+ 4. Put the check in ci - a step running `npx marver context check` on a full-history checkout
111
+ (`fetch-depth: 0`) - and make it pass before you hand over.
112
+
113
+ **Context that exists but is scattered** - specs out of date, nobody sure what is live - is a job
114
+ for `context/playbooks/reorganize-context/PLAYBOOK.md`, not for a rewrite.
115
+
116
+ ## The check
117
+
118
+ `npx marver context check` - availability only in `shipped.md`, a level and a citation on every
119
+ evidence cell, citations and links that resolve, the index under budget and equal to the map,
120
+ audiences, feedback and contract states, playbook freshness; with `--base <ref>` (or in ci on a pull
121
+ request) a contract change or a stated reason for every mapped change. Exit 0 pass, 1 fail, 2 cannot
122
+ determine (fetch full history - `fetch-depth: 0`). `npx marver context index` rewrites the index's
123
+ generated table after you change the map.
@@ -0,0 +1,72 @@
1
+ ---
2
+ name: publish-canvas
3
+ description: Use this when the canvas should be on a URL - a first publish, a republish after changes, or moving it to a new host. Steps with checks, the traps that have bitten, how to roll back.
4
+ last_success: none
5
+ depends_on:
6
+ - design/publish.json
7
+ - design/config.ts
8
+ - Dockerfile
9
+ - package.json
10
+ - "service: the canvas host (Railway or any volume-capable host)"
11
+ - "account: the host, and Marver Sign In or the canvas password"
12
+ ---
13
+
14
+ # Publish a canvas
15
+
16
+ The recipe. `design/instructions/publish.md` is the reference - every variable, both gates, the
17
+ host contract; read it when a step here is not enough. Marver maintains this file; record your
18
+ project's values (which account, which domain, which gate) and every run in the log at the end.
19
+
20
+ ## Prerequisites
21
+
22
+ - `@marver-design/marver` resolves from the registry in `package.json` - a `link:` or `file:`
23
+ dependency cannot ride to a remote host.
24
+ - The three decisions are made: **which boards ship** (`design/publish.json` - publishing is
25
+ default-closed), **who gets in** (Marver Sign In by default; the canvas password when nothing
26
+ external may be in the sign-in path), **collaboration** (a persistent volume for comments, or none).
27
+ - The canvas has a name: `share: { name: "..." }` in `design/config.ts`.
28
+
29
+ ## Steps
30
+
31
+ 1. **Build locally first.** `npx marver build`.
32
+ *Check:* it lists exactly the boards you meant, `textures:` reports the hi-fi frames asleep, and
33
+ no warning names a deck with no slide frames.
34
+ 2. **Check what ships.** `npx marver context check` when the project has `context/`.
35
+ *Check:* no `team` or `restricted` file is reachable from a published board; a board's status
36
+ ships only where its publish row says `"showStatus": true`.
37
+ 3. **Commit the host files** - a `Dockerfile` (it gives the build a browser for the glass
38
+ textures) and a `.dockerignore` with `node_modules`, `design/.dist`, `design/.local`.
39
+ *Check:* `design/.dist` is gitignored and not in the upload.
40
+ 4. **Create the service once:** a persistent volume, `MARVER_DATA_DIR` on it, the gate's variables,
41
+ `MARVER_TRUSTED_PROXY=1` behind a proxy, `MARVER_CLI_TOKEN` generated (`openssl rand -hex 24`),
42
+ never chosen. *Check:* the host lists the volume mounted at the path the variable names.
43
+ 5. **Deploy** - on Railway, `railway up` from the repository root.
44
+ *Check:* the deploy succeeds, and the URL answers with the gate.
45
+ 6. **Walk it.** Sign in as the owner, open each published board, leave a comment, reload.
46
+ *Check:* the comment survives the reload, and a second account sees only what its rights allow.
47
+ 7. **Write the run** in the log below - date, revision, URL, what was checked.
48
+
49
+ ## Traps
50
+
51
+ - **The upload is the working tree.** Untracked files and full-resolution images go up with it;
52
+ `marver build` ships images at full size, and a large `design/assets/` can pass the host's upload
53
+ cap. Keep heavy originals out of published frames, or out of the upload.
54
+ - **A repository linked to another project uploads there.** Check which project and service the
55
+ host CLI is linked to before `up` - a canvas deployed onto the app's own service replaces the app.
56
+ - **An ephemeral data directory loses every comment** on the next deploy. The volume is not optional
57
+ once collaboration is on.
58
+ - **Run one instance.** The comment log is single-writer.
59
+ - **Fonts only on your laptop** render differently in the build container; the glass textures then
60
+ rest live. Bundle fonts with the project.
61
+ - **A republish never clobbers comments** - the server unions the seeded logs on boot - but tabs
62
+ already open keep the old textures until they reload.
63
+
64
+ ## Rolling back
65
+
66
+ Redeploy the previous revision the same way. Comments live on the volume, not in the build, so a
67
+ rollback keeps them; a board removed from `publish.json` hides its threads until it ships again.
68
+
69
+ ## Run log
70
+
71
+ | Date | Revision | Who | URL | Outcome |
72
+ |---|---|---|---|---|
@@ -0,0 +1,232 @@
1
+ ---
2
+ name: reorganize-context
3
+ description: Reorganize an existing project's accumulated context - specs, changelogs, roadmaps, transcripts, research, boards, agent memory - so a fresh agent can answer what is available to users, how it works, why, and what is next, from the repo alone. Builds a shipped record with evidence levels, writes current contracts for a bounded first batch, adds a checker, reconciles only what that batch covers, and proves the result with a blind before/after eval. Works on a branch and never deletes. Use when specs are out of date, nobody can say what is on production, or context is scattered.
4
+ last_success: none
5
+ depends_on: []
6
+ ---
7
+
8
+ # Reorganize a project's context
9
+
10
+ `design/instructions/context.md` defines the target - the files under `context/`, what each holds,
11
+ the evidence levels. This is how an agent gets an existing project there. Marver maintains it: it
12
+ was first run on a client project in October 2026 (16 of 16 questions answered before and after,
13
+ with 40% fewer tokens after), and each rule exists because that project, or an earlier draft of this
14
+ playbook, broke it. Record your project's runs in the run log at the end.
15
+
16
+ **The order matters: truth first, then authority, then files.** Moving documents before you know
17
+ which one is right only relocates the contradiction.
18
+
19
+ ## Done means
20
+
21
+ A fresh agent with no memory reads `context/INDEX.md` and at most two more files per question, and
22
+ answers correctly - or says "unknown" with the reason - what each production service runs, whether a
23
+ capability is available and to whom, how it works, why, and what is next. Every artifact that
24
+ existed before has a recorded disposition. A blind before/after eval proves it.
25
+
26
+ ## Hard rules
27
+
28
+ 1. **Worktree and branch.** Never main. Never touch the original checkout's uncommitted changes.
29
+ 2. **Never delete.** Originals are preserved; superseded ones get a header pointing at their
30
+ replacement. Delete only proven disposable output, listed in the report.
31
+ 3. **Code says what is implemented, not what is intended.** When code and an accepted requirement
32
+ disagree, record both on the exception list - do not pick a winner silently.
33
+ 4. **Evidence has levels.** `confirmed`: observed in a system of record. `reported`: a written claim,
34
+ cited - a changelog line, an ops note, agent memory. `unknown`: neither. Never upgrade a level
35
+ because a claim sounds sure. Commit ancestry says a change is inside a revision, not that the
36
+ revision is running.
37
+ 5. **Stable paths stay stable.** Frame ids, board and scene slugs, URLs, comment anchors and asset
38
+ paths are identities. Do not move them in this pass.
39
+ 6. **Read-only outside the repo.** You may read deploy records (`gh`, a provider's status command).
40
+ You never deploy, migrate, query a production database, or edit agent memory. A read that needs
41
+ credentials you were not given is an exception-list item, not a workaround.
42
+ 7. **No secrets; mind the audience.** Never copy credentials. Record sensitive files by path and
43
+ hash only.
44
+ 8. **History is not rewritten.** Old changelog entries stay as written.
45
+ 9. **Bounded scope.** This pass rewrites only what the first batch of capabilities needs. Everything
46
+ else is inventoried and marked `deferred` - never superseded on a guess.
47
+
48
+ ## Phase 0 - baseline, in this order
49
+
50
+ 1. **The original checkout first.** In `context/reorg/start.md` (written later into the worktree,
51
+ from notes kept outside it): the revision, every dirty and untracked path with the hash of its
52
+ **working** file (`shasum`), and the tracked hashes (`git ls-files -s`).
53
+ 2. **The worktree.** `git worktree add ../<repo>-context -b context/reorganize <revision>` and work
54
+ only there from now on. Note that the original's dirty files are not in it.
55
+ 3. **Preimages.** Hash every context-bearing file in the worktree. This list is what Phase 9
56
+ compares against: untouched files must match; edited files must appear on the edit list with
57
+ their preimage hash and the reason.
58
+ 4. **The eval, blind.** Follow `eval.md` (beside this file). The questions go into the evaluated
59
+ agent's prompt; the **answer key never enters the worktree**. Run the before pass now.
60
+
61
+ ## Phase 1 - inventory, bounded
62
+
63
+ `context/reorg/inventory.md`:
64
+ - **Every** context-bearing path at file or folder level - docs, plans, changelogs, roadmaps,
65
+ feedback lists, READMEs inside apps, research, `design/` (boards, briefs, sticky notes, content
66
+ frames, comments), asset indexes, sync logs, `CLAUDE.md` / `AGENTS.md`, CI and scripts that
67
+ encode procedure, and the agent's project memory (Claude Code:
68
+ `~/.claude/projects/<path-slug>/memory/`). Size, last change, one-line role.
69
+ - **Section by section** only for documents that touch the first batch (Phase 4).
70
+ - Everything else gets disposition `deferred`.
71
+
72
+ Classes: current contract · proposed · historical evidence · decision · release/validation evidence ·
73
+ operational procedure · receipt · generated view · convention · tool state.
74
+
75
+ Flag, with file and line: status lines in prose; supersession inside a current file ("section 20
76
+ wins"); double homes ("the board is canonical, this file is its digest"); conventions pointing at
77
+ paths that no longer exist; design status phrased as product status; hand-kept counts.
78
+
79
+ Start `context/reorg/exceptions.md` now and keep adding to it in every phase.
80
+
81
+ ## Phase 2 - the shipped record
82
+
83
+ `context/shipped.md`, two tables, every cell with an evidence level and a citation.
84
+
85
+ **Services × environments** - discover the services first (deploy workflow, Dockerfiles, provider
86
+ config), then for each: revision, deployed at, how (pipeline / by hand), evidence level + citation,
87
+ migrations applied.
88
+ - `confirmed` needs the run that deployed **that service** and succeeded - the job inside the deploy
89
+ workflow, or the provider's own deploy status - not an environment record, which some pipelines
90
+ mark "deployed" even when the run stood down.
91
+ - Hand deploys leave only prose: `reported`, cite the line.
92
+ - A health endpoint that reports no revision cannot confirm one - note the gap.
93
+ - An API you cannot reach: `unknown`, with the error.
94
+
95
+ **Capabilities** - implemented (code reference); verified (a test or recorded walk: revision,
96
+ environment, date, scope - a test file's existence is not a run); available (which environment, and
97
+ for which tenants - a flag or setting is separate evidence from the deploy, usually `reported`
98
+ unless someone with access confirms it); contract link.
99
+
100
+ ## Phase 3 - the implementation map
101
+
102
+ `context/map.json`, machine-readable because the checker reads it:
103
+
104
+ ```json
105
+ { "capabilities": { "quote-intake": { "paths": ["apps/quote/app/**"], "tests": [] } },
106
+ "excluded": [{ "path": "apps/run/app/preview/**", "reason": "generated from the canvas" }] }
107
+ ```
108
+
109
+ Enumerate from the code: routes and actions, domain commands, tables and SQL functions, migrations
110
+ (every mechanism - a second app may keep its own schema), queued jobs and schedules, integrations,
111
+ config, tests. Map or exclude each, with a reason. Shared files may map to several capabilities. The
112
+ map is routing, not proof that anything is documented.
113
+
114
+ ## Phase 4 - contracts, the first batch
115
+
116
+ Pick 5-7 capabilities where the evidence showed contradictions or no contract. For each,
117
+ `context/product/<capability>.md`:
118
+
119
+ ```yaml
120
+ state: current
121
+ capability: quote-intake
122
+ audience: team # team | publishable | restricted
123
+ reviewed: { revision: 606ad4e, scope: "apps/quote, db schema", method: "read code and tests" }
124
+ ```
125
+
126
+ Then: what it does and its limits; who can use it and when; how it works; what it touches -
127
+ frontend, backend, database, workers, integrations; invariants; code and test references; decisions;
128
+ its lo-fi and hi-fi frames and the PRs that built it; what is still proposed. Read the code to write
129
+ it. Apply the plan's amendments - one answer per rule. Shared rules go once in `context/system/`.
130
+
131
+ The contracts are independent: one agent each, in parallel, from one shared brief (the rules, the
132
+ format, "no availability, no status lines", and a report back of every code-vs-document conflict and
133
+ the exact sections it now covers, `full` or `part`). Then a second model reviews the batch against
134
+ the code before anything relies on it - on its first run it found 18 overstated or mis-cited claims.
135
+ A shared file maps to every contract whose behaviour it changes, or the pull-request rule misses it.
136
+
137
+ ## Phase 5 - the checker, before reconciling anything
138
+
139
+ `scripts/context-check.mjs` (or the project's language), wired into existing CI. Inputs: front
140
+ matter, `context/map.json`, `context/shipped.md`, the publish configuration and what it pulls in, and
141
+ on a pull request the changed paths (`git diff --name-only <base>...<head>`, base and head from the
142
+ CI event) and the PR body (from the event payload).
143
+
144
+ Audiences decide where a file may go: `restricted` is never tracked by git and never published;
145
+ `team` may be committed in a private repo but never published; `publishable` may appear in a
146
+ publish. "Published" means everything a canvas build carries: the boards in its publish config, the
147
+ frames on them, and every markdown import, image, note and asset those frames pull in.
148
+
149
+ | Rule | Result |
150
+ |---|---|
151
+ | a `state: current` doc has a status line (`^\s*\**status\**\s*:`) - lines, not domain words like a bank account's `verified` | fail |
152
+ | a current doc carries supersession markers ("wins where it differs", "overrides section") | fail |
153
+ | a shipped cell has no evidence level, or `confirmed` / `reported` with no citation | fail |
154
+ | a dead relative link, `?raw` import or asset reference under `context/` or in a frame that renders it | fail |
155
+ | `INDEX.md` over 800 words | fail |
156
+ | a `restricted` file tracked by git, or a `team` or `restricted` file reachable from what a publish carries | fail |
157
+ | a PR changes mapped paths without its contract, and without `no-contract-change: <capability> - <reason>` in the PR body | fail on pull requests |
158
+ | an `unknown` evidence cell | pass - honest unknowns are the right answer |
159
+ | a rule it cannot evaluate (no git history, API unreachable) | exit 2, "cannot determine", printed |
160
+
161
+ Exit 0 pass, 1 fail, 2 cannot determine. In CI, 2 blocks like 1 - fix the cause (fetch full history,
162
+ grant the read). During this pass, Phase 6 may start on exit 0, or on exit 2 when every
163
+ "cannot determine" is on the exception list with its reason.
164
+
165
+ ## Phase 6 - reconcile, first batch only
166
+
167
+ For each document that the first batch's contracts now supersede, **only the sections they cover**:
168
+
169
+ | Material | Action |
170
+ |---|---|
171
+ | amendment-stacked spec | its covered sections get "Superseded by `context/product/x.md`"; the file stays as history; uncovered sections stay `deferred` |
172
+ | mixed spec | covered contract → `product/`; review diary stays; open items → backlog |
173
+ | dated feedback lists | covered items get a resolution: contract + shipped row, or `open` |
174
+ | roadmap in two homes | one editable source; the board renders it (`?raw`) |
175
+ | README with contradictions | fix the covered contradiction; keep run instructions by the app |
176
+ | changelog | untouched; its release lines are cited from `shipped.md` |
177
+ | sync log | keep cursors and pending failures; drop "agent memory" as a target |
178
+ | agent memory | project facts the batch needs land in their repo home; `context/reorg/memory-trim.md` proposes the rest |
179
+ | wrong binding instructions | correct them - an agent will obey them |
180
+
181
+ Diff duplicate prose before merging; check inbound references after every move.
182
+
183
+ **Line citations move when you insert.** Every note added to an old document shifts the lines that
184
+ contracts cite. Either reconcile before the contracts cite those documents, or remap every citation
185
+ once, mechanically, from the base revision's lines to the new ones - and verify by content (the cited
186
+ line still says what the citation claims). Never run a remap twice. Prefer line-neutral edits where a
187
+ document is cited heavily.
188
+
189
+ ## Phase 7 - entry point and index
190
+
191
+ - Root `AGENTS.md` is the shared entry point; tool files route to it. Keep existing binding routes
192
+ (for example to `design/AGENTS.md`).
193
+ - `context/INDEX.md`, **at most 800 words**, counted by the checker: the product in five lines; the
194
+ capabilities with their contract and shipped row; where each role lives, including healthy homes
195
+ left in place; the rules (status only in `shipped.md`; a behaviour change updates its contract or
196
+ says `no-contract-change`; every deploy writes the record; plans fold into contracts when they
197
+ land; changelog lines start with kind and capability: `[product:quote]`, `[design:board]`,
198
+ `[infra]`, `[decision 41]`). Generated parts sit between `<!-- generated -->` fences.
199
+
200
+ ## Phase 8 - the exception list, asked once
201
+
202
+ Everything collected since Phase 1: `unknown` availability that matters, code-vs-intent conflicts,
203
+ disputed capability boundaries, audience and retention calls, reads that need credentials. Ask in
204
+ one numbered batch now - or immediately, only if an item blocks the next step.
205
+
206
+ | The human decides | The agent decides from evidence |
207
+ |---|---|
208
+ | intended behaviour where code and an accepted requirement conflict | what the code implements |
209
+ | access, audiences, retention, redaction | inventory, hashes, references |
210
+ | priorities, commitments, pilot scope | what is done vs still proposed |
211
+ | disputed capability boundaries | the initial grouping (reversible) |
212
+
213
+ ## Phase 9 - verify
214
+
215
+ `context/reorg/report.md`:
216
+
217
+ | Check | Passing result |
218
+ |---|---|
219
+ | Preservation | untouched files match their preimage; every edited file is on the edit list with preimage and reason |
220
+ | Disposition | every inventoried path is unchanged, edited, superseded (covered sections only) or `deferred` |
221
+ | The checker | exit 0, or exit 2 with each "cannot determine" explained |
222
+ | Release honesty | every availability claim carries a level and a citation, or says `unknown` |
223
+ | Reconciliation | no Phase 1 contradiction inside the first batch survives as competing current text |
224
+ | The eval | the after pass beats the before pass, per `eval.md` |
225
+
226
+ Then stop. The report lists what moved, what was decided without asking, the open exceptions, what
227
+ was deferred, and the memory-trim proposal. The human reviews the branch; nothing merges on its own.
228
+
229
+ ## Run log
230
+
231
+ | Date | Revision | Who | Batch | Outcome |
232
+ |---|---|---|---|---|
@@ -0,0 +1,93 @@
1
+ # Context eval - can a fresh agent answer from the repo alone?
2
+
3
+ The reorganization is judged by one thing: a fresh agent, with no memory and no chat history,
4
+ answering real questions about the product correctly and cheaply. Run it **before** and **after**,
5
+ the same way both times.
6
+
7
+ This file is the protocol and the questions. The **answer key is a separate file**
8
+ (`answer-key-<project>.md`) and never enters the worktree the agent reads.
9
+
10
+ ## Isolation - check before each pass
11
+
12
+ - The agent runs in the worktree (`../<repo>-context`), a fresh process per question.
13
+ - No agent memory for that path: Claude Code keys memory by project path, so
14
+ `~/.claude/projects/<worktree-slug>/memory/` must not exist. Codex keeps none by default.
15
+ - No hook injects project context for that path (a vault feed keyed by path, a knowledge inbox).
16
+ - The answer key, the reorg notes kept outside the worktree, and earlier pass results are not in
17
+ the worktree. Before the first pass, `context/` does not exist yet.
18
+ - **A clean profile.** Global instructions also count as injected context: Claude Code loads
19
+ `~/.claude/CLAUDE.md`, Codex loads its global `AGENTS.md`. Launch with an empty config home and
20
+ only credentials in it - Claude Code: `CLAUDE_CONFIG_DIR=<empty dir>` with an API key; Codex:
21
+ `CODEX_HOME=<dir holding only auth>`. Confirm from the transcript's first event that no user or
22
+ global instruction file was loaded.
23
+ - *Learned on the first run:* without an API key, Claude Code cannot run clean - `--bare` needs one,
24
+ and a plain run loads the auto-memory that worktrees of one repository share. Codex with a
25
+ `CODEX_HOME` holding a copy of `auth.json` works; check `codex features list` shows `memories`
26
+ off, and keep each pass's session rollouts out of that home before the next pass. Copy the auth
27
+ only when its `last_refresh` is recent, and delete the copy after.
28
+ - **Read-only.** Claude Code with read tools only (`--allowedTools Read,Grep,Glob`); Codex with
29
+ `--sandbox read-only`.
30
+ - Same CLI, same model, same prompt, both passes.
31
+
32
+ ## The prompt
33
+
34
+ > Answer this question using only files in this repository: <question>. Give the answer, the
35
+ > files you used, and your confidence. If the repository cannot settle it, say "cannot tell" and
36
+ > why. Do not run the app, query a database, call an API, or read outside this directory.
37
+
38
+ ## The measurement - one metric, per question
39
+
40
+ Run each question as its own invocation with a machine-readable transcript (Claude Code:
41
+ `claude -p "<prompt>" --output-format stream-json --verbose`; Codex: `codex exec --json`). From the
42
+ transcript record:
43
+
44
+ - **result** - correct / partial / wrong / cannot-tell, judged against the answer key;
45
+ - **confident wrong** - a wrong answer given with high confidence, counted separately;
46
+ - **bytes read** - the summed size of every tool result returned to the agent (file reads, searches,
47
+ listings). Measure what the model saw, not what the command printed: Codex truncates long output
48
+ (each result carries its `original_token_count`), so read the session rollout's tool outputs, not
49
+ `exec --json`'s `aggregated_output`. Record the session's **input tokens** beside it;
50
+ - **tool calls**;
51
+ - **out of bounds** - any read outside the worktree. One voids that answer.
52
+
53
+ "Cannot tell" with the right reason counts as correct where the key says the repository genuinely
54
+ does not hold the answer.
55
+
56
+ ## Pass
57
+
58
+ - no confident wrong answer after;
59
+ - more correct answers after than before - or, if the before pass was already perfect, equal;
60
+ - fewer total bytes read after.
61
+
62
+ A capable agent may already answer everything before - the first run's before pass was 16 / 16 at
63
+ medium effort, the traps notwithstanding. Then the pass is decided on cost, and a weaker or cheaper
64
+ profile is worth a second pair of passes to see whether the reorganization moves correctness.
65
+
66
+ ## The six kinds of question
67
+
68
+ Write 10-16, at least one of each, from Phase 2 evidence, with the key signed off by the human where
69
+ Phase 8 asked:
70
+
71
+ 1. **Shipped** - is X available, where, for whom, since when?
72
+ 2. **Behaviour** - how does surface Y work today?
73
+ 3. **Interfaces** - what runs behind action Z: jobs, mails, data written?
74
+ 4. **Why** - why was decision N taken, and does it still apply?
75
+ 5. **Next** - what is being built next, and what counts as done?
76
+ 6. **Receipts** - what did someone say about T, and where is it recorded?
77
+
78
+ Include **traps**: questions where a stale spec, a design frame, or a wrong instruction gives a
79
+ confident wrong answer today. They are the point.
80
+
81
+ ## Your project's questions
82
+
83
+ Write them before the before pass, from Phase 2's evidence, and keep the key out of the worktree:
84
+
85
+ | # | Kind | Question | Trap |
86
+ |---|---|---|---|
87
+
88
+ ## Report format
89
+
90
+ | # | Before: result | Before: bytes / calls | After: result | After: bytes / calls |
91
+ |---|---|---|---|---|
92
+
93
+ Write the before column before Phase 1 begins - not from memory afterwards.