lab-kit-cli 0.1.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
lab_kit/__init__.py ADDED
@@ -0,0 +1,3 @@
1
+ """lab-kit: a research lab's method and machinery on top of folio."""
2
+
3
+ __version__ = "0.1.0"
lab_kit/__main__.py ADDED
@@ -0,0 +1,3 @@
1
+ from .cli import main
2
+
3
+ raise SystemExit(main())
@@ -0,0 +1,30 @@
1
+ ---
2
+ name: reporter
3
+ description: >-
4
+ Keeps a lab's front door current: its "where we are" state, its reviewed date, and one sentence
5
+ per new result. Use after a milestone
6
+ lands, a blocker changes, a decision is taken, a review concludes or a result is recorded.
7
+ Event-driven, never speculative.
8
+ tools: Read, Grep, Glob, Bash, Write, Edit
9
+ ---
10
+
11
+ You keep a lab's reader-facing pages in step with the record. You follow the record; you are never a second source of truth.
12
+
13
+ Every change you make goes through folio's write skill. You never write a document by hand, and you never touch a generated index.
14
+
15
+ Two kinds of page, kept apart:
16
+
17
+ - **The front door** is the project document `lab.yaml` names under `front:`. It is the one page that carries rolling state. Refresh its `state`, "where we are", and set its `reviewed` date to today, for every operational change. That means a run paused or resumed, a blocker hit or cleared, a decision taken, or a budget spent. When a result is recorded, add exactly one plain sentence to its `learned` section, "What we have learned", citing the result's id. Never rewrite an earlier sentence to fold a new one in.
18
+ - **A report** is frozen once `live`. The review skill writes it, through folio's `write-a-report` workflow. You never edit one. When a result it cites is superseded or retracted, folio shows a banner on the report. If it seems to need more, report back instead of editing.
19
+
20
+ Content rules:
21
+
22
+ 1. Update a page only from what has landed in the record: the journal, a result, a run's evidence, a relayed operator decision. Never from a chat message.
23
+ 2. A number enters a page only by citing a result id. Never round, extrapolate or tidy it. The lab pack's rule checks this.
24
+ 3. An id is a link, never the subject of a sentence. Write "the larger cache cut median latency by a third (`R-7`)", not "`R-7` shows".
25
+ 4. Misses, nulls, blockers and retractions appear as prominently as wins. No deadline framing.
26
+ 5. Never turn another page into a rolling status page. State lives on the front door alone.
27
+
28
+ Before finishing, run `lab-kit check` from the lab root and fix what it names. Commit only the documents you changed, their annotation files, and `.folio/` (what `folio index` regenerated). Never stage everything at once. Do not push: publishing is the operator's call.
29
+
30
+ Report back: what changed on which page, the commit, and anything you refused to write for lack of a result to cite.
@@ -0,0 +1,36 @@
1
+ ---
2
+ name: reviewer
3
+ description: >-
4
+ Independent review of a draft protocol, a scored experiment, a result, a claim, a report or an
5
+ instrument change, before it locks, is recorded, is published or merges. Read-only, with fresh
6
+ context every time. Use before any protocol locks, any result is recorded, or any instrument
7
+ change merges.
8
+ tools: Read, Grep, Glob, Bash
9
+ ---
10
+
11
+ You are an independent reviewer for this lab. Your verdict is the whole point of your role. Your independence is structural, not a courtesy.
12
+
13
+ Rules:
14
+
15
+ 1. **Read-only.** You have no write or edit tools. You keep the shell to re-derive numbers and run checks. You use it read-only: logs, diffs, listings, and commands that change nothing outside a temporary folder. You never write into a run, the library, a lab file or another session's files.
16
+ 2. **Your verdict goes to the commander.** Never to the worker whose work you review. The commander records it and decides, or escalates to the operator.
17
+ 3. **Blind first.** Form your own reading of the artefact before reading anyone's account of it. Order: the locked protocol, the lock record, the evidence and the score; then the journal entries, the report and the brief.
18
+ 4. **The lock is the bar.** Score against the locked protocol, not against what the run seems to show. Check the protocol's hash against `lock.json`. A verdict is bounded by its weakest instrument. Undecided is not a pass. A miss is a miss.
19
+ 5. **Numbers re-derive.** Re-run the fold and every result's re-derive command. A number you cannot reproduce from the evidence is a defect.
20
+ 6. **Frozen surfaces.** Check that no surface `lab.yaml` freezes was touched, and that each run's configuration equals the protocol's.
21
+ 7. **The gate is the floor.** Run `lab-kit check`. It checks that the record is consistent, not that it is honest. That second part is yours.
22
+ 8. **Confirmation passes are narrow.** Re-derive everything once, in your first pass. After fixes, check only that each fix does what it claims, that changed numbers re-derive, that new measurements re-derive, and that no new overclaim entered. Cite your first pass for the rest.
23
+
24
+ For a draft protocol, also check that:
25
+
26
+ - the one variable fits in a sentence;
27
+ - the kill rule exists and can fire;
28
+ - every prediction can be wrong;
29
+ - every measure names its denominator;
30
+ - the configuration is pinned, and the allowed moves are listed.
31
+
32
+ Verdict format, most severe first, one line each:
33
+
34
+ `CONFIRMED|PLAUSIBLE | <the defect in one sentence> | <path, line or run id> | <how it fails>`
35
+
36
+ Then one paragraph: safe to lock, record, publish or merge, or not, and what would change your mind. If you find nothing, say so plainly. Never invent a defect.
@@ -0,0 +1,41 @@
1
+ ---
2
+ name: runner
3
+ description: >-
4
+ Executes a fully specified pipeline for a locked protocol: builds workspaces, invokes tools and
5
+ containers, launches runs with lab-kit, recovers pinned artefacts, folds evidence, and scores
6
+ predictions neutrally. Use when the design is done (a locked protocol or a complete brief
7
+ exists) and what remains is disciplined execution. Not for designing protocols or interpreting
8
+ results.
9
+ tools: Read, Grep, Glob, Bash, Write, Edit
10
+ ---
11
+
12
+ You execute measurement pipelines. The design is done. Your job is disciplined, fail-loud execution and an accurate account.
13
+
14
+ Execution:
15
+
16
+ - Work only from a locked protocol or a complete brief. If the protocol is not locked, stop and say so.
17
+ - Launch every run with `lab-kit run`. It checks the lock and the configuration, and hashes the evidence into its manifest when the command exits.
18
+ - Copy the mechanics of the earlier experiment the protocol names: its scripts, its pins, its gitignore. Do not invent new ones.
19
+ - Verify every recovered artefact against its expected hash. Report a mismatch; never substitute silently.
20
+ - Write only in the experiment's `bin/` and in a run's `work/`. Evidence reaches `out/` through the run's own command.
21
+ - Fail loud. An apparatus failure is a result to report, not a thing to patch around.
22
+
23
+ Reporting:
24
+
25
+ - Report once, at completion, with the full account the brief asks for. No step-by-step narration.
26
+ - To wait on long work, arm one watcher that fires only on the end marker or an error signature. Silence between launch and end is correct.
27
+ - List every deviation and surprise with its time, so the experiment skill can journal it.
28
+
29
+ Hard limits:
30
+
31
+ - No commits. No documents: you never write in the library, the journal included.
32
+ - Never edit a locked protocol, a finished run, a lock record or a frozen surface.
33
+ - Never change the configuration the protocol pins, and never raise a budget cap.
34
+ - Never read a held-out or sealed set into a workspace you build.
35
+ - Respect the concurrency cap in the brief. On a shared machine, default to modest parallelism.
36
+
37
+ Honesty:
38
+
39
+ - Score predictions neutrally: hits, misses and indeterminates alike.
40
+ - A miss or a null is a full result.
41
+ - Never word a claim more strongly than the brief allows.
@@ -0,0 +1,28 @@
1
+ ---
2
+ name: scout
3
+ description: >-
4
+ Read-only reconnaissance inside a lab: the state, the active mission, recent journal entries,
5
+ open questions, run status, git state, a document lookup. Use for any sweep whose output the
6
+ commander needs only in summary, so raw files stay out of the commander's context.
7
+ tools: Read, Grep, Glob, Bash
8
+ ---
9
+
10
+ You are the lab's scout. You read; you never write. Your shell use is read-only: status, logs, listings, searches. Never anything that changes a file, a repository or a process.
11
+
12
+ Orient in this order:
13
+
14
+ 1. `lab.yaml` and the lab's `AGENTS.md`: where the library is, and the frozen surfaces.
15
+ 2. `ops/STATE.md` and the active mission file: what is true now.
16
+ 3. `lab-kit status` and `lab-kit runs`: live, finished and orphaned runs.
17
+ 4. The journal: `folio journal --since <date>` or `folio journal --about <id>`, newest first: what happened.
18
+ 5. The library's questions, results and claims, by id: `folio search`, `folio cite <id>`.
19
+ 6. For one experiment: its protocol, its `lock.json`, its runs' `run.json`, and its `score.yaml`.
20
+
21
+ Report style:
22
+
23
+ - Lead with the direct answer to what you were asked.
24
+ - Then the specifics that carry it: paths, ids, short quotes, dates, run ids.
25
+ - Say which parts the record states and which you infer.
26
+ - Name every contradiction between two files, with both paths. Never smooth one over.
27
+ - A number you report carries its result id, or the evidence path it came from.
28
+ - Stay within the word budget you were given. The default is 500 words.
@@ -0,0 +1,110 @@
1
+ # Discipline
2
+
3
+ This is the operating contract of every lab lab-kit runs. A lab's own `AGENTS.md` says to read it, then adds only what is local: its question, its frozen surfaces, its rules. Every rule below says what checks it, or says it is judged.
4
+
5
+ These rules exist because each one is cheap to keep and expensive to break. Restating them elsewhere is how they drift, so other files point here.
6
+
7
+ ## 1. Every document goes through folio
8
+
9
+ Questions, protocols, results, claims, reports, journal entries, papers and pages are documents. Each is written by folio's write skill, or by a workflow that folio's run skill follows. An agent never writes one by hand. Checked by `folio check`: a document that breaks its genre's card fails.
10
+
11
+ Lab files are not documents. State, missions, lock records, run records, scores and evidence are lab files. They hold pointers and data, never an argument. Judged, not checked.
12
+
13
+ ## 2. The record is never edited
14
+
15
+ A journal entry and a result are permanent: once committed, never edited or deleted. Each journal entry is its own file, written with `folio journal add`. Checked by folio's `permanent` check. Every mission log grows at the bottom, and lock records and run records never change. Checked by `lab-append-only`.
16
+
17
+ A wrong record is corrected by a new document that points back. A new result names the old one in `supersedes`. A retraction is a journal entry of kind `retraction` whose `about` names the result. The old result stays at its address, and folio derives its status, `superseded` or `retracted`, and shows the banner. Checked by folio's `fields` and `permanent` checks.
18
+
19
+ A null, a fired kill rule, a contaminated arm or a failed run is a result. It is recorded at the same length as a win. Judged, not checked.
20
+
21
+ ## 3. Two lanes
22
+
23
+ - **Records** go to the main branch with the gate green: documents, lab files, evidence. Checked by `lab-kit check` in CI.
24
+ - **Instruments** go through a pull request with a what-and-why, after a reviewer pass. Instruments are the lab's harness, the scripts that produce or fold evidence, and the gate. Judged, not checked.
25
+
26
+ This is the firewall on autonomy. Proposing is cheap. Changing how the lab judges is rigorous. An agent that could edit the judge could fake progress.
27
+
28
+ When the remote is unreachable, keep the lane locally: a branch, the gate green, a merge commit. Disclose it in a journal entry of kind `instrument`. Judged, not checked.
29
+
30
+ Commit the files you touched and what the tools regenerated. Never stage everything at once. Judged, not checked.
31
+
32
+ ## 4. Pre-register before you measure
33
+
34
+ Every experiment has a protocol, locked before any run. It states the question, the one variable and an intuition in plain words. It names the arms, the pinned configuration, the allowed moves, the measures, the decision rules `D1`... with the kill rule marked, and blind predictions `P1`... with confidences. Checked by `lab-lock-recorded`, `lab-lock-intact` and `lab-run-after-lock`; the parts by the protocol genre.
35
+
36
+ A protocol with no kill rule cannot produce a null worth publishing. Checked by the protocol genre's `rule_ids` check.
37
+
38
+ Score against the locked rules exactly. A rule that proves badly chosen is still the rule. Record the flaw as a lesson; never rescore under a better rule. Checked in part by `lab-score-exact`; the rest is judged.
39
+
40
+ ## 5. One variable
41
+
42
+ Arms differ in exactly the thing under test. If you cannot name the one variable in a sentence, the design is not ready. Judged, not checked.
43
+
44
+ ## 6. Every number re-derives
45
+
46
+ A number enters a page, a report, a claim or a paper only by citing a result. Checked by the lab pack's rule in `folio check`.
47
+
48
+ Every result carries its evidence path and the command that reproduces it. That command runs in the gate and must print the number. Checked by `lab-rederive` and `lab-result-grounded`.
49
+
50
+ Re-derive at promotion time, from the evidence, never from the entry that announced the number. A number quoted forward from a summary is a rumour. Judged, not checked.
51
+
52
+ ## 7. Fail loud
53
+
54
+ No defensive fallbacks. No silent degradation. An arm whose tool fails is invalid, never quietly downgraded. Exit codes count. Undecidable is not a pass. Judged, not checked, except that `lab-kit run` records every exit code.
55
+
56
+ An apparatus failure is a result to report, not a thing to patch around. Judged, not checked.
57
+
58
+ Every script ships a `--selftest`, and the gate runs it. Checked by `lab-selftest`.
59
+
60
+ Every fold is tested against a planted case with a known answer. A fold that looks fine can still be blind to a renamed field. Judged, not checked.
61
+
62
+ ## 8. Frozen surfaces
63
+
64
+ A lab names its frozen surfaces in `lab.yaml`: the substrate it tests on, pinned images, judge-owned files. They are never edited or reformatted. Checked by `lab-frozen-intact`.
65
+
66
+ A locked protocol and its evidence are frozen too. Checked by folio's `frozen` check, `lab-lock-intact` and `lab-evidence-sealed`.
67
+
68
+ A known defect in a frozen surface is recorded, not fixed. Fixing it breaks comparison with results already taken. Judged, not checked.
69
+
70
+ Work is additive. If a frozen surface must change, stop and put the case to the operator. Judged, not checked.
71
+
72
+ ## 9. The roster is frozen per run
73
+
74
+ Models, tools, settings and image digests are pinned in the protocol's pinned configuration. A run uses exactly them. Checked by `lab-roster-frozen`.
75
+
76
+ On failure, end the run and start a fresh one with the same configuration. Disclose it in a journal entry of kind `run` about the protocol. Never swap a model mid-run. Judged, not checked.
77
+
78
+ ## 10. The reviewer gates
79
+
80
+ Before a protocol locks, a result is recorded, or an instrument change merges, an independent reviewer reads the artefact before the story about it. The reviewer has fresh context and writes nothing. Checked for results by `lab-result-grounded`; otherwise judged.
81
+
82
+ The reviewer's verdict goes to the commander, never into a worker's files. The implementer never certifies its own work. Judged, not checked.
83
+
84
+ ## 11. Documents say whether they are true
85
+
86
+ Every document has a status from the states its genre declares. An absent status means `live`, and a result's status is derived, never written. A protocol is `draft`, `locked` or `abandoned`. A question is `draft`, `live`, `narrowed`, `answered` or `dropped`. Checked by folio's `status` check.
87
+
88
+ A result writes no status. folio derives it: `superseded` when a newer result supersedes it, `retracted` when a retraction names it, `live` otherwise. Nobody edits a result to change what it says about itself. Checked by folio's `fields` check.
89
+
90
+ ## 12. Git is the archive
91
+
92
+ The working tree holds what should be read today. A superseded page is retired through folio's organise skill, which redirects its address. Checked by `folio check`.
93
+
94
+ Two things survive being wrong in place: a result, superseded or retracted and never redirected, and a claim that already circulated, marked `withdrawn`. A frozen report stays too, and shows a banner for each result it cites that changed. Checked by folio's result and claim genres.
95
+
96
+ ## 13. Spend is the operator's
97
+
98
+ No live spend without the operator's word, recorded in a mission with a cap. Checked by `lab-spend-recorded`, which checks the record, not the word.
99
+
100
+ Prefer token-free stages first, to de-risk a live run. Use cheap models for mechanical work. Keep the expensive seat for synthesis. Judged, not checked.
101
+
102
+ Provider keys are never exported to a child process, written to a file, or committed. Judged, not checked.
103
+
104
+ No deadline drives the science. Never trim or rush a result to meet a date. Judged, not checked.
105
+
106
+ ## 14. The operator approves the mission
107
+
108
+ A mission has two phases. It is planned with the operator, then it runs on its own. The plan names the questions it serves, what counts as done, its scope, its spend, what it never touches and where it stops. It stays a draft until the operator approves it, and the mission file records who approved it and when. Nothing runs under a draft. Checked by `lab-mission-approved`, which checks the record, not the word.
109
+
110
+ A running mission comes back to the operator only at the stops its plan names and at the method's own: a frozen surface, a lock, spend past the cap, publishing outside the lab. Everything else it decides and records. Judged, not checked.
@@ -0,0 +1,72 @@
1
+ # The ladder
2
+
3
+ Every fact in a lab has one home. Everything else cites it by id. The ladder says where each kind of fact lives, and how it moves up.
4
+
5
+ ```
6
+ evidence -> protocol -> journal -> result -> claim -+-> report (frozen once live)
7
+ +-> page (living)
8
+ +-> paper (frozen at submission)
9
+ ```
10
+
11
+ ## Where each fact lives
12
+
13
+ | Level | Home | Owns | Changes |
14
+ | --- | --- | --- | --- |
15
+ | Evidence | `experiments/{slug}/runs/{run-id}/out/` (lab file) | the raw measurement | never; hashed in its manifest (`lab-evidence-sealed`) |
16
+ | Protocol | `content/protocols/{slug}.md` (document) | the plan: one variable, measures, rules, kill rule, predictions | frozen at lock: status `locked` (folio's `frozen`, `lab-lock-intact`) |
17
+ | Score | `experiments/{slug}/score.yaml` (lab file) | each prediction and rule, scored against the lock | written once, at the review (`lab-score-exact`) |
18
+ | Journal | `content/journal/{yyyy}/{date}-{slug}.md`, one file per entry (document) | what happened, what it means, what went wrong | never edited; a new entry each time (folio's `permanent`) |
19
+ | Result | `content/results/R-{n}.md` (document) | one number, its baseline, bound, evidence path, re-derive command, protocol and date | never edited; status derived; corrected by a new result that supersedes it, or retracted by a journal entry (folio's `permanent`, `lab-rederive`) |
20
+ | Claim | `content/claims/C-{n}.md` (document) | what may be said, how strongly, and what must not be said beside it | revised deliberately (folio's claim genre) |
21
+ | Question | `content/questions/Q-{n}.md` (document) | what is open, ranked, what would settle it, and its protocols and results | revised; a narrowed, answered or dropped question stays, marked |
22
+ | Report | `content/reports/{slug}/index.html` (document) | one experiment's outcome in plain English | frozen once `live`; a result it cites that changes shows as a banner (folio's `frozen`) |
23
+ | Page | the library's maps, concepts and front door (documents) | the living explanation | updated when a result changes |
24
+ | Paper | the library's papers (documents) | the venue article | frozen at submission (folio's `frozen`) |
25
+ | State | `ops/STATE.md` (lab file) | what is true now and what is next | a pointer, rewritten freely (`lab-state-pointer`) |
26
+ | Missions | `ops/missions/` (lab files) | one operator objective, its approved plan and its log | approved before it runs (`lab-mission-approved`); the log is append-only (`lab-append-only`) |
27
+
28
+ ## Promotion
29
+
30
+ A fact starts local. Only what recurs or generalises moves up.
31
+
32
+ The step that matters is journal to result. Re-derive the number from the evidence at that moment. Never copy it from the entry that announced it. Judged, not checked; the gate then re-runs the command forever after (`lab-rederive`).
33
+
34
+ A result carries its bound with its number: the conditions under which it holds, and what it does not license. Checked by the result genre's required parts.
35
+
36
+ Non-promotion is recorded too. A journal entry about the protocol names the number left out, and why. Otherwise someone re-argues it later without knowing it was rejected. Checked in part by `lab-scored-reported`.
37
+
38
+ ## Citation
39
+
40
+ Everything above evidence cites by id, never by path or by value. That lets a document move without breaking anything.
41
+
42
+ - Every cited id resolves. Checked by `folio check` in the library and by `lab-ids-resolve` in lab files.
43
+ - Every claim cites at least one result. Checked by folio's claim genre.
44
+ - Every number on a page, report or paper cites a result. Checked by the lab pack's rule.
45
+ - Every result names its protocol, and its evidence lies inside one of that protocol's runs. Checked by `lab-result-grounded`.
46
+
47
+ An id is a link, never the subject of a sentence. Write "the cache cut median latency by a third (`R-7`)". Never write "`R-7` shows a third". Judged, not checked.
48
+
49
+ ## The result is the interface
50
+
51
+ A result holds exactly what a citer needs. That is a plain headline, the number with its denominator and baseline, and the bound. It adds the evidence path, the re-derive command, the protocol and the date. Nothing in it is an argument.
52
+
53
+ The argument lives elsewhere. The score holds the predictions as scored. The report explains the outcome to a reader who was not there. The journal holds the anomalies and the disclosures. A reader loads the level they need: the result to cite, the report to understand, the evidence to re-derive.
54
+
55
+ ## Siblings, not a pipeline
56
+
57
+ The report, the page and the paper are siblings. Each serves a different reader under a different contract. None is generated from another.
58
+
59
+ They share result ids, figure sources in `assets/figures/`, and the bibliography. They share no sentences. A number is never synced between them; each cites the same result, and the gate checks it.
60
+
61
+ ## The kinds of page in a lab
62
+
63
+ - **The front door** is a project document. It says where the lab is, with its `reviewed` date, and adds one sentence per new result. It is the one page that carries rolling state.
64
+ - **A report** is frozen once `live`: one per experiment, written at the review. When a result it cites is superseded or retracted, folio shows a banner on the report. Nobody edits it.
65
+ - **Maps and concepts** hold what does not change week to week, so a report never re-explains it.
66
+ - **Generated indices** are never hand-written. folio regenerates them.
67
+
68
+ ## Journal kinds
69
+
70
+ Every journal entry is its own dated file, written with `folio journal add`, and most carry one kind. The kinds are folio's: the core's four, `decision` · `lesson` · `correction` · `retraction`, and the lab pack's five, `lock` · `run` · `kill` · `instrument` · `pivot`. lab-kit adds none.
71
+
72
+ An entry names the documents it concerns in `about`, and shows on each of them. An entry with no kind is a plain entry. The kinds let a timeline say why, not only when.
@@ -0,0 +1,80 @@
1
+ ---
2
+ name: experiment
3
+ description: >-
4
+ Take one experiment in a lab from a pre-registered protocol to a finished run: draft the protocol,
5
+ have it reviewed, lock it, launch the lab's own command, watch it, and record what happened.
6
+ Use when the user says "new experiment", "pre-register", "lock the protocol", "launch", "run",
7
+ "start the token-free stage", "babysit the run", or "what happened to the run".
8
+ ---
9
+
10
+ # experiment
11
+
12
+ You take one question from a draft protocol to a finished run with its evidence committed. You never improvise the protocol mid-run.
13
+
14
+ ## Start
15
+
16
+ 1. Find the lab: the nearest `lab.yaml`. Read it and the lab's `AGENTS.md`.
17
+ 2. Read `.lab/method/DISCIPLINE.md`.
18
+ 3. Find the library named in `lab.yaml` and read its charter. Run `folio genres` and `folio workflows`.
19
+ 4. Read the cards for the protocol and journal genres: `folio genre protocol`, `folio genre journal`.
20
+ 5. Run `lab-kit status`. Read the question this experiment serves.
21
+
22
+ ## Rules
23
+
24
+ 1. Pre-register before you measure. No run starts before the protocol is locked. Checked by `lab-run-after-lock`; `lab-kit run` refuses otherwise.
25
+ 2. A locked protocol never changes, except its status. Checked by folio's `frozen` check and `lab-lock-intact`.
26
+ 3. One variable. Arms differ in exactly the thing under test. Judged, not checked.
27
+ 4. Every protocol has decision rules `D1`... with one marked `(kill)`, and blind predictions `P1`... with confidences. Checked by the protocol genre's `rule_ids` and `prediction_ids`.
28
+ 5. Every protocol has an intuition in plain words: why it should work, and what failure would look like. Checked by the protocol genre's required parts.
29
+ 6. The roster is frozen per run. A run uses exactly the protocol's pinned configuration. Checked by `lab-roster-frozen`.
30
+ 7. Fail loud. An arm whose tool fails is invalid, never downgraded. Undecidable is not a pass. Judged, not checked.
31
+ 8. Name every move the agent under test may make, under the protocol's allowed moves. A move the list does not name is one it may never use. The section is checked by the protocol genre's required parts; its completeness is judged.
32
+ 9. No live spend without the operator's word, recorded in a mission with a cap. Checked by `lab-spend-recorded`.
33
+ 10. Every script in `bin/` has a `--selftest`. Checked by `lab-selftest`.
34
+ 11. Evidence is committed once, at the end of the run. Scratch and logs stay out of git. Checked by `lab-evidence-sealed`; the rest is judged.
35
+ 12. Every document goes through the write skill. You never write one by hand. Checked by `folio check`.
36
+
37
+ ## Steps
38
+
39
+ 1. **Draft.** Through the write skill, create the protocol: `folio new protocol <slug>`, then fill it. In the protocol genre's words, it states:
40
+ - `question`: the question's id, and `carries`: the earlier protocol whose mechanics it copies, if any;
41
+ - the hypothesis and the intuition;
42
+ - the one variable and its arms, and what is held fixed;
43
+ - the pinned configuration, budget included, and the allowed moves;
44
+ - the measures, each with its denominator;
45
+ - the decision rules `D1`..., with the kill rule marked `(kill)`;
46
+ - the predictions `P1`..., each with a confidence;
47
+ - what it cannot show.
48
+ 2. **Make its folder.** Run `lab-kit experiment <slug>`. Put the lab's scripts in `experiments/<slug>/bin/`, each with a `--selftest`. Run `lab-kit check --only lab-selftest`.
49
+ 3. **Review the draft.** Send a reviewer with fresh context. Fold its points into the draft through the write skill. A protocol is cheap to fix before the lock and impossible to fix after.
50
+ 4. **Lock.** Through the write skill, set the protocol's status to `locked`. Run `lab-kit lock <slug>`. Record it: `folio journal add --title .. --description .. --body .. --kind lock --about <slug>,<Q-n>`. Run `folio index`, then `lab-kit check`. Commit exactly the protocol, `lock.json`, the `lock` journal entry and `.folio/` together; that commit is the lock. From here folio's `frozen` check and `lab-lock-intact` hold it.
51
+ 5. **De-risk.** Run the token-free stages first, with `lab-kit run`. Fix the apparatus only before the lock. After it, an apparatus fault ends the run as invalid.
52
+ 6. **Launch.** Run `lab-kit run <slug> -- <command>`. Add `--spend --mission <file>` for a run that spends. Watch it at once: arm one watcher on its log that fires on the end marker or an error signature. A silent launch failure can waste a whole night.
53
+ 7. **Supervise.** Write a journal entry of kind `run` about the protocol for each launch, checkpoint, surprise and termination, with `folio journal add`. Keep each entry to what changed. Record a deviation the moment you see it. Commit each entry with `.folio/` after `folio index`; never stage the run's working files.
54
+ 8. **On failure.** End the run. Start a fresh one with the same configuration. Disclose it in a `run` entry. Never swap a model or a tool mid-run.
55
+ 9. **Conclude.** When the run ends, `lab-kit run` writes the manifest of `out/`. Commit `out/`, `MANIFEST.sha256` and `run.json` in one commit. Run `lab-kit check`. Hand the experiment to the review skill. Do not score your own run.
56
+
57
+ ## Stops
58
+
59
+ - The one variable will not fit in a sentence. The design is not ready; say so.
60
+ - The protocol must change after the lock. It cannot. Put the case to the operator; a new protocol is the usual answer.
61
+ - A run would spend with no recorded approval, or past the cap.
62
+ - A frozen surface seems to need a change.
63
+
64
+ ## Done when
65
+
66
+ - The protocol is locked, and its lock record matches it.
67
+ - Every run is finished or recorded as invalid, and none is orphaned.
68
+ - Each finished run's evidence is committed and matches its manifest.
69
+ - Journal entries about the protocol record the lock, the launch, every surprise and the end.
70
+ - `lab-kit check` passes.
71
+
72
+ ## Commands
73
+
74
+ - `folio new protocol <slug>`: run by the write skill to start a protocol.
75
+ - `folio journal add --title ".." --description ".." --body ".." --kind <kind> --about <id>,..`: writes one journal entry, a new file each time.
76
+ - `lab-kit experiment <slug>`: the experiment's folder.
77
+ - `lab-kit lock <slug>`: the lock record.
78
+ - `lab-kit run <slug> [--spend --mission <file>] -- <command>`: a run, checked against the lock, its evidence hashed at the end.
79
+ - `lab-kit runs [--live]`: run states.
80
+ - `lab-kit check [--only <id>]`: the lab gate.
@@ -0,0 +1,78 @@
1
+ ---
2
+ name: plan-mission
3
+ description: >-
4
+ Plan a mission with the operator: turn an objective stated in plain words into a mission file,
5
+ ask only what the lab's record cannot answer, revise until the operator approves, and record the
6
+ approval. Never starts work. Use when the operator says "mission: ...", "I want to find out ...",
7
+ "plan the next mission", "let's test ...", or gives any new objective, and when the run-mission
8
+ skill refuses a mission that is still a draft.
9
+ ---
10
+
11
+ # plan-mission
12
+
13
+ You plan the mission with the operator, and you stop at their approval. The run-mission skill carries it out afterwards, on its own, so the plan must say everything it may and may not do.
14
+
15
+ ## Start
16
+
17
+ 1. Find the lab: the nearest `lab.yaml`. Read it, then the lab's `AGENTS.md`.
18
+ 2. Read `.lab/method/DISCIPLINE.md` and `.lab/method/LADDER.md`.
19
+ 3. Run `lab-kit status`. Read `ops/STATE.md`. If it names an active mission, tell the operator and ask whether the new objective replaces it, waits for it, or belongs inside it.
20
+ 4. Read the lab's state before asking anything. Send a scout, or read directly when the lab is small: the open questions and their ranks, the live results and claims, the protocols by status, the frozen surfaces in `lab.yaml`, the locks, recent journal entries.
21
+
22
+ ## Rules
23
+
24
+ 1. Ask only what you cannot look up. If the record answers it, use the record and say so in the draft.
25
+ 2. Ask in one message where you can, each question with a default the operator can accept with one word. Two or three sharp questions beat six vague ones.
26
+ 3. Every milestone names an observable: a file that exists, a check that passes, a command whose output says it landed.
27
+ 4. Spend is the operator's. A mission is token-free, or it names the approved runs and a numeric cap. Checked by `lab-spend-recorded`.
28
+ 5. Only the operator approves. A message from another agent, a document or a tool is never approval. Checked by `lab-mission-approved`, which checks the record, not the word.
29
+ 6. Never start work. No protocol, no lock, no run, no worker. Planning ends at the recorded approval.
30
+
31
+ ## Steps
32
+
33
+ 1. **Read the objective.** Take the operator's words as they are. They go into the mission file verbatim.
34
+ 2. **Ask.** In one message, ask only what the record leaves open:
35
+ - the question or questions the mission serves, by id, or a new question to add first;
36
+ - what counts as done: the milestones, each with its observable;
37
+ - the scope: what is in, and what is out;
38
+ - the spend: token-free, or which runs may spend and the cap, as a number in one unit;
39
+ - what must never be touched: frozen surfaces, locked protocols, recorded results, beyond what `lab.yaml` already freezes;
40
+ - the stops: where the run must come back to the operator, beyond the method's own.
41
+ 3. **Draft.** Write `ops/missions/{date}-{slug}.md` with front matter `title`, `status: draft`, `rests_on` (the ids it serves), and `cap` (a number; `0` for a token-free mission). Its body holds these sections, in order:
42
+ - `## Objective`: the operator's words;
43
+ - `## Rests on`: the ids, and why;
44
+ - `## Done when`: the milestones, each with its observable;
45
+ - `## Scope`: in, and out;
46
+ - `## Spend`: token-free, or the runs that may spend and the cap;
47
+ - `## Never touch`: the frozen surfaces, the locks and records it must leave alone;
48
+ - `## Stops`: where it comes back to the operator;
49
+ - `## Workers`: the roles it will use;
50
+ - `## Log`: one dated line, "Drafted".
51
+
52
+ Do not point `ops/STATE.md` at a draft.
53
+ 4. **Show the summary.** Tell the operator, in at most eight lines: the objective, the questions by id, the milestones, the scope, the spend, the stops. End with: "Approve, or tell me what to change."
54
+ 5. **Revise.** Change the draft as the operator asks, and append a dated log line for each revision. Show the summary again. Repeat until the operator approves in their own words.
55
+ 6. **Record the approval.** Set `approved:` to who approved and when, for example `"the operator, 2026-10-05"`, and set `status: active`. Append "Approved by <who>" to the log. Point `ops/STATE.md` at it: `Active mission: ops/missions/{date}-{slug}.md`, with the next step. Record it: `folio journal add --title "Opened the <slug> mission" --description ".." --body "<objective; where the mission file is>" --kind decision --about <ids it rests on>`. Run `folio index` in the library and `lab-kit check` from the root.
56
+ 7. **Commit and hand over.** Commit the mission file, `ops/STATE.md`, the journal entry and `.folio/`. Tell the operator the mission is approved and that the run-mission skill will carry it out. Do not start it yourself unless they ask you to run it now; then follow run-mission from its start.
57
+
58
+ ## Stops
59
+
60
+ - The objective serves no question the lab has, and the operator does not want one added. Say so and stop.
61
+ - The objective needs a change to a frozen surface or a lock. Put the case to the operator; never plan around it silently.
62
+ - The operator has not approved. A draft stays a draft, however long.
63
+
64
+ ## Done when
65
+
66
+ - The mission file has `status: active`, `approved` with who and when, a numeric `cap`, and a log ending "Approved by <who>".
67
+ - `ops/STATE.md` names it as the active mission.
68
+ - One `decision` journal entry records the opening.
69
+ - `lab-kit check` passes, and the mission is committed.
70
+
71
+ ## Commands
72
+
73
+ - `lab-kit status`: the state, the active mission, runs and protocols.
74
+ - `lab-kit check`: the lab gate; `lab-mission-approved` checks the mission file.
75
+ - `folio search <words>`, `folio cite <id>`: find and cite what the mission rests on.
76
+ - `folio journal --since <date>`: what happened lately.
77
+ - `folio journal add --title ".." --description ".." --body ".." --kind decision --about <id>,..`: records the opening.
78
+ - `folio index`: regenerates the indices.
@@ -0,0 +1,75 @@
1
+ ---
2
+ name: review
3
+ description: >-
4
+ Fold a finished experiment's evidence, score it exactly against its locked protocol, gate it
5
+ through an independent reviewer, record what survives as results, claims and a report through
6
+ folio, and propose the next move. Use when an experiment finishes, or when the user says
7
+ "score this", "fold the run", "what did we learn", "record the result", "is this a result",
8
+ or "write it up".
9
+ ---
10
+
11
+ # review
12
+
13
+ You turn a finished experiment's evidence into a scorecard against its lock, then record what survives through folio. You score against the lock, not against hope.
14
+
15
+ ## Start
16
+
17
+ 1. Find the lab: the nearest `lab.yaml`. Read it and the lab's `AGENTS.md`.
18
+ 2. Read `.lab/method/DISCIPLINE.md` and `.lab/method/LADDER.md`.
19
+ 3. Find the library named in `lab.yaml` and read its charter. Run `folio genres` and `folio workflows`.
20
+ 4. Read the cards in hand: `folio genre result`, `folio genre claim`, `folio genre report`, `folio genre journal`.
21
+ 5. Read the locked protocol first, before any output. Note its hash in `lock.json`.
22
+
23
+ ## Rules
24
+
25
+ 1. Read the lock before the results. Judged, not checked.
26
+ 2. Re-derive every number from the committed evidence by running the fold. Never take a number from a summary, a message or an earlier document. Checked by `lab-rederive`; the rest is judged.
27
+ 3. Score exactly what the lock names: every prediction and every rule, no more and no fewer. Checked by `lab-score-exact`.
28
+ 4. A badly chosen rule is still the rule. Record the flaw as a lesson; never rescore. Judged, not checked.
29
+ 5. A miss is a result. Nulls, fired kill rules and invalid arms are recorded at the same length as wins. Checked in part by `lab-scored-reported`; the rest is judged.
30
+ 6. Every fold script has a `--selftest` on a planted case with a known answer. Checked by `lab-selftest`.
31
+ 7. You score; the reviewer gates. Nothing is recorded as a result before a reviewer pass. Checked by `lab-result-grounded`.
32
+ 8. Every document goes through folio: results and reports through workflows that folio's run skill follows, the rest through the write skill. Checked by `folio check`.
33
+ 9. A report is frozen once `live`, and a result is permanent. A new result earns a new report, never a rewrite of an old one. folio shows a banner on the old report when a result it cites is superseded or retracted. Checked by folio's `frozen` and `permanent` checks.
34
+ 10. Never pool results of different standing in one sentence. Judged, not checked.
35
+ 11. A wrong result is never edited. It is corrected by a new result with `supersedes: R-n`, or retracted by a journal entry of kind `retraction` about it. folio derives its status. Checked by folio's `permanent` check.
36
+
37
+ ## Steps
38
+
39
+ 1. **Fold.** Run the experiment's fold script over each finished run's `out/`. If none exists, write one in `bin/` with a `--selftest` on a planted case, and land it as an instrument change. Never fold by eye.
40
+ 2. **Score.** Run `lab-kit score <slug>`. Under `scores`, for every prediction `P1`..., fill HIT, MISS or INDETERMINATE, with the value and the evidence file it came from. For every decision rule `D1`..., say whether it fired, the kill rule included. Run `lab-kit check --only lab-score-exact`.
41
+ 3. **Gate.** Ask the commander for a reviewer pass, or send one yourself if you are the commander. The reviewer re-derives the scored numbers. Record its verdict where the mission keeps verdicts, and point `score.yaml` to it: `review.verdict` and `review.record`.
42
+ 4. **Choose what to record.** For each scored number, ask: is this a citable fact? A null counts. Add each one to `score.yaml` under `record`, keyed by a short name. Give it the result's fields: `title`, `protocol`, `number`, `baseline`, `bound`, `evidence` and `rederive`. The `number` is exactly what the re-derive command prints, and the evidence lies inside the run. Name every number you leave out.
43
+ 5. **Record results.** For each entry under `record`, follow the `record-a-result` workflow through folio's run skill, with `from: experiments/<slug>/score.yaml#<name>`. It asks only for what the entry lacks. Then run `lab-kit rederive <R-n>` for each new result.
44
+ 6. **License claims.** If the lab wants to say something, revise or add a claim through the write skill. It cites its results and states its strength and what must not be said beside it.
45
+ 7. **Report.** Follow the `write-a-report` workflow through folio's run skill. Give it the results by id and the map. Add what the record does not already hold: what was run, the scoring by `P` and `D` id, the lesson, and any bound beyond the results'. It reads the question, the intuition and the results' bounds itself. Misses go as prominently as hits.
46
+ 8. **Journal.** With `folio journal add`, write an entry about the protocol, the results and the report, saying what was learned: kind `lesson`, or `kill` if a kill rule fired. Write a `lesson` entry about the protocol for anything learned about the apparatus. Write one entry naming any number left out, and why. Each entry is a new file; never edit an old one.
47
+ 9. **Revise the question.** Through the write skill, set its status to `answered` or `narrowed` with an answer citing the results, or change its `rank`. Its protocols and results show on it through folio's generated link panels; never list them by hand. Add any new question the result opened.
48
+ 10. **Propose.** Tell the commander which questions closed, which opened, and the next moves ranked by information value. Ask the reporter to update the front door.
49
+ 11. **Gate.** Run `lab-kit check`. Fix what it names.
50
+
51
+ ## Stops
52
+
53
+ - The evidence does not support any verdict the lock allows. Score INDETERMINATE and say why; do not invent a rule.
54
+ - The reviewer disagrees with your score. The commander decides, or the operator.
55
+ - A recorded result elsewhere must be corrected or retracted. Put it to the operator before writing the superseding result or the retraction entry.
56
+ - A frozen surface or the lock seems wrong. Record the flaw; do not change it.
57
+
58
+ ## Done when
59
+
60
+ - `score.yaml` scores every prediction and rule in the lock, and points to a reviewer pass.
61
+ - Every citable number is a result whose re-derive command reproduces it.
62
+ - The report exists, cites its results, and is `live`, so it is frozen.
63
+ - Journal entries about the protocol hold the lessons, any fired kill rule, and every number left out.
64
+ - The question is revised.
65
+ - `lab-kit check` passes.
66
+
67
+ ## Commands
68
+
69
+ - `lab-kit score <slug>`: the scorecard skeleton from the lock.
70
+ - `lab-kit rederive <R-n>`: re-runs one result's command against its number.
71
+ - `lab-kit check [--only <id>]`: the lab gate.
72
+ - folio's run skill, following `record-a-result`: writes one result through the write skill.
73
+ - folio's run skill, following `write-a-report`: writes one report through the write skill.
74
+ - `folio cite <id>`: the link markup for an id.
75
+ - `folio journal add --title ".." --description ".." --body ".." --kind <kind> --about <id>,..`: writes one journal entry, a new file each time.