brevet 0.3.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (56) hide show
  1. brevet-0.3.0/ABOUT.md +434 -0
  2. brevet-0.3.0/BENCHMARK.md +116 -0
  3. brevet-0.3.0/CHANGELOG.md +172 -0
  4. brevet-0.3.0/CITATION.cff +25 -0
  5. brevet-0.3.0/GLOSSARY.md +56 -0
  6. brevet-0.3.0/LICENSE +201 -0
  7. brevet-0.3.0/MANIFEST.in +3 -0
  8. brevet-0.3.0/PKG-INFO +222 -0
  9. brevet-0.3.0/README.md +183 -0
  10. brevet-0.3.0/README_PYPI.md +175 -0
  11. brevet-0.3.0/SECURITY.md +30 -0
  12. brevet-0.3.0/SPEC.md +137 -0
  13. brevet-0.3.0/brevet/__init__.py +75 -0
  14. brevet-0.3.0/brevet/adapters.py +328 -0
  15. brevet-0.3.0/brevet/approvals.py +644 -0
  16. brevet-0.3.0/brevet/assist.py +92 -0
  17. brevet-0.3.0/brevet/canonical.py +106 -0
  18. brevet-0.3.0/brevet/chap_bridge.py +203 -0
  19. brevet-0.3.0/brevet/chap_evidence.py +433 -0
  20. brevet-0.3.0/brevet/cli.py +533 -0
  21. brevet-0.3.0/brevet/delta.py +256 -0
  22. brevet-0.3.0/brevet/demo.py +152 -0
  23. brevet-0.3.0/brevet/evals.py +67 -0
  24. brevet-0.3.0/brevet/evidence.py +73 -0
  25. brevet-0.3.0/brevet/identity.py +34 -0
  26. brevet-0.3.0/brevet/ledger.py +144 -0
  27. brevet-0.3.0/brevet/lifecycle.py +409 -0
  28. brevet-0.3.0/brevet/mcp_server.py +527 -0
  29. brevet-0.3.0/brevet/models.py +302 -0
  30. brevet-0.3.0/brevet/playground.py +458 -0
  31. brevet-0.3.0/brevet/runner.py +92 -0
  32. brevet-0.3.0/brevet/shell.py +295 -0
  33. brevet-0.3.0/brevet/workdir.py +82 -0
  34. brevet-0.3.0/brevet.egg-info/PKG-INFO +222 -0
  35. brevet-0.3.0/brevet.egg-info/SOURCES.txt +54 -0
  36. brevet-0.3.0/brevet.egg-info/dependency_links.txt +1 -0
  37. brevet-0.3.0/brevet.egg-info/entry_points.txt +2 -0
  38. brevet-0.3.0/brevet.egg-info/requires.txt +23 -0
  39. brevet-0.3.0/brevet.egg-info/top_level.txt +1 -0
  40. brevet-0.3.0/pyproject.toml +68 -0
  41. brevet-0.3.0/schemas/agent_manifest.schema.json +154 -0
  42. brevet-0.3.0/schemas/capabilities_lock.schema.json +33 -0
  43. brevet-0.3.0/schemas/capability_object.schema.json +165 -0
  44. brevet-0.3.0/schemas/recall_notice.schema.json +37 -0
  45. brevet-0.3.0/schemas/release_record.schema.json +37 -0
  46. brevet-0.3.0/setup.cfg +4 -0
  47. brevet-0.3.0/tests/test_adapters.py +131 -0
  48. brevet-0.3.0/tests/test_approvals.py +301 -0
  49. brevet-0.3.0/tests/test_chap_evidence.py +234 -0
  50. brevet-0.3.0/tests/test_chap_evidence_missing_source.py +25 -0
  51. brevet-0.3.0/tests/test_core.py +143 -0
  52. brevet-0.3.0/tests/test_hardening.py +442 -0
  53. brevet-0.3.0/tests/test_playground.py +78 -0
  54. brevet-0.3.0/tests/test_release_artifacts.py +62 -0
  55. brevet-0.3.0/tests/test_runner_assist_mcp.py +204 -0
  56. brevet-0.3.0/tests/test_schemas.py +73 -0
brevet-0.3.0/ABOUT.md ADDED
@@ -0,0 +1,434 @@
1
+ # About Brevet
2
+
3
+ The detail behind the [README](README.md): how the governed evolution loop
4
+ works, how Brevet is built and how the repository is organised. For
5
+ definitions of every term, see the [GLOSSARY](GLOSSARY.md). For the rules any
6
+ implementation must follow, see [SPEC.md](SPEC.md).
7
+
8
+ ## Why Brevet exists
9
+
10
+ Organisations already know how to manage change when people learn on the
11
+ job. When someone wants to change how a task is done, they write the change
12
+ down, a named person approves it, the procedure gets a new version number,
13
+ and a bad change can be rolled back. Those steps are how trust in a team
14
+ extends beyond the people who know each other personally.
15
+
16
+ AI agents now learn on the job too. They save memories, add skills and edit
17
+ their own instructions, but those changes skip every one of those steps.
18
+ That leaves three questions without answers:
19
+
20
+ 1. **What has the agent learned?** Not the model's weights, but the rules,
21
+ skills, memories and tool permissions it has picked up since it was
22
+ deployed.
23
+ 2. **Who approved it?** A learned behaviour changed a real decision. Which
24
+ accountable person said it could be used, and on what evidence?
25
+ 3. **How do we take it back?** A learned rule is wrong. Where does it live,
26
+ which versions included it, and what shows it is no longer in use?
27
+
28
+ Brevet gives an agent's learning the same steps a team's procedures already
29
+ have. It is not an agent framework and does not replace one. It wraps the
30
+ framework you already use and keeps the records around it.
31
+
32
+ ## The governed evolution loop
33
+
34
+ The loop has seven stages. Each one appends its envelopes to the evidence
35
+ chain.
36
+
37
+ | Stage | What happens | Who acts | What is recorded |
38
+ |---|---|---|---|
39
+ | **Work** `run()` | The agent drafts under one signed harness. | the agent | `brevet.task`, `brevet.artefact` |
40
+ | **Override** `record_final()` | An expert corrects the draft; the difference and the reason become an override. | an expert | `brevet.override` |
41
+ | **Dream** `dream()` | Offline, recurring overrides become candidate capabilities and eval cases, with no authority. | Brevet | `brevet.candidate` |
42
+ | **Dawn** `dawn()` | Each candidate is promoted, held, rejected or sent back for re-elicitation. | a named human or mission group | `brevet.promotion` |
43
+ | **Evals** `evaluate()` | Overrides replay as tests. The conservative gate needs neither half (held-in or held-out) to get worse and at least one to improve. | Brevet | `brevet.eval_run` |
44
+ | **Release** `release()` | Promoted capabilities ship in a signed release with its `capabilities.lock`. | a named human or mission group | `brevet.release` |
45
+ | **Recall** `recall()` | A capability is withdrawn, and every release that shipped it is flagged. | a named human or mission group | `brevet.recall` |
46
+
47
+ The names follow the loop's day-and-night rhythm, which the paper calls the
48
+ circadian contract. Awake, the agent works under one signed version and does
49
+ not change itself. While it sleeps, the dream cycle mines what the day's
50
+ overrides imply. At dawn, people review what the night proposed, and only what
51
+ they promote reaches the agent, through a signed release.
52
+
53
+ Overrides come in two kinds. A **refining** override keeps the agent's decision
54
+ and changes its wording. A **substituting** override reaches a different
55
+ decision, as when *minor* becomes *major*. The dream cycle treats them as soft
56
+ and hard signals.
57
+
58
+ The dream cycle adds only what is new, so it can run every night. A group of
59
+ overrides that already produced a candidate is not proposed again, whatever
60
+ happened to that candidate since. When more overrides join the group, the new
61
+ candidate supersedes the pending one, and a candidate whose content matches a
62
+ recalled capability is never proposed.
63
+
64
+ ## The authority ladder
65
+
66
+ <p align="center">
67
+ <picture>
68
+ <source media="(prefers-color-scheme: dark)" srcset="docs/assets/levels-dark.svg">
69
+ <img src="docs/assets/levels-light.svg" alt="The authority ladder: Evidence holds candidates with no operational authority; promotion at the dawn gate by a named human raises a capability to Advisory, where it may inform drafts a human still checks; promotion by the mission group raises it to Controlled, where it may drive actions directly. A recalled capability leaves future releases and every release that shipped it is flagged." width="860">
70
+ </picture>
71
+ </p>
72
+
73
+ Every capability holds one authority layer. It enters at **Evidence**, with no
74
+ operational authority. Promotion at the dawn gate raises it to **Advisory**,
75
+ where it may inform what the agent drafts while a human still checks each
76
+ result. **Controlled**, where it may drive actions directly, needs promotion
77
+ by a mission group. Only `human:` and `mission_group:` identities can decide;
78
+ every other namespace, including `agent:`, `model:` and `dream:`, is refused.
79
+ Rejection is not deletion: rejected candidates stay on the evidence chain, so
80
+ "was this ever proposed, and why did we say no?" always has an answer.
81
+
82
+ ## Signed approvals
83
+
84
+ An identity such as `human:alice@example.com` is a name anyone can type,
85
+ including an agent that can call Brevet's tools. Signed approvals make each
86
+ decision provable: once a workspace registers its first approver, every dawn
87
+ decision, release and recall must be signed with a key that an agent does
88
+ not hold.
89
+
90
+ ```console
91
+ brevet approver add --identity human:alice@example.com --group mission_group:quality_team
92
+ brevet dawn --decide cap_123:promote --approver mission_group:quality_team
93
+ brevet approve req_7f3a9c2e41d0 # shows the request, asks for the passphrase, signs
94
+ ```
95
+
96
+ - **Keys stay with people.** `brevet approver add` creates an Ed25519 key,
97
+ encrypted with a passphrase, in `~/.config/brevet/approvers/` (or
98
+ `BREVET_APPROVER_DIR`), outside every workspace.
99
+ - **Request, sign, apply.** Anyone may request a decision, an agent over MCP
100
+ included; nothing changes until the people it needs sign it with
101
+ `brevet approve`. The signed request is bound to the exact decision: the
102
+ capability's content hash for a promotion, the lock's digest for a release.
103
+ A release whose promotions changed after signing is refused, and a signed
104
+ request can be applied only once.
105
+ - **Mission groups sign through their members.** A decision taken as
106
+ `mission_group:<name>` needs signatures from that group's members, one by
107
+ default or more with `brevet approver threshold`.
108
+ - **The register is on the evidence chain.** Each registration, revocation
109
+ and threshold is a `brevet.approver` envelope. The first approver registers
110
+ themselves; every later change needs an existing approver's signature, and
111
+ a new key signs its own registration. `brevet verify` replays the register
112
+ and checks every signature as it stood at the time.
113
+
114
+ A signature proves that the holder of a registered key signed. Binding each
115
+ key to a verified person, and keeping it safe, is left to the deployment.
116
+
117
+ ## The worked example
118
+
119
+ [examples/pump_vibration.py](examples/pump_vibration.py) runs the README's
120
+ example from the first override to the recall. Running it prints:
121
+
122
+ ```text
123
+ 1. Work and override: 4 overrides recorded from 5 drafts (1 draft accepted as it was).
124
+ 2. Dream: 1 candidate capability, Evidence layer (no authority yet):
125
+ In equipment_triage work involving 'vibration-during-cleaning', reviewers changed
126
+ the agent's decision 4 times. Their reason: Vibration during cleaning is an early
127
+ sign of seal wear. Proposed rule: when this situation applies, raise it explicitly
128
+ and follow the reviewers' decision.
129
+ 3. Dawn: rejected an approval from dream:nightly (machine identities cannot promote).
130
+ Dawn: promoted to Advisory by mission_group:quality_team.
131
+ 4. Evals: 0/4 passed before, 4/4 after. Conservative gate: pass.
132
+ 5. Release: 0.2.0 signed; capabilities.lock lists 1 promoted capability and its approver.
133
+ 6. Recall: capability recalled; releases flagged: 0.2.0.
134
+ 7. Verify: evidence chain intact.
135
+ Records are in <temporary folder>
136
+ ```
137
+
138
+ ## What is in the box
139
+
140
+ - **The harness as data.** `agent.yaml`, the manifest, holds the agent's
141
+ instructions, tools, loop policy, memory bindings and safety settings in
142
+ five layers. It is versioned and signed at every release.
143
+ - **One capability object for everything learned.** Prompt rules, skills,
144
+ tool bindings, eval cases, escalation rules and memory bindings share one
145
+ structure, one authority ladder and one recall mechanism.
146
+ - **Override-compiled evals.** `agent.evaluate()` replays the overrides as
147
+ tests, so no separate labelling project is needed. Cases are sorted by
148
+ content hash and dealt alternately into the held-in and held-out halves.
149
+ - **A capability bill of materials for every release.** `capabilities.lock`
150
+ lists what the release knows, where each capability came from, the hash of
151
+ its exact content and who approved it. The release signature covers the
152
+ lock's digest.
153
+ - **Optional model assist.** `brevet dream --assist ollama` lets a local
154
+ model word candidates more clearly. Its output is always a draft, every use
155
+ is logged, and Brevet works fully without it.
156
+ - **CHAP mirroring.** Set `ledger: chap:<workspace>` in `agent.yaml`, install
157
+ the extra with `pip install "brevet[chap]"`, and every evidence envelope is
158
+ also mirrored to a CHAP coordinator running alongside, with its own database
159
+ in the working directory. To use a coordinator on another machine, write
160
+ `ledger: chap:<workspace>@<url>`; envelopes queue in an outbox while the
161
+ connection is down. The local evidence chain remains the record of
162
+ reference.
163
+
164
+ ## Supported frameworks
165
+
166
+ <p align="center">
167
+ <picture>
168
+ <source media="(prefers-color-scheme: dark)" srcset="docs/assets/wrap-dark.svg">
169
+ <img src="docs/assets/wrap-light.svg" alt="brevet.wrap(agent) places the agent, unchanged, inside four records: the signed manifest, the hash-linked evidence chain, the dawn gate, and capabilities.lock." width="860">
170
+ </picture>
171
+ </p>
172
+
173
+ `brevet.wrap()` works out which framework built the agent from the agent
174
+ object itself. The framework keeps running the agent; Brevet keeps the
175
+ records.
176
+
177
+ | Framework | How to wrap it |
178
+ |---|---|
179
+ | LangGraph | `brevet.wrap(compiled_graph)` |
180
+ | Claude Agent SDK | `brevet.wrap(sdk_client)` |
181
+ | DeepAgents | `brevet.wrap(deep_agent)` |
182
+ | AutoGen (AgentChat) | `brevet.wrap(agent_or_team)` |
183
+ | LlamaIndex | `brevet.wrap(agent_or_engine)` |
184
+ | Pydantic AI | `brevet.wrap(pydantic_agent)` |
185
+ | Google Agent Development Kit | `brevet.wrap(runner)` |
186
+ | CrewAI | `brevet.wrap(crew)` |
187
+ | OpenAI Agents SDK | `brevet.wrap(agent)` |
188
+ | any Python function | `brevet.wrap(fn)` |
189
+
190
+ Another framework needs one adapter class, registered with
191
+ `brevet.register_adapter`:
192
+
193
+ ```python
194
+ @brevet.register_adapter("myfw", prefixes=("myfw",))
195
+ class MyAdapter(brevet.BaseAdapter):
196
+ def invoke(self, task, context):
197
+ return self.target.do(task), [{"step": "do"}]
198
+ ```
199
+
200
+ No framework is ever a required dependency: adapters check the shape of the
201
+ object they are given, and the tests use stand-ins rather than the real
202
+ frameworks. Outside the shadow channel, a wrapped agent runs only while its
203
+ manifest signature verifies, so an edit made after a release stops the agent
204
+ instead of running unrecorded.
205
+
206
+ ## The MCP server
207
+
208
+ `brevet mcp` serves the loop to any MCP client, such as Claude Desktop, Claude
209
+ Code or Cursor. Clients may start servers from any directory, so point
210
+ `BREVET_HOME` at a folder that holds `agent.yaml` and `.brevet/`:
211
+
212
+ ```json
213
+ {"mcpServers": {"brevet": {
214
+ "command": "uvx", "args": ["brevet", "mcp"],
215
+ "env": {"BREVET_HOME": "/path/to/workspace"}}}}
216
+ ```
217
+
218
+ | Tool | What it does |
219
+ |---|---|
220
+ | `brevet_record` | records a draft and the expert's final; a difference becomes an override |
221
+ | `brevet_chap_ingest` | imports CHAP review verdicts as overrides |
222
+ | `brevet_dream` | runs the dream cycle |
223
+ | `brevet_dawn_pending` | lists the candidates awaiting a dawn decision |
224
+ | `brevet_dawn_decide` | records one dawn decision under a human or mission-group identity |
225
+ | `brevet_release` | signs and records a release, with the conservative gate for trial and production |
226
+ | `brevet_recall` | recalls a capability and flags the releases that shipped it |
227
+ | `brevet_active` | serves the governed rules of the latest release, after checking the chain, the lock, the signature and each rule's content |
228
+ | `brevet_status` | reports the version, capabilities by layer, the dawn queue and chain health |
229
+ | `brevet_verify` | replays the evidence chain |
230
+
231
+ Read-only tools carry the MCP read-only hint, and the dawn, release and recall
232
+ tools carry the destructive hint, so clients can ask before running them. When
233
+ the workspace requires signed approvals, those three tools return a request
234
+ and the command that signs it, and nothing changes until a person signs.
235
+ Sessions record corrections only when asked, unless the workspace owner turns
236
+ on automatic capture with `BREVET_AUTO_CAPTURE=1` or
237
+ `runtime_safety.evidence.auto_capture: true` in the manifest.
238
+
239
+ ## Commands
240
+
241
+ | Command | What it does |
242
+ |---|---|
243
+ | `brevet demo` | runs the whole loop once on synthetic data, offline |
244
+ | `brevet playground` | steps through the loop in a browser |
245
+ | `brevet init` | writes a starter `agent.yaml` and `.brevet/` |
246
+ | `brevet dream` | mines recurring overrides into candidates |
247
+ | `brevet dawn` | lists the dawn queue, or records one decision with `--decide` |
248
+ | `brevet release` | passes the conservative gate, then signs and records a release |
249
+ | `brevet recall` | recalls a capability and flags every release that shipped it |
250
+ | `brevet verify` | replays the evidence chain |
251
+ | `brevet status` | shows the version, capabilities by layer and chain health |
252
+ | `brevet chap-ingest` | imports CHAP review verdicts as overrides |
253
+ | `brevet mcp` | serves the loop to an MCP client |
254
+ | `brevet approver` | registers approvers, revokes them and sets mission-group thresholds |
255
+ | `brevet approve` | lists the requests waiting for signatures, or signs them |
256
+
257
+ ## Governing what Claude itself learns
258
+
259
+ Claude Desktop and Cowork already learn between sessions through memory, saved
260
+ skills and standing instructions. [examples/claude-cowork](examples/claude-cowork)
261
+ puts that learning under the governed evolution loop in a few minutes. Your
262
+ corrections become overrides, you promote candidates at dawn, and the governed
263
+ rules Claude receives come only from a signed release. In Claude these
264
+ controls detect rather than prevent: Claude's own memory and skills keep
265
+ working outside Brevet, so the example makes them visible and reviewable
266
+ rather than impossible.
267
+
268
+ ## Where Brevet fits
269
+
270
+ <p align="center">
271
+ <picture>
272
+ <source media="(prefers-color-scheme: dark)" srcset="docs/assets/suite-dark.svg">
273
+ <img src="docs/assets/suite-light.svg" alt="CHAP records what happened, Metis captures what experts know, and Brevet governs how the agent changes" width="860">
274
+ </picture>
275
+ </p>
276
+
277
+ | Project | Question it answers | What it owns |
278
+ |---|---|---|
279
+ | [CHAP](https://github.com/BrightbeamAI/chap) | What happened between people and agents? | the record of tasks, drafts, reviews and decisions |
280
+ | [Metis](https://github.com/BrightbeamAI/metis) | What do our experts know that is not written down? | experts' know-how, captured as governed memory |
281
+ | **Brevet** | How does the agent change, and on whose approval? | approvals, capability state and recalls |
282
+
283
+ The three are designed to work together, and each owns its own records.
284
+ Brevet's evidence envelopes follow CHAP's envelope model, so Brevet writes
285
+ CHAP-compatible evidence and can mirror it to a live CHAP coordinator;
286
+ `brevet chap-ingest` imports CHAP review decisions as overrides. Brevet's
287
+ capability object extends the tuple Metis uses for tacit fragments, adding a
288
+ `kind` field. For `memory_fragment` capabilities, Metis stays the system of
289
+ record; Brevet records only the binding and its approval.
290
+
291
+ ## What Brevet is not
292
+
293
+ Brevet is not an agent framework. It does not let an agent improve itself
294
+ without oversight: candidates have no authority until a human or mission
295
+ group promotes them, and approvals under machine identities are refused. It
296
+ is not a monitoring tool either. The manifest declares whose overrides the
297
+ dream cycle may learn from (its consent scope), and Evidence-layer material
298
+ never influences the agent; the dream cycle does not yet filter overrides by
299
+ that declaration, so a deployment applies it at capture.
300
+
301
+ ## Project status
302
+
303
+ Brevet implements the whole loop and keeps every record. Some protections
304
+ are left to the system you deploy it in, and the [paper](README.md#citation)
305
+ sets them out in full:
306
+
307
+ - **Verified approvers.** With signed approvals, each decision is signed by
308
+ a registered key. Making sure that key belongs to the named person, through
309
+ identity proofing, custody and recovery, is the deployment's job; CHAP
310
+ participant keys or an organisation's single sign-on can supply it. Without
311
+ registered approvers, Brevet records the identity given and refuses
312
+ anything other than `human:` and `mission_group:`.
313
+ - **Recall that reaches running agents.** A recalled capability, and any
314
+ capability with identical content, is left out of every later release, and
315
+ the releases that shipped it are flagged. Confirming that running agents
316
+ have stopped using it needs checks where the agent runs.
317
+ - **An anchored evidence chain.** Replay detects edits that break the chain.
318
+ Detecting a wholesale rewrite, or a chain cut short at the end, needs the
319
+ latest chain hash held outside the machine, which a CHAP coordinator can
320
+ hold.
321
+ - **Measured eval deltas.** The conservative gate checks the deltas passed to
322
+ `release()`. `agent.evaluate()` measures them, but nothing yet ties a release
323
+ to the eval runs that produced its numbers.
324
+
325
+ ## How the repository is organised
326
+
327
+ ```
328
+ brevet/
329
+ ├── brevet/ the Python package
330
+ │ ├── shell.py brevet.wrap() and the agent object it returns
331
+ │ ├── adapters.py framework adapters and the adapter registry
332
+ │ ├── evidence.py harvests overrides from drafts and finals
333
+ │ ├── delta.py the dream cycle: mines overrides into candidates
334
+ │ ├── lifecycle.py the dawn gate, releases and recall
335
+ │ ├── evals.py override-compiled evals; the conservative gate
336
+ │ ├── runner.py runs the evals before and after a change
337
+ │ ├── ledger.py the hash-linked evidence chain
338
+ │ ├── canonical.py content hashing and Ed25519 signing
339
+ │ ├── identity.py which identities may decide
340
+ │ ├── approvals.py approver keys, the register and signed approvals
341
+ │ ├── workdir.py workspace location, permissions and file locks
342
+ │ ├── models.py data models matching the schemas
343
+ │ ├── assist.py optional local model assist
344
+ │ ├── chap_bridge.py mirrors envelopes to a CHAP coordinator
345
+ │ ├── chap_evidence.py imports CHAP review decisions as overrides
346
+ │ ├── mcp_server.py the loop as MCP tools
347
+ │ ├── playground.py the loop, stage by stage, in a browser
348
+ │ ├── demo.py the end-to-end demonstration
349
+ │ └── cli.py the `brevet` command
350
+ ├── schemas/ the five JSON Schemas that define the records
351
+ ├── examples/
352
+ │ ├── pump_vibration.py the worked example from the README
353
+ │ └── claude-cowork/ governing what Claude itself learns
354
+ ├── docs/
355
+ │ ├── demo.html the interactive tour
356
+ │ └── assets/ the diagrams, in light and dark versions
357
+ ├── scripts/
358
+ │ ├── make_diagrams.py regenerates the diagrams
359
+ │ └── make_pypi_readme.py writes README_PYPI.md for PyPI
360
+ ├── tests/ the test suite; runs offline
361
+ ├── server.json the MCP Registry entry
362
+ └── README.md · ABOUT.md · GLOSSARY.md · SPEC.md · BENCHMARK.md
363
+ ```
364
+
365
+ ## What Brevet stores, and where
366
+
367
+ `brevet.wrap()` keeps everything in a working directory, `.brevet/` by default
368
+ (change it with `workdir=`). Brevet creates it readable only by you, with a
369
+ `.gitignore` that keeps it out of version control:
370
+
371
+ ```
372
+ .brevet/
373
+ ├── agent.yaml the manifest, if none was supplied
374
+ ├── capabilities.lock the latest release's lock, beside the manifest
375
+ ├── ledger.jsonl the evidence chain: one envelope per line
376
+ ├── capabilities.jsonl the capability store; latest line per id wins
377
+ ├── keys/brevet_ed25519.pem the signing key, created at the first release
378
+ ├── approvals.jsonl decisions waiting for approver signatures
379
+ ├── chap_cursor.json how far each CHAP source has been imported
380
+ ├── chap.db an embedded CHAP coordinator's store, if used
381
+ ├── chap_outbox.jsonl envelopes waiting to be mirrored to CHAP
382
+ └── *.lock short-lived locks that let processes share files
383
+ ```
384
+
385
+ When a manifest path is given, `agent.yaml` and `capabilities.lock` live
386
+ beside it instead. Never commit `.brevet/`: it holds a private key and the
387
+ evidence of real work. Approver keys never live here; each person keeps
388
+ theirs in `~/.config/brevet/approvers/`.
389
+
390
+ ## Evidence envelopes
391
+
392
+ Every envelope is appended, never edited, and hash-linked to the one before
393
+ it: `chain_hash = sha256(encode(envelope) ‖ prev_hash)`. Appends take a file
394
+ lock, so several processes can share one chain.
395
+
396
+ | Envelope | Appended when |
397
+ |---|---|
398
+ | `brevet.task` | the agent is given a task |
399
+ | `brevet.artefact` | the agent produces a draft |
400
+ | `brevet.override` | an expert's correction is recorded |
401
+ | `brevet.candidate` | the dream cycle proposes a candidate |
402
+ | `brevet.model_assist` | a local model helps word a candidate |
403
+ | `brevet.promotion` | a decision is made at the dawn gate |
404
+ | `brevet.eval_run` | the evals run |
405
+ | `brevet.release` | a new version is released |
406
+ | `brevet.recall` | a capability is recalled |
407
+ | `brevet.approver` | an approver is registered or revoked, or a group threshold changes |
408
+
409
+ ## The benchmark
410
+
411
+ Research on self-improving agents mostly measures one thing: whether the
412
+ agent got better. Teams running agents in production need four measures,
413
+ and [BENCHMARK.md](BENCHMARK.md) proposes a governed-adaptation benchmark
414
+ that reports all four side by side: improvement, regression discipline,
415
+ lineage completeness and recall compliance. No implementation of the
416
+ benchmark exists yet.
417
+
418
+ ## Design principles
419
+
420
+ - **Authority is granted, never grabbed.** Every increase in a
421
+ capability's authority is a recorded decision by a named human or mission
422
+ group, signed by them once the workspace registers approvers.
423
+ - **The running agent does not change itself.** It runs one released
424
+ version. Learning happens between versions, where it can be reviewed.
425
+ - **Overrides are evidence, not truth.** Experts' corrections are the best
426
+ available record of judgement, but they can be mistaken, so candidates are
427
+ reviewed at dawn before they gain authority.
428
+ - **No network needed.** The tests, the demonstration and the whole loop run
429
+ offline. Model assist and CHAP mirroring are optional and fail safely.
430
+ - **Rejection is not deletion.** Rejected and recalled capabilities stay on
431
+ the evidence chain with their lineage intact.
432
+ - **Prove it from the chain.** Promotions, releases and recalls are all
433
+ envelopes on the evidence chain, so anyone holding it can replay and check
434
+ them.
@@ -0,0 +1,116 @@
1
+ # The Governed-Adaptation Benchmark (design specification)
2
+
3
+ Research on self-improving agents mostly measures one thing: whether the
4
+ agent got better. Teams that run agents in production need four
5
+ properties, and improvement is only the first. This document specifies a
6
+ benchmark that scores all four side by side. **No implementation of it
7
+ exists yet**; this is its design, as specified in the paper.
8
+
9
+ ## What a submission is given
10
+
11
+ A submission is any system that adapts an agent over time, whether or not it
12
+ uses Brevet. It receives:
13
+
14
+ - **a task stream**, partitioned by incident before any mining, so that
15
+ related cases never end up on both sides of a split. The stream includes
16
+ accepted work as well as corrected work;
17
+ - **an override stream** in CHAP's envelope format (the difference, the
18
+ rationale, tags and `intent_preserved`), replayed on a schedule. It
19
+ supplies human judgement where no automatic verifier exists;
20
+ - **an initial signed harness** `H0`;
21
+ - **revocation events** issued mid-stream, including one consent withdrawal
22
+ and one recall of a capability that an endpoint has already cached or that
23
+ a later candidate was derived from;
24
+ - **at least one pair of conflicting candidates** whose applicability
25
+ conditions overlap, so the scoring can see whether review detects the
26
+ interaction or promotes both; and
27
+ - **a consent policy** over override sources.
28
+
29
+ The submission runs its loop and produces a harness lineage `H0 … HN`, each
30
+ version with its lockfile, its promotion records and its evidence log.
31
+ Proposal generation and release selection may not see the final test split.
32
+
33
+ ## The four scored axes
34
+
35
+ 1. **Improvement.** The held-out pass-rate change across the lineage: how
36
+ much better the agent gets. This is the axis current research reports.
37
+ 2. **Regression discipline.** How often, and how badly, a change made
38
+ individual tasks or risk groups worse and still survived promotion.
39
+ Gains and losses are reported separately, because an average can hide a
40
+ loss behind a gain.
41
+ 3. **Lineage completeness.** For a sample of changes between two harness
42
+ versions, whether the records link each change to its capability, the
43
+ evidence for it, the eval run that tested it and the authenticated
44
+ approver. Audit replay scores completeness. Showing that a capability
45
+ actually *caused* a change in behaviour needs an extra ablation.
46
+ 4. **Recall compliance.** For each revocation event: how long until every
47
+ endpoint acknowledges the recall, whether the capability's content hash is
48
+ excluded from later lockfiles, and whether task-level checks find it still
49
+ in use, including through derived capabilities and cached material.
50
+ Inventory exclusion and operational cessation are scored separately.
51
+
52
+ Results are a profile, not a single number:
53
+ `(δperf, regressions, lineage %, recall)`. Submissions also report how many
54
+ candidates were rejected, held or blocked at each gate, so the cost of
55
+ governance is visible next to its benefit. A system that improves faster but
56
+ fails recall is not automatically better where revocation is required.
57
+
58
+ ## Record formats
59
+
60
+ A submission reads override-stream events and revocation events, and writes
61
+ one profile record per run. Representative values:
62
+
63
+ ```jsonc
64
+ // input: one override-stream event (CHAP envelope format)
65
+ { "seq": 214, "group_id": "incident_0031",
66
+ "envelope": { "kind": "brevet.override",
67
+ "body": { "participant": "human:qa@site", "intent_preserved": false,
68
+ "diff": [ ... ], "tags": ["vibration-cip-underrated"],
69
+ "rationale": "Seal-wear precursor; treat as major." } } }
70
+
71
+ // input: one revocation event, issued mid-stream
72
+ { "event": "revocation", "issued_at_seq": 305,
73
+ "target_digest": "sha256:9c41...", "reason_class": "consent_withdrawn",
74
+ "notes": "target also used to derive a later candidate" }
75
+
76
+ // output: the profile record for one run
77
+ { "delta_perf": 0.12,
78
+ "regressions": { "count": 2, "worst_group": "risk:sterility" },
79
+ "lineage_pct": 96.0,
80
+ "recall": { "ack_latency_events": 41, "digest_excluded": true,
81
+ "continued_use_detected": 1 },
82
+ "candidates": { "promoted": 9, "rejected": 4, "held": 3, "blocked_at_gate": 2 } }
83
+ ```
84
+
85
+ ## Systems to compare
86
+
87
+ All under the same data and compute budget:
88
+
89
+ - **A harness optimiser** in the style of Self-Harness
90
+ ([arXiv:2606.09498](https://arxiv.org/abs/2606.09498)).
91
+ - **A frozen, no-adaptation baseline.** It is not automatically perfect on
92
+ recall: a capability it started with can still be recalled later.
93
+ - **A simple version-and-approval workflow**, in which every change is
94
+ versioned and human-approved, but nothing is mined from overrides,
95
+ compiled into evals or recalled by content hash.
96
+ - **The Brevet reference loop** in this repository.
97
+
98
+ Every system's scores are measured, not assigned in advance. Ablations
99
+ should remove human promotion, paired regression checks or recall
100
+ propagation one at a time, to show what each contributes.
101
+
102
+ ## Data
103
+
104
+ A first release could contain synthetic override corpora built from worked
105
+ scenarios such as deviation triage, batch quality and shift handover,
106
+ published in CHAP's envelope format. Each override should carry its
107
+ rationale and the limits of where it applies, and the data should include
108
+ disagreements between reviewers and corrections that were later reversed.
109
+ Later releases could add consented, anonymised records from real operations.
110
+
111
+ ## What this benchmark is not
112
+
113
+ It does not measure raw capability (SWE-bench and Terminal-Bench do that) or
114
+ memory recall (LoCoMo and LongMemEval do that). It measures whether an
115
+ agent's changes over time are governable: improved without hidden
116
+ regressions, traced to their evidence and approver, and recalled when wrong.