brevet 0.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- brevet-0.3.0/ABOUT.md +434 -0
- brevet-0.3.0/BENCHMARK.md +116 -0
- brevet-0.3.0/CHANGELOG.md +172 -0
- brevet-0.3.0/CITATION.cff +25 -0
- brevet-0.3.0/GLOSSARY.md +56 -0
- brevet-0.3.0/LICENSE +201 -0
- brevet-0.3.0/MANIFEST.in +3 -0
- brevet-0.3.0/PKG-INFO +222 -0
- brevet-0.3.0/README.md +183 -0
- brevet-0.3.0/README_PYPI.md +175 -0
- brevet-0.3.0/SECURITY.md +30 -0
- brevet-0.3.0/SPEC.md +137 -0
- brevet-0.3.0/brevet/__init__.py +75 -0
- brevet-0.3.0/brevet/adapters.py +328 -0
- brevet-0.3.0/brevet/approvals.py +644 -0
- brevet-0.3.0/brevet/assist.py +92 -0
- brevet-0.3.0/brevet/canonical.py +106 -0
- brevet-0.3.0/brevet/chap_bridge.py +203 -0
- brevet-0.3.0/brevet/chap_evidence.py +433 -0
- brevet-0.3.0/brevet/cli.py +533 -0
- brevet-0.3.0/brevet/delta.py +256 -0
- brevet-0.3.0/brevet/demo.py +152 -0
- brevet-0.3.0/brevet/evals.py +67 -0
- brevet-0.3.0/brevet/evidence.py +73 -0
- brevet-0.3.0/brevet/identity.py +34 -0
- brevet-0.3.0/brevet/ledger.py +144 -0
- brevet-0.3.0/brevet/lifecycle.py +409 -0
- brevet-0.3.0/brevet/mcp_server.py +527 -0
- brevet-0.3.0/brevet/models.py +302 -0
- brevet-0.3.0/brevet/playground.py +458 -0
- brevet-0.3.0/brevet/runner.py +92 -0
- brevet-0.3.0/brevet/shell.py +295 -0
- brevet-0.3.0/brevet/workdir.py +82 -0
- brevet-0.3.0/brevet.egg-info/PKG-INFO +222 -0
- brevet-0.3.0/brevet.egg-info/SOURCES.txt +54 -0
- brevet-0.3.0/brevet.egg-info/dependency_links.txt +1 -0
- brevet-0.3.0/brevet.egg-info/entry_points.txt +2 -0
- brevet-0.3.0/brevet.egg-info/requires.txt +23 -0
- brevet-0.3.0/brevet.egg-info/top_level.txt +1 -0
- brevet-0.3.0/pyproject.toml +68 -0
- brevet-0.3.0/schemas/agent_manifest.schema.json +154 -0
- brevet-0.3.0/schemas/capabilities_lock.schema.json +33 -0
- brevet-0.3.0/schemas/capability_object.schema.json +165 -0
- brevet-0.3.0/schemas/recall_notice.schema.json +37 -0
- brevet-0.3.0/schemas/release_record.schema.json +37 -0
- brevet-0.3.0/setup.cfg +4 -0
- brevet-0.3.0/tests/test_adapters.py +131 -0
- brevet-0.3.0/tests/test_approvals.py +301 -0
- brevet-0.3.0/tests/test_chap_evidence.py +234 -0
- brevet-0.3.0/tests/test_chap_evidence_missing_source.py +25 -0
- brevet-0.3.0/tests/test_core.py +143 -0
- brevet-0.3.0/tests/test_hardening.py +442 -0
- brevet-0.3.0/tests/test_playground.py +78 -0
- brevet-0.3.0/tests/test_release_artifacts.py +62 -0
- brevet-0.3.0/tests/test_runner_assist_mcp.py +204 -0
- brevet-0.3.0/tests/test_schemas.py +73 -0
brevet-0.3.0/ABOUT.md
ADDED
|
@@ -0,0 +1,434 @@
|
|
|
1
|
+
# About Brevet
|
|
2
|
+
|
|
3
|
+
The detail behind the [README](README.md): how the governed evolution loop
|
|
4
|
+
works, how Brevet is built and how the repository is organised. For
|
|
5
|
+
definitions of every term, see the [GLOSSARY](GLOSSARY.md). For the rules any
|
|
6
|
+
implementation must follow, see [SPEC.md](SPEC.md).
|
|
7
|
+
|
|
8
|
+
## Why Brevet exists
|
|
9
|
+
|
|
10
|
+
Organisations already know how to manage change when people learn on the
|
|
11
|
+
job. When someone wants to change how a task is done, they write the change
|
|
12
|
+
down, a named person approves it, the procedure gets a new version number,
|
|
13
|
+
and a bad change can be rolled back. Those steps are how trust in a team
|
|
14
|
+
extends beyond the people who know each other personally.
|
|
15
|
+
|
|
16
|
+
AI agents now learn on the job too. They save memories, add skills and edit
|
|
17
|
+
their own instructions, but those changes skip every one of those steps.
|
|
18
|
+
That leaves three questions without answers:
|
|
19
|
+
|
|
20
|
+
1. **What has the agent learned?** Not the model's weights, but the rules,
|
|
21
|
+
skills, memories and tool permissions it has picked up since it was
|
|
22
|
+
deployed.
|
|
23
|
+
2. **Who approved it?** A learned behaviour changed a real decision. Which
|
|
24
|
+
accountable person said it could be used, and on what evidence?
|
|
25
|
+
3. **How do we take it back?** A learned rule is wrong. Where does it live,
|
|
26
|
+
which versions included it, and what shows it is no longer in use?
|
|
27
|
+
|
|
28
|
+
Brevet gives an agent's learning the same steps a team's procedures already
|
|
29
|
+
have. It is not an agent framework and does not replace one. It wraps the
|
|
30
|
+
framework you already use and keeps the records around it.
|
|
31
|
+
|
|
32
|
+
## The governed evolution loop
|
|
33
|
+
|
|
34
|
+
The loop has seven stages. Each one appends its envelopes to the evidence
|
|
35
|
+
chain.
|
|
36
|
+
|
|
37
|
+
| Stage | What happens | Who acts | What is recorded |
|
|
38
|
+
|---|---|---|---|
|
|
39
|
+
| **Work** `run()` | The agent drafts under one signed harness. | the agent | `brevet.task`, `brevet.artefact` |
|
|
40
|
+
| **Override** `record_final()` | An expert corrects the draft; the difference and the reason become an override. | an expert | `brevet.override` |
|
|
41
|
+
| **Dream** `dream()` | Offline, recurring overrides become candidate capabilities and eval cases, with no authority. | Brevet | `brevet.candidate` |
|
|
42
|
+
| **Dawn** `dawn()` | Each candidate is promoted, held, rejected or sent back for re-elicitation. | a named human or mission group | `brevet.promotion` |
|
|
43
|
+
| **Evals** `evaluate()` | Overrides replay as tests. The conservative gate needs neither half (held-in or held-out) to get worse and at least one to improve. | Brevet | `brevet.eval_run` |
|
|
44
|
+
| **Release** `release()` | Promoted capabilities ship in a signed release with its `capabilities.lock`. | a named human or mission group | `brevet.release` |
|
|
45
|
+
| **Recall** `recall()` | A capability is withdrawn, and every release that shipped it is flagged. | a named human or mission group | `brevet.recall` |
|
|
46
|
+
|
|
47
|
+
The names follow the loop's day-and-night rhythm, which the paper calls the
|
|
48
|
+
circadian contract. Awake, the agent works under one signed version and does
|
|
49
|
+
not change itself. While it sleeps, the dream cycle mines what the day's
|
|
50
|
+
overrides imply. At dawn, people review what the night proposed, and only what
|
|
51
|
+
they promote reaches the agent, through a signed release.
|
|
52
|
+
|
|
53
|
+
Overrides come in two kinds. A **refining** override keeps the agent's decision
|
|
54
|
+
and changes its wording. A **substituting** override reaches a different
|
|
55
|
+
decision, as when *minor* becomes *major*. The dream cycle treats them as soft
|
|
56
|
+
and hard signals.
|
|
57
|
+
|
|
58
|
+
The dream cycle adds only what is new, so it can run every night. A group of
|
|
59
|
+
overrides that already produced a candidate is not proposed again, whatever
|
|
60
|
+
happened to that candidate since. When more overrides join the group, the new
|
|
61
|
+
candidate supersedes the pending one, and a candidate whose content matches a
|
|
62
|
+
recalled capability is never proposed.
|
|
63
|
+
|
|
64
|
+
## The authority ladder
|
|
65
|
+
|
|
66
|
+
<p align="center">
|
|
67
|
+
<picture>
|
|
68
|
+
<source media="(prefers-color-scheme: dark)" srcset="docs/assets/levels-dark.svg">
|
|
69
|
+
<img src="docs/assets/levels-light.svg" alt="The authority ladder: Evidence holds candidates with no operational authority; promotion at the dawn gate by a named human raises a capability to Advisory, where it may inform drafts a human still checks; promotion by the mission group raises it to Controlled, where it may drive actions directly. A recalled capability leaves future releases and every release that shipped it is flagged." width="860">
|
|
70
|
+
</picture>
|
|
71
|
+
</p>
|
|
72
|
+
|
|
73
|
+
Every capability holds one authority layer. It enters at **Evidence**, with no
|
|
74
|
+
operational authority. Promotion at the dawn gate raises it to **Advisory**,
|
|
75
|
+
where it may inform what the agent drafts while a human still checks each
|
|
76
|
+
result. **Controlled**, where it may drive actions directly, needs promotion
|
|
77
|
+
by a mission group. Only `human:` and `mission_group:` identities can decide;
|
|
78
|
+
every other namespace, including `agent:`, `model:` and `dream:`, is refused.
|
|
79
|
+
Rejection is not deletion: rejected candidates stay on the evidence chain, so
|
|
80
|
+
"was this ever proposed, and why did we say no?" always has an answer.
|
|
81
|
+
|
|
82
|
+
## Signed approvals
|
|
83
|
+
|
|
84
|
+
An identity such as `human:alice@example.com` is a name anyone can type,
|
|
85
|
+
including an agent that can call Brevet's tools. Signed approvals make each
|
|
86
|
+
decision provable: once a workspace registers its first approver, every dawn
|
|
87
|
+
decision, release and recall must be signed with a key that an agent does
|
|
88
|
+
not hold.
|
|
89
|
+
|
|
90
|
+
```console
|
|
91
|
+
brevet approver add --identity human:alice@example.com --group mission_group:quality_team
|
|
92
|
+
brevet dawn --decide cap_123:promote --approver mission_group:quality_team
|
|
93
|
+
brevet approve req_7f3a9c2e41d0 # shows the request, asks for the passphrase, signs
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
- **Keys stay with people.** `brevet approver add` creates an Ed25519 key,
|
|
97
|
+
encrypted with a passphrase, in `~/.config/brevet/approvers/` (or
|
|
98
|
+
`BREVET_APPROVER_DIR`), outside every workspace.
|
|
99
|
+
- **Request, sign, apply.** Anyone may request a decision, an agent over MCP
|
|
100
|
+
included; nothing changes until the people it needs sign it with
|
|
101
|
+
`brevet approve`. The signed request is bound to the exact decision: the
|
|
102
|
+
capability's content hash for a promotion, the lock's digest for a release.
|
|
103
|
+
A release whose promotions changed after signing is refused, and a signed
|
|
104
|
+
request can be applied only once.
|
|
105
|
+
- **Mission groups sign through their members.** A decision taken as
|
|
106
|
+
`mission_group:<name>` needs signatures from that group's members, one by
|
|
107
|
+
default or more with `brevet approver threshold`.
|
|
108
|
+
- **The register is on the evidence chain.** Each registration, revocation
|
|
109
|
+
and threshold is a `brevet.approver` envelope. The first approver registers
|
|
110
|
+
themselves; every later change needs an existing approver's signature, and
|
|
111
|
+
a new key signs its own registration. `brevet verify` replays the register
|
|
112
|
+
and checks every signature as it stood at the time.
|
|
113
|
+
|
|
114
|
+
A signature proves that the holder of a registered key signed. Binding each
|
|
115
|
+
key to a verified person, and keeping it safe, is left to the deployment.
|
|
116
|
+
|
|
117
|
+
## The worked example
|
|
118
|
+
|
|
119
|
+
[examples/pump_vibration.py](examples/pump_vibration.py) runs the README's
|
|
120
|
+
example from the first override to the recall. Running it prints:
|
|
121
|
+
|
|
122
|
+
```text
|
|
123
|
+
1. Work and override: 4 overrides recorded from 5 drafts (1 draft accepted as it was).
|
|
124
|
+
2. Dream: 1 candidate capability, Evidence layer (no authority yet):
|
|
125
|
+
In equipment_triage work involving 'vibration-during-cleaning', reviewers changed
|
|
126
|
+
the agent's decision 4 times. Their reason: Vibration during cleaning is an early
|
|
127
|
+
sign of seal wear. Proposed rule: when this situation applies, raise it explicitly
|
|
128
|
+
and follow the reviewers' decision.
|
|
129
|
+
3. Dawn: rejected an approval from dream:nightly (machine identities cannot promote).
|
|
130
|
+
Dawn: promoted to Advisory by mission_group:quality_team.
|
|
131
|
+
4. Evals: 0/4 passed before, 4/4 after. Conservative gate: pass.
|
|
132
|
+
5. Release: 0.2.0 signed; capabilities.lock lists 1 promoted capability and its approver.
|
|
133
|
+
6. Recall: capability recalled; releases flagged: 0.2.0.
|
|
134
|
+
7. Verify: evidence chain intact.
|
|
135
|
+
Records are in <temporary folder>
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
## What is in the box
|
|
139
|
+
|
|
140
|
+
- **The harness as data.** `agent.yaml`, the manifest, holds the agent's
|
|
141
|
+
instructions, tools, loop policy, memory bindings and safety settings in
|
|
142
|
+
five layers. It is versioned and signed at every release.
|
|
143
|
+
- **One capability object for everything learned.** Prompt rules, skills,
|
|
144
|
+
tool bindings, eval cases, escalation rules and memory bindings share one
|
|
145
|
+
structure, one authority ladder and one recall mechanism.
|
|
146
|
+
- **Override-compiled evals.** `agent.evaluate()` replays the overrides as
|
|
147
|
+
tests, so no separate labelling project is needed. Cases are sorted by
|
|
148
|
+
content hash and dealt alternately into the held-in and held-out halves.
|
|
149
|
+
- **A capability bill of materials for every release.** `capabilities.lock`
|
|
150
|
+
lists what the release knows, where each capability came from, the hash of
|
|
151
|
+
its exact content and who approved it. The release signature covers the
|
|
152
|
+
lock's digest.
|
|
153
|
+
- **Optional model assist.** `brevet dream --assist ollama` lets a local
|
|
154
|
+
model word candidates more clearly. Its output is always a draft, every use
|
|
155
|
+
is logged, and Brevet works fully without it.
|
|
156
|
+
- **CHAP mirroring.** Set `ledger: chap:<workspace>` in `agent.yaml`, install
|
|
157
|
+
the extra with `pip install "brevet[chap]"`, and every evidence envelope is
|
|
158
|
+
also mirrored to a CHAP coordinator running alongside, with its own database
|
|
159
|
+
in the working directory. To use a coordinator on another machine, write
|
|
160
|
+
`ledger: chap:<workspace>@<url>`; envelopes queue in an outbox while the
|
|
161
|
+
connection is down. The local evidence chain remains the record of
|
|
162
|
+
reference.
|
|
163
|
+
|
|
164
|
+
## Supported frameworks
|
|
165
|
+
|
|
166
|
+
<p align="center">
|
|
167
|
+
<picture>
|
|
168
|
+
<source media="(prefers-color-scheme: dark)" srcset="docs/assets/wrap-dark.svg">
|
|
169
|
+
<img src="docs/assets/wrap-light.svg" alt="brevet.wrap(agent) places the agent, unchanged, inside four records: the signed manifest, the hash-linked evidence chain, the dawn gate, and capabilities.lock." width="860">
|
|
170
|
+
</picture>
|
|
171
|
+
</p>
|
|
172
|
+
|
|
173
|
+
`brevet.wrap()` works out which framework built the agent from the agent
|
|
174
|
+
object itself. The framework keeps running the agent; Brevet keeps the
|
|
175
|
+
records.
|
|
176
|
+
|
|
177
|
+
| Framework | How to wrap it |
|
|
178
|
+
|---|---|
|
|
179
|
+
| LangGraph | `brevet.wrap(compiled_graph)` |
|
|
180
|
+
| Claude Agent SDK | `brevet.wrap(sdk_client)` |
|
|
181
|
+
| DeepAgents | `brevet.wrap(deep_agent)` |
|
|
182
|
+
| AutoGen (AgentChat) | `brevet.wrap(agent_or_team)` |
|
|
183
|
+
| LlamaIndex | `brevet.wrap(agent_or_engine)` |
|
|
184
|
+
| Pydantic AI | `brevet.wrap(pydantic_agent)` |
|
|
185
|
+
| Google Agent Development Kit | `brevet.wrap(runner)` |
|
|
186
|
+
| CrewAI | `brevet.wrap(crew)` |
|
|
187
|
+
| OpenAI Agents SDK | `brevet.wrap(agent)` |
|
|
188
|
+
| any Python function | `brevet.wrap(fn)` |
|
|
189
|
+
|
|
190
|
+
Another framework needs one adapter class, registered with
|
|
191
|
+
`brevet.register_adapter`:
|
|
192
|
+
|
|
193
|
+
```python
|
|
194
|
+
@brevet.register_adapter("myfw", prefixes=("myfw",))
|
|
195
|
+
class MyAdapter(brevet.BaseAdapter):
|
|
196
|
+
def invoke(self, task, context):
|
|
197
|
+
return self.target.do(task), [{"step": "do"}]
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
No framework is ever a required dependency: adapters check the shape of the
|
|
201
|
+
object they are given, and the tests use stand-ins rather than the real
|
|
202
|
+
frameworks. Outside the shadow channel, a wrapped agent runs only while its
|
|
203
|
+
manifest signature verifies, so an edit made after a release stops the agent
|
|
204
|
+
instead of running unrecorded.
|
|
205
|
+
|
|
206
|
+
## The MCP server
|
|
207
|
+
|
|
208
|
+
`brevet mcp` serves the loop to any MCP client, such as Claude Desktop, Claude
|
|
209
|
+
Code or Cursor. Clients may start servers from any directory, so point
|
|
210
|
+
`BREVET_HOME` at a folder that holds `agent.yaml` and `.brevet/`:
|
|
211
|
+
|
|
212
|
+
```json
|
|
213
|
+
{"mcpServers": {"brevet": {
|
|
214
|
+
"command": "uvx", "args": ["brevet", "mcp"],
|
|
215
|
+
"env": {"BREVET_HOME": "/path/to/workspace"}}}}
|
|
216
|
+
```
|
|
217
|
+
|
|
218
|
+
| Tool | What it does |
|
|
219
|
+
|---|---|
|
|
220
|
+
| `brevet_record` | records a draft and the expert's final; a difference becomes an override |
|
|
221
|
+
| `brevet_chap_ingest` | imports CHAP review verdicts as overrides |
|
|
222
|
+
| `brevet_dream` | runs the dream cycle |
|
|
223
|
+
| `brevet_dawn_pending` | lists the candidates awaiting a dawn decision |
|
|
224
|
+
| `brevet_dawn_decide` | records one dawn decision under a human or mission-group identity |
|
|
225
|
+
| `brevet_release` | signs and records a release, with the conservative gate for trial and production |
|
|
226
|
+
| `brevet_recall` | recalls a capability and flags the releases that shipped it |
|
|
227
|
+
| `brevet_active` | serves the governed rules of the latest release, after checking the chain, the lock, the signature and each rule's content |
|
|
228
|
+
| `brevet_status` | reports the version, capabilities by layer, the dawn queue and chain health |
|
|
229
|
+
| `brevet_verify` | replays the evidence chain |
|
|
230
|
+
|
|
231
|
+
Read-only tools carry the MCP read-only hint, and the dawn, release and recall
|
|
232
|
+
tools carry the destructive hint, so clients can ask before running them. When
|
|
233
|
+
the workspace requires signed approvals, those three tools return a request
|
|
234
|
+
and the command that signs it, and nothing changes until a person signs.
|
|
235
|
+
Sessions record corrections only when asked, unless the workspace owner turns
|
|
236
|
+
on automatic capture with `BREVET_AUTO_CAPTURE=1` or
|
|
237
|
+
`runtime_safety.evidence.auto_capture: true` in the manifest.
|
|
238
|
+
|
|
239
|
+
## Commands
|
|
240
|
+
|
|
241
|
+
| Command | What it does |
|
|
242
|
+
|---|---|
|
|
243
|
+
| `brevet demo` | runs the whole loop once on synthetic data, offline |
|
|
244
|
+
| `brevet playground` | steps through the loop in a browser |
|
|
245
|
+
| `brevet init` | writes a starter `agent.yaml` and `.brevet/` |
|
|
246
|
+
| `brevet dream` | mines recurring overrides into candidates |
|
|
247
|
+
| `brevet dawn` | lists the dawn queue, or records one decision with `--decide` |
|
|
248
|
+
| `brevet release` | passes the conservative gate, then signs and records a release |
|
|
249
|
+
| `brevet recall` | recalls a capability and flags every release that shipped it |
|
|
250
|
+
| `brevet verify` | replays the evidence chain |
|
|
251
|
+
| `brevet status` | shows the version, capabilities by layer and chain health |
|
|
252
|
+
| `brevet chap-ingest` | imports CHAP review verdicts as overrides |
|
|
253
|
+
| `brevet mcp` | serves the loop to an MCP client |
|
|
254
|
+
| `brevet approver` | registers approvers, revokes them and sets mission-group thresholds |
|
|
255
|
+
| `brevet approve` | lists the requests waiting for signatures, or signs them |
|
|
256
|
+
|
|
257
|
+
## Governing what Claude itself learns
|
|
258
|
+
|
|
259
|
+
Claude Desktop and Cowork already learn between sessions through memory, saved
|
|
260
|
+
skills and standing instructions. [examples/claude-cowork](examples/claude-cowork)
|
|
261
|
+
puts that learning under the governed evolution loop in a few minutes. Your
|
|
262
|
+
corrections become overrides, you promote candidates at dawn, and the governed
|
|
263
|
+
rules Claude receives come only from a signed release. In Claude these
|
|
264
|
+
controls detect rather than prevent: Claude's own memory and skills keep
|
|
265
|
+
working outside Brevet, so the example makes them visible and reviewable
|
|
266
|
+
rather than impossible.
|
|
267
|
+
|
|
268
|
+
## Where Brevet fits
|
|
269
|
+
|
|
270
|
+
<p align="center">
|
|
271
|
+
<picture>
|
|
272
|
+
<source media="(prefers-color-scheme: dark)" srcset="docs/assets/suite-dark.svg">
|
|
273
|
+
<img src="docs/assets/suite-light.svg" alt="CHAP records what happened, Metis captures what experts know, and Brevet governs how the agent changes" width="860">
|
|
274
|
+
</picture>
|
|
275
|
+
</p>
|
|
276
|
+
|
|
277
|
+
| Project | Question it answers | What it owns |
|
|
278
|
+
|---|---|---|
|
|
279
|
+
| [CHAP](https://github.com/BrightbeamAI/chap) | What happened between people and agents? | the record of tasks, drafts, reviews and decisions |
|
|
280
|
+
| [Metis](https://github.com/BrightbeamAI/metis) | What do our experts know that is not written down? | experts' know-how, captured as governed memory |
|
|
281
|
+
| **Brevet** | How does the agent change, and on whose approval? | approvals, capability state and recalls |
|
|
282
|
+
|
|
283
|
+
The three are designed to work together, and each owns its own records.
|
|
284
|
+
Brevet's evidence envelopes follow CHAP's envelope model, so Brevet writes
|
|
285
|
+
CHAP-compatible evidence and can mirror it to a live CHAP coordinator;
|
|
286
|
+
`brevet chap-ingest` imports CHAP review decisions as overrides. Brevet's
|
|
287
|
+
capability object extends the tuple Metis uses for tacit fragments, adding a
|
|
288
|
+
`kind` field. For `memory_fragment` capabilities, Metis stays the system of
|
|
289
|
+
record; Brevet records only the binding and its approval.
|
|
290
|
+
|
|
291
|
+
## What Brevet is not
|
|
292
|
+
|
|
293
|
+
Brevet is not an agent framework. It does not let an agent improve itself
|
|
294
|
+
without oversight: candidates have no authority until a human or mission
|
|
295
|
+
group promotes them, and approvals under machine identities are refused. It
|
|
296
|
+
is not a monitoring tool either. The manifest declares whose overrides the
|
|
297
|
+
dream cycle may learn from (its consent scope), and Evidence-layer material
|
|
298
|
+
never influences the agent; the dream cycle does not yet filter overrides by
|
|
299
|
+
that declaration, so a deployment applies it at capture.
|
|
300
|
+
|
|
301
|
+
## Project status
|
|
302
|
+
|
|
303
|
+
Brevet implements the whole loop and keeps every record. Some protections
|
|
304
|
+
are left to the system you deploy it in, and the [paper](README.md#citation)
|
|
305
|
+
sets them out in full:
|
|
306
|
+
|
|
307
|
+
- **Verified approvers.** With signed approvals, each decision is signed by
|
|
308
|
+
a registered key. Making sure that key belongs to the named person, through
|
|
309
|
+
identity proofing, custody and recovery, is the deployment's job; CHAP
|
|
310
|
+
participant keys or an organisation's single sign-on can supply it. Without
|
|
311
|
+
registered approvers, Brevet records the identity given and refuses
|
|
312
|
+
anything other than `human:` and `mission_group:`.
|
|
313
|
+
- **Recall that reaches running agents.** A recalled capability, and any
|
|
314
|
+
capability with identical content, is left out of every later release, and
|
|
315
|
+
the releases that shipped it are flagged. Confirming that running agents
|
|
316
|
+
have stopped using it needs checks where the agent runs.
|
|
317
|
+
- **An anchored evidence chain.** Replay detects edits that break the chain.
|
|
318
|
+
Detecting a wholesale rewrite, or a chain cut short at the end, needs the
|
|
319
|
+
latest chain hash held outside the machine, which a CHAP coordinator can
|
|
320
|
+
hold.
|
|
321
|
+
- **Measured eval deltas.** The conservative gate checks the deltas passed to
|
|
322
|
+
`release()`. `agent.evaluate()` measures them, but nothing yet ties a release
|
|
323
|
+
to the eval runs that produced its numbers.
|
|
324
|
+
|
|
325
|
+
## How the repository is organised
|
|
326
|
+
|
|
327
|
+
```
|
|
328
|
+
brevet/
|
|
329
|
+
├── brevet/ the Python package
|
|
330
|
+
│ ├── shell.py brevet.wrap() and the agent object it returns
|
|
331
|
+
│ ├── adapters.py framework adapters and the adapter registry
|
|
332
|
+
│ ├── evidence.py harvests overrides from drafts and finals
|
|
333
|
+
│ ├── delta.py the dream cycle: mines overrides into candidates
|
|
334
|
+
│ ├── lifecycle.py the dawn gate, releases and recall
|
|
335
|
+
│ ├── evals.py override-compiled evals; the conservative gate
|
|
336
|
+
│ ├── runner.py runs the evals before and after a change
|
|
337
|
+
│ ├── ledger.py the hash-linked evidence chain
|
|
338
|
+
│ ├── canonical.py content hashing and Ed25519 signing
|
|
339
|
+
│ ├── identity.py which identities may decide
|
|
340
|
+
│ ├── approvals.py approver keys, the register and signed approvals
|
|
341
|
+
│ ├── workdir.py workspace location, permissions and file locks
|
|
342
|
+
│ ├── models.py data models matching the schemas
|
|
343
|
+
│ ├── assist.py optional local model assist
|
|
344
|
+
│ ├── chap_bridge.py mirrors envelopes to a CHAP coordinator
|
|
345
|
+
│ ├── chap_evidence.py imports CHAP review decisions as overrides
|
|
346
|
+
│ ├── mcp_server.py the loop as MCP tools
|
|
347
|
+
│ ├── playground.py the loop, stage by stage, in a browser
|
|
348
|
+
│ ├── demo.py the end-to-end demonstration
|
|
349
|
+
│ └── cli.py the `brevet` command
|
|
350
|
+
├── schemas/ the five JSON Schemas that define the records
|
|
351
|
+
├── examples/
|
|
352
|
+
│ ├── pump_vibration.py the worked example from the README
|
|
353
|
+
│ └── claude-cowork/ governing what Claude itself learns
|
|
354
|
+
├── docs/
|
|
355
|
+
│ ├── demo.html the interactive tour
|
|
356
|
+
│ └── assets/ the diagrams, in light and dark versions
|
|
357
|
+
├── scripts/
|
|
358
|
+
│ ├── make_diagrams.py regenerates the diagrams
|
|
359
|
+
│ └── make_pypi_readme.py writes README_PYPI.md for PyPI
|
|
360
|
+
├── tests/ the test suite; runs offline
|
|
361
|
+
├── server.json the MCP Registry entry
|
|
362
|
+
└── README.md · ABOUT.md · GLOSSARY.md · SPEC.md · BENCHMARK.md
|
|
363
|
+
```
|
|
364
|
+
|
|
365
|
+
## What Brevet stores, and where
|
|
366
|
+
|
|
367
|
+
`brevet.wrap()` keeps everything in a working directory, `.brevet/` by default
|
|
368
|
+
(change it with `workdir=`). Brevet creates it readable only by you, with a
|
|
369
|
+
`.gitignore` that keeps it out of version control:
|
|
370
|
+
|
|
371
|
+
```
|
|
372
|
+
.brevet/
|
|
373
|
+
├── agent.yaml the manifest, if none was supplied
|
|
374
|
+
├── capabilities.lock the latest release's lock, beside the manifest
|
|
375
|
+
├── ledger.jsonl the evidence chain: one envelope per line
|
|
376
|
+
├── capabilities.jsonl the capability store; latest line per id wins
|
|
377
|
+
├── keys/brevet_ed25519.pem the signing key, created at the first release
|
|
378
|
+
├── approvals.jsonl decisions waiting for approver signatures
|
|
379
|
+
├── chap_cursor.json how far each CHAP source has been imported
|
|
380
|
+
├── chap.db an embedded CHAP coordinator's store, if used
|
|
381
|
+
├── chap_outbox.jsonl envelopes waiting to be mirrored to CHAP
|
|
382
|
+
└── *.lock short-lived locks that let processes share files
|
|
383
|
+
```
|
|
384
|
+
|
|
385
|
+
When a manifest path is given, `agent.yaml` and `capabilities.lock` live
|
|
386
|
+
beside it instead. Never commit `.brevet/`: it holds a private key and the
|
|
387
|
+
evidence of real work. Approver keys never live here; each person keeps
|
|
388
|
+
theirs in `~/.config/brevet/approvers/`.
|
|
389
|
+
|
|
390
|
+
## Evidence envelopes
|
|
391
|
+
|
|
392
|
+
Every envelope is appended, never edited, and hash-linked to the one before
|
|
393
|
+
it: `chain_hash = sha256(encode(envelope) ‖ prev_hash)`. Appends take a file
|
|
394
|
+
lock, so several processes can share one chain.
|
|
395
|
+
|
|
396
|
+
| Envelope | Appended when |
|
|
397
|
+
|---|---|
|
|
398
|
+
| `brevet.task` | the agent is given a task |
|
|
399
|
+
| `brevet.artefact` | the agent produces a draft |
|
|
400
|
+
| `brevet.override` | an expert's correction is recorded |
|
|
401
|
+
| `brevet.candidate` | the dream cycle proposes a candidate |
|
|
402
|
+
| `brevet.model_assist` | a local model helps word a candidate |
|
|
403
|
+
| `brevet.promotion` | a decision is made at the dawn gate |
|
|
404
|
+
| `brevet.eval_run` | the evals run |
|
|
405
|
+
| `brevet.release` | a new version is released |
|
|
406
|
+
| `brevet.recall` | a capability is recalled |
|
|
407
|
+
| `brevet.approver` | an approver is registered or revoked, or a group threshold changes |
|
|
408
|
+
|
|
409
|
+
## The benchmark
|
|
410
|
+
|
|
411
|
+
Research on self-improving agents mostly measures one thing: whether the
|
|
412
|
+
agent got better. Teams running agents in production need four measures,
|
|
413
|
+
and [BENCHMARK.md](BENCHMARK.md) proposes a governed-adaptation benchmark
|
|
414
|
+
that reports all four side by side: improvement, regression discipline,
|
|
415
|
+
lineage completeness and recall compliance. No implementation of the
|
|
416
|
+
benchmark exists yet.
|
|
417
|
+
|
|
418
|
+
## Design principles
|
|
419
|
+
|
|
420
|
+
- **Authority is granted, never grabbed.** Every increase in a
|
|
421
|
+
capability's authority is a recorded decision by a named human or mission
|
|
422
|
+
group, signed by them once the workspace registers approvers.
|
|
423
|
+
- **The running agent does not change itself.** It runs one released
|
|
424
|
+
version. Learning happens between versions, where it can be reviewed.
|
|
425
|
+
- **Overrides are evidence, not truth.** Experts' corrections are the best
|
|
426
|
+
available record of judgement, but they can be mistaken, so candidates are
|
|
427
|
+
reviewed at dawn before they gain authority.
|
|
428
|
+
- **No network needed.** The tests, the demonstration and the whole loop run
|
|
429
|
+
offline. Model assist and CHAP mirroring are optional and fail safely.
|
|
430
|
+
- **Rejection is not deletion.** Rejected and recalled capabilities stay on
|
|
431
|
+
the evidence chain with their lineage intact.
|
|
432
|
+
- **Prove it from the chain.** Promotions, releases and recalls are all
|
|
433
|
+
envelopes on the evidence chain, so anyone holding it can replay and check
|
|
434
|
+
them.
|
|
@@ -0,0 +1,116 @@
|
|
|
1
|
+
# The Governed-Adaptation Benchmark (design specification)
|
|
2
|
+
|
|
3
|
+
Research on self-improving agents mostly measures one thing: whether the
|
|
4
|
+
agent got better. Teams that run agents in production need four
|
|
5
|
+
properties, and improvement is only the first. This document specifies a
|
|
6
|
+
benchmark that scores all four side by side. **No implementation of it
|
|
7
|
+
exists yet**; this is its design, as specified in the paper.
|
|
8
|
+
|
|
9
|
+
## What a submission is given
|
|
10
|
+
|
|
11
|
+
A submission is any system that adapts an agent over time, whether or not it
|
|
12
|
+
uses Brevet. It receives:
|
|
13
|
+
|
|
14
|
+
- **a task stream**, partitioned by incident before any mining, so that
|
|
15
|
+
related cases never end up on both sides of a split. The stream includes
|
|
16
|
+
accepted work as well as corrected work;
|
|
17
|
+
- **an override stream** in CHAP's envelope format (the difference, the
|
|
18
|
+
rationale, tags and `intent_preserved`), replayed on a schedule. It
|
|
19
|
+
supplies human judgement where no automatic verifier exists;
|
|
20
|
+
- **an initial signed harness** `H0`;
|
|
21
|
+
- **revocation events** issued mid-stream, including one consent withdrawal
|
|
22
|
+
and one recall of a capability that an endpoint has already cached or that
|
|
23
|
+
a later candidate was derived from;
|
|
24
|
+
- **at least one pair of conflicting candidates** whose applicability
|
|
25
|
+
conditions overlap, so the scoring can see whether review detects the
|
|
26
|
+
interaction or promotes both; and
|
|
27
|
+
- **a consent policy** over override sources.
|
|
28
|
+
|
|
29
|
+
The submission runs its loop and produces a harness lineage `H0 … HN`, each
|
|
30
|
+
version with its lockfile, its promotion records and its evidence log.
|
|
31
|
+
Proposal generation and release selection may not see the final test split.
|
|
32
|
+
|
|
33
|
+
## The four scored axes
|
|
34
|
+
|
|
35
|
+
1. **Improvement.** The held-out pass-rate change across the lineage: how
|
|
36
|
+
much better the agent gets. This is the axis current research reports.
|
|
37
|
+
2. **Regression discipline.** How often, and how badly, a change made
|
|
38
|
+
individual tasks or risk groups worse and still survived promotion.
|
|
39
|
+
Gains and losses are reported separately, because an average can hide a
|
|
40
|
+
loss behind a gain.
|
|
41
|
+
3. **Lineage completeness.** For a sample of changes between two harness
|
|
42
|
+
versions, whether the records link each change to its capability, the
|
|
43
|
+
evidence for it, the eval run that tested it and the authenticated
|
|
44
|
+
approver. Audit replay scores completeness. Showing that a capability
|
|
45
|
+
actually *caused* a change in behaviour needs an extra ablation.
|
|
46
|
+
4. **Recall compliance.** For each revocation event: how long until every
|
|
47
|
+
endpoint acknowledges the recall, whether the capability's content hash is
|
|
48
|
+
excluded from later lockfiles, and whether task-level checks find it still
|
|
49
|
+
in use, including through derived capabilities and cached material.
|
|
50
|
+
Inventory exclusion and operational cessation are scored separately.
|
|
51
|
+
|
|
52
|
+
Results are a profile, not a single number:
|
|
53
|
+
`(δperf, regressions, lineage %, recall)`. Submissions also report how many
|
|
54
|
+
candidates were rejected, held or blocked at each gate, so the cost of
|
|
55
|
+
governance is visible next to its benefit. A system that improves faster but
|
|
56
|
+
fails recall is not automatically better where revocation is required.
|
|
57
|
+
|
|
58
|
+
## Record formats
|
|
59
|
+
|
|
60
|
+
A submission reads override-stream events and revocation events, and writes
|
|
61
|
+
one profile record per run. Representative values:
|
|
62
|
+
|
|
63
|
+
```jsonc
|
|
64
|
+
// input: one override-stream event (CHAP envelope format)
|
|
65
|
+
{ "seq": 214, "group_id": "incident_0031",
|
|
66
|
+
"envelope": { "kind": "brevet.override",
|
|
67
|
+
"body": { "participant": "human:qa@site", "intent_preserved": false,
|
|
68
|
+
"diff": [ ... ], "tags": ["vibration-cip-underrated"],
|
|
69
|
+
"rationale": "Seal-wear precursor; treat as major." } } }
|
|
70
|
+
|
|
71
|
+
// input: one revocation event, issued mid-stream
|
|
72
|
+
{ "event": "revocation", "issued_at_seq": 305,
|
|
73
|
+
"target_digest": "sha256:9c41...", "reason_class": "consent_withdrawn",
|
|
74
|
+
"notes": "target also used to derive a later candidate" }
|
|
75
|
+
|
|
76
|
+
// output: the profile record for one run
|
|
77
|
+
{ "delta_perf": 0.12,
|
|
78
|
+
"regressions": { "count": 2, "worst_group": "risk:sterility" },
|
|
79
|
+
"lineage_pct": 96.0,
|
|
80
|
+
"recall": { "ack_latency_events": 41, "digest_excluded": true,
|
|
81
|
+
"continued_use_detected": 1 },
|
|
82
|
+
"candidates": { "promoted": 9, "rejected": 4, "held": 3, "blocked_at_gate": 2 } }
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
## Systems to compare
|
|
86
|
+
|
|
87
|
+
All under the same data and compute budget:
|
|
88
|
+
|
|
89
|
+
- **A harness optimiser** in the style of Self-Harness
|
|
90
|
+
([arXiv:2606.09498](https://arxiv.org/abs/2606.09498)).
|
|
91
|
+
- **A frozen, no-adaptation baseline.** It is not automatically perfect on
|
|
92
|
+
recall: a capability it started with can still be recalled later.
|
|
93
|
+
- **A simple version-and-approval workflow**, in which every change is
|
|
94
|
+
versioned and human-approved, but nothing is mined from overrides,
|
|
95
|
+
compiled into evals or recalled by content hash.
|
|
96
|
+
- **The Brevet reference loop** in this repository.
|
|
97
|
+
|
|
98
|
+
Every system's scores are measured, not assigned in advance. Ablations
|
|
99
|
+
should remove human promotion, paired regression checks or recall
|
|
100
|
+
propagation one at a time, to show what each contributes.
|
|
101
|
+
|
|
102
|
+
## Data
|
|
103
|
+
|
|
104
|
+
A first release could contain synthetic override corpora built from worked
|
|
105
|
+
scenarios such as deviation triage, batch quality and shift handover,
|
|
106
|
+
published in CHAP's envelope format. Each override should carry its
|
|
107
|
+
rationale and the limits of where it applies, and the data should include
|
|
108
|
+
disagreements between reviewers and corrections that were later reversed.
|
|
109
|
+
Later releases could add consented, anonymised records from real operations.
|
|
110
|
+
|
|
111
|
+
## What this benchmark is not
|
|
112
|
+
|
|
113
|
+
It does not measure raw capability (SWE-bench and Terminal-Bench do that) or
|
|
114
|
+
memory recall (LoCoMo and LongMemEval do that). It measures whether an
|
|
115
|
+
agent's changes over time are governable: improved without hidden
|
|
116
|
+
regressions, traced to their evidence and approver, and recalled when wrong.
|