monogate-capcard-cli 1.0.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- monogate_capcard_cli-1.0.0/LICENSE +17 -0
- monogate_capcard_cli-1.0.0/PKG-INFO +663 -0
- monogate_capcard_cli-1.0.0/README.md +640 -0
- monogate_capcard_cli-1.0.0/capcard_cli/__init__.py +55 -0
- monogate_capcard_cli-1.0.0/capcard_cli/attest.py +444 -0
- monogate_capcard_cli-1.0.0/capcard_cli/bundle.py +403 -0
- monogate_capcard_cli-1.0.0/capcard_cli/cli.py +1530 -0
- monogate_capcard_cli-1.0.0/capcard_cli/playbook.py +561 -0
- monogate_capcard_cli-1.0.0/capcard_cli/registry.py +484 -0
- monogate_capcard_cli-1.0.0/capcard_cli/training.py +463 -0
- monogate_capcard_cli-1.0.0/capcard_cli/verify.py +901 -0
- monogate_capcard_cli-1.0.0/monogate_capcard_cli.egg-info/PKG-INFO +663 -0
- monogate_capcard_cli-1.0.0/monogate_capcard_cli.egg-info/SOURCES.txt +24 -0
- monogate_capcard_cli-1.0.0/monogate_capcard_cli.egg-info/dependency_links.txt +1 -0
- monogate_capcard_cli-1.0.0/monogate_capcard_cli.egg-info/entry_points.txt +14 -0
- monogate_capcard_cli-1.0.0/monogate_capcard_cli.egg-info/requires.txt +3 -0
- monogate_capcard_cli-1.0.0/monogate_capcard_cli.egg-info/top_level.txt +1 -0
- monogate_capcard_cli-1.0.0/pyproject.toml +58 -0
- monogate_capcard_cli-1.0.0/setup.cfg +4 -0
- monogate_capcard_cli-1.0.0/tests/test_attest.py +864 -0
- monogate_capcard_cli-1.0.0/tests/test_cli.py +626 -0
- monogate_capcard_cli-1.0.0/tests/test_phase7.py +555 -0
- monogate_capcard_cli-1.0.0/tests/test_playbook.py +422 -0
- monogate_capcard_cli-1.0.0/tests/test_registry.py +306 -0
- monogate_capcard_cli-1.0.0/tests/test_training.py +529 -0
- monogate_capcard_cli-1.0.0/tests/test_verify.py +647 -0
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
Apache License
|
|
2
|
+
Version 2.0, January 2004
|
|
3
|
+
http://www.apache.org/licenses/
|
|
4
|
+
|
|
5
|
+
Copyright 2026 Mosa Creates LLC
|
|
6
|
+
|
|
7
|
+
Licensed under the Apache License, Version 2.0 (the "License");
|
|
8
|
+
you may not use this file except in compliance with the License.
|
|
9
|
+
You may obtain a copy of the License at
|
|
10
|
+
|
|
11
|
+
http://www.apache.org/licenses/LICENSE-2.0
|
|
12
|
+
|
|
13
|
+
Unless required by applicable law or agreed to in writing, software
|
|
14
|
+
distributed under the License is distributed on an "AS IS" BASIS,
|
|
15
|
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
16
|
+
See the License for the specific language governing permissions and
|
|
17
|
+
limitations under the License.
|
|
@@ -0,0 +1,663 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: monogate-capcard-cli
|
|
3
|
+
Version: 1.0.0
|
|
4
|
+
Summary: Persistent identity + verification CLIs for Monogate CapCard agents
|
|
5
|
+
Author-email: Monogate <almaguer1986@gmail.com>
|
|
6
|
+
License: Apache-2.0
|
|
7
|
+
Project-URL: Homepage, https://capcard.ai
|
|
8
|
+
Project-URL: Source, https://github.com/agent-maestro/capcard-cli
|
|
9
|
+
Keywords: monogate,capcard,agents,verification,trust
|
|
10
|
+
Classifier: Development Status :: 3 - Alpha
|
|
11
|
+
Classifier: Intended Audience :: Developers
|
|
12
|
+
Classifier: License :: OSI Approved :: Apache Software License
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
17
|
+
Requires-Python: >=3.10
|
|
18
|
+
Description-Content-Type: text/markdown
|
|
19
|
+
License-File: LICENSE
|
|
20
|
+
Provides-Extra: dev
|
|
21
|
+
Requires-Dist: pytest>=8; extra == "dev"
|
|
22
|
+
Dynamic: license-file
|
|
23
|
+
|
|
24
|
+
# capcard-cli
|
|
25
|
+
|
|
26
|
+
Persistent identity + weighted-trust verification + knowledge
|
|
27
|
+
transfer + cross-agent attestation + RL training-data extraction +
|
|
28
|
+
cross-registry wire format for Monogate CapCard agents. All seven
|
|
29
|
+
phases of the [CapCard: Verified Agents](https://capcard.ai)
|
|
30
|
+
roadmap — turns anonymous Claude Code sessions into named agents
|
|
31
|
+
with stable IDs + generation lineage, evidence-backed verification
|
|
32
|
+
records scored on a Lean-proof-to-self-report scale, structured
|
|
33
|
+
session-end playbooks the next agent on a related task can read, a
|
|
34
|
+
"builder posts / verifier signs / both trust scores update"
|
|
35
|
+
cross-attestation flow so trust is earned, a training-data export
|
|
36
|
+
pipeline that turns the accumulated playbook archive into shaped-
|
|
37
|
+
reward episodes for an RL agent training loop, and a stable JSON
|
|
38
|
+
wire format with idempotent merge for cross-registry sync.
|
|
39
|
+
|
|
40
|
+
> **All phases shipped (1.0.0).** Twelve binaries on PATH:
|
|
41
|
+
> `capcard-register` (with `--generation` / `--parent` /
|
|
42
|
+
> `--pubkey`), `capcard-status`, `capcard-verify` (with
|
|
43
|
+
> `--evidence-type` and `--needs-attestation`), `capcard-attest`
|
|
44
|
+
> (`--pass` / `--reject` / `--dispute`), `capcard-attestations`
|
|
45
|
+
> (`--pending` / `--by` / `--on`), `capcard-recompute-trust` (with
|
|
46
|
+
> `--iterative` and `--dispute-decay-days`), `capcard-playbook`
|
|
47
|
+
> (`generate` / `show` / `search`), `capcard-export`,
|
|
48
|
+
> `capcard-tactics`, `capcard-stats`, `capcard-bundle`
|
|
49
|
+
> (`export` / `import`), `capcard-fingerprint`. Trust is the
|
|
50
|
+
> weighted formula `sum(w_i) / sum(|w_i|)` clamped to `[0, 1]`,
|
|
51
|
+
> with cross-attestation gating contribution when the builder
|
|
52
|
+
> requests it. Disputes resolve by quorum (≥2 distinct
|
|
53
|
+
> non-disputers agreeing); per-verifier votes collapse; stale
|
|
54
|
+
> dispute markers can be aged out with `--dispute-decay-days`.
|
|
55
|
+
> Phase 5 turns the playbook archive into shaped-reward episodes
|
|
56
|
+
> for an RL pipeline. Phase 7 ships a versioned JSON wire format
|
|
57
|
+
> ready for a future network gossip layer (the actual gossip + ZK
|
|
58
|
+
> proofs are an out-of-CLI research project — see
|
|
59
|
+
> [`docs/BACKLOG.md`](docs/BACKLOG.md)).
|
|
60
|
+
|
|
61
|
+
## Install
|
|
62
|
+
|
|
63
|
+
```bash
|
|
64
|
+
git clone https://github.com/agent-maestro/capcard-cli # once on PyPI: pip install monogate-capcard-cli
|
|
65
|
+
cd capcard-cli
|
|
66
|
+
pip install -e .
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
The package depends on nothing outside the Python stdlib — installs
|
|
70
|
+
into any Claude Code session's Python without dragging in a wider
|
|
71
|
+
requirements graph.
|
|
72
|
+
|
|
73
|
+
## Use
|
|
74
|
+
|
|
75
|
+
### Register at session start
|
|
76
|
+
|
|
77
|
+
```bash
|
|
78
|
+
capcard-register --name "Frontier C Researcher" \
|
|
79
|
+
--role researcher \
|
|
80
|
+
--task "Galois dimension proof"
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
Roles: `researcher`, `builder`, `tooling`, `verifier` (any other
|
|
84
|
+
string is accepted with a warning to stderr — extensible without
|
|
85
|
+
a registry migration).
|
|
86
|
+
|
|
87
|
+
The first registration with a given `--name` allocates an
|
|
88
|
+
`agent_id` slug (`researcher-001`, `builder-001`, etc.). Subsequent
|
|
89
|
+
calls with the same name return the same ID and update
|
|
90
|
+
`current_task` + `last_active`.
|
|
91
|
+
|
|
92
|
+
### Update mid-session
|
|
93
|
+
|
|
94
|
+
```bash
|
|
95
|
+
capcard-status --agent researcher-001 \
|
|
96
|
+
--task "writing arXiv update" \
|
|
97
|
+
--status working
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
Status values: `working`, `idle`, `blocked`, `done`, `failed`
|
|
101
|
+
(also extensible).
|
|
102
|
+
|
|
103
|
+
### List
|
|
104
|
+
|
|
105
|
+
```bash
|
|
106
|
+
capcard list
|
|
107
|
+
capcard list --json # machine-readable
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
### Verify a unit of work
|
|
111
|
+
|
|
112
|
+
`capcard-verify` records one outcome with cryptographic evidence,
|
|
113
|
+
appends a row to the verification log, and updates the agent's
|
|
114
|
+
trust counters. The agent owns the framing — `capcard-verify` is
|
|
115
|
+
how they sign their claim into the log.
|
|
116
|
+
|
|
117
|
+
```bash
|
|
118
|
+
# Treat --evidence as a literal artifact (you've already collected it)
|
|
119
|
+
capcard-verify --agent researcher-001 \
|
|
120
|
+
--evidence "lake build clean (12/12 modules)" \
|
|
121
|
+
--result proof-closed \
|
|
122
|
+
--theorem tanh_monotone
|
|
123
|
+
|
|
124
|
+
# Have capcard-verify execute the evidence command and capture the result
|
|
125
|
+
capcard-verify --agent builder-001 \
|
|
126
|
+
--evidence "pytest -q" \
|
|
127
|
+
--result feature-shipped \
|
|
128
|
+
--run
|
|
129
|
+
|
|
130
|
+
# Honest-negative: investigation that returned a negative finding.
|
|
131
|
+
# Counts as a verified outcome — does NOT decrease trust.
|
|
132
|
+
capcard-verify --agent researcher-001 \
|
|
133
|
+
--evidence "C-246 NOTES.md: Phi homomorphism is trivial" \
|
|
134
|
+
--result honest-negative
|
|
135
|
+
|
|
136
|
+
# Failure: shipped broken code, broke CI, false-claim correction.
|
|
137
|
+
# Decreases trust.
|
|
138
|
+
capcard-verify --agent builder-001 \
|
|
139
|
+
--evidence "CI red on main: type signature regression" \
|
|
140
|
+
--result failure
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
The result types and their effect on the simple counter pair (the
|
|
144
|
+
weighted trust score uses the `--evidence-type` table below — see
|
|
145
|
+
"Phase 4: weighted trust"):
|
|
146
|
+
|
|
147
|
+
| `--result` | Counter bumped | Trust direction |
|
|
148
|
+
|------------------------|----------------------|------------------------|
|
|
149
|
+
| `proof-closed` | `verified_outputs` | up |
|
|
150
|
+
| `feature-shipped` | `verified_outputs` | up |
|
|
151
|
+
| `theorem-established` | `verified_outputs` | up |
|
|
152
|
+
| `bug-fixed` | `verified_outputs` | up |
|
|
153
|
+
| `honest-negative` | `verified_outputs` | up (per spec) |
|
|
154
|
+
| `failure` | `failed_outputs` | down |
|
|
155
|
+
|
|
156
|
+
### Phase 4: weighted trust
|
|
157
|
+
|
|
158
|
+
`capcard-verify` accepts an optional `--evidence-type` that drives
|
|
159
|
+
the weighted trust formula. Heavier weights mean stronger evidence:
|
|
160
|
+
a Lean proof checked by `lake build` is unforgeable (1.0), a self-
|
|
161
|
+
report is just the agent's word (0.2). Negative weights model
|
|
162
|
+
verifiable harm.
|
|
163
|
+
|
|
164
|
+
| `--evidence-type` | Weight | When to use |
|
|
165
|
+
|-------------------|-------:|---------------------------------------------------|
|
|
166
|
+
| `lean-proof` | +1.0 | Lean proof passed `lake build` cleanly |
|
|
167
|
+
| `bfs-closure` | +0.9 | BFS proof engine closed a sorry; lake-verified |
|
|
168
|
+
| `test-suite` | +0.7 | Full test suite passed |
|
|
169
|
+
| `ci-green` | +0.6 | CI green (depends on test quality) |
|
|
170
|
+
| `empirical` | +0.5 | N/N empirical hold; suggestive, not proven |
|
|
171
|
+
| `honest-negative` | +0.5 | Investigated; reported negative finding honestly |
|
|
172
|
+
| `self-report` | +0.2 | No independent verification |
|
|
173
|
+
| `ci-broken` | -0.3 | Agent broke CI |
|
|
174
|
+
| `false-claim` | -0.5 | Agent's earlier claim was corrected |
|
|
175
|
+
|
|
176
|
+
```bash
|
|
177
|
+
# 47/47 spot-check is empirical, not a proof — be explicit
|
|
178
|
+
capcard-verify --agent researcher-001 \
|
|
179
|
+
--evidence "47/47 holds across [501, 2000]" \
|
|
180
|
+
--result theorem-established \
|
|
181
|
+
--evidence-type empirical
|
|
182
|
+
|
|
183
|
+
# BFS engine closed a Discovered/ sorry
|
|
184
|
+
capcard-verify --agent tooling-001 \
|
|
185
|
+
--evidence "tools/sweep closed exp_strict_mono_at_zero" \
|
|
186
|
+
--result proof-closed \
|
|
187
|
+
--evidence-type bfs-closure
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
When `--evidence-type` is omitted, the type is inferred from the
|
|
191
|
+
result + evidence keywords (`lake build` → `lean-proof`, `pytest …
|
|
192
|
+
passed` → `test-suite`, "ci"/"broke"/"build" tokens in a `failure`
|
|
193
|
+
→ `ci-broken`, etc.). Pass it explicitly when the heuristic is
|
|
194
|
+
wrong — the conservative default is `self-report` (0.2) for
|
|
195
|
+
positive results without obvious signal.
|
|
196
|
+
|
|
197
|
+
**Trust formula:** `trust = sum(w_i) / sum(|w_i|)` over all
|
|
198
|
+
verifications, clamped to `[0, 1]` from below. An empty history
|
|
199
|
+
returns 0.0 ("no track record"). The cached `trust_score` field on
|
|
200
|
+
the registry row is updated by every `record_verification` call;
|
|
201
|
+
operators can rebuild it from the log via `recompute_trust(agent_id)`
|
|
202
|
+
after a weight-table change.
|
|
203
|
+
|
|
204
|
+
`--run` shells out via `shlex.split` (no shell interpolation —
|
|
205
|
+
quote your arguments) and captures `exit_code`, wall-clock
|
|
206
|
+
`runtime_s`, full `stdout` + `stderr`, and a SHA-256 digest of
|
|
207
|
+
the combined output. `--timeout-s` defaults to 600s. If
|
|
208
|
+
`--run` finishes with a non-zero exit but you claimed a positive
|
|
209
|
+
result, the record is written as you stated it but a note is
|
|
210
|
+
attached so a Phase 6 verifier can spot the mismatch.
|
|
211
|
+
|
|
212
|
+
When `--run` is *not* set, `--evidence` is stored verbatim and
|
|
213
|
+
the digest is `sha256(evidence)`.
|
|
214
|
+
|
|
215
|
+
## What it stores
|
|
216
|
+
|
|
217
|
+
A single JSON file at `~/.agent-registry.json` (override via
|
|
218
|
+
`CAPCARD_REGISTRY_PATH`):
|
|
219
|
+
|
|
220
|
+
```json
|
|
221
|
+
{
|
|
222
|
+
"version": 1,
|
|
223
|
+
"agents": {
|
|
224
|
+
"researcher-001": {
|
|
225
|
+
"agent_id": "researcher-001",
|
|
226
|
+
"name": "Frontier C Researcher",
|
|
227
|
+
"role": "researcher",
|
|
228
|
+
"status": "working",
|
|
229
|
+
"current_task": "Galois dimension proof",
|
|
230
|
+
"total_tasks": 0,
|
|
231
|
+
"verified_outputs": 0,
|
|
232
|
+
"failed_outputs": 0,
|
|
233
|
+
"trust_score": 0.0,
|
|
234
|
+
"created": "2026-05-04T13:30:00Z",
|
|
235
|
+
"last_active": "2026-05-04T13:30:00Z",
|
|
236
|
+
"playbook_path": null
|
|
237
|
+
}
|
|
238
|
+
}
|
|
239
|
+
}
|
|
240
|
+
```
|
|
241
|
+
|
|
242
|
+
Writes are atomic (temp-file-then-rename) and serialized via
|
|
243
|
+
`flock` on a sidecar `~/.agent-registry.json.lock`. Two parallel
|
|
244
|
+
`capcard-register` calls cannot lose updates.
|
|
245
|
+
|
|
246
|
+
### Phase 6: cross-agent attestation
|
|
247
|
+
|
|
248
|
+
By default `capcard-verify` self-attests — the agent's claim
|
|
249
|
+
contributes to their own trust immediately. Phase 6 adds a
|
|
250
|
+
two-step path so a verifier can confirm or reject another agent's
|
|
251
|
+
work before it counts.
|
|
252
|
+
|
|
253
|
+
```bash
|
|
254
|
+
# Builder records a pending verification (--needs-attestation).
|
|
255
|
+
# The record lands in the 'pending' state and contributes 0 to
|
|
256
|
+
# trust until cross-attested.
|
|
257
|
+
capcard-verify --agent builder-001 \
|
|
258
|
+
--evidence "pytest -q && cargo test --release" \
|
|
259
|
+
--result feature-shipped \
|
|
260
|
+
--evidence-type test-suite \
|
|
261
|
+
--needs-attestation
|
|
262
|
+
# Output ends with:
|
|
263
|
+
# id: 6f00d02d4244 ← capture this
|
|
264
|
+
# attestation: pending (--needs-attestation set)
|
|
265
|
+
|
|
266
|
+
# A verifier (must be a different agent) signs off:
|
|
267
|
+
capcard-attest --verification-id 6f00d02d4244 \
|
|
268
|
+
--verifier-agent verifier-001 \
|
|
269
|
+
--pass \
|
|
270
|
+
--evidence "reviewed diff + ran pytest myself" \
|
|
271
|
+
--note "looks clean"
|
|
272
|
+
```
|
|
273
|
+
|
|
274
|
+
Outcomes:
|
|
275
|
+
|
|
276
|
+
| flag | effect on builder trust | effect on verifier trust |
|
|
277
|
+
|-------------|----------------------------------------------------------------------------------------------------------|-------------------------------------------|
|
|
278
|
+
| `--pass` | record contributes `base_weight × max(0.5, verifier_trust)` to the numerator (denom stays at base_weight) | +0.6 weight (review work is real work) |
|
|
279
|
+
| `--reject` | record treated as false-claim against builder (-0.5 weight) | +0.6 weight |
|
|
280
|
+
| `--dispute` | record held in `disputed` state (contributes 0) until quorum resolves it (Phase 6.5) | +0.6 weight |
|
|
281
|
+
|
|
282
|
+
Guards:
|
|
283
|
+
|
|
284
|
+
- **No self-grading** — `--verifier-agent` must be different from
|
|
285
|
+
the builder. Same-agent attestation returns exit 2.
|
|
286
|
+
- **Cross-attesting a self-attested record** returns exit 2 — the
|
|
287
|
+
builder must opt in via `--needs-attestation`.
|
|
288
|
+
- **Verifier must be registered** — needed so the trust snapshot
|
|
289
|
+
has a real value to reference.
|
|
290
|
+
|
|
291
|
+
The append-only attestation log lives at
|
|
292
|
+
`~/.capcard/attestations.jsonl` (override with
|
|
293
|
+
`CAPCARD_ATTESTATION_LOG_PATH`). Overturning a previous decision
|
|
294
|
+
means recording a *new* attestation, not editing the old row.
|
|
295
|
+
|
|
296
|
+
### Phase 6.5: dispute resolution + trust hygiene
|
|
297
|
+
|
|
298
|
+
The state machine resolves disputes by **quorum** instead of holding
|
|
299
|
+
disputed records open forever. Two rules:
|
|
300
|
+
|
|
301
|
+
1. **Per-verifier vote collapse.** A verifier who casts multiple
|
|
302
|
+
attestations on the same record contributes one current vote —
|
|
303
|
+
their latest stance. Flip-flopping doesn't manufacture extra
|
|
304
|
+
votes (the append-only log preserves the audit trail for
|
|
305
|
+
`capcard-attestations --on <vid>`).
|
|
306
|
+
2. **Dispute quorum.** A `dispute` marker is overridden when ≥2
|
|
307
|
+
distinct non-disputer verifiers' current votes agree on the same
|
|
308
|
+
verdict and that verdict has a strict majority among non-disputer
|
|
309
|
+
votes. Examples:
|
|
310
|
+
|
|
311
|
+
| non-disputer current votes | result |
|
|
312
|
+
|----------------------------|-------------|
|
|
313
|
+
| pass | disputed |
|
|
314
|
+
| pass + pass | confirmed |
|
|
315
|
+
| pass + reject | disputed |
|
|
316
|
+
| pass + pass + reject | confirmed |
|
|
317
|
+
| pass + pass + reject + reject | disputed |
|
|
318
|
+
|
|
319
|
+
```bash
|
|
320
|
+
# Read-only views of the attestation log + pending queue.
|
|
321
|
+
capcard-attestations --pending # awaiting cross-attestation
|
|
322
|
+
capcard-attestations --by verifier-001 # what this verifier signed
|
|
323
|
+
capcard-attestations --on 6f00d02d4244 # full audit trail on one record
|
|
324
|
+
```
|
|
325
|
+
|
|
326
|
+
**Trust hygiene.** Cross-attestation snapshots the verifier's
|
|
327
|
+
trust *at attestation time* into the row. If the verifier's trust
|
|
328
|
+
later changes (e.g. they made a string of bad calls), builder
|
|
329
|
+
records they confirmed keep the old modulator. Run
|
|
330
|
+
`capcard-recompute-trust` to propagate the change:
|
|
331
|
+
|
|
332
|
+
```bash
|
|
333
|
+
# Default: re-walk the log using snapshots (cheap).
|
|
334
|
+
capcard-recompute-trust --all
|
|
335
|
+
|
|
336
|
+
# Hygiene mode: re-walk using each verifier's *current* cached
|
|
337
|
+
# trust as the modulator. With --all we do a 2-pass converge so
|
|
338
|
+
# verifier scores stabilise before they're used as modulators on
|
|
339
|
+
# builder rows.
|
|
340
|
+
capcard-recompute-trust --all --use-current-verifier-trust
|
|
341
|
+
```
|
|
342
|
+
|
|
343
|
+
The library API (`capcard_cli.verify.recompute_trust`,
|
|
344
|
+
`recompute_all_trust`) takes the same `use_current_verifier_trust`
|
|
345
|
+
flag — useful for daemons that want to schedule periodic hygiene
|
|
346
|
+
runs without shelling out.
|
|
347
|
+
|
|
348
|
+
### Phase 5: training-data extraction + agent lineage
|
|
349
|
+
|
|
350
|
+
The CLI doesn't train RL agents — it makes the CapCard archive
|
|
351
|
+
*trainable*. Three primitives:
|
|
352
|
+
|
|
353
|
+
```bash
|
|
354
|
+
# Lineage. Generation auto-increments from --parent.
|
|
355
|
+
capcard-register --name bfs-engine-v2 --role researcher \
|
|
356
|
+
--parent researcher-001
|
|
357
|
+
# (researcher-002 → generation 2, parent_id=researcher-001)
|
|
358
|
+
|
|
359
|
+
# Walk the parent chain, root → leaf:
|
|
360
|
+
capcard list --by-lineage researcher-002
|
|
361
|
+
# Or every descendant of an ancestor:
|
|
362
|
+
capcard list --descendants-of researcher-001
|
|
363
|
+
|
|
364
|
+
# Training-data export — one shaped-reward episode per playbook.
|
|
365
|
+
capcard-export --role researcher --format jsonl > episodes.jsonl
|
|
366
|
+
capcard-export --since 2026-04-01T00:00:00Z --out training.jsonl
|
|
367
|
+
|
|
368
|
+
# Tactic effectiveness across every playbook on disk, ranked by
|
|
369
|
+
# (trust-weighted worked) - 0.5 * (trust-weighted failed).
|
|
370
|
+
capcard-tactics --on-goal "0 <= exp x" --limit 5
|
|
371
|
+
|
|
372
|
+
# Aggregate per-agent stats + evidence-type histogram. Useful to
|
|
373
|
+
# see whether a generation is graduating from self-report to
|
|
374
|
+
# lean-proof / test-suite over time.
|
|
375
|
+
capcard-stats --role researcher --by-generation
|
|
376
|
+
```
|
|
377
|
+
|
|
378
|
+
**Reward shaping** (`reward_for_playbook`):
|
|
379
|
+
|
|
380
|
+
base = clamp(pb.trust_delta, -1, 1)
|
|
381
|
+
bonus = +0.2 if any verification in the playbook window has
|
|
382
|
+
result = "honest-negative" with non-empty evidence
|
|
383
|
+
penalty = -0.3 if any verification in window has
|
|
384
|
+
evidence_type = "ci-broken"
|
|
385
|
+
reward = clamp(base + bonus + penalty, -1, 1)
|
|
386
|
+
|
|
387
|
+
Bonuses and penalties stack independently; a session that ships an
|
|
388
|
+
honest negative AND broke CI nets to `trust_delta - 0.1`. The
|
|
389
|
+
clamp at the end keeps the RL pipeline's reward bounded.
|
|
390
|
+
|
|
391
|
+
**Library API** (`capcard_cli.training`) — importable from RL
|
|
392
|
+
training scripts:
|
|
393
|
+
|
|
394
|
+
| function | purpose |
|
|
395
|
+
|--------------------------------|-----------------------------------------------------------------|
|
|
396
|
+
| `reward_for_playbook(pb)` | shaped scalar reward in `[-1, 1]` |
|
|
397
|
+
| `episode_for_playbook(pb)` | flatten one playbook into an `Episode` |
|
|
398
|
+
| `export_episodes(role=, ...)` | iterator over every playbook, optionally filtered |
|
|
399
|
+
| `tactic_score(t, on_goal=)` | scalar score for one tactic |
|
|
400
|
+
| `rank_tactics(on_goal=, ...)` | top-N tactics ranked by trust-weighted success |
|
|
401
|
+
| `agent_stats(role=, since=)` | per-agent rollup (sessions, cumulative_reward, avg_reward) |
|
|
402
|
+
| `evidence_type_histogram(...)` | rollup of evidence types from the verification log |
|
|
403
|
+
|
|
404
|
+
Generation lineage is stored on the `AgentRecord` itself —
|
|
405
|
+
`generation: int` (defaults to 1, or `parent.generation + 1`) and
|
|
406
|
+
`parent_id: str | None`. Lineage is locked in on first registration
|
|
407
|
+
for a given `--name`; subsequent re-registrations don't change it
|
|
408
|
+
(fork a new lineage with a new name instead).
|
|
409
|
+
|
|
410
|
+
### Phase 7: cross-registry wire format
|
|
411
|
+
|
|
412
|
+
The CLI doesn't run the network gossip layer (that's a separate
|
|
413
|
+
project — see [`docs/BACKLOG.md`](docs/BACKLOG.md)). What it does
|
|
414
|
+
provide is a **stable, versioned JSON wire format** + an
|
|
415
|
+
idempotent merge importer, ready for a future gossip layer to
|
|
416
|
+
serialize over.
|
|
417
|
+
|
|
418
|
+
```bash
|
|
419
|
+
# Snapshot the local state into one JSON document.
|
|
420
|
+
capcard-bundle export --out snapshot.json --source-id "lab-monogate"
|
|
421
|
+
|
|
422
|
+
# Merge into another registry. Idempotent on primary keys
|
|
423
|
+
# (agent_id, verification_id, attestation_id, (agent_id, session_id)).
|
|
424
|
+
# Local state is authoritative — conflicts are logged and skipped,
|
|
425
|
+
# never overwritten.
|
|
426
|
+
capcard-bundle import snapshot.json
|
|
427
|
+
capcard-bundle import snapshot.json --dry-run # preview only
|
|
428
|
+
capcard-bundle import snapshot.json --json # ImportSummary as JSON
|
|
429
|
+
|
|
430
|
+
# Verify a verification record's stored output digest matches a
|
|
431
|
+
# fresh re-run of the original evidence command. Poor-man's
|
|
432
|
+
# "did this happen" check without ZK proofs.
|
|
433
|
+
capcard-fingerprint --verification-id 6f00d02d4244 --rerun
|
|
434
|
+
# → MATCH (exit 0) | MISMATCH (exit 3) | unknown id (exit 1)
|
|
435
|
+
```
|
|
436
|
+
|
|
437
|
+
The bundle format is a single JSON object with `bundle_version`,
|
|
438
|
+
`created`, `source_id`, and arrays of agents / verifications /
|
|
439
|
+
attestations / playbooks. After import, every agent's
|
|
440
|
+
`trust_score` is recomputed under the same lock so newly-merged
|
|
441
|
+
verifications and attestations propagate into the cached score.
|
|
442
|
+
|
|
443
|
+
`AgentRecord` carries an optional `pubkey: str | None` field
|
|
444
|
+
(`capcard-register --pubkey "..."`). It's forward-reserved — the
|
|
445
|
+
stdlib-only constraint means the CLI doesn't sign anything yet.
|
|
446
|
+
The field is here so a Phase 7+ network protocol that *does* sign
|
|
447
|
+
can import bundles without a schema migration.
|
|
448
|
+
|
|
449
|
+
### Phase 6.7: dispute time-decay + iterative trust converge
|
|
450
|
+
|
|
451
|
+
Two hygiene knobs on `capcard-recompute-trust` close the open
|
|
452
|
+
edges from the Phase 6.5 review:
|
|
453
|
+
|
|
454
|
+
```bash
|
|
455
|
+
# Stale disputes (latest dispute attestation older than N days)
|
|
456
|
+
# are ignored, letting the underlying verdict surface. Useful when
|
|
457
|
+
# nobody followed up on a contested record.
|
|
458
|
+
capcard-recompute-trust --all --dispute-decay-days 7
|
|
459
|
+
|
|
460
|
+
# Iterative converge to a fixed point. Default 2-pass already
|
|
461
|
+
# handles the common case; --iterative is for verifier↔verifier
|
|
462
|
+
# review chains that could oscillate. Stops when the L-infinity
|
|
463
|
+
# per-agent delta drops below --epsilon (default 0.001) or
|
|
464
|
+
# --max-iter (default 10) is hit.
|
|
465
|
+
capcard-recompute-trust --all --use-current-verifier-trust \
|
|
466
|
+
--iterative --epsilon 0.0001
|
|
467
|
+
```
|
|
468
|
+
|
|
469
|
+
Both flags are off by default so the conservative Phase 6.5
|
|
470
|
+
semantics still apply unless the operator opts in.
|
|
471
|
+
|
|
472
|
+
### Generate a session-end playbook
|
|
473
|
+
|
|
474
|
+
When the agent finishes a session (or hits a meaningful checkpoint),
|
|
475
|
+
`capcard-playbook generate` writes a structured summary to disk and
|
|
476
|
+
points the agent's `playbook_path` at the new file. The trust delta,
|
|
477
|
+
verification window, and result_summary are computed automatically
|
|
478
|
+
from the verification log since the agent's previous playbook.
|
|
479
|
+
|
|
480
|
+
```bash
|
|
481
|
+
capcard-playbook generate --agent researcher-001 \
|
|
482
|
+
--result theorem-established \
|
|
483
|
+
--lesson "P(tuple) undefined at depth >= 2" \
|
|
484
|
+
--lesson "OfScientific vs OfNat blocks proofs" \
|
|
485
|
+
--tactic-worked "exact exp_nonneg _||0 <= exp x" \
|
|
486
|
+
--tactic-failed "trivial||0 <= exp x" \
|
|
487
|
+
--artifact "machlib:9589804" \
|
|
488
|
+
--open-question "depth-2 unfold under products?"
|
|
489
|
+
```
|
|
490
|
+
|
|
491
|
+
`--tactic-worked` / `--tactic-failed` use a literal `||` separator
|
|
492
|
+
between tactic and goal-shape label (the goal label is what the BFS
|
|
493
|
+
engine keys on — keep it consistent across sessions).
|
|
494
|
+
|
|
495
|
+
For bulk input, write the same fields to a JSON file and pass
|
|
496
|
+
`--from-file`. Per-flag values are merged on top:
|
|
497
|
+
|
|
498
|
+
```bash
|
|
499
|
+
cat > playbook.json <<'JSON'
|
|
500
|
+
{
|
|
501
|
+
"lessons": ["P(tuple) undefined at depth >= 2"],
|
|
502
|
+
"tactics_that_worked": [
|
|
503
|
+
{"tactic": "exact exp_nonneg _", "on_goal": "0 <= exp x"}
|
|
504
|
+
],
|
|
505
|
+
"tactics_that_failed": [
|
|
506
|
+
{"tactic": "trivial", "on_goal": "0 <= exp x"}
|
|
507
|
+
],
|
|
508
|
+
"artifacts": ["machlib:9589804"],
|
|
509
|
+
"open_questions": ["depth-2 unfold under products?"]
|
|
510
|
+
}
|
|
511
|
+
JSON
|
|
512
|
+
|
|
513
|
+
capcard-playbook generate --agent researcher-001 \
|
|
514
|
+
--result theorem-established \
|
|
515
|
+
--from-file playbook.json
|
|
516
|
+
```
|
|
517
|
+
|
|
518
|
+
### Search prior playbooks (transfer)
|
|
519
|
+
|
|
520
|
+
When a new agent picks up a related task, search the playbook
|
|
521
|
+
archive for what's already been tried — the highest-leverage
|
|
522
|
+
transfer is the list of tactics that *failed*, so future agents
|
|
523
|
+
don't repeat them.
|
|
524
|
+
|
|
525
|
+
```bash
|
|
526
|
+
# Find related work — score is the count of distinct query terms
|
|
527
|
+
# matched across task / lessons / tactics / open_questions.
|
|
528
|
+
capcard-playbook search --task "Galois dimension proof for chain order"
|
|
529
|
+
|
|
530
|
+
# Filter to one role + tighten the score floor
|
|
531
|
+
capcard-playbook search --task "exp_strict_mono on negative inputs" \
|
|
532
|
+
--role researcher --min-score 2
|
|
533
|
+
|
|
534
|
+
# Machine-readable for piping into another agent's context
|
|
535
|
+
capcard-playbook search --task "..." --json | jq '.[0].playbook'
|
|
536
|
+
```
|
|
537
|
+
|
|
538
|
+
The transfer mechanism is **explicit, not magic** — the calling
|
|
539
|
+
agent runs the search and decides what to load. (Hidden auto-load
|
|
540
|
+
on `capcard-register` would silently change session behavior, which
|
|
541
|
+
we don't want.)
|
|
542
|
+
|
|
543
|
+
### Show / inspect
|
|
544
|
+
|
|
545
|
+
```bash
|
|
546
|
+
capcard-playbook show --agent researcher-001 # newest playbook
|
|
547
|
+
capcard-playbook show --path /path/to/playbook.json # specific file
|
|
548
|
+
```
|
|
549
|
+
|
|
550
|
+
### Verification log
|
|
551
|
+
|
|
552
|
+
`capcard-verify` appends one JSONL row per call to
|
|
553
|
+
`~/.capcard/verifications.jsonl` (override via
|
|
554
|
+
`CAPCARD_VERIFY_LOG_PATH`):
|
|
555
|
+
|
|
556
|
+
```json
|
|
557
|
+
{"agent_id":"researcher-001","evidence":"lake build","exit_code":null,
|
|
558
|
+
"notes":null,"output_digest":"d41d8cd9...","ran":false,
|
|
559
|
+
"result":"proof-closed","runtime_s":null,"task":"Galois dimension proof",
|
|
560
|
+
"theorem":"tanh_monotone","timestamp":"2026-05-04T13:30:00Z"}
|
|
561
|
+
```
|
|
562
|
+
|
|
563
|
+
The log is append-only and serialized under its own flock
|
|
564
|
+
(`~/.capcard/verifications.jsonl.lock`). The registry stores the
|
|
565
|
+
rolled-up counters; the log keeps the per-event detail Phases 3
|
|
566
|
+
(playbook), 6 (cross-verify), and 7 (ZK fingerprints) need.
|
|
567
|
+
|
|
568
|
+
## Wire it into a Claude Code session
|
|
569
|
+
|
|
570
|
+
Add to your shell rc (or per-project `.envrc`):
|
|
571
|
+
|
|
572
|
+
```bash
|
|
573
|
+
# At session start, register and capture the assigned agent_id
|
|
574
|
+
export CAPCARD_AGENT_ID=$(capcard-register \
|
|
575
|
+
--name "$USER on $(hostname -s)" \
|
|
576
|
+
--role researcher \
|
|
577
|
+
--json | python -c 'import json,sys; print(json.load(sys.stdin)["agent_id"])')
|
|
578
|
+
|
|
579
|
+
# Then any time you switch tasks
|
|
580
|
+
capcard-status --agent "$CAPCARD_AGENT_ID" --task "$NEW_TASK_DESC"
|
|
581
|
+
|
|
582
|
+
# When you finish a unit of work, log the outcome with evidence
|
|
583
|
+
capcard-verify --agent "$CAPCARD_AGENT_ID" \
|
|
584
|
+
--evidence "lake build" \
|
|
585
|
+
--result proof-closed \
|
|
586
|
+
--theorem tanh_monotone \
|
|
587
|
+
--run
|
|
588
|
+
```
|
|
589
|
+
|
|
590
|
+
Once the Command Center's `/agents` page is updated to read the
|
|
591
|
+
extended `blackwell-agent` `/agents` endpoint (which already merges
|
|
592
|
+
the registry with live tmux sessions as of this commit), the
|
|
593
|
+
registered agent appears with its persistent ID, role, and trust
|
|
594
|
+
score alongside the anonymous tmux sessions on the same machine.
|
|
595
|
+
|
|
596
|
+
## What's coming next
|
|
597
|
+
|
|
598
|
+
Per the [vision doc](https://capcard.ai/whitepaper) :
|
|
599
|
+
|
|
600
|
+
- **Phase 2 (`capcard-verify`)** — **shipped in 0.2.0**: log work
|
|
601
|
+
outcomes with cryptographic evidence (`lake build` passing,
|
|
602
|
+
`pytest 87 passed`, CI red). Bumps `verified_outputs` /
|
|
603
|
+
`failed_outputs` per the result-type contract; appends a row to
|
|
604
|
+
`~/.capcard/verifications.jsonl`.
|
|
605
|
+
- **Phase 3 (`capcard-playbook`)** — **shipped in 0.3.0**: emit a
|
|
606
|
+
structured summary at session end (lessons / tactics_that_worked
|
|
607
|
+
/ tactics_that_failed / open_questions); next agent inherits via
|
|
608
|
+
`capcard-playbook search`.
|
|
609
|
+
- **Phase 4 — weighted trust** — **shipped in 0.4.0**: replaces the
|
|
610
|
+
simple `verified / total` ratio with the weight table above.
|
|
611
|
+
`--evidence-type` declares what kind of proof the agent has;
|
|
612
|
+
trust is the weighted average clamped to `[0, 1]`.
|
|
613
|
+
- **Phase 6 — cross-agent attestation** — **shipped in 0.5.0**:
|
|
614
|
+
`--needs-attestation` + `capcard-attest`. Builder = verifier
|
|
615
|
+
rejected; verifier trust modulates the builder's contribution
|
|
616
|
+
with a 0.5 floor.
|
|
617
|
+
- **Phase 6.5 — dispute resolution + trust hygiene** — **shipped
|
|
618
|
+
in 0.6.0**: per-verifier vote collapse, dispute quorum
|
|
619
|
+
resolution (≥2 distinct non-disputers agreeing override the
|
|
620
|
+
dispute), `capcard-attestations` (`--pending` / `--by` /
|
|
621
|
+
`--on`), `capcard-recompute-trust` with optional
|
|
622
|
+
`--use-current-verifier-trust` mode for propagating verifier-
|
|
623
|
+
trust changes back into builder rows.
|
|
624
|
+
- **Phase 5 — RL training-data extraction** — **shipped in 0.7.0**:
|
|
625
|
+
generation/parent_id on AgentRecord, `capcard-export` (shaped-
|
|
626
|
+
reward episodes), `capcard-tactics` (trust-weighted ranking),
|
|
627
|
+
`capcard-stats` (per-agent rollup + evidence histogram). The CLI
|
|
628
|
+
doesn't train models — it makes the archive trainable.
|
|
629
|
+
- **Phase 6.7 — trust hygiene** — **shipped in 1.0.0**: time-decay
|
|
630
|
+
for stuck dispute markers, iterative fixed-point converge in
|
|
631
|
+
`recompute_all_trust`. Both opt-in via flags on
|
|
632
|
+
`capcard-recompute-trust`.
|
|
633
|
+
- **Phase 7 — wire format + fingerprint verify** — **shipped in
|
|
634
|
+
1.0.0**: `capcard-bundle export/import` (versioned, idempotent
|
|
635
|
+
merge), `capcard-fingerprint --rerun` for spot-checking imported
|
|
636
|
+
output digests, `pubkey: str | None` field on AgentRecord. The
|
|
637
|
+
*actual* network gossip layer + ZK proofs are an independent
|
|
638
|
+
research project — see [`docs/BACKLOG.md`](docs/BACKLOG.md).
|
|
639
|
+
|
|
640
|
+
The capcard-cli pipeline is feature-complete for the whitepaper
|
|
641
|
+
roadmap as of 1.0.0. Future work (signing, gossip, ZK
|
|
642
|
+
fingerprints, network resilience) belongs to the network/protocol
|
|
643
|
+
side of the project, not this CLI.
|
|
644
|
+
|
|
645
|
+
## Testing
|
|
646
|
+
|
|
647
|
+
```bash
|
|
648
|
+
pip install -e '.[dev]'
|
|
649
|
+
python -m pytest
|
|
650
|
+
```
|
|
651
|
+
|
|
652
|
+
244 tests cover: schema, atomic writes, file locking under 20-process
|
|
653
|
+
contention, idempotency, role-migration discipline, evidence-weight
|
|
654
|
+
math, playbook windowing, attestation state machine + dispute
|
|
655
|
+
quorum resolution + time-decay, two-pass and iterative trust
|
|
656
|
+
converge, lineage walkers, reward shaping (incl. clamps and
|
|
657
|
+
bonus/penalty stacking), tactic ranking across playbooks, episode
|
|
658
|
+
export filters, bundle round-trip + idempotent merge + conflict
|
|
659
|
+
preservation, fingerprint match/mismatch, and CLI surface.
|
|
660
|
+
|
|
661
|
+
## License
|
|
662
|
+
|
|
663
|
+
Apache-2.0 (matches the rest of the Monogate stack).
|