monogate-capcard-cli 1.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (26) hide show
  1. monogate_capcard_cli-1.0.0/LICENSE +17 -0
  2. monogate_capcard_cli-1.0.0/PKG-INFO +663 -0
  3. monogate_capcard_cli-1.0.0/README.md +640 -0
  4. monogate_capcard_cli-1.0.0/capcard_cli/__init__.py +55 -0
  5. monogate_capcard_cli-1.0.0/capcard_cli/attest.py +444 -0
  6. monogate_capcard_cli-1.0.0/capcard_cli/bundle.py +403 -0
  7. monogate_capcard_cli-1.0.0/capcard_cli/cli.py +1530 -0
  8. monogate_capcard_cli-1.0.0/capcard_cli/playbook.py +561 -0
  9. monogate_capcard_cli-1.0.0/capcard_cli/registry.py +484 -0
  10. monogate_capcard_cli-1.0.0/capcard_cli/training.py +463 -0
  11. monogate_capcard_cli-1.0.0/capcard_cli/verify.py +901 -0
  12. monogate_capcard_cli-1.0.0/monogate_capcard_cli.egg-info/PKG-INFO +663 -0
  13. monogate_capcard_cli-1.0.0/monogate_capcard_cli.egg-info/SOURCES.txt +24 -0
  14. monogate_capcard_cli-1.0.0/monogate_capcard_cli.egg-info/dependency_links.txt +1 -0
  15. monogate_capcard_cli-1.0.0/monogate_capcard_cli.egg-info/entry_points.txt +14 -0
  16. monogate_capcard_cli-1.0.0/monogate_capcard_cli.egg-info/requires.txt +3 -0
  17. monogate_capcard_cli-1.0.0/monogate_capcard_cli.egg-info/top_level.txt +1 -0
  18. monogate_capcard_cli-1.0.0/pyproject.toml +58 -0
  19. monogate_capcard_cli-1.0.0/setup.cfg +4 -0
  20. monogate_capcard_cli-1.0.0/tests/test_attest.py +864 -0
  21. monogate_capcard_cli-1.0.0/tests/test_cli.py +626 -0
  22. monogate_capcard_cli-1.0.0/tests/test_phase7.py +555 -0
  23. monogate_capcard_cli-1.0.0/tests/test_playbook.py +422 -0
  24. monogate_capcard_cli-1.0.0/tests/test_registry.py +306 -0
  25. monogate_capcard_cli-1.0.0/tests/test_training.py +529 -0
  26. monogate_capcard_cli-1.0.0/tests/test_verify.py +647 -0
@@ -0,0 +1,17 @@
1
+ Apache License
2
+ Version 2.0, January 2004
3
+ http://www.apache.org/licenses/
4
+
5
+ Copyright 2026 Mosa Creates LLC
6
+
7
+ Licensed under the Apache License, Version 2.0 (the "License");
8
+ you may not use this file except in compliance with the License.
9
+ You may obtain a copy of the License at
10
+
11
+ http://www.apache.org/licenses/LICENSE-2.0
12
+
13
+ Unless required by applicable law or agreed to in writing, software
14
+ distributed under the License is distributed on an "AS IS" BASIS,
15
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
16
+ See the License for the specific language governing permissions and
17
+ limitations under the License.
@@ -0,0 +1,663 @@
1
+ Metadata-Version: 2.4
2
+ Name: monogate-capcard-cli
3
+ Version: 1.0.0
4
+ Summary: Persistent identity + verification CLIs for Monogate CapCard agents
5
+ Author-email: Monogate <almaguer1986@gmail.com>
6
+ License: Apache-2.0
7
+ Project-URL: Homepage, https://capcard.ai
8
+ Project-URL: Source, https://github.com/agent-maestro/capcard-cli
9
+ Keywords: monogate,capcard,agents,verification,trust
10
+ Classifier: Development Status :: 3 - Alpha
11
+ Classifier: Intended Audience :: Developers
12
+ Classifier: License :: OSI Approved :: Apache Software License
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Programming Language :: Python :: 3.10
15
+ Classifier: Programming Language :: Python :: 3.11
16
+ Classifier: Programming Language :: Python :: 3.12
17
+ Requires-Python: >=3.10
18
+ Description-Content-Type: text/markdown
19
+ License-File: LICENSE
20
+ Provides-Extra: dev
21
+ Requires-Dist: pytest>=8; extra == "dev"
22
+ Dynamic: license-file
23
+
24
+ # capcard-cli
25
+
26
+ Persistent identity + weighted-trust verification + knowledge
27
+ transfer + cross-agent attestation + RL training-data extraction +
28
+ cross-registry wire format for Monogate CapCard agents. All seven
29
+ phases of the [CapCard: Verified Agents](https://capcard.ai)
30
+ roadmap — turns anonymous Claude Code sessions into named agents
31
+ with stable IDs + generation lineage, evidence-backed verification
32
+ records scored on a Lean-proof-to-self-report scale, structured
33
+ session-end playbooks the next agent on a related task can read, a
34
+ "builder posts / verifier signs / both trust scores update"
35
+ cross-attestation flow so trust is earned, a training-data export
36
+ pipeline that turns the accumulated playbook archive into shaped-
37
+ reward episodes for an RL agent training loop, and a stable JSON
38
+ wire format with idempotent merge for cross-registry sync.
39
+
40
+ > **All phases shipped (1.0.0).** Twelve binaries on PATH:
41
+ > `capcard-register` (with `--generation` / `--parent` /
42
+ > `--pubkey`), `capcard-status`, `capcard-verify` (with
43
+ > `--evidence-type` and `--needs-attestation`), `capcard-attest`
44
+ > (`--pass` / `--reject` / `--dispute`), `capcard-attestations`
45
+ > (`--pending` / `--by` / `--on`), `capcard-recompute-trust` (with
46
+ > `--iterative` and `--dispute-decay-days`), `capcard-playbook`
47
+ > (`generate` / `show` / `search`), `capcard-export`,
48
+ > `capcard-tactics`, `capcard-stats`, `capcard-bundle`
49
+ > (`export` / `import`), `capcard-fingerprint`. Trust is the
50
+ > weighted formula `sum(w_i) / sum(|w_i|)` clamped to `[0, 1]`,
51
+ > with cross-attestation gating contribution when the builder
52
+ > requests it. Disputes resolve by quorum (≥2 distinct
53
+ > non-disputers agreeing); per-verifier votes collapse; stale
54
+ > dispute markers can be aged out with `--dispute-decay-days`.
55
+ > Phase 5 turns the playbook archive into shaped-reward episodes
56
+ > for an RL pipeline. Phase 7 ships a versioned JSON wire format
57
+ > ready for a future network gossip layer (the actual gossip + ZK
58
+ > proofs are an out-of-CLI research project — see
59
+ > [`docs/BACKLOG.md`](docs/BACKLOG.md)).
60
+
61
+ ## Install
62
+
63
+ ```bash
64
+ git clone https://github.com/agent-maestro/capcard-cli # once on PyPI: pip install monogate-capcard-cli
65
+ cd capcard-cli
66
+ pip install -e .
67
+ ```
68
+
69
+ The package depends on nothing outside the Python stdlib — installs
70
+ into any Claude Code session's Python without dragging in a wider
71
+ requirements graph.
72
+
73
+ ## Use
74
+
75
+ ### Register at session start
76
+
77
+ ```bash
78
+ capcard-register --name "Frontier C Researcher" \
79
+ --role researcher \
80
+ --task "Galois dimension proof"
81
+ ```
82
+
83
+ Roles: `researcher`, `builder`, `tooling`, `verifier` (any other
84
+ string is accepted with a warning to stderr — extensible without
85
+ a registry migration).
86
+
87
+ The first registration with a given `--name` allocates an
88
+ `agent_id` slug (`researcher-001`, `builder-001`, etc.). Subsequent
89
+ calls with the same name return the same ID and update
90
+ `current_task` + `last_active`.
91
+
92
+ ### Update mid-session
93
+
94
+ ```bash
95
+ capcard-status --agent researcher-001 \
96
+ --task "writing arXiv update" \
97
+ --status working
98
+ ```
99
+
100
+ Status values: `working`, `idle`, `blocked`, `done`, `failed`
101
+ (also extensible).
102
+
103
+ ### List
104
+
105
+ ```bash
106
+ capcard list
107
+ capcard list --json # machine-readable
108
+ ```
109
+
110
+ ### Verify a unit of work
111
+
112
+ `capcard-verify` records one outcome with cryptographic evidence,
113
+ appends a row to the verification log, and updates the agent's
114
+ trust counters. The agent owns the framing — `capcard-verify` is
115
+ how they sign their claim into the log.
116
+
117
+ ```bash
118
+ # Treat --evidence as a literal artifact (you've already collected it)
119
+ capcard-verify --agent researcher-001 \
120
+ --evidence "lake build clean (12/12 modules)" \
121
+ --result proof-closed \
122
+ --theorem tanh_monotone
123
+
124
+ # Have capcard-verify execute the evidence command and capture the result
125
+ capcard-verify --agent builder-001 \
126
+ --evidence "pytest -q" \
127
+ --result feature-shipped \
128
+ --run
129
+
130
+ # Honest-negative: investigation that returned a negative finding.
131
+ # Counts as a verified outcome — does NOT decrease trust.
132
+ capcard-verify --agent researcher-001 \
133
+ --evidence "C-246 NOTES.md: Phi homomorphism is trivial" \
134
+ --result honest-negative
135
+
136
+ # Failure: shipped broken code, broke CI, false-claim correction.
137
+ # Decreases trust.
138
+ capcard-verify --agent builder-001 \
139
+ --evidence "CI red on main: type signature regression" \
140
+ --result failure
141
+ ```
142
+
143
+ The result types and their effect on the simple counter pair (the
144
+ weighted trust score uses the `--evidence-type` table below — see
145
+ "Phase 4: weighted trust"):
146
+
147
+ | `--result` | Counter bumped | Trust direction |
148
+ |------------------------|----------------------|------------------------|
149
+ | `proof-closed` | `verified_outputs` | up |
150
+ | `feature-shipped` | `verified_outputs` | up |
151
+ | `theorem-established` | `verified_outputs` | up |
152
+ | `bug-fixed` | `verified_outputs` | up |
153
+ | `honest-negative` | `verified_outputs` | up (per spec) |
154
+ | `failure` | `failed_outputs` | down |
155
+
156
+ ### Phase 4: weighted trust
157
+
158
+ `capcard-verify` accepts an optional `--evidence-type` that drives
159
+ the weighted trust formula. Heavier weights mean stronger evidence:
160
+ a Lean proof checked by `lake build` is unforgeable (1.0), a self-
161
+ report is just the agent's word (0.2). Negative weights model
162
+ verifiable harm.
163
+
164
+ | `--evidence-type` | Weight | When to use |
165
+ |-------------------|-------:|---------------------------------------------------|
166
+ | `lean-proof` | +1.0 | Lean proof passed `lake build` cleanly |
167
+ | `bfs-closure` | +0.9 | BFS proof engine closed a sorry; lake-verified |
168
+ | `test-suite` | +0.7 | Full test suite passed |
169
+ | `ci-green` | +0.6 | CI green (depends on test quality) |
170
+ | `empirical` | +0.5 | N/N empirical hold; suggestive, not proven |
171
+ | `honest-negative` | +0.5 | Investigated; reported negative finding honestly |
172
+ | `self-report` | +0.2 | No independent verification |
173
+ | `ci-broken` | -0.3 | Agent broke CI |
174
+ | `false-claim` | -0.5 | Agent's earlier claim was corrected |
175
+
176
+ ```bash
177
+ # 47/47 spot-check is empirical, not a proof — be explicit
178
+ capcard-verify --agent researcher-001 \
179
+ --evidence "47/47 holds across [501, 2000]" \
180
+ --result theorem-established \
181
+ --evidence-type empirical
182
+
183
+ # BFS engine closed a Discovered/ sorry
184
+ capcard-verify --agent tooling-001 \
185
+ --evidence "tools/sweep closed exp_strict_mono_at_zero" \
186
+ --result proof-closed \
187
+ --evidence-type bfs-closure
188
+ ```
189
+
190
+ When `--evidence-type` is omitted, the type is inferred from the
191
+ result + evidence keywords (`lake build` → `lean-proof`, `pytest …
192
+ passed` → `test-suite`, "ci"/"broke"/"build" tokens in a `failure`
193
+ → `ci-broken`, etc.). Pass it explicitly when the heuristic is
194
+ wrong — the conservative default is `self-report` (0.2) for
195
+ positive results without obvious signal.
196
+
197
+ **Trust formula:** `trust = sum(w_i) / sum(|w_i|)` over all
198
+ verifications, clamped to `[0, 1]` from below. An empty history
199
+ returns 0.0 ("no track record"). The cached `trust_score` field on
200
+ the registry row is updated by every `record_verification` call;
201
+ operators can rebuild it from the log via `recompute_trust(agent_id)`
202
+ after a weight-table change.
203
+
204
+ `--run` shells out via `shlex.split` (no shell interpolation —
205
+ quote your arguments) and captures `exit_code`, wall-clock
206
+ `runtime_s`, full `stdout` + `stderr`, and a SHA-256 digest of
207
+ the combined output. `--timeout-s` defaults to 600s. If
208
+ `--run` finishes with a non-zero exit but you claimed a positive
209
+ result, the record is written as you stated it but a note is
210
+ attached so a Phase 6 verifier can spot the mismatch.
211
+
212
+ When `--run` is *not* set, `--evidence` is stored verbatim and
213
+ the digest is `sha256(evidence)`.
214
+
215
+ ## What it stores
216
+
217
+ A single JSON file at `~/.agent-registry.json` (override via
218
+ `CAPCARD_REGISTRY_PATH`):
219
+
220
+ ```json
221
+ {
222
+ "version": 1,
223
+ "agents": {
224
+ "researcher-001": {
225
+ "agent_id": "researcher-001",
226
+ "name": "Frontier C Researcher",
227
+ "role": "researcher",
228
+ "status": "working",
229
+ "current_task": "Galois dimension proof",
230
+ "total_tasks": 0,
231
+ "verified_outputs": 0,
232
+ "failed_outputs": 0,
233
+ "trust_score": 0.0,
234
+ "created": "2026-05-04T13:30:00Z",
235
+ "last_active": "2026-05-04T13:30:00Z",
236
+ "playbook_path": null
237
+ }
238
+ }
239
+ }
240
+ ```
241
+
242
+ Writes are atomic (temp-file-then-rename) and serialized via
243
+ `flock` on a sidecar `~/.agent-registry.json.lock`. Two parallel
244
+ `capcard-register` calls cannot lose updates.
245
+
246
+ ### Phase 6: cross-agent attestation
247
+
248
+ By default `capcard-verify` self-attests — the agent's claim
249
+ contributes to their own trust immediately. Phase 6 adds a
250
+ two-step path so a verifier can confirm or reject another agent's
251
+ work before it counts.
252
+
253
+ ```bash
254
+ # Builder records a pending verification (--needs-attestation).
255
+ # The record lands in the 'pending' state and contributes 0 to
256
+ # trust until cross-attested.
257
+ capcard-verify --agent builder-001 \
258
+ --evidence "pytest -q && cargo test --release" \
259
+ --result feature-shipped \
260
+ --evidence-type test-suite \
261
+ --needs-attestation
262
+ # Output ends with:
263
+ # id: 6f00d02d4244 ← capture this
264
+ # attestation: pending (--needs-attestation set)
265
+
266
+ # A verifier (must be a different agent) signs off:
267
+ capcard-attest --verification-id 6f00d02d4244 \
268
+ --verifier-agent verifier-001 \
269
+ --pass \
270
+ --evidence "reviewed diff + ran pytest myself" \
271
+ --note "looks clean"
272
+ ```
273
+
274
+ Outcomes:
275
+
276
+ | flag | effect on builder trust | effect on verifier trust |
277
+ |-------------|----------------------------------------------------------------------------------------------------------|-------------------------------------------|
278
+ | `--pass` | record contributes `base_weight × max(0.5, verifier_trust)` to the numerator (denom stays at base_weight) | +0.6 weight (review work is real work) |
279
+ | `--reject` | record treated as false-claim against builder (-0.5 weight) | +0.6 weight |
280
+ | `--dispute` | record held in `disputed` state (contributes 0) until quorum resolves it (Phase 6.5) | +0.6 weight |
281
+
282
+ Guards:
283
+
284
+ - **No self-grading** — `--verifier-agent` must be different from
285
+ the builder. Same-agent attestation returns exit 2.
286
+ - **Cross-attesting a self-attested record** returns exit 2 — the
287
+ builder must opt in via `--needs-attestation`.
288
+ - **Verifier must be registered** — needed so the trust snapshot
289
+ has a real value to reference.
290
+
291
+ The append-only attestation log lives at
292
+ `~/.capcard/attestations.jsonl` (override with
293
+ `CAPCARD_ATTESTATION_LOG_PATH`). Overturning a previous decision
294
+ means recording a *new* attestation, not editing the old row.
295
+
296
+ ### Phase 6.5: dispute resolution + trust hygiene
297
+
298
+ The state machine resolves disputes by **quorum** instead of holding
299
+ disputed records open forever. Two rules:
300
+
301
+ 1. **Per-verifier vote collapse.** A verifier who casts multiple
302
+ attestations on the same record contributes one current vote —
303
+ their latest stance. Flip-flopping doesn't manufacture extra
304
+ votes (the append-only log preserves the audit trail for
305
+ `capcard-attestations --on <vid>`).
306
+ 2. **Dispute quorum.** A `dispute` marker is overridden when ≥2
307
+ distinct non-disputer verifiers' current votes agree on the same
308
+ verdict and that verdict has a strict majority among non-disputer
309
+ votes. Examples:
310
+
311
+ | non-disputer current votes | result |
312
+ |----------------------------|-------------|
313
+ | pass | disputed |
314
+ | pass + pass | confirmed |
315
+ | pass + reject | disputed |
316
+ | pass + pass + reject | confirmed |
317
+ | pass + pass + reject + reject | disputed |
318
+
319
+ ```bash
320
+ # Read-only views of the attestation log + pending queue.
321
+ capcard-attestations --pending # awaiting cross-attestation
322
+ capcard-attestations --by verifier-001 # what this verifier signed
323
+ capcard-attestations --on 6f00d02d4244 # full audit trail on one record
324
+ ```
325
+
326
+ **Trust hygiene.** Cross-attestation snapshots the verifier's
327
+ trust *at attestation time* into the row. If the verifier's trust
328
+ later changes (e.g. they made a string of bad calls), builder
329
+ records they confirmed keep the old modulator. Run
330
+ `capcard-recompute-trust` to propagate the change:
331
+
332
+ ```bash
333
+ # Default: re-walk the log using snapshots (cheap).
334
+ capcard-recompute-trust --all
335
+
336
+ # Hygiene mode: re-walk using each verifier's *current* cached
337
+ # trust as the modulator. With --all we do a 2-pass converge so
338
+ # verifier scores stabilise before they're used as modulators on
339
+ # builder rows.
340
+ capcard-recompute-trust --all --use-current-verifier-trust
341
+ ```
342
+
343
+ The library API (`capcard_cli.verify.recompute_trust`,
344
+ `recompute_all_trust`) takes the same `use_current_verifier_trust`
345
+ flag — useful for daemons that want to schedule periodic hygiene
346
+ runs without shelling out.
347
+
348
+ ### Phase 5: training-data extraction + agent lineage
349
+
350
+ The CLI doesn't train RL agents — it makes the CapCard archive
351
+ *trainable*. Three primitives:
352
+
353
+ ```bash
354
+ # Lineage. Generation auto-increments from --parent.
355
+ capcard-register --name bfs-engine-v2 --role researcher \
356
+ --parent researcher-001
357
+ # (researcher-002 → generation 2, parent_id=researcher-001)
358
+
359
+ # Walk the parent chain, root → leaf:
360
+ capcard list --by-lineage researcher-002
361
+ # Or every descendant of an ancestor:
362
+ capcard list --descendants-of researcher-001
363
+
364
+ # Training-data export — one shaped-reward episode per playbook.
365
+ capcard-export --role researcher --format jsonl > episodes.jsonl
366
+ capcard-export --since 2026-04-01T00:00:00Z --out training.jsonl
367
+
368
+ # Tactic effectiveness across every playbook on disk, ranked by
369
+ # (trust-weighted worked) - 0.5 * (trust-weighted failed).
370
+ capcard-tactics --on-goal "0 <= exp x" --limit 5
371
+
372
+ # Aggregate per-agent stats + evidence-type histogram. Useful to
373
+ # see whether a generation is graduating from self-report to
374
+ # lean-proof / test-suite over time.
375
+ capcard-stats --role researcher --by-generation
376
+ ```
377
+
378
+ **Reward shaping** (`reward_for_playbook`):
379
+
380
+ base = clamp(pb.trust_delta, -1, 1)
381
+ bonus = +0.2 if any verification in the playbook window has
382
+ result = "honest-negative" with non-empty evidence
383
+ penalty = -0.3 if any verification in window has
384
+ evidence_type = "ci-broken"
385
+ reward = clamp(base + bonus + penalty, -1, 1)
386
+
387
+ Bonuses and penalties stack independently; a session that ships an
388
+ honest negative AND broke CI nets to `trust_delta - 0.1`. The
389
+ clamp at the end keeps the RL pipeline's reward bounded.
390
+
391
+ **Library API** (`capcard_cli.training`) — importable from RL
392
+ training scripts:
393
+
394
+ | function | purpose |
395
+ |--------------------------------|-----------------------------------------------------------------|
396
+ | `reward_for_playbook(pb)` | shaped scalar reward in `[-1, 1]` |
397
+ | `episode_for_playbook(pb)` | flatten one playbook into an `Episode` |
398
+ | `export_episodes(role=, ...)` | iterator over every playbook, optionally filtered |
399
+ | `tactic_score(t, on_goal=)` | scalar score for one tactic |
400
+ | `rank_tactics(on_goal=, ...)` | top-N tactics ranked by trust-weighted success |
401
+ | `agent_stats(role=, since=)` | per-agent rollup (sessions, cumulative_reward, avg_reward) |
402
+ | `evidence_type_histogram(...)` | rollup of evidence types from the verification log |
403
+
404
+ Generation lineage is stored on the `AgentRecord` itself —
405
+ `generation: int` (defaults to 1, or `parent.generation + 1`) and
406
+ `parent_id: str | None`. Lineage is locked in on first registration
407
+ for a given `--name`; subsequent re-registrations don't change it
408
+ (fork a new lineage with a new name instead).
409
+
410
+ ### Phase 7: cross-registry wire format
411
+
412
+ The CLI doesn't run the network gossip layer (that's a separate
413
+ project — see [`docs/BACKLOG.md`](docs/BACKLOG.md)). What it does
414
+ provide is a **stable, versioned JSON wire format** + an
415
+ idempotent merge importer, ready for a future gossip layer to
416
+ serialize over.
417
+
418
+ ```bash
419
+ # Snapshot the local state into one JSON document.
420
+ capcard-bundle export --out snapshot.json --source-id "lab-monogate"
421
+
422
+ # Merge into another registry. Idempotent on primary keys
423
+ # (agent_id, verification_id, attestation_id, (agent_id, session_id)).
424
+ # Local state is authoritative — conflicts are logged and skipped,
425
+ # never overwritten.
426
+ capcard-bundle import snapshot.json
427
+ capcard-bundle import snapshot.json --dry-run # preview only
428
+ capcard-bundle import snapshot.json --json # ImportSummary as JSON
429
+
430
+ # Verify a verification record's stored output digest matches a
431
+ # fresh re-run of the original evidence command. Poor-man's
432
+ # "did this happen" check without ZK proofs.
433
+ capcard-fingerprint --verification-id 6f00d02d4244 --rerun
434
+ # → MATCH (exit 0) | MISMATCH (exit 3) | unknown id (exit 1)
435
+ ```
436
+
437
+ The bundle format is a single JSON object with `bundle_version`,
438
+ `created`, `source_id`, and arrays of agents / verifications /
439
+ attestations / playbooks. After import, every agent's
440
+ `trust_score` is recomputed under the same lock so newly-merged
441
+ verifications and attestations propagate into the cached score.
442
+
443
+ `AgentRecord` carries an optional `pubkey: str | None` field
444
+ (`capcard-register --pubkey "..."`). It's forward-reserved — the
445
+ stdlib-only constraint means the CLI doesn't sign anything yet.
446
+ The field is here so a Phase 7+ network protocol that *does* sign
447
+ can import bundles without a schema migration.
448
+
449
+ ### Phase 6.7: dispute time-decay + iterative trust converge
450
+
451
+ Two hygiene knobs on `capcard-recompute-trust` close the open
452
+ edges from the Phase 6.5 review:
453
+
454
+ ```bash
455
+ # Stale disputes (latest dispute attestation older than N days)
456
+ # are ignored, letting the underlying verdict surface. Useful when
457
+ # nobody followed up on a contested record.
458
+ capcard-recompute-trust --all --dispute-decay-days 7
459
+
460
+ # Iterative converge to a fixed point. Default 2-pass already
461
+ # handles the common case; --iterative is for verifier↔verifier
462
+ # review chains that could oscillate. Stops when the L-infinity
463
+ # per-agent delta drops below --epsilon (default 0.001) or
464
+ # --max-iter (default 10) is hit.
465
+ capcard-recompute-trust --all --use-current-verifier-trust \
466
+ --iterative --epsilon 0.0001
467
+ ```
468
+
469
+ Both flags are off by default so the conservative Phase 6.5
470
+ semantics still apply unless the operator opts in.
471
+
472
+ ### Generate a session-end playbook
473
+
474
+ When the agent finishes a session (or hits a meaningful checkpoint),
475
+ `capcard-playbook generate` writes a structured summary to disk and
476
+ points the agent's `playbook_path` at the new file. The trust delta,
477
+ verification window, and result_summary are computed automatically
478
+ from the verification log since the agent's previous playbook.
479
+
480
+ ```bash
481
+ capcard-playbook generate --agent researcher-001 \
482
+ --result theorem-established \
483
+ --lesson "P(tuple) undefined at depth >= 2" \
484
+ --lesson "OfScientific vs OfNat blocks proofs" \
485
+ --tactic-worked "exact exp_nonneg _||0 <= exp x" \
486
+ --tactic-failed "trivial||0 <= exp x" \
487
+ --artifact "machlib:9589804" \
488
+ --open-question "depth-2 unfold under products?"
489
+ ```
490
+
491
+ `--tactic-worked` / `--tactic-failed` use a literal `||` separator
492
+ between tactic and goal-shape label (the goal label is what the BFS
493
+ engine keys on — keep it consistent across sessions).
494
+
495
+ For bulk input, write the same fields to a JSON file and pass
496
+ `--from-file`. Per-flag values are merged on top:
497
+
498
+ ```bash
499
+ cat > playbook.json <<'JSON'
500
+ {
501
+ "lessons": ["P(tuple) undefined at depth >= 2"],
502
+ "tactics_that_worked": [
503
+ {"tactic": "exact exp_nonneg _", "on_goal": "0 <= exp x"}
504
+ ],
505
+ "tactics_that_failed": [
506
+ {"tactic": "trivial", "on_goal": "0 <= exp x"}
507
+ ],
508
+ "artifacts": ["machlib:9589804"],
509
+ "open_questions": ["depth-2 unfold under products?"]
510
+ }
511
+ JSON
512
+
513
+ capcard-playbook generate --agent researcher-001 \
514
+ --result theorem-established \
515
+ --from-file playbook.json
516
+ ```
517
+
518
+ ### Search prior playbooks (transfer)
519
+
520
+ When a new agent picks up a related task, search the playbook
521
+ archive for what's already been tried — the highest-leverage
522
+ transfer is the list of tactics that *failed*, so future agents
523
+ don't repeat them.
524
+
525
+ ```bash
526
+ # Find related work — score is the count of distinct query terms
527
+ # matched across task / lessons / tactics / open_questions.
528
+ capcard-playbook search --task "Galois dimension proof for chain order"
529
+
530
+ # Filter to one role + tighten the score floor
531
+ capcard-playbook search --task "exp_strict_mono on negative inputs" \
532
+ --role researcher --min-score 2
533
+
534
+ # Machine-readable for piping into another agent's context
535
+ capcard-playbook search --task "..." --json | jq '.[0].playbook'
536
+ ```
537
+
538
+ The transfer mechanism is **explicit, not magic** — the calling
539
+ agent runs the search and decides what to load. (Hidden auto-load
540
+ on `capcard-register` would silently change session behavior, which
541
+ we don't want.)
542
+
543
+ ### Show / inspect
544
+
545
+ ```bash
546
+ capcard-playbook show --agent researcher-001 # newest playbook
547
+ capcard-playbook show --path /path/to/playbook.json # specific file
548
+ ```
549
+
550
+ ### Verification log
551
+
552
+ `capcard-verify` appends one JSONL row per call to
553
+ `~/.capcard/verifications.jsonl` (override via
554
+ `CAPCARD_VERIFY_LOG_PATH`):
555
+
556
+ ```json
557
+ {"agent_id":"researcher-001","evidence":"lake build","exit_code":null,
558
+ "notes":null,"output_digest":"d41d8cd9...","ran":false,
559
+ "result":"proof-closed","runtime_s":null,"task":"Galois dimension proof",
560
+ "theorem":"tanh_monotone","timestamp":"2026-05-04T13:30:00Z"}
561
+ ```
562
+
563
+ The log is append-only and serialized under its own flock
564
+ (`~/.capcard/verifications.jsonl.lock`). The registry stores the
565
+ rolled-up counters; the log keeps the per-event detail Phases 3
566
+ (playbook), 6 (cross-verify), and 7 (ZK fingerprints) need.
567
+
568
+ ## Wire it into a Claude Code session
569
+
570
+ Add to your shell rc (or per-project `.envrc`):
571
+
572
+ ```bash
573
+ # At session start, register and capture the assigned agent_id
574
+ export CAPCARD_AGENT_ID=$(capcard-register \
575
+ --name "$USER on $(hostname -s)" \
576
+ --role researcher \
577
+ --json | python -c 'import json,sys; print(json.load(sys.stdin)["agent_id"])')
578
+
579
+ # Then any time you switch tasks
580
+ capcard-status --agent "$CAPCARD_AGENT_ID" --task "$NEW_TASK_DESC"
581
+
582
+ # When you finish a unit of work, log the outcome with evidence
583
+ capcard-verify --agent "$CAPCARD_AGENT_ID" \
584
+ --evidence "lake build" \
585
+ --result proof-closed \
586
+ --theorem tanh_monotone \
587
+ --run
588
+ ```
589
+
590
+ Once the Command Center's `/agents` page is updated to read the
591
+ extended `blackwell-agent` `/agents` endpoint (which already merges
592
+ the registry with live tmux sessions as of this commit), the
593
+ registered agent appears with its persistent ID, role, and trust
594
+ score alongside the anonymous tmux sessions on the same machine.
595
+
596
+ ## What's coming next
597
+
598
+ Per the [vision doc](https://capcard.ai/whitepaper) :
599
+
600
+ - **Phase 2 (`capcard-verify`)** — **shipped in 0.2.0**: log work
601
+ outcomes with cryptographic evidence (`lake build` passing,
602
+ `pytest 87 passed`, CI red). Bumps `verified_outputs` /
603
+ `failed_outputs` per the result-type contract; appends a row to
604
+ `~/.capcard/verifications.jsonl`.
605
+ - **Phase 3 (`capcard-playbook`)** — **shipped in 0.3.0**: emit a
606
+ structured summary at session end (lessons / tactics_that_worked
607
+ / tactics_that_failed / open_questions); next agent inherits via
608
+ `capcard-playbook search`.
609
+ - **Phase 4 — weighted trust** — **shipped in 0.4.0**: replaces the
610
+ simple `verified / total` ratio with the weight table above.
611
+ `--evidence-type` declares what kind of proof the agent has;
612
+ trust is the weighted average clamped to `[0, 1]`.
613
+ - **Phase 6 — cross-agent attestation** — **shipped in 0.5.0**:
614
+ `--needs-attestation` + `capcard-attest`. Builder = verifier
615
+ rejected; verifier trust modulates the builder's contribution
616
+ with a 0.5 floor.
617
+ - **Phase 6.5 — dispute resolution + trust hygiene** — **shipped
618
+ in 0.6.0**: per-verifier vote collapse, dispute quorum
619
+ resolution (≥2 distinct non-disputers agreeing override the
620
+ dispute), `capcard-attestations` (`--pending` / `--by` /
621
+ `--on`), `capcard-recompute-trust` with optional
622
+ `--use-current-verifier-trust` mode for propagating verifier-
623
+ trust changes back into builder rows.
624
+ - **Phase 5 — RL training-data extraction** — **shipped in 0.7.0**:
625
+ generation/parent_id on AgentRecord, `capcard-export` (shaped-
626
+ reward episodes), `capcard-tactics` (trust-weighted ranking),
627
+ `capcard-stats` (per-agent rollup + evidence histogram). The CLI
628
+ doesn't train models — it makes the archive trainable.
629
+ - **Phase 6.7 — trust hygiene** — **shipped in 1.0.0**: time-decay
630
+ for stuck dispute markers, iterative fixed-point converge in
631
+ `recompute_all_trust`. Both opt-in via flags on
632
+ `capcard-recompute-trust`.
633
+ - **Phase 7 — wire format + fingerprint verify** — **shipped in
634
+ 1.0.0**: `capcard-bundle export/import` (versioned, idempotent
635
+ merge), `capcard-fingerprint --rerun` for spot-checking imported
636
+ output digests, `pubkey: str | None` field on AgentRecord. The
637
+ *actual* network gossip layer + ZK proofs are an independent
638
+ research project — see [`docs/BACKLOG.md`](docs/BACKLOG.md).
639
+
640
+ The capcard-cli pipeline is feature-complete for the whitepaper
641
+ roadmap as of 1.0.0. Future work (signing, gossip, ZK
642
+ fingerprints, network resilience) belongs to the network/protocol
643
+ side of the project, not this CLI.
644
+
645
+ ## Testing
646
+
647
+ ```bash
648
+ pip install -e '.[dev]'
649
+ python -m pytest
650
+ ```
651
+
652
+ 244 tests cover: schema, atomic writes, file locking under 20-process
653
+ contention, idempotency, role-migration discipline, evidence-weight
654
+ math, playbook windowing, attestation state machine + dispute
655
+ quorum resolution + time-decay, two-pass and iterative trust
656
+ converge, lineage walkers, reward shaping (incl. clamps and
657
+ bonus/penalty stacking), tactic ranking across playbooks, episode
658
+ export filters, bundle round-trip + idempotent merge + conflict
659
+ preservation, fingerprint match/mismatch, and CLI surface.
660
+
661
+ ## License
662
+
663
+ Apache-2.0 (matches the rest of the Monogate stack).