@tangle-network/agent-eval 0.94.0 → 0.95.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (140) hide show
  1. package/CHANGELOG.md +32 -0
  2. package/README.md +44 -30
  3. package/dist/adapters/http.d.ts +8 -7
  4. package/dist/adapters/http.js.map +1 -1
  5. package/dist/adapters/langchain.d.ts +3 -2
  6. package/dist/adapters/otel.d.ts +5 -4
  7. package/dist/analyst/index.d.ts +11 -31
  8. package/dist/analyst/index.js +5 -65
  9. package/dist/analyst/index.js.map +1 -1
  10. package/dist/{analyze-runs-B6Ljo_dI.d.ts → analyze-runs-DtT6F_6T.d.ts} +3 -3
  11. package/dist/belief-state/index.d.ts +4 -3
  12. package/dist/benchmarks/index.d.ts +3 -2
  13. package/dist/campaign/index.d.ts +727 -616
  14. package/dist/campaign/index.js +1863 -1316
  15. package/dist/campaign/index.js.map +1 -1
  16. package/dist/{chunk-2K6UUZ7P.js → chunk-2T4EZACH.js} +1 -1
  17. package/dist/chunk-2T4EZACH.js.map +1 -0
  18. package/dist/{chunk-CTBHKLEU.js → chunk-77T4STFI.js} +59 -86
  19. package/dist/chunk-77T4STFI.js.map +1 -0
  20. package/dist/{chunk-EGPMSBEZ.js → chunk-7QTQKIDD.js} +178 -177
  21. package/dist/chunk-7QTQKIDD.js.map +1 -0
  22. package/dist/{chunk-MIFZUPEK.js → chunk-AQ5WQAIV.js} +21 -6
  23. package/dist/chunk-AQ5WQAIV.js.map +1 -0
  24. package/dist/chunk-DJWX3GVS.js +81 -0
  25. package/dist/chunk-DJWX3GVS.js.map +1 -0
  26. package/dist/{chunk-S6OZEZQK.js → chunk-HMA63UEO.js} +37 -9
  27. package/dist/{chunk-S6OZEZQK.js.map → chunk-HMA63UEO.js.map} +1 -1
  28. package/dist/{chunk-TBDR6PAI.js → chunk-IZCEK2HR.js} +2 -2
  29. package/dist/{chunk-SD2YFWQQ.js → chunk-KKWJD5E6.js} +20 -20
  30. package/dist/chunk-KKWJD5E6.js.map +1 -0
  31. package/dist/{chunk-KWRRMR3J.js → chunk-LO6IOIJ2.js} +10 -10
  32. package/dist/chunk-LO6IOIJ2.js.map +1 -0
  33. package/dist/{chunk-E4GH6USR.js → chunk-NZEQVRH5.js} +2 -2
  34. package/dist/chunk-NZEQVRH5.js.map +1 -0
  35. package/dist/{chunk-MPQWFX6Y.js → chunk-PSWWQXHF.js} +13 -88
  36. package/dist/chunk-PSWWQXHF.js.map +1 -0
  37. package/dist/{chunk-Q5LIB7BC.js → chunk-S4SYLDFX.js} +2 -2
  38. package/dist/chunk-S4SYLDFX.js.map +1 -0
  39. package/dist/{chunk-KW53MSA5.js → chunk-X74V6ESX.js} +2 -2
  40. package/dist/{chunk-QMUEXQJS.js → chunk-YBIGNSCZ.js} +81 -4
  41. package/dist/chunk-YBIGNSCZ.js.map +1 -0
  42. package/dist/{chunk-2KNZHH3P.js → chunk-Z6L6YSU6.js} +2 -2
  43. package/dist/{code-agent-session-BO8nCnv3.d.ts → code-agent-session-CPHRCb4-.d.ts} +1 -1
  44. package/dist/contract/index.d.ts +91 -43
  45. package/dist/contract/index.js +127 -17
  46. package/dist/contract/index.js.map +1 -1
  47. package/dist/{control-D6qwHXIR.d.ts → control-Doncu-B_.d.ts} +2 -2
  48. package/dist/control.d.ts +3 -2
  49. package/dist/control.js +2 -2
  50. package/dist/{corpus-B8A4BDR3.d.ts → corpus-D4YW9UoJ.d.ts} +1 -1
  51. package/dist/{default-registry-6dhErQbs.d.ts → default-registry-GyE8X5SP.d.ts} +3 -3
  52. package/dist/diagnose.d.ts +4 -3
  53. package/dist/diagnose.js +1 -1
  54. package/dist/{run-improvement-loop-DBahB8Ax.d.ts → gepa-C1NCIZ9o.d.ts} +117 -130
  55. package/dist/hosted/index.d.ts +5 -4
  56. package/dist/{index-Bx3gZ8xl.d.ts → index-_Y4oNOOb.d.ts} +1 -1
  57. package/dist/index.d.ts +76 -81
  58. package/dist/index.js +66 -31
  59. package/dist/index.js.map +1 -1
  60. package/dist/{insight-report-DWl3z9tl.d.ts → insight-report-BnRjTibG.d.ts} +1 -1
  61. package/dist/{kind-factory-0BhLSI27.d.ts → kind-factory-X3eDYbKn.d.ts} +2 -3
  62. package/dist/matrix/index.d.ts +1 -1
  63. package/dist/meta-eval/index.d.ts +3 -2
  64. package/dist/multishot/index.d.ts +4 -4
  65. package/dist/multishot/index.js.map +1 -1
  66. package/dist/openapi.json +1 -1
  67. package/dist/{pre-registration-mAnCugl9.d.ts → pre-registration-nfUdc9EQ.d.ts} +2 -42
  68. package/dist/{provenance-P-bCL2Fo.d.ts → provenance-CncDq9qE.d.ts} +26 -41
  69. package/dist/{release-report-BEbWmVYj.d.ts → release-report-pidWUMZ2.d.ts} +2 -2
  70. package/dist/reporting.d.ts +5 -4
  71. package/dist/{researcher-B0C2_fVO.d.ts → researcher-Jr8ME1dZ.d.ts} +2 -2
  72. package/dist/rl.d.ts +516 -515
  73. package/dist/rl.js +612 -612
  74. package/dist/rl.js.map +1 -1
  75. package/dist/{rubric-predictive-validity-Cy_W-hWZ.d.ts → rubric-predictive-validity-C2hDKM8Z.d.ts} +1 -1
  76. package/dist/{run-campaign-7WNXMDSN.js → run-campaign-WXY7KI67.js} +2 -2
  77. package/dist/{run-record-e7vj1uZQ.d.ts → run-record-CP2ObebC.d.ts} +14 -18
  78. package/dist/{runtime-trajectory-BDgfGZSr.d.ts → runtime-trajectory-BOUUjI0y.d.ts} +1 -1
  79. package/dist/{semantic-concept-judge-B9MgmBnM.d.ts → semantic-concept-judge-DSBB2Cfp.d.ts} +2 -2
  80. package/dist/{summary-report-BDOFevaT.d.ts → summary-report-CInXwsza.d.ts} +1 -1
  81. package/dist/testing-C21CHsq2.d.ts +20 -0
  82. package/dist/testing.d.ts +1 -0
  83. package/dist/testing.js +8 -0
  84. package/dist/testing.js.map +1 -0
  85. package/dist/traces.d.ts +26 -10
  86. package/dist/traces.js +41 -11
  87. package/dist/{types-Ce17tDlG.d.ts → types-B5x54y6n.d.ts} +1 -1
  88. package/dist/{types-mn5Aqk7x.d.ts → types-BUxNaJ8c.d.ts} +2 -4
  89. package/dist/{types-BU-7W85F.d.ts → types-DQRY8ZT-.d.ts} +60 -58
  90. package/dist/workflow/index.d.ts +5 -4
  91. package/dist/workflow/index.js +1 -1
  92. package/docs/campaign-proposers.md +170 -0
  93. package/docs/concepts.md +8 -4
  94. package/docs/customer-journeys.md +15 -13
  95. package/docs/design/loop-taxonomy.md +34 -66
  96. package/docs/distributed-driver.md +14 -14
  97. package/docs/feature-guide.md +1 -1
  98. package/docs/hosted-ingest-spec.md +2 -3
  99. package/docs/multi-shot-optimization.md +8 -8
  100. package/docs/product-eval-adoption.md +1 -1
  101. package/docs/self-improvement-map.md +33 -29
  102. package/package.json +8 -14
  103. package/dist/chunk-2K6UUZ7P.js.map +0 -1
  104. package/dist/chunk-CTBHKLEU.js.map +0 -1
  105. package/dist/chunk-E4GH6USR.js.map +0 -1
  106. package/dist/chunk-EGPMSBEZ.js.map +0 -1
  107. package/dist/chunk-KWRRMR3J.js.map +0 -1
  108. package/dist/chunk-MIFZUPEK.js.map +0 -1
  109. package/dist/chunk-MPQWFX6Y.js.map +0 -1
  110. package/dist/chunk-Q5LIB7BC.js.map +0 -1
  111. package/dist/chunk-QMUEXQJS.js.map +0 -1
  112. package/dist/chunk-SD2YFWQQ.js.map +0 -1
  113. package/docs/design/external-agent-wedge.md +0 -89
  114. package/docs/design/phase-d-rfc.md +0 -125
  115. package/docs/design/phase4-consumer-migration.md +0 -70
  116. package/docs/design/primitives-integration-spec.md +0 -393
  117. package/docs/design/product-self-improvement-loop.md +0 -146
  118. package/docs/design/self-improvement-engine.md +0 -140
  119. package/docs/design/self-improvement-protocol.md +0 -223
  120. package/docs/design/self-improvement-roadmap.md +0 -106
  121. package/docs/design/substrate-gaps.md +0 -118
  122. package/docs/phase-b-pairing-kit.md +0 -188
  123. package/docs/phase-b-runbook.md +0 -176
  124. package/docs/pilot/README.md +0 -62
  125. package/docs/pilot/customer-checklist.md +0 -90
  126. package/docs/pilot/integration-foreign-stack.md +0 -296
  127. package/docs/pilot/integration-tangle-stack.md +0 -248
  128. package/docs/pilot/one-pager.md +0 -161
  129. package/docs/pilot/sample-insight-report.json +0 -172
  130. package/docs/quickstart-external.md +0 -229
  131. package/docs/research/belief-state-agent-eval-roadmap.md +0 -593
  132. package/docs/research/research-roadmap.md +0 -205
  133. package/docs/specs/driver-honest-spec.md +0 -251
  134. package/docs/specs/hermes-self-improvement-audit.md +0 -93
  135. package/docs/specs/profile-versioning.md +0 -291
  136. package/docs/three-package-architecture.md +0 -168
  137. /package/dist/{chunk-TBDR6PAI.js.map → chunk-IZCEK2HR.js.map} +0 -0
  138. /package/dist/{chunk-KW53MSA5.js.map → chunk-X74V6ESX.js.map} +0 -0
  139. /package/dist/{chunk-2KNZHH3P.js.map → chunk-Z6L6YSU6.js.map} +0 -0
  140. /package/dist/{run-campaign-7WNXMDSN.js.map → run-campaign-WXY7KI67.js.map} +0 -0
@@ -1,291 +0,0 @@
1
- # Profile versioning — closing the offline/online drift gap
2
-
3
- **Status:** Architecture design. Greenfield, replace existing primitives in place. No V2 suffix.
4
- **Owner:** spans agent-eval + agent-runtime + agent-knowledge + sandbox SDK.
5
- **Tracking:** task #98.
6
- **Date:** 2026-05-27.
7
-
8
- ## Architecture in one diagram — symmetric fork
9
-
10
- Neither writer is privileged. Both branches are first-class. When they reconverge, the substrate's job is to BENCHMARK the branches and propose what to keep — not to be the authority.
11
-
12
- ```
13
- AgentProfile lineage
14
- ╱ ╲
15
- ╱ ╲
16
- harness branch substrate branch
17
- (per-turn writes) (selfImprove diff)
18
- ╲ ╱
19
- ╲ ╱
20
- DIVERGENCE EVENT
21
-
22
-
23
- benchmark both branches
24
- against the same held-out
25
-
26
- ┌────────┼────────┐
27
- ▼ ▼ ▼
28
- ship-harness ship-substrate merge
29
-
30
-
31
- inconclusive → expand
32
- corpus / human review
33
- ```
34
-
35
- The substrate becomes a peer, not an owner. The gate verdict names *which* branch won, not just "ship."
36
-
37
- ## What we are fixing
38
-
39
- Two writers, same state, no coordination:
40
-
41
- - **Harness writer** — Hermes-style per-turn `spawn_background_review_thread`, agent-runtime's runLoop, any future in-sandbox self-modification. Online, continuous, fires every turn.
42
- - **Substrate writer** — `selfImprove()` running offline against a frozen snapshot, producing a winner with held-out gate confidence. Batch, fires per campaign.
43
-
44
- Failure modes today:
45
-
46
- 1. **Lost update.** Substrate ships a winner. Harness's per-turn updates since baseline evaporate.
47
- 2. **Stale eval.** Substrate's lift CI is `winner vs P₀`. Production is at `P_h`. The CI says nothing about `winner vs P_h`.
48
- 3. **Gate becomes a lie.** `gateDecision: ship` against `P₀` looks legitimate. Consumer ships. Regresses against `P_h`. Detection fails because metrics moved too.
49
-
50
- ## The minimum design
51
-
52
- Single concept, single operation, content-addressable.
53
-
54
- ### `AgentProfile` is a versioned, content-addressable object
55
-
56
- ```typescript
57
- // src/profile/types.ts
58
-
59
- export interface AgentProfileVersion {
60
- /** Content-hash of the materialised profile state. */
61
- hash: string
62
- /** Parent in the lineage, null for the genesis profile. */
63
- parentHash: string | null
64
- /** Who wrote this version. */
65
- source: 'harness' | 'substrate' | 'human'
66
- /** When. */
67
- timestamp: number
68
- /** Human-readable label, optional. */
69
- label?: string
70
- }
71
-
72
- export type ProfileDiff =
73
- | { kind: 'patch'; edits: ProfileEdit[] }
74
- | { kind: 'replace'; content: MutableSurface }
75
-
76
- export interface ProfileEdit {
77
- /** Which surface inside the profile this edit targets. */
78
- surface: 'systemPrompt' | 'skill' | 'tool' | 'mcp' | 'subagent' | 'modelByRole'
79
- /** Surface-scoped identifier — skillName, toolName, mcpId, subagentId, role. */
80
- surfaceId?: string
81
- op: 'append' | 'insert_after' | 'replace' | 'delete'
82
- target?: string
83
- content: string
84
- /** Support count from multi-trial evidence. */
85
- supportCount?: number
86
- /** Source classification for the merge/rank stage. */
87
- sourceType?: 'failure' | 'success'
88
- }
89
- ```
90
-
91
- That's the whole substrate type surface. Two types. No interface explosion.
92
-
93
- ### `RunRecord` carries the version it was captured at
94
-
95
- Replace the existing `commitSha` / `promptHash` / `configHash` triple with a single canonical hash. Greenfield, no compat shim:
96
-
97
- ```typescript
98
- // src/run-record.ts — IN-PLACE replacement
99
- export interface RunRecord {
100
- // ... existing fields ...
101
- /** Content-hash of the AgentProfileVersion that produced this run. */
102
- agentProfileHash: string
103
- }
104
- ```
105
-
106
- `commitSha`, `promptHash`, `configHash` become *inputs* to `hashProfile()`, not separate fields.
107
-
108
- ### `selfImprove()` returns a diff, and the gate becomes 4-way
109
-
110
- Replace the current return shape. Greenfield, in place:
111
-
112
- ```typescript
113
- // src/contract/self-improve.ts — IN-PLACE replacement
114
- export interface SelfImproveResult {
115
- /** What we measured against. */
116
- baselineHash: string
117
- /** What we recommend applying. */
118
- diff: ProfileDiff
119
- /** Hash of `applyDiff(baseline, diff)` — verifiable by consumer. */
120
- winningHash: string
121
- /** Statistical evidence — paired bootstrap CI vs baseline. */
122
- lift: LiftInsight
123
- /** Substrate verdict — see DriftGateDecision below. */
124
- gateDecision: DriftGateDecision
125
- insight: InsightReport
126
- }
127
-
128
- export type DriftGateDecision =
129
- | { kind: 'ship-substrate'; reason: string; vs?: 'baseline' | 'harness-live' }
130
- | { kind: 'ship-harness'; reason: string }
131
- | { kind: 'merge'; mergedDiff: ProfileDiff; reason: string }
132
- | { kind: 'inconclusive'; reason: string }
133
- ```
134
-
135
- When the substrate runs WITHOUT `driftPolicy: benchmark-branches`, only `ship-substrate` / `inconclusive` (or the equivalent `hold` framing) are possible. When `benchmark-branches` is on, all four kinds may surface.
136
-
137
- The substrate is now explicit: *"this diff is statistically valid against `baselineHash`. Whether to apply it to your live state is your call — and we'll tell you what we found when we compared branches."*
138
-
139
- ### The opt-in drift policy
140
-
141
- ```typescript
142
- selfImprove({
143
- // ... existing
144
- driftPolicy?:
145
- | { kind: 'ignore' } // default — assume single-writer
146
- | { kind: 'reject-on-drift' } // cheap safety mode
147
- | { kind: 'benchmark-branches'; benchmarkBudget: { generations, populationSize } }
148
- })
149
- ```
150
-
151
- - **`ignore`** is the default. Same as today. Zero overhead for consumers whose sandbox harness doesn't self-modify.
152
- - **`reject-on-drift`** is the cheap safety mode. Substrate notices `currentHash != baselineHash` at apply time and refuses to ship. Tells the consumer "your profile drifted; re-run selfImprove against current state."
153
- - **`benchmark-branches`** is the full thing — only used when the harness DOES self-modify (Hermes per-turn, Claude Code with skill creation, Codex with user-prompted skill edits, agent-builder RL bridge, any future autonomous improvement loop). Costs an extra mini-campaign. Returns the 4-way `DriftGateDecision`.
154
-
155
- ### Generalises past Hermes
156
-
157
- Any in-sandbox profile mutation appends to the same profile log, regardless of trigger:
158
-
159
- - Hermes-style autonomous (per-turn `background_review` fork)
160
- - Claude/Codex user-prompted ("hey, create a skill for X")
161
- - agent-runtime's runLoop self-modifying its prompt addendum
162
- - RL-style policy parameter updates
163
- - Manual user edits via `skill_manage` commands
164
-
165
- The substrate doesn't care WHY the harness wrote. It just sees: live profile is at hash X, my baseline was Y. Same merge protocol applies.
166
-
167
- ### Conflict resolution — the four cases
168
-
169
- For the `benchmark-branches` policy, the substrate handles four cases:
170
-
171
- 1. **No conflict.** Edits target different surfaces (substrate edited `systemPrompt`, harness wrote a new `skill/X.md`). Auto-merge into a combined candidate, benchmark merged vs each branch.
172
-
173
- 2. **Orthogonal edits to the same surface.** Both touched `systemPrompt` but different H2 sections (subsumed by `GepaDriverConstraints.preserveSections`). Auto-merge by union of edits, benchmark.
174
-
175
- 3. **Semantic duplication.** Substrate proposed a new skill `summarize-pr`; harness already created `pr-summarizer` (similar purpose, different file). Substrate runs a similarity-detection step: embed both, threshold cosine similarity, surface as a "duplicate-likely" finding. Resolution: head-to-head benchmark with both → keep the winner → archive the loser.
176
-
177
- 4. **Direct same-region conflict.** Both edited the same paragraph. Three resolution paths the substrate offers:
178
- - **Head-to-head**: run both branches, pick the winner.
179
- - **LLM-mediated merge**: prompt an LLM with both candidate edits + the held-out failure trials, ask for a synthesis that addresses both. Benchmark the synthesis.
180
- - **Human review**: surface the diff with `requires-resolution: true` and stop.
181
-
182
- ### Sandbox-side merge protocol
183
-
184
- ```typescript
185
- // agent-runtime exports:
186
- export async function getCurrentProfileVersion(): Promise<AgentProfileVersion>
187
- export async function applyDiff(diff: ProfileDiff): Promise<ApplyResult>
188
-
189
- export type ApplyResult =
190
- | { ok: true; newHash: string }
191
- | { ok: false; reason: 'conflict'; ancestor: string; ours: string; theirs: string }
192
- | { ok: false; reason: 'stale-baseline'; expected: string; actual: string }
193
- ```
194
-
195
- Sandbox keeps an append-only profile log at `~/.tangle/profile-log.jsonl`. Every harness write appends an entry. Every substrate-proposed apply appends or returns conflict.
196
-
197
- ### The merge algorithm (3-way, surface-scoped)
198
-
199
- When substrate proposes `diff(baselineHash → winningHash)` but live state is at `currentHash != baselineHash`:
200
-
201
- 1. **Walk the lineage** — find common ancestor of `baselineHash` and `currentHash`. If `baselineHash` IS an ancestor of `currentHash`, we have a clean rebase target.
202
- 2. **Per-surface 3-way merge** — for each `ProfileEdit` in the diff:
203
- - If the targeted surface (skillName, toolName, etc.) hasn't been touched in `currentHash` lineage since `baselineHash` → apply.
204
- - If touched but the textual edit is on a different region → apply (no conflict).
205
- - If touched on the same region → return `conflict` with ancestor/ours/theirs for the human or substrate to resolve.
206
- 3. **Re-eval recommendation** — if non-trivial conflicts, recommend `selfImprove()` re-run against `currentHash` rather than blind merge.
207
-
208
- The consumer chooses: rebase + re-eval (statistically clean), force merge (skip re-eval, ship-at-own-risk), or reject (substrate's proposal is too stale).
209
-
210
- ## How this changes the substrate flow
211
-
212
- ```
213
- Today:
214
- ingest_baseline_P0 → eval → winner W → consumer ships W (regardless of drift)
215
-
216
- Tomorrow:
217
- ingest_baseline_hashed → eval → {baselineHash, diff, winningHash, lift, gate}
218
-
219
- sandbox.applyDiff(diff) → ok | conflict | stale-baseline
220
-
221
- if stale-baseline: substrate re-eval against currentHash
222
- if conflict: substrate proposes targeted resolution OR human reviews
223
- if ok: profile log gets a new entry, substrate notified
224
- ```
225
-
226
- ## What changes per package
227
-
228
- | Package | Files | Change |
229
- |---|---|---|
230
- | **agent-eval** | `src/profile/types.ts` (new) | `AgentProfileVersion`, `ProfileDiff`, `ProfileEdit` |
231
- | | `src/profile/hash.ts` (new) | `hashProfile()` — content-hash of the materialised state |
232
- | | `src/profile/diff.ts` (new) | `diffProfiles(a, b)`, `applyDiff(profile, diff)`, `threeWayMerge(ancestor, ours, theirs)` |
233
- | | `src/run-record.ts` | REPLACE `commitSha`/`promptHash`/`configHash` triple with `agentProfileHash` (greenfield) |
234
- | | `src/contract/self-improve.ts` | REPLACE `SelfImproveResult` to return `{baselineHash, diff, winningHash, lift, gateDecision, insight}` |
235
- | | `src/contract/analyze-runs.ts` | Add `agentProfileLineage` section to `InsightReport` — what versions ran, drift detected |
236
- | **agent-runtime** | `src/profile/log.ts` (new) | Append-only `~/.tangle/profile-log.jsonl`. `appendVersion()`, `readLineage()`, `findCommonAncestor()` |
237
- | | `src/profile/api.ts` (new) | `getCurrentProfileVersion()`, `applyDiff()` |
238
- | | `src/loops/run-loop.ts` | Every harness-side write to skills/memory/prompt-addendum appends to profile log |
239
- | **agent-knowledge** | `src/skills/version.ts` (new) | Skills become independently versioned objects; profile references them by `skillSetHash` |
240
- | **sandbox** | `src/agent-profile.ts` | Expose `getCurrentProfileVersion()` over the SDK |
241
-
242
- ## What the gate semantics become
243
-
244
- `defaultProductionGate` today: "is the candidate statistically better than the baseline?"
245
-
246
- `defaultProductionGate` tomorrow: same question, scoped to the baseline. The consumer (sandbox / human / hosted-tier) decides whether to apply, given the answer + the current live state.
247
-
248
- We do NOT downgrade our paired-bootstrap CI. That's our edge over SkillOpt and Hermes. We just stop pretending the ship verdict is a deployment decision — it's a measurement.
249
-
250
- ## The forcing function (task C from the audit)
251
-
252
- Before we commit weeks to this implementation, set up the empirical case:
253
-
254
- 1. Run Hermes on top of our sandbox.
255
- 2. Hermes' per-turn loop mutates skills.
256
- 3. Run `selfImprove()` against the baseline at sandbox boot.
257
- 4. Observe `gateDecision: ship` produce a winner that, when applied to the now-drifted live state, regresses.
258
- 5. Capture the actual lift CI gap between `winner vs baseline` and `winner vs live`.
259
-
260
- If that gap is small (< MDE), profile-versioning is over-engineering. If it's large, this work is critical. We should know the number, not the intuition.
261
-
262
- ## Phasing
263
-
264
- ### Phase 0 — forcing function (1 week)
265
- Hermes-on-sandbox drift experiment. Real numbers on the gap. Either proves this work is needed or kills it.
266
-
267
- ### Phase 1 — types + hashing (3 days)
268
- `AgentProfileVersion`, `ProfileDiff`, `ProfileEdit`. `hashProfile()`. `diffProfiles()`. `applyDiff()`. Pure functions, fully tested, no integration yet.
269
-
270
- ### Phase 2 — substrate-side rewire (5 days)
271
- Replace `RunRecord` triple with `agentProfileHash`. Replace `SelfImproveResult` shape. Update `analyzeRuns` to detect lineage drift. Update tests + all 6 consumer products.
272
-
273
- ### Phase 3 — sandbox + runtime (1 week)
274
- Profile log primitive in agent-runtime. `getCurrentProfileVersion()` + `applyDiff()` API. Sandbox SDK surface. Three-way merge for surface-scoped edits.
275
-
276
- ### Phase 4 — agent-knowledge skill versioning (3 days)
277
- Skills become independently versioned. `skillSetHash` referenced from profile.
278
-
279
- ### Phase 5 — Hermes adapter (3 days)
280
- Bridge: Hermes' `~/.hermes/skills/` write events → our profile log via a runtime hook.
281
-
282
- Total: ~3 weeks of focused work. Phase 0 in this session if Drew greenlights.
283
-
284
- ## Source pointers
285
-
286
- - Task: #98
287
- - Related audit: `docs/specs/hermes-self-improvement-audit.md`
288
- - Related spec: `docs/specs/driver-honest-spec.md`
289
- - Current pre-versioning `RunRecord`: `src/run-record.ts`
290
- - Current pre-versioning `SelfImproveResult`: `src/contract/self-improve.ts`
291
- - Current gate: `src/campaign/gates/default-production-gate.ts`
@@ -1,168 +0,0 @@
1
- # Three-package architecture: agent-eval × agent-knowledge × agent-runtime
2
-
3
- The Tangle agent stack splits responsibilities across three TypeScript
4
- packages with explicit, narrow contracts. This doc is the reference for how
5
- they fit together — what each owns, what each consumes from the others, and
6
- the canonical data shapes that move between them.
7
-
8
- ## Why three packages
9
-
10
- Each one has a single, defensible job. Combining them was a real temptation
11
- (less version drift, fewer registries) and we said no on purpose:
12
-
13
- - **`@tangle-network/agent-eval`** owns measurement, optimization, and the
14
- RL bridge. It has no opinion about *what* the agent does or *how* it runs;
15
- it has strong opinions about whether the answer is good and how to make it
16
- better.
17
- - **`@tangle-network/agent-knowledge`** owns the data side: source-grounded
18
- knowledge graphs, source citations, eval-gated knowledge growth, knowledge
19
- readiness scoring. It is domain-agnostic — legal, tax, coding, research
20
- workflows define their own policies on top of it.
21
- - **`@tangle-network/agent-runtime`** owns the *execution* side: the task
22
- lifecycle, knowledge-readiness gating, control-loop orchestration,
23
- streaming session kernels. It does not own domain policy, models, tools,
24
- or UI; it standardizes the lifecycle and delegates domain behavior to
25
- adapters.
26
-
27
- Each package can be reasoned about independently. Each can be replaced
28
- without rewriting the others.
29
-
30
- ## The data interchange — `RunRecord`, `Scenario`, `KnowledgeBundle`
31
-
32
- These three types travel between the packages and tie the architecture
33
- together.
34
-
35
- ### `RunRecord` (owned by agent-eval)
36
-
37
- Every measurable thing — a campaign cell, an optimization trial, a
38
- production rollout, a deployment outcome — projects to a `RunRecord`. It
39
- carries identity (`runId`, `experimentId`, `candidateId`, `seed`,
40
- `scenarioId`), provenance (`commitSha`, `model`, `promptHash`, `configHash`),
41
- cost (`costUsd`, `tokenUsage`), and the outcome (per-split scores +
42
- free-form `raw` metric bag).
43
-
44
- agent-knowledge consumes `RunRecord[]` for release reporting and
45
- optimization analysis. agent-runtime exposes hooks for projecting its own
46
- task results into `RunRecord` shape. Every consumer of agent-eval's
47
- campaign / RL primitives produces `RunRecord[]`.
48
-
49
- ### `Scenario` (currently each owner defines its own)
50
-
51
- agent-eval's `runEvalCampaign` takes
52
- `{ scenarioId: string; tags?: Record<string,string> }`. agent-knowledge
53
- defines richer scenario types for knowledge-base optimization. agent-runtime
54
- takes `TaskSpec` which is one task at a time, not a scenario set.
55
-
56
- This is a known minor friction; not load-bearing yet. When it becomes one,
57
- `Scenario` will get promoted to a shared interface.
58
-
59
- ### `KnowledgeBundle` (owned by agent-knowledge)
60
-
61
- agent-knowledge produces `KnowledgeBundle` (a versioned graph of source
62
- citations + generated content) and `KnowledgeReadinessReport` (gap
63
- analysis). agent-eval's `KnowledgeRequirement` / `KnowledgeBundle` types
64
- are imported from agent-eval into agent-knowledge — agent-knowledge
65
- **adapts** its richer types to agent-eval's wire types, not the other way
66
- around. The wire types are the contract; the rich types are agent-knowledge's
67
- internal model.
68
-
69
- ## Dependency direction
70
-
71
- ```
72
- ┌────────────────────┐
73
- │ agent-runtime │
74
- │ (executor) │
75
- └─────────┬──────────┘
76
-
77
- ▼ imports
78
- ┌────────────────────┐
79
- │ agent-eval │
80
- │ (measurement) │
81
- └────────────────────┘
82
-
83
- │ imports
84
- ┌─────────┴──────────┐
85
- │ agent-knowledge │
86
- │ (data side) │
87
- └────────────────────┘
88
- ```
89
-
90
- **Both** agent-runtime and agent-knowledge import agent-eval. agent-eval
91
- imports neither. This is deliberate: agent-eval is the leaf — its API is
92
- the bottleneck, so its surface stays narrow and stable.
93
-
94
- ## What each package contributes to the auto-research loop
95
-
96
- ```
97
- ┌────────────────────┐ ┌────────────────────┐
98
- │ agent-knowledge │ ────► │ agent-eval │
99
- │ │ │ │
100
- │ - scenario sets │ │ - runEvalCampaign │
101
- │ - knowledge bundle │ │ - capture integrity│
102
- │ - readiness gates │ │ - researchReport │
103
- │ - source citations │ │ - replayCampaign │
104
- │ │ │ - sequential │
105
- │ produces: │ │ - RL bridge │
106
- │ KnowledgeBundle │ │ - preferences │
107
- │ Scenario │ │ - off-policy │
108
- └────────────────────┘ │ - tournament │
109
- │ │
110
- │ produces: │
111
- │ RunRecord[] │
112
- │ PreferenceTriple │
113
- │ etc. │
114
- └─────────▲──────────┘
115
-
116
- ┌────────────────────┐ │
117
- │ agent-runtime │ ──────────────────┘
118
- │ │
119
- │ - runAgentTask │
120
- │ - runAgentControl │
121
- │ - readiness gating │
122
- │ - SSE / sessions │
123
- │ │
124
- │ produces: │
125
- │ ControlRunResult │
126
- │ SSE events │
127
- └────────────────────┘
128
- ```
129
-
130
- agent-knowledge brings the *what* (scenarios, knowledge, source data).
131
- agent-runtime brings the *how to run it once* (task lifecycle, control
132
- loop). agent-eval brings the *measurement and improvement* (campaign,
133
- report, RL bridge).
134
-
135
- ## Cross-package contracts
136
-
137
- | From → To | Type | What it carries |
138
- |---|---|---|
139
- | agent-knowledge → agent-eval | `RunRecord` | (consumed via `runImprovementLoop` for knowledge-base optimization) |
140
- | agent-knowledge → agent-eval | `KnowledgeReadinessReport`, `KnowledgeBundle`, `KnowledgeRequirement` | (re-exported from agent-eval; agent-knowledge populates) |
141
- | agent-knowledge → agent-eval | `ControlRuntimeConfig<KnowledgeBaseCandidate>` | (knowledge research adapter) |
142
- | agent-runtime → agent-eval | `runAgentControlLoop`, `scoreKnowledgeReadiness`, `blockingKnowledgeEval` | (consumed; agent-runtime calls these in its task lifecycle) |
143
- | agent-runtime → agent-eval | `RunRecord`, `TraceStore`, `ControlRunResult`, `ControlStep` | (re-exported types; agent-runtime adapters projects into these) |
144
- | agent-eval ↘ neither package | (no upstream imports) | |
145
-
146
- ## Known gaps in the contracts
147
-
148
- 1. **Shared `Scenario` interface.** Each package has its own scenario
149
- shape. agent-eval will promote a minimal `Scenario` to shared use when
150
- the second consumer needs it.
151
- 2. **agent-knowledge and agent-runtime pin older agent-eval minors.**
152
- Until both bump to current, `RunRecord`'s `scenarioId` field won't be
153
- populated by their existing run records and `RawProviderSink`
154
- integration is per-consumer rather than automatic.
155
- 3. **No first-class trace-analyst hook in agent-runtime.** agent-runtime's
156
- `runAgentTask` emits traces but doesn't auto-execute the trace analyst
157
- on completion the way `runEvalCampaign` does. A `onRunComplete` hook
158
- on agent-runtime would close this.
159
-
160
- ## Versioning policy
161
-
162
- Each package versions independently. The minor-version axis carries
163
- breaking changes; agent-eval's minor versions are tied to major
164
- methodological shifts.
165
-
166
- When agent-eval ships a minor, agent-knowledge and agent-runtime get a
167
- follow-up PR to consume the new surface. The follow-up is tracked as a
168
- deliberate change, not a passive caret pickup.