@tangle-network/agent-eval 0.125.0 → 0.126.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (82) hide show
  1. package/CHANGELOG.md +62 -35
  2. package/README.md +270 -189
  3. package/dist/analyst/index.d.ts +15 -145
  4. package/dist/analyst/index.js +33 -47
  5. package/dist/analyst/index.js.map +1 -1
  6. package/dist/benchmarks/index.d.ts +45 -162
  7. package/dist/benchmarks/index.js +8 -9
  8. package/dist/campaign/index.d.ts +3674 -5393
  9. package/dist/campaign/index.js +21 -95
  10. package/dist/{chunk-R226UZOI.js → chunk-474LBSOX.js} +2 -2
  11. package/dist/{chunk-HM6V7F3M.js → chunk-FO7HEH76.js} +3 -3
  12. package/dist/chunk-IILEIWGW.js +635 -0
  13. package/dist/chunk-IILEIWGW.js.map +1 -0
  14. package/dist/{chunk-EQUK3RFS.js → chunk-J5SQWP6Y.js} +8 -5
  15. package/dist/chunk-J5SQWP6Y.js.map +1 -0
  16. package/dist/{chunk-W5B3ZGP3.js → chunk-KE2VWPZX.js} +8 -6
  17. package/dist/{chunk-W5B3ZGP3.js.map → chunk-KE2VWPZX.js.map} +1 -1
  18. package/dist/{chunk-DT7OXY3C.js → chunk-LUNF2SEL.js} +538 -851
  19. package/dist/chunk-LUNF2SEL.js.map +1 -0
  20. package/dist/chunk-NGUYT5CI.js +4637 -0
  21. package/dist/chunk-NGUYT5CI.js.map +1 -0
  22. package/dist/{chunk-QFQZ3U3X.js → chunk-OCFJACJU.js} +2 -2
  23. package/dist/{chunk-GID26AN4.js → chunk-P22LJ3Y2.js} +4 -6
  24. package/dist/{chunk-GID26AN4.js.map → chunk-P22LJ3Y2.js.map} +1 -1
  25. package/dist/{chunk-SJT4OBVL.js → chunk-SDPM6554.js} +3 -3
  26. package/dist/{chunk-D5JZ7UDZ.js → chunk-UCLVDLCH.js} +136 -50
  27. package/dist/chunk-UCLVDLCH.js.map +1 -0
  28. package/dist/chunk-VMUENW6F.js +7274 -0
  29. package/dist/chunk-VMUENW6F.js.map +1 -0
  30. package/dist/{chunk-JKDNAOF5.js → chunk-W4L6C2XT.js} +2 -2
  31. package/dist/chunk-WGXIEX7P.js +116 -0
  32. package/dist/chunk-WGXIEX7P.js.map +1 -0
  33. package/dist/{chunk-GRCDRKII.js → chunk-WS3NZZQQ.js} +58 -20
  34. package/dist/chunk-WS3NZZQQ.js.map +1 -0
  35. package/dist/cli.js +3 -3
  36. package/dist/contract/index.d.ts +3220 -3094
  37. package/dist/contract/index.js +173 -42
  38. package/dist/contract/index.js.map +1 -1
  39. package/dist/control.js +2 -3
  40. package/dist/fuzz.d.ts +14 -1
  41. package/dist/fuzz.js +1 -1
  42. package/dist/hosted/index.d.ts +8 -1
  43. package/dist/index.d.ts +71 -687
  44. package/dist/index.js +178 -497
  45. package/dist/index.js.map +1 -1
  46. package/dist/openapi.json +1 -1
  47. package/dist/rl.d.ts +5 -100
  48. package/dist/rl.js +4 -5
  49. package/dist/rl.js.map +1 -1
  50. package/dist/{run-campaign-I3JXKVAK.js → run-campaign-LVFKZCEU.js} +3 -3
  51. package/dist/traces.js +2 -3
  52. package/dist/wire/index.d.ts +14 -1
  53. package/dist/wire/index.js +3 -3
  54. package/docs/campaign-proposers.md +363 -168
  55. package/docs/design/loop-taxonomy.md +142 -190
  56. package/docs/design.md +1 -1
  57. package/docs/distributed-driver.md +8 -11
  58. package/docs/feature-guide.md +20 -19
  59. package/docs/knowledge-readiness.md +2 -5
  60. package/docs/multi-shot-optimization.md +35 -27
  61. package/docs/rollout.md +5 -5
  62. package/package.json +4 -4
  63. package/dist/chunk-A62YMFWA.js +0 -9269
  64. package/dist/chunk-A62YMFWA.js.map +0 -1
  65. package/dist/chunk-A6GT67HT.js +0 -550
  66. package/dist/chunk-A6GT67HT.js.map +0 -1
  67. package/dist/chunk-D5JZ7UDZ.js.map +0 -1
  68. package/dist/chunk-DT7OXY3C.js.map +0 -1
  69. package/dist/chunk-EQUK3RFS.js.map +0 -1
  70. package/dist/chunk-GC4ATIKK.js +0 -317
  71. package/dist/chunk-GC4ATIKK.js.map +0 -1
  72. package/dist/chunk-GRCDRKII.js.map +0 -1
  73. package/dist/chunk-LOW3U7JZ.js +0 -328
  74. package/dist/chunk-LOW3U7JZ.js.map +0 -1
  75. package/dist/chunk-PMITBABE.js +0 -3841
  76. package/dist/chunk-PMITBABE.js.map +0 -1
  77. /package/dist/{chunk-R226UZOI.js.map → chunk-474LBSOX.js.map} +0 -0
  78. /package/dist/{chunk-HM6V7F3M.js.map → chunk-FO7HEH76.js.map} +0 -0
  79. /package/dist/{chunk-QFQZ3U3X.js.map → chunk-OCFJACJU.js.map} +0 -0
  80. /package/dist/{chunk-SJT4OBVL.js.map → chunk-SDPM6554.js.map} +0 -0
  81. /package/dist/{chunk-JKDNAOF5.js.map → chunk-W4L6C2XT.js.map} +0 -0
  82. /package/dist/{run-campaign-I3JXKVAK.js.map → run-campaign-LVFKZCEU.js.map} +0 -0
package/CHANGELOG.md CHANGED
@@ -6,49 +6,76 @@ All notable changes to `@tangle-network/agent-eval` and its sibling `agent-eval-
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [0.126.1] - 2026-07-24 - optimizer lifecycle integrity
10
+
11
+ ### Changed
12
+
13
+ - `compareOptimizationMethods()` uses one caller-visible spend limit across optimizer models, train and selection evaluation, and final scoring.
14
+ - Candidate surfaces are detached before every optimizer and scoring callback.
15
+ - SkillOpt receives complete non-secret model settings automatically while provider credentials remain in the local Node proxy.
16
+
17
+ ### Fixed
18
+
19
+ - Aborting a comparison now stops official optimizer subprocess groups, callback work, and model proxy requests.
20
+ - Resumed final scoring retains the original cost identity and reports cumulative spend instead of resetting cached calls to zero.
21
+ - Model-backed GEPA fails before search when the full upstream dependency set is missing instead of silently returning the baseline after swallowed reflection errors.
22
+
23
+ ## [0.126.0] - 2026-07-24 - official optimizer engines
24
+
25
+ ### Added
26
+
27
+ - `skillOptOptimizationMethod()` delegates skill optimization to Microsoft's official `ReflACTTrainer`.
28
+ - `gepaOptimizationMethod()` delegates sequential, adaptive, best-of, vote, and Omni recipes to the official GEPA package.
29
+ - `engineModules` lets callers register custom engines through GEPA's public engine registry.
30
+ - `DspyJudgeMetric` exposes an `agent-eval-rpc` judge as a native DSPy metric.
31
+ - `externalTextOptimizationMethod()` adapts another optimizer to the same text or component-surface contract without adding a named local algorithm.
32
+ - Optimizer runs record optimizer and bridge package identity, source commits and source-tree hashes, Python runtime, custom engine module hashes, recipe, configuration, evaluation count, tokens, cost, and resume compatibility.
33
+
9
34
  ### Changed
10
35
 
11
- - `selfImprove({ budget: { candidateConcurrency } })` exposes the existing `runOptimization()` control for scoring candidate campaigns in parallel; it remains opt-in and defaults to one candidate campaign at a time.
12
- - `llmPolicyEditProposer()` and `projectPolicyEditHistory()` accept `scenarioOrder: 'input'` when controlled comparisons must preserve first-occurrence caller order; ranked evidence selection remains the default.
13
- - `callLlmJson()` accepts `jsonPayloadMode: 'exact'` when callers must reject fenced, prose-wrapped, or multi-root responses instead of extracting a JSON value.
14
- - `llmPolicyEditProposer({ redactCurrentSurfaceForModel })` can remove credentials and unrelated fields from the current surface sent to the model while applying validated edits to the complete original surface.
15
- - `CostLedger.listPending()` exposes immutable pending paid calls and distinguishes calls that are active, late after cancellation, or interrupted by a prior process so durable workflows can reconcile exact reservations before resuming.
16
- - `traceAnalystProposer()` accepts an opt-in `resolvePriorFindings` callback that forwards canonical prior findings into the existing analyst registry.
17
- - Trace-analysis actors are instructed to emit one executable JavaScript program per turn, report named turn-limit exhaustion, and preserve per-analyst failure details when a proposer produces no findings.
18
- - Keep one live tip per lineage track when another track branches or merges from it, and compare those track tips when building the frontier.
19
- - Pass track identity, operation, vision, ancestry, and proposer choice to candidate generation so independent tracks can pursue distinct strategies.
20
- - Route named tracks to caller-supplied proposers, with the default proposer as fallback, and make heuristic branches inherit their parent track's proposer.
36
+ - `selfImprove()` requires an explicit `OptimizationMethod` or caller-owned `SurfaceProposer`.
37
+ - `selfImprove({ budget: { candidateConcurrency } })` can score independent candidate campaigns concurrently.
38
+ - `callLlmJson()` accepts `jsonPayloadMode: 'exact'` to reject fenced, prose-wrapped, or multi-root responses.
39
+ - `CostLedger.listPending()` distinguishes active, late, and interrupted paid calls so durable runs can reconcile reservations before resuming.
40
+ - GEPA and SkillOpt share one OpenAI-compatible optimizer model configuration with request, output, response-size, timeout, and dollar limits.
41
+ - Model credentials remain in the Node process.
42
+ Official Python libraries receive only a temporary loopback endpoint and credential.
43
+ - `compareOptimizationMethods()` owns pairwise-disjoint train, selection, and final test sets.
44
+ Methods never receive final test rows, and final test scoring starts only after every method completes.
45
+ - Independent methods and candidate campaigns can run concurrently under the same exact cost accounting.
46
+ - Resume files use atomic writes, exclusive run locks, and one content-derived identity shared by Node and Python.
47
+ - `runCampaign()` gives cancelled dispatches five seconds to stop by default.
48
+ `dispatchShutdownTimeoutMs` changes that bound.
49
+ - AppWorld and GSM8K comparisons reject unequal candidate-evaluation limits and record every model, split, limit, rate, actual call count, and cost basis.
50
+ - CI installs GEPA and SkillOpt from exact official source commits for compatibility testing.
51
+ Users install either optimizer separately when they need its bridge.
52
+ DSPy remains a separate optional environment because DSPy 3.2.1 currently requires an older GEPA release.
21
53
 
22
54
  ### Fixed
23
55
 
24
- - `llmPolicyEditProposer()` asks the model only for semantic edit choices, deterministically binds caller-owned policy fields, and requires one exact whole-response JSON value before admission.
25
- - `compareOptimizationMethods()` owns three non-empty, pairwise-disjoint scenario sets.
26
- Methods receive independent copies of train and selection data, every method finishes before final test scoring starts, and final test data is absent from `OptimizationMethodInput`.
27
- Built-in GEPA, SkillOpt, and FAPO methods read the shared baseline, runner, judges, directories, and execution settings from the comparison call instead of duplicating them in each method config.
28
- Final test scoring uses one shared spend limit.
29
- Reused cost ledgers isolate comparison receipts from unrelated calls, and final test scenarios are copied for each measured surface.
30
- Results report optimization, test, and total cost with accounting completeness and reasons for unknown charges.
31
- Per-method results include the paired scenario values used to compute lift.
32
- Cost orders a tied-lift group only when every method in that group has complete accounting.
33
- Bootstrap intervals adjust across all method contrasts, and the default resample count increases when needed to represent the adjusted tails.
34
- `optimizationConcurrency` can run independent methods in parallel.
35
- - `runSkillOpt()` validates every numeric control before scoring, enforces `patchesPerEpoch` for custom proposers, and no longer mutates the caller's `runDir`.
36
- - Include the complete required response shape in `llmPolicyEditProposer` model instructions, and add `LlmClientOptions.jsonSchemaTransport: 'json-object'` for providers that do not implement native JSON Schema enforcement.
37
- - `compositeProposer` restores each member's original labels when replaying history, so stateful members do not repeat candidates whose labels were decorated for provenance.
38
- - `InsightReport.interRater.kappa` now reports quadratic weighted kappa instead of Pearson correlation.
39
- Read `interRater.pearson` for the previous correlation measure; `icc` and `spearman` are now reported separately.
40
- - Compute contextual-bandit doubly robust estimates with separate logged-action and target-policy value terms, expose how many rows use DR versus IPS or the deprecated scalar path, and carry both values through belief-state records.
56
+ - Reserve worst-case optimizer spend before each provider call and settle it from provider-reported cost or complete token usage.
57
+ - Preserve cached input, cache creation, reasoning, and output token classes without double counting.
58
+ - Price cache-read tokens at the configured cache rate when providers omit billed cost.
59
+ - Reject incomplete usage, conflicting token details, hidden provider endpoints, hidden credentials, unsupported streaming, oversized input or output, and process descendants that outlive any bridge exit.
60
+ - Abort active candidate evaluations when an optimizer exits or times out, and bind cached scores to the exact evaluation identity.
61
+ - Require explicit trust and local ownership checks before restoring GEPA pickle state.
62
+ - Hash packaged prompts and other behavior files in optimizer identity, and report failed model attempts separately from successful calls.
63
+ - Reject native Windows optimizer subprocesses rather than claiming process-tree cleanup that cannot be proven; Linux and WSL use verified POSIX process groups.
64
+ - Abort and drain sibling campaign lanes after the first lane failure.
65
+ - Delay campaign results until cancelled dispatches and their paid calls settle, so late receipts cannot change reported cost after return.
66
+ - Preserve both the primary operation error and any subprocess, temporary-directory, callback, or model-proxy cleanup error.
67
+ - Keep optimizer, evaluation, and final test costs separate when a shared cost ledger contains unrelated receipts.
68
+ - `LlmClientOptions.jsonSchemaTransport: 'json-object'` supports providers that do not implement native JSON Schema enforcement.
69
+ - `InsightReport.interRater.kappa` reports quadratic weighted kappa.
70
+ `interRater.pearson` preserves the previous correlation measure, with ICC and Spearman reported separately.
71
+ - Contextual-bandit doubly robust estimates keep logged-action and target-policy value terms separate and report which rows use DR, IPS, or the deprecated scalar path.
41
72
 
42
73
  ### Breaking
43
74
 
44
- - `compareProposers()` is replaced by `compareOptimizationMethods()`.
45
- Rename `proposers` to `methods`, `ProposerEntry` to `OptimizationMethod`, and the built-in `*Entry` factories to `*Method`.
46
- Pass `trainScenarios`, `selectionScenarios`, and `testScenarios`; the ambiguous `holdoutScenarios` option is rejected.
47
- `OptimizerEntryConfig` is replaced by `BuiltinOptimizationMethodConfig`, which contains method-specific settings only.
48
- Move shared method execution settings to `optimizationRunOptions` on the comparison call.
49
- Read `optimizationCost`, `testCost`, and `totalCost`; each includes `totalCostUsd`, `accountingComplete`, and `incompleteReasons`.
50
- - `runSkillOpt({ holdoutScenarios })` fails closed because those rows are adaptively reused.
51
- Pass `selectionScenarios`; selection result fields now use `Selection` instead of `Holdout`, and `lift` is now `selectionLift`.
75
+ - Remove the local GEPA, SkillOpt, ACE, FAPO, HALO, policy-edit, memory-curation, trace-analyst, evolutionary, and composite proposer implementations.
76
+ Use the official GEPA or SkillOpt method, `externalTextOptimizationMethod()`, or pass a caller-owned `SurfaceProposer`.
77
+ - Remove `runSkillOpt()`, built-in optimization method factories, lineage loops, skill-patch parsing, analyst policy editing, and the experimental distillation workflow.
78
+ - Remove stale examples and generated comparison results tied to the deleted local implementations.
52
79
 
53
80
  ## [0.125.0] — 2026-07-24 — Claude Code supervision reader and path-bound policy edits
54
81