@tangle-network/agent-app 0.47.25 → 0.48.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/eval-architect/SKILL.md +34 -31
- package/.claude/skills/improve-conductor/SKILL.md +46 -41
- package/.claude/skills/measurement-validation/SKILL.md +40 -34
- package/README.md +7 -0
- package/dist/eval-campaign/index.d.ts +4 -4
- package/dist/eval-campaign/index.js +3 -2
- package/dist/eval-campaign/index.js.map +1 -1
- package/dist/eval-campaign/trust-gate.d.ts +10 -33
- package/package.json +15 -15
|
@@ -1,44 +1,47 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: eval-architect
|
|
3
|
-
description: Build
|
|
3
|
+
description: Build or repair evaluations that score the agent's actual deliverable through the production path, with independent controls and visible missing evidence.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
# Eval
|
|
6
|
+
# Eval architect
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
Build the measurement around the product's required outcome and actual execution path.
|
|
9
|
+
Reuse maintained Eval contracts and existing product checks before creating another scorer.
|
|
9
10
|
|
|
10
|
-
|
|
11
|
+
## Locate the deliverable
|
|
11
12
|
|
|
12
|
-
|
|
13
|
+
1. Inspect real runs to find the output channel: replies, validated tool calls, persisted artifacts, application state, or rendered UI.
|
|
14
|
+
Score the channel that carries the required outcome.
|
|
15
|
+
2. Define the completion boundary for the task.
|
|
16
|
+
For work that accumulates across turns, evaluate the completed artifact and retain intermediate evidence needed to explain failures.
|
|
17
|
+
3. Trace every consumer of the score, including completion checks, optimization selection, and release decisions.
|
|
18
|
+
When the output channel changes, update every affected consumer.
|
|
13
19
|
|
|
14
|
-
|
|
20
|
+
## Build the checks
|
|
15
21
|
|
|
16
|
-
|
|
22
|
+
1. Map each requirement to observable evidence and an explicit failure condition.
|
|
23
|
+
Use answer keys when available; otherwise use independently justified constraints, executable checks, or calibrated judgment.
|
|
24
|
+
Keep unsupported requirements and missing evidence visible.
|
|
25
|
+
2. Establish a simple baseline through the same entrypoint as the candidate.
|
|
26
|
+
Investigate surprising scores instead of assuming either the scorer or the agent caused them.
|
|
27
|
+
3. Separate training, candidate selection, and final comparison evidence where the improvement claim requires those partitions.
|
|
28
|
+
Preserve scenario identities and shared source-unit mappings across baseline and candidate runs.
|
|
29
|
+
4. Define critical failure checks separately from aggregate quality.
|
|
30
|
+
A favorable composite must not erase a failure that violates the product's requirements.
|
|
31
|
+
5. Preserve scorer identity, actual cost receipts, execution failures, and diagnostic artifacts.
|
|
32
|
+
Change `judgeVersion` when an ensemble scorer's configuration changes.
|
|
17
33
|
|
|
18
|
-
##
|
|
34
|
+
## Prove the measurement
|
|
19
35
|
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
36
|
+
Run known positive and negative examples through the complete scoring path.
|
|
37
|
+
Perturb a real deliverable so required behavior improves or regresses, then check that the score detects each change.
|
|
38
|
+
Confirm that absent output and evaluator failure remain distinguishable from measured low quality.
|
|
39
|
+
Report case coverage, detectable failures, uncertainty, and any requirements the evaluation cannot assess.
|
|
40
|
+
A training gain without a final-comparison gain needs diagnosis; it does not identify the cause by itself.
|
|
24
41
|
|
|
25
|
-
##
|
|
42
|
+
## Then consider
|
|
26
43
|
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
## Self-test (prove the metric works before trusting it)
|
|
33
|
-
|
|
34
|
-
- **Baseline sanity:** run it. Is the score non-zero and plausible for a competent agent? A near-zero baseline usually means you're scoring the wrong channel, not that the agent is terrible.
|
|
35
|
-
- **The mutation test (the one that catches the empty-string bug):** hand-edit the produced artifact to be *obviously better* and *obviously worse*. Does the score move in the right direction and magnitude? A metric that doesn't move under obvious changes is measuring the wrong thing.
|
|
36
|
-
- **Audit EVERY scoring surface together.** Completion, quality, and the optimizer's own scorer all read *something*. When the deliverable's channel moves, all of them that read the old channel silently zero. (Session: completion + quality were fixed; the optimizer's own scorer was missed and only found by tracing. Three surfaces — enumerate them, don't assume one.)
|
|
37
|
-
|
|
38
|
-
## Evolves-by
|
|
39
|
-
|
|
40
|
-
When a later optimization shows lift on *training* but none on *held-out*, your eval was overfittable or gameable — add the gap it missed as a new judgment rule. The architect's judgment surface is itself optimized by the meta-eval *"did evals built this way yield real held-out lift, no critical regression?"* See `skill-evolution`.
|
|
41
|
-
|
|
42
|
-
## Fleet as dogfood
|
|
43
|
-
|
|
44
|
-
legal / tax / gtm / creative / insurance each put their deliverable in a *different* channel — filings, forms, published copy, rendered artifacts, routed proposals. The skill is general precisely because it forces you to *locate* the channel for the product in front of you rather than hardcode "the reply text."
|
|
44
|
+
| Condition | Skill |
|
|
45
|
+
|---|---|
|
|
46
|
+
| The evaluation path executes and needs calibration or comparison checks | `measurement-validation` with the baseline and control results |
|
|
47
|
+
| The validated measurement supports a candidate search | `surface-evolution` with the target surface, evidence, and resource limits |
|
|
@@ -1,45 +1,50 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: improve-conductor
|
|
3
|
-
description:
|
|
3
|
+
description: Drive a requested optimization from its target and resource limits through measured candidate selection, evidence review, and the product's promotion policy.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
# Improve
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
2.
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
##
|
|
44
|
-
|
|
45
|
-
|
|
6
|
+
# Improve conductor
|
|
7
|
+
|
|
8
|
+
Own the requested outcome, actual spend, and the decision supported by the evidence.
|
|
9
|
+
Distinguish building an evaluation, searching for candidates, and proving an improvement.
|
|
10
|
+
|
|
11
|
+
## Establish the work
|
|
12
|
+
|
|
13
|
+
1. Recover the target, required outcome, existing evaluation, and authorization from the user's request and product context.
|
|
14
|
+
Ask only for consequential information that the available evidence cannot resolve.
|
|
15
|
+
2. Inspect the current baseline and the system's failure cases.
|
|
16
|
+
Choose a surface change, capability change, or architecture experiment according to the observed limitation.
|
|
17
|
+
3. Check that the evaluation can detect the required behavior through the production entrypoint.
|
|
18
|
+
If it cannot, build that measurement before making improvement claims.
|
|
19
|
+
Record measurement work as measurement work, including its actual cost.
|
|
20
|
+
4. Set resource limits and a stopping rule for the selected experiment.
|
|
21
|
+
Estimate cost from the actual execution path and retain uncertainty about additional calls, retries, and candidate evaluations.
|
|
22
|
+
More spend does not guarantee a useful candidate or a conclusive result.
|
|
23
|
+
|
|
24
|
+
## Run and decide
|
|
25
|
+
|
|
26
|
+
1. Use the maintained optimization method and execution path already available to the product.
|
|
27
|
+
Preserve training, selection, and final-comparison boundaries along with the registered observation units.
|
|
28
|
+
2. Retain candidate artifacts, scorer identity, paired raw evidence, failures, and complete attempt costs.
|
|
29
|
+
Missing usage remains an explicit accounting gap.
|
|
30
|
+
3. Read the shared deciding statistic and the producer's actual verdict.
|
|
31
|
+
Distinguish missing evidence, a measured failure, an inconclusive comparison, and a result that meets the release policy.
|
|
32
|
+
4. Investigate surprising gains and null results with controls that isolate the suspected mechanism.
|
|
33
|
+
Use a footprint control when the claim concerns content versus added context; it is not a universal release prerequisite.
|
|
34
|
+
5. Promote only through the product's authorized decision path after its required checks pass.
|
|
35
|
+
A promising exploratory result can justify another scoped experiment without establishing an improvement.
|
|
36
|
+
|
|
37
|
+
## Report
|
|
38
|
+
|
|
39
|
+
State what changed, the baseline comparison, deciding interval, independent-unit count, actual costs, and the verdict's reasons.
|
|
40
|
+
Explain whether the run stopped because it reached its registered criterion, exhausted its budget, or could not capture valid evidence.
|
|
41
|
+
A proposed follow-up must state which uncertainty it could resolve; extra budget alone does not promise confirmation.
|
|
42
|
+
|
|
43
|
+
## Then consider
|
|
44
|
+
|
|
45
|
+
| Condition | Skill |
|
|
46
|
+
|---|---|
|
|
47
|
+
| The product has no usable evaluation path | `eval-bootstrap` with the observed deliverable and missing checks |
|
|
48
|
+
| A scorer or output-channel defect prevents assessment | `eval-architect` with the failed case and execution evidence |
|
|
49
|
+
| Existing measurements need calibration or comparison review | `measurement-validation` with the baseline and retained results |
|
|
50
|
+
| The supported next experiment changes an existing surface | `surface-evolution` with the target, acceptance criteria, and resource limits |
|
|
@@ -1,38 +1,44 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: measurement-validation
|
|
3
|
-
description:
|
|
3
|
+
description: Check whether an evaluation supports an optimization or release decision through calibrated scoring, independent comparison units, and complete paired evidence.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
# Measurement
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
2.
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
6
|
+
# Measurement validation
|
|
7
|
+
|
|
8
|
+
Check the measurement through the product's actual evaluation path before interpreting an optimization result.
|
|
9
|
+
Separate exploratory evidence from evidence that satisfies a release policy.
|
|
10
|
+
|
|
11
|
+
## Establish the measurement
|
|
12
|
+
|
|
13
|
+
1. State the user outcome, target population, independent comparison unit, and smallest useful effect.
|
|
14
|
+
Record the scoring revision, selection procedure, resource limits, and stopping rule before comparing candidates.
|
|
15
|
+
2. Run a simple baseline and independent positive and negative controls through the scorer.
|
|
16
|
+
Use `auditEvaluator` from `@tangle-network/agent-eval/meta-eval` when the scorer needs an accuracy audit.
|
|
17
|
+
Report false acceptances, false rejections, and missing observations against the product's requirements.
|
|
18
|
+
3. Check repeatability and choose enough independent units to resolve the intended effect.
|
|
19
|
+
Use `powerPreflight` to guide sampling and budget decisions.
|
|
20
|
+
An underpowered result can guide exploration; it cannot certify an improvement.
|
|
21
|
+
4. Keep final comparison cases separate from training and candidate selection.
|
|
22
|
+
Register shared source units when several scenarios come from the same source.
|
|
23
|
+
Additional repetitions do not create new independent units.
|
|
24
|
+
|
|
25
|
+
## Assess the result
|
|
26
|
+
|
|
27
|
+
1. Inspect raw baseline and candidate cells, judge failures, pairing, and the registered unit mapping.
|
|
28
|
+
Missing or asymmetric evidence must remain visible and must block a promotion claim.
|
|
29
|
+
2. Use Eval's `heldOutGate` or `heldoutSignificance` for the shared promotion decision.
|
|
30
|
+
Report its deciding interval, independent-unit count, eligibility, and vetoes.
|
|
31
|
+
The bootstrap diagnostic may differ from the deciding interval for binary outcomes.
|
|
32
|
+
3. Use App's `trustVerdicts` to check rater agreement and surviving-judge coverage when an ensemble is present.
|
|
33
|
+
Agreement alone does not establish accuracy against independent controls or authorize release.
|
|
34
|
+
4. Investigate null or surprising results before assigning a cause.
|
|
35
|
+
A control intervention supports a causal explanation only when it isolates the proposed mechanism.
|
|
36
|
+
5. Retain actual costs, failures, exclusions, and uncertainty with the result.
|
|
37
|
+
Apply the product's release policy to the deciding evidence and record what the evidence cannot establish.
|
|
38
|
+
|
|
39
|
+
## Then consider
|
|
40
|
+
|
|
41
|
+
| Condition | Skill |
|
|
42
|
+
|---|---|
|
|
43
|
+
| Scoring or case coverage cannot test the required outcome | `eval-architect` with the failed controls and missing coverage |
|
|
44
|
+
| Measurement supports a scoped optimization experiment | `improve-conductor` with the baseline, registered comparison, and resource limits |
|
package/README.md
CHANGED
|
@@ -151,6 +151,13 @@ The **complete, always-current reference** — every published subpath, its expo
|
|
|
151
151
|
- [`/runtime`](src/runtime) — the bounded tool loop; the same loop drives a sandbox agent, a Worker, or an in-browser copilot behind one `streamTurn` seam.
|
|
152
152
|
- [`/trace`](src/trace) — bounded stage timing, turn waterfalls, and mission traces over an injected telemetry carrier.
|
|
153
153
|
|
|
154
|
+
**Evaluate changes**
|
|
155
|
+
|
|
156
|
+
- [`/eval-campaign`](docs/api/eval-campaign.md) composes Eval campaigns, optimization methods, and ensemble judges.
|
|
157
|
+
Change `judgeVersion` when the scorer configuration changes.
|
|
158
|
+
Record paid judge calls through the supplied `costLedger`, `costPhase`, and `costTags`.
|
|
159
|
+
`trustVerdicts` checks agreement and coverage; use Eval's `auditEvaluator` to test accuracy against independent controls.
|
|
160
|
+
|
|
154
161
|
**The server chat vertical** ([`examples/chat-app.md`](./examples/chat-app.md))
|
|
155
162
|
- [`/chat-routes`](src/chat-routes) — `createChatTurnRoutes`: auth → store → streaming turn with buffered replay → uploads → sidecar question answering, assembled. Plus `runDetachedTurn` for autonomous turns a browser can still watch live.
|
|
156
163
|
- [`/chat-store`](src/chat-store) · [`/interactions`](src/interactions) · [`/plans`](src/plans) — persistence, human-in-the-loop asks, and the durable plan projection.
|
|
@@ -27,6 +27,8 @@ import type { JudgeConfig, Scenario } from '@tangle-network/agent-eval/campaign'
|
|
|
27
27
|
export interface EnsembleJudgeConfig<TArtifact, TScenario extends Scenario, D extends string> {
|
|
28
28
|
/** Judge name — appears in traces and scorecards. */
|
|
29
29
|
name: string;
|
|
30
|
+
/** Scoring revision for campaign caches. Change it when rubric, model, or callback settings change. */
|
|
31
|
+
judgeVersion?: JudgeConfig<TArtifact, TScenario>['judgeVersion'];
|
|
30
32
|
/** Stable-ordered rubric dimensions. Drives the `JudgeDimension` list AND the
|
|
31
33
|
* reducer keys, so a judge that omits a dimension scores it 0 (never silently
|
|
32
34
|
* dropped). */
|
|
@@ -38,11 +40,9 @@ export interface EnsembleJudgeConfig<TArtifact, TScenario extends Scenario, D ex
|
|
|
38
40
|
* `{ model, perDimension: null }` to record a judge failure WITHOUT killing
|
|
39
41
|
* the ensemble; throw only on an unrecoverable error (the whole rep is then
|
|
40
42
|
* treated as a failed judge).
|
|
43
|
+
* Use the supplied costLedger, costPhase, and costTags to record paid judge calls.
|
|
41
44
|
*/
|
|
42
|
-
scoreOne: (input: {
|
|
43
|
-
artifact: TArtifact;
|
|
44
|
-
scenario: TScenario;
|
|
45
|
-
signal: AbortSignal;
|
|
45
|
+
scoreOne: (input: Parameters<JudgeConfig<TArtifact, TScenario>['score']>[0] & {
|
|
46
46
|
rep: number;
|
|
47
47
|
}) => Promise<JudgeVerdict<D>>;
|
|
48
48
|
/** Independent judge calls per artifact, reduced by `aggregateJudgeVerdicts`.
|
|
@@ -115,10 +115,11 @@ function buildEnsembleJudge(cfg) {
|
|
|
115
115
|
}
|
|
116
116
|
return {
|
|
117
117
|
name: cfg.name,
|
|
118
|
+
...cfg.judgeVersion === void 0 ? {} : { judgeVersion: cfg.judgeVersion },
|
|
118
119
|
dimensions: cfg.rubric.map((key) => ({ key, description: cfg.describe?.(key) ?? key })),
|
|
119
|
-
async score(
|
|
120
|
+
async score(input) {
|
|
120
121
|
const settled = await Promise.allSettled(
|
|
121
|
-
Array.from({ length: reps }, (_, rep) => cfg.scoreOne({
|
|
122
|
+
Array.from({ length: reps }, (_, rep) => cfg.scoreOne({ ...input, rep }))
|
|
122
123
|
);
|
|
123
124
|
const verdicts = settled.map(
|
|
124
125
|
(r, rep) => r.status === "fulfilled" ? r.value : { model: `${cfg.name}-rep${rep}`, perDimension: null, rationale: String(r.reason) }
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"sources":["../../src/eval-campaign/index.ts","../../src/eval-campaign/trust-gate.ts"],"sourcesContent":["/**\n * Eval-campaign — the app-shell's curated surface for a product's\n * self-improvement loop, NOT a reimplementation.\n *\n * The loop ENGINE lives in `@tangle-network/agent-eval` (a peer dependency):\n * `selfImprove` already owns execution, scoring, data separation, release\n * decisions, durable provenance, and hosted ingest. Candidate search is always\n * explicit: pass an official optimization `method`, or pass a caller-owned\n * `SurfaceProposer`.\n *\n * This module adds the one piece `selfImprove` does not own and which every\n * multi-model product re-hand-rolls — the ensemble judge:\n *\n * {@link buildEnsembleJudge} — turn a per-rubric `scoreOne` into a\n * `JudgeConfig` that fans out N uncorrelated judge calls and reduces them via\n * the substrate's `aggregateJudgeVerdicts` (survivor-mean, inter-rater spread,\n * fail-loud on all-failed). A product writes its rubric + one judge call; the\n * fan-out, partial-failure handling, and composite are the scaffold's.\n *\n * Everything else is a curated re-export so a product has ONE eval import:\n * `selfImprove` + release policies + optimization methods + their types. See\n * `.claude/skills/eval-campaign/SKILL.md` for the wiring contract.\n */\n\nimport {\n aggregateJudgeVerdicts,\n type JudgeVerdict,\n} from '@tangle-network/agent-eval'\nimport type {\n JudgeConfig,\n JudgeScore,\n Scenario,\n} from '@tangle-network/agent-eval/campaign'\n\n/** Config for {@link buildEnsembleJudge}. `D` = the rubric's dimension union. */\nexport interface EnsembleJudgeConfig<TArtifact, TScenario extends Scenario, D extends string> {\n /** Judge name — appears in traces and scorecards. */\n name: string\n /** Stable-ordered rubric dimensions. Drives the `JudgeDimension` list AND the\n * reducer keys, so a judge that omits a dimension scores it 0 (never silently\n * dropped). */\n rubric: readonly D[]\n /**\n * Score ONE artifact on the rubric → a raw per-dimension verdict. Called\n * `judgeReps` times per artifact; vary the model by `rep` for an uncorrelated\n * ensemble (judges that share a base model share its bias). Return\n * `{ model, perDimension: null }` to record a judge failure WITHOUT killing\n * the ensemble; throw only on an unrecoverable error (the whole rep is then\n * treated as a failed judge).\n */\n scoreOne: (input: {\n artifact: TArtifact\n scenario: TScenario\n signal: AbortSignal\n rep: number\n }) => Promise<JudgeVerdict<D>>\n /** Independent judge calls per artifact, reduced by `aggregateJudgeVerdicts`.\n * Default 1. Raise (with model variety in `scoreOne`) for inter-rater bands. */\n judgeReps?: number\n /** Per-dimension composite weights. Default: uniform over `rubric`. A partial\n * map selects-and-weights exactly the named dimensions. */\n weights?: Partial<Record<D, number>>\n /** Optional human-readable dimension descriptions. Default: the key itself. */\n describe?: (dim: D) => string\n}\n\n/**\n * Build a `JudgeConfig` whose `score()` fans out `judgeReps` independent\n * `scoreOne` calls and reduces them with the substrate's\n * `aggregateJudgeVerdicts`. A single judge call failing does NOT fail the cell\n * (it is recorded and dropped); only ALL judges failing throws — which the\n * campaign records as a failed cell, never a silent zero.\n *\n * Pass the result straight to `selfImprove({ judge })` (or `runCampaign`).\n */\nexport function buildEnsembleJudge<TArtifact, TScenario extends Scenario, D extends string>(\n cfg: EnsembleJudgeConfig<TArtifact, TScenario, D>,\n): JudgeConfig<TArtifact, TScenario> {\n const reps = cfg.judgeReps ?? 1\n if (reps < 1) {\n throw new Error(`buildEnsembleJudge: judgeReps must be >= 1 (got ${reps})`)\n }\n if (cfg.rubric.length === 0) {\n throw new Error('buildEnsembleJudge: rubric is empty')\n }\n return {\n name: cfg.name,\n dimensions: cfg.rubric.map((key) => ({ key, description: cfg.describe?.(key) ?? key })),\n async score({ artifact, scenario, signal }): Promise<JudgeScore> {\n const settled = await Promise.allSettled(\n Array.from({ length: reps }, (_, rep) => cfg.scoreOne({ artifact, scenario, signal, rep })),\n )\n const verdicts: JudgeVerdict<D>[] = settled.map((r, rep) =>\n r.status === 'fulfilled'\n ? r.value\n : { model: `${cfg.name}-rep${rep}`, perDimension: null, rationale: String(r.reason) },\n )\n // Throws iff EVERY rep failed → the campaign records a failed cell.\n const agg = aggregateJudgeVerdicts(verdicts, cfg.rubric, cfg.weights)\n return { composite: agg.composite, dimensions: agg.perDimension, notes: agg.rationale }\n },\n }\n}\n\n// ── Trust gate — the after-gate (\"is this result allowed to be believed\") ────\n// One level up from `aggregateJudgeVerdicts`: it audits the raters ACROSS items\n// and reports whether the composites are believable before a lift is reported.\nexport {\n trustVerdicts,\n type TrustItem,\n type TrustThresholds,\n type TrustVerdict,\n} from './trust-gate'\n\n// ── Curated re-exports — the one eval import for a product loop ──────────────\n// The loop engine + gates + drivers + the ensemble reducer, so a product wires\n// its self-improvement loop from a single module instead of reaching across\n// three agent-eval subpaths. All DOWNWARD imports (agent-app consumes the\n// substrate); the layering rule is preserved.\n\nexport { aggregateJudgeVerdicts } from '@tangle-network/agent-eval'\nexport type {\n EnsembleAggregate,\n JudgeVerdict,\n RunRecord,\n} from '@tangle-network/agent-eval'\nexport {\n compareOptimizationMethods,\n defaultProductionGate,\n externalTextOptimizationMethod,\n gepaOptimizationMethod,\n paretoSignificanceGate,\n runCampaign,\n skillOptOptimizationMethod,\n} from '@tangle-network/agent-eval/campaign'\nexport type {\n CampaignResult,\n CompareOptimizationMethodsOptions,\n DispatchContext,\n ExternalTextOptimizationMethodConfig,\n Gate,\n GepaOptimizationMethodConfig,\n JudgeConfig,\n JudgeDimension,\n JudgeScore,\n LabeledScenarioStore,\n MutableSurface,\n OptimizationMethod,\n OptimizationMethodResult,\n Scenario,\n SkillOptOptimizationMethodConfig,\n SurfaceProposer,\n} from '@tangle-network/agent-eval/campaign'\nexport { selfImprove } from '@tangle-network/agent-eval/contract'\nexport type {\n SelfImproveBudget,\n SelfImproveOptions,\n SelfImproveResult,\n} from '@tangle-network/agent-eval/contract'\n","/**\n * Trust gate — decides whether an ensemble's scores are allowed to be BELIEVED,\n * one level up from {@link aggregateJudgeVerdicts} (which only reduces ONE\n * artifact's raters to a composite). A composite is a number; this is the check\n * that the number means anything. It is the code \"Enforced by\" for the\n * measurement-validation skill's after-gate (\"is this result allowed to be\n * believed\").\n *\n * Three checks, each fail-loud and named in `trustReasons`:\n * (1) inter-rater reliability over the corpus ≥ `irrFloor` — raters that\n * disagree no better than chance carry no signal to optimize against.\n * (2) per-item rater spread ≤ `spreadCeiling` — for EACH item, raters must\n * converge on THAT item.\n * (3) surviving raters per item ≥ `minSurvivors` — a mean over one or two\n * raters is an anecdote, not an ensemble.\n *\n * CRITICAL metric semantics — per-item spread is rater disagreement about the\n * SAME item: `max(score) − min(score)` across the raters that scored THAT item\n * (max over its dimensions), never pooled across different items or across the\n * baseline/candidate sides. Pooling reads a genuine quality gap BETWEEN items as\n * \"the raters split\" and so trips the gate exactly when the finding is largest —\n * the failure mode the after-gate exists to prevent. The corpus IRR (check 1)\n * leans on the substrate's `interRaterReliability`, whose expected-disagreement\n * denominator already pools across items, so genuine item-to-item variation\n * RAISES reliability rather than lowering it.\n */\n\nimport {\n interRaterReliability,\n type JudgeScore,\n type JudgeVerdict,\n} from '@tangle-network/agent-eval'\n\n/** One item's raters: the per-judge verdicts {@link aggregateJudgeVerdicts}\n * reduces, tagged with the item they scored so spread stays within-item. */\nexport interface TrustItem<D extends string = string> {\n /** Stable item identifier — surfaces in `perItemSpread` and `trustReasons`. */\n itemId: string\n /** The raters' verdicts for THIS item (one per judge call). A failed judge\n * (`perDimension: null`) is dropped before spread/IRR, never folded as 0. */\n verdicts: readonly JudgeVerdict<D>[]\n}\n\n/** Thresholds for {@link trustVerdicts}. All overridable; defaults are the\n * conservative after-gate bar. */\nexport interface TrustThresholds {\n /** Minimum corpus inter-rater reliability (Krippendorff-style α). Below this\n * the raters agree no better than chance. Default 0.2. */\n irrFloor?: number\n /** Maximum per-item rater spread (`max − min` over a single item's surviving\n * raters, across its dimensions). Above this the raters split ON THAT ITEM.\n * Default 0.5. */\n spreadCeiling?: number\n /** Minimum surviving (non-failed) raters required per item. Default 3. */\n minSurvivors?: number\n}\n\n/** Result of the trust gate. `trustworthy` iff every check passed; `trustReasons`\n * is empty iff `trustworthy`. */\nexport interface TrustVerdict {\n /** True iff IRR ≥ floor AND every item's spread ≤ ceiling AND every item has\n * ≥ `minSurvivors` surviving raters. */\n trustworthy: boolean\n /** One entry per FAILED check, each naming its number + the offending value.\n * Empty iff `trustworthy`. */\n trustReasons: string[]\n /** Corpus inter-rater reliability actually measured (the check-1 value). */\n interRaterReliability: number\n /** Per-item spread (`max − min` over surviving raters, max over dimensions),\n * keyed by `itemId`. The check-2 input, surfaced for drill-down. */\n perItemSpread: Record<string, number>\n}\n\nconst DEFAULT_IRR_FLOOR = 0.2\nconst DEFAULT_SPREAD_CEILING = 0.5\nconst DEFAULT_MIN_SURVIVORS = 3\n\n/** Surviving (non-failed) verdicts for an item — those with a real\n * `perDimension` map. A failed judge carries no scores and is excluded from\n * every statistic (it is NOT a zero rater). */\nfunction survivors<D extends string>(item: TrustItem<D>): JudgeVerdict<D>[] {\n return item.verdicts.filter((v) => v.perDimension !== null)\n}\n\n/**\n * Within-item rater spread: for each dimension, `max − min` across the item's\n * surviving raters; the item's spread is the max over its dimensions (the worst\n * dimension the raters split on). Pooled ONLY within this one item — never\n * across items — so a quality gap between items cannot inflate it.\n */\nfunction itemSpread<D extends string>(survivorVerdicts: JudgeVerdict<D>[]): number {\n if (survivorVerdicts.length < 2) return 0\n const dims = new Set<string>()\n for (const v of survivorVerdicts) {\n for (const d of Object.keys(v.perDimension as Record<string, number>)) dims.add(d)\n }\n let worst = 0\n for (const d of dims) {\n let min = Infinity\n let max = -Infinity\n for (const v of survivorVerdicts) {\n const score = (v.perDimension as Record<string, number>)[d]\n if (score === undefined) continue\n if (score < min) min = score\n if (score > max) max = score\n }\n if (max > -Infinity && max - min > worst) worst = max - min\n }\n return worst\n}\n\n/**\n * Decide whether an ensemble's per-item verdicts are trustworthy enough to\n * believe a lift computed from them. Pure: no LLM, no I/O, no clock, no random —\n * the same `items` + `thresholds` always yield the same verdict.\n *\n * Sibling to {@link aggregateJudgeVerdicts}: that reduces ONE item's raters to a\n * composite; this audits the raters ACROSS items and reports whether the\n * composites are believable. Run it on the corpus of held-out items before\n * reporting any lift over their scores.\n *\n * @throws if `items` is empty — an empty corpus has no measurable trust, and a\n * silent `trustworthy: true` over zero evidence is the exact lie the gate\n * exists to refuse.\n */\nexport function trustVerdicts<D extends string>(\n items: readonly TrustItem<D>[],\n thresholds: TrustThresholds = {},\n): TrustVerdict {\n if (items.length === 0) {\n throw new Error('trustVerdicts: items is empty — no evidence to trust')\n }\n const irrFloor = thresholds.irrFloor ?? DEFAULT_IRR_FLOOR\n const spreadCeiling = thresholds.spreadCeiling ?? DEFAULT_SPREAD_CEILING\n const minSurvivors = thresholds.minSurvivors ?? DEFAULT_MIN_SURVIVORS\n\n // Rater-major JudgeScore series for the substrate's IRR. Each item's surviving\n // raters are assigned a stable column index so the same rater across items\n // lines up; per (item, dimension) one JudgeScore per rater, in item-then-\n // dimension order — the layout interRaterReliability chunks back into items.\n const maxRaters = items.reduce((m, it) => Math.max(m, survivors(it).length), 0)\n const raterSeries: JudgeScore[][] = Array.from({ length: maxRaters }, () => [])\n const perItemSpread: Record<string, number> = {}\n const splitItems: Array<{ itemId: string; spread: number }> = []\n const starvedItems: Array<{ itemId: string; n: number }> = []\n\n for (const item of items) {\n const surv = survivors(item)\n if (surv.length < minSurvivors) starvedItems.push({ itemId: item.itemId, n: surv.length })\n\n const spread = itemSpread(surv)\n perItemSpread[item.itemId] = spread\n if (spread > spreadCeiling) splitItems.push({ itemId: item.itemId, spread })\n\n if (surv.length >= 2) {\n const dims = Array.from(\n new Set(surv.flatMap((v) => Object.keys(v.perDimension as Record<string, number>))),\n ).sort()\n surv.forEach((v, raterIdx) => {\n // raterIdx < surv.length ≤ maxRaters = raterSeries.length, so the column\n // always exists; the ??= keeps the access provably defined for the type.\n const column = (raterSeries[raterIdx] ??= [])\n const pd = v.perDimension as Record<string, number>\n for (const d of dims) {\n const score = pd[d]\n if (score === undefined) continue\n column.push({\n judgeName: v.model,\n dimension: `${item.itemId}::${d}`,\n score,\n reasoning: v.rationale ?? '',\n })\n }\n })\n }\n }\n\n const irr = interRaterReliability(raterSeries)\n\n const trustReasons: string[] = []\n if (irr < irrFloor) {\n trustReasons.push(`(1) IRR ${round(irr)} < ${irrFloor}`)\n }\n for (const { itemId, spread } of splitItems) {\n trustReasons.push(`(2) item ${itemId} spread ${round(spread)} > ${spreadCeiling} — raters split`)\n }\n for (const { itemId, n } of starvedItems) {\n trustReasons.push(`(3) item ${itemId}: ${n} surviving raters < ${minSurvivors}`)\n }\n\n return {\n trustworthy: trustReasons.length === 0,\n trustReasons,\n interRaterReliability: irr,\n perItemSpread,\n }\n}\n\n/** Round to 2 decimals for stable, readable reason strings. */\nfunction round(n: number): number {\n return Math.round(n * 100) / 100\n}\n"],"mappings":";AAwBA;AAAA,EACE;AAAA,OAEK;;;ACAP;AAAA,EACE;AAAA,OAGK;AA0CP,IAAM,oBAAoB;AAC1B,IAAM,yBAAyB;AAC/B,IAAM,wBAAwB;AAK9B,SAAS,UAA4B,MAAuC;AAC1E,SAAO,KAAK,SAAS,OAAO,CAAC,MAAM,EAAE,iBAAiB,IAAI;AAC5D;AAQA,SAAS,WAA6B,kBAA6C;AACjF,MAAI,iBAAiB,SAAS,EAAG,QAAO;AACxC,QAAM,OAAO,oBAAI,IAAY;AAC7B,aAAW,KAAK,kBAAkB;AAChC,eAAW,KAAK,OAAO,KAAK,EAAE,YAAsC,EAAG,MAAK,IAAI,CAAC;AAAA,EACnF;AACA,MAAI,QAAQ;AACZ,aAAW,KAAK,MAAM;AACpB,QAAI,MAAM;AACV,QAAI,MAAM;AACV,eAAW,KAAK,kBAAkB;AAChC,YAAM,QAAS,EAAE,aAAwC,CAAC;AAC1D,UAAI,UAAU,OAAW;AACzB,UAAI,QAAQ,IAAK,OAAM;AACvB,UAAI,QAAQ,IAAK,OAAM;AAAA,IACzB;AACA,QAAI,MAAM,aAAa,MAAM,MAAM,MAAO,SAAQ,MAAM;AAAA,EAC1D;AACA,SAAO;AACT;AAgBO,SAAS,cACd,OACA,aAA8B,CAAC,GACjB;AACd,MAAI,MAAM,WAAW,GAAG;AACtB,UAAM,IAAI,MAAM,2DAAsD;AAAA,EACxE;AACA,QAAM,WAAW,WAAW,YAAY;AACxC,QAAM,gBAAgB,WAAW,iBAAiB;AAClD,QAAM,eAAe,WAAW,gBAAgB;AAMhD,QAAM,YAAY,MAAM,OAAO,CAAC,GAAG,OAAO,KAAK,IAAI,GAAG,UAAU,EAAE,EAAE,MAAM,GAAG,CAAC;AAC9E,QAAM,cAA8B,MAAM,KAAK,EAAE,QAAQ,UAAU,GAAG,MAAM,CAAC,CAAC;AAC9E,QAAM,gBAAwC,CAAC;AAC/C,QAAM,aAAwD,CAAC;AAC/D,QAAM,eAAqD,CAAC;AAE5D,aAAW,QAAQ,OAAO;AACxB,UAAM,OAAO,UAAU,IAAI;AAC3B,QAAI,KAAK,SAAS,aAAc,cAAa,KAAK,EAAE,QAAQ,KAAK,QAAQ,GAAG,KAAK,OAAO,CAAC;AAEzF,UAAM,SAAS,WAAW,IAAI;AAC9B,kBAAc,KAAK,MAAM,IAAI;AAC7B,QAAI,SAAS,cAAe,YAAW,KAAK,EAAE,QAAQ,KAAK,QAAQ,OAAO,CAAC;AAE3E,QAAI,KAAK,UAAU,GAAG;AACpB,YAAM,OAAO,MAAM;AAAA,QACjB,IAAI,IAAI,KAAK,QAAQ,CAAC,MAAM,OAAO,KAAK,EAAE,YAAsC,CAAC,CAAC;AAAA,MACpF,EAAE,KAAK;AACP,WAAK,QAAQ,CAAC,GAAG,aAAa;AAG5B,cAAM,SAAU,YAAY,QAAQ,MAAM,CAAC;AAC3C,cAAM,KAAK,EAAE;AACb,mBAAW,KAAK,MAAM;AACpB,gBAAM,QAAQ,GAAG,CAAC;AAClB,cAAI,UAAU,OAAW;AACzB,iBAAO,KAAK;AAAA,YACV,WAAW,EAAE;AAAA,YACb,WAAW,GAAG,KAAK,MAAM,KAAK,CAAC;AAAA,YAC/B;AAAA,YACA,WAAW,EAAE,aAAa;AAAA,UAC5B,CAAC;AAAA,QACH;AAAA,MACF,CAAC;AAAA,IACH;AAAA,EACF;AAEA,QAAM,MAAM,sBAAsB,WAAW;AAE7C,QAAM,eAAyB,CAAC;AAChC,MAAI,MAAM,UAAU;AAClB,iBAAa,KAAK,WAAW,MAAM,GAAG,CAAC,MAAM,QAAQ,EAAE;AAAA,EACzD;AACA,aAAW,EAAE,QAAQ,OAAO,KAAK,YAAY;AAC3C,iBAAa,KAAK,YAAY,MAAM,WAAW,MAAM,MAAM,CAAC,MAAM,aAAa,sBAAiB;AAAA,EAClG;AACA,aAAW,EAAE,QAAQ,EAAE,KAAK,cAAc;AACxC,iBAAa,KAAK,YAAY,MAAM,KAAK,CAAC,uBAAuB,YAAY,EAAE;AAAA,EACjF;AAEA,SAAO;AAAA,IACL,aAAa,aAAa,WAAW;AAAA,IACrC;AAAA,IACA,uBAAuB;AAAA,IACvB;AAAA,EACF;AACF;AAGA,SAAS,MAAM,GAAmB;AAChC,SAAO,KAAK,MAAM,IAAI,GAAG,IAAI;AAC/B;;;ADjFA,SAAS,0BAAAA,+BAA8B;AAMvC;AAAA,EACE;AAAA,EACA;AAAA,EACA;AAAA,EACA;AAAA,EACA;AAAA,EACA;AAAA,EACA;AAAA,OACK;AAmBP,SAAS,mBAAmB;AA9ErB,SAAS,mBACd,KACmC;AACnC,QAAM,OAAO,IAAI,aAAa;AAC9B,MAAI,OAAO,GAAG;AACZ,UAAM,IAAI,MAAM,mDAAmD,IAAI,GAAG;AAAA,EAC5E;AACA,MAAI,IAAI,OAAO,WAAW,GAAG;AAC3B,UAAM,IAAI,MAAM,qCAAqC;AAAA,EACvD;AACA,SAAO;AAAA,IACL,MAAM,IAAI;AAAA,IACV,YAAY,IAAI,OAAO,IAAI,CAAC,SAAS,EAAE,KAAK,aAAa,IAAI,WAAW,GAAG,KAAK,IAAI,EAAE;AAAA,IACtF,MAAM,MAAM,EAAE,UAAU,UAAU,OAAO,GAAwB;AAC/D,YAAM,UAAU,MAAM,QAAQ;AAAA,QAC5B,MAAM,KAAK,EAAE,QAAQ,KAAK,GAAG,CAAC,GAAG,QAAQ,IAAI,SAAS,EAAE,UAAU,UAAU,QAAQ,IAAI,CAAC,CAAC;AAAA,MAC5F;AACA,YAAM,WAA8B,QAAQ;AAAA,QAAI,CAAC,GAAG,QAClD,EAAE,WAAW,cACT,EAAE,QACF,EAAE,OAAO,GAAG,IAAI,IAAI,OAAO,GAAG,IAAI,cAAc,MAAM,WAAW,OAAO,EAAE,MAAM,EAAE;AAAA,MACxF;AAEA,YAAM,MAAM,uBAAuB,UAAU,IAAI,QAAQ,IAAI,OAAO;AACpE,aAAO,EAAE,WAAW,IAAI,WAAW,YAAY,IAAI,cAAc,OAAO,IAAI,UAAU;AAAA,IACxF;AAAA,EACF;AACF;","names":["aggregateJudgeVerdicts"]}
|
|
1
|
+
{"version":3,"sources":["../../src/eval-campaign/index.ts","../../src/eval-campaign/trust-gate.ts"],"sourcesContent":["/**\n * Eval-campaign — the app-shell's curated surface for a product's\n * self-improvement loop, NOT a reimplementation.\n *\n * The loop ENGINE lives in `@tangle-network/agent-eval` (a peer dependency):\n * `selfImprove` already owns execution, scoring, data separation, release\n * decisions, durable provenance, and hosted ingest. Candidate search is always\n * explicit: pass an official optimization `method`, or pass a caller-owned\n * `SurfaceProposer`.\n *\n * This module adds the one piece `selfImprove` does not own and which every\n * multi-model product re-hand-rolls — the ensemble judge:\n *\n * {@link buildEnsembleJudge} — turn a per-rubric `scoreOne` into a\n * `JudgeConfig` that fans out N uncorrelated judge calls and reduces them via\n * the substrate's `aggregateJudgeVerdicts` (survivor-mean, inter-rater spread,\n * fail-loud on all-failed). A product writes its rubric + one judge call; the\n * fan-out, partial-failure handling, and composite are the scaffold's.\n *\n * Everything else is a curated re-export so a product has ONE eval import:\n * `selfImprove` + release policies + optimization methods + their types. See\n * `.claude/skills/eval-campaign/SKILL.md` for the wiring contract.\n */\n\nimport {\n aggregateJudgeVerdicts,\n type JudgeVerdict,\n} from '@tangle-network/agent-eval'\nimport type {\n JudgeConfig,\n JudgeScore,\n Scenario,\n} from '@tangle-network/agent-eval/campaign'\n\n/** Config for {@link buildEnsembleJudge}. `D` = the rubric's dimension union. */\nexport interface EnsembleJudgeConfig<TArtifact, TScenario extends Scenario, D extends string> {\n /** Judge name — appears in traces and scorecards. */\n name: string\n /** Scoring revision for campaign caches. Change it when rubric, model, or callback settings change. */\n judgeVersion?: JudgeConfig<TArtifact, TScenario>['judgeVersion']\n /** Stable-ordered rubric dimensions. Drives the `JudgeDimension` list AND the\n * reducer keys, so a judge that omits a dimension scores it 0 (never silently\n * dropped). */\n rubric: readonly D[]\n /**\n * Score ONE artifact on the rubric → a raw per-dimension verdict. Called\n * `judgeReps` times per artifact; vary the model by `rep` for an uncorrelated\n * ensemble (judges that share a base model share its bias). Return\n * `{ model, perDimension: null }` to record a judge failure WITHOUT killing\n * the ensemble; throw only on an unrecoverable error (the whole rep is then\n * treated as a failed judge).\n * Use the supplied costLedger, costPhase, and costTags to record paid judge calls.\n */\n scoreOne: (input: Parameters<JudgeConfig<TArtifact, TScenario>['score']>[0] & {\n rep: number\n }) => Promise<JudgeVerdict<D>>\n /** Independent judge calls per artifact, reduced by `aggregateJudgeVerdicts`.\n * Default 1. Raise (with model variety in `scoreOne`) for inter-rater bands. */\n judgeReps?: number\n /** Per-dimension composite weights. Default: uniform over `rubric`. A partial\n * map selects-and-weights exactly the named dimensions. */\n weights?: Partial<Record<D, number>>\n /** Optional human-readable dimension descriptions. Default: the key itself. */\n describe?: (dim: D) => string\n}\n\n/**\n * Build a `JudgeConfig` whose `score()` fans out `judgeReps` independent\n * `scoreOne` calls and reduces them with the substrate's\n * `aggregateJudgeVerdicts`. A single judge call failing does NOT fail the cell\n * (it is recorded and dropped); only ALL judges failing throws — which the\n * campaign records as a failed cell, never a silent zero.\n *\n * Pass the result straight to `selfImprove({ judge })` (or `runCampaign`).\n */\nexport function buildEnsembleJudge<TArtifact, TScenario extends Scenario, D extends string>(\n cfg: EnsembleJudgeConfig<TArtifact, TScenario, D>,\n): JudgeConfig<TArtifact, TScenario> {\n const reps = cfg.judgeReps ?? 1\n if (reps < 1) {\n throw new Error(`buildEnsembleJudge: judgeReps must be >= 1 (got ${reps})`)\n }\n if (cfg.rubric.length === 0) {\n throw new Error('buildEnsembleJudge: rubric is empty')\n }\n return {\n name: cfg.name,\n ...(cfg.judgeVersion === undefined ? {} : { judgeVersion: cfg.judgeVersion }),\n dimensions: cfg.rubric.map((key) => ({ key, description: cfg.describe?.(key) ?? key })),\n async score(input): Promise<JudgeScore> {\n const settled = await Promise.allSettled(\n Array.from({ length: reps }, (_, rep) => cfg.scoreOne({ ...input, rep })),\n )\n const verdicts: JudgeVerdict<D>[] = settled.map((r, rep) =>\n r.status === 'fulfilled'\n ? r.value\n : { model: `${cfg.name}-rep${rep}`, perDimension: null, rationale: String(r.reason) },\n )\n // Throws iff EVERY rep failed → the campaign records a failed cell.\n const agg = aggregateJudgeVerdicts(verdicts, cfg.rubric, cfg.weights)\n return { composite: agg.composite, dimensions: agg.perDimension, notes: agg.rationale }\n },\n }\n}\n\n// Agreement across items is distinct from evaluator accuracy against independent controls.\n// Consumers can opt into auditEvaluator from agent-eval/meta-eval for that separate evidence.\nexport {\n trustVerdicts,\n type TrustItem,\n type TrustThresholds,\n type TrustVerdict,\n} from './trust-gate'\n\n// ── Curated re-exports — the one eval import for a product loop ──────────────\n// The loop engine + gates + drivers + the ensemble reducer, so a product wires\n// its self-improvement loop from a single module instead of reaching across\n// three agent-eval subpaths. All DOWNWARD imports (agent-app consumes the\n// substrate); the layering rule is preserved.\n\nexport { aggregateJudgeVerdicts } from '@tangle-network/agent-eval'\nexport type {\n EnsembleAggregate,\n JudgeVerdict,\n RunRecord,\n} from '@tangle-network/agent-eval'\nexport {\n compareOptimizationMethods,\n defaultProductionGate,\n externalTextOptimizationMethod,\n gepaOptimizationMethod,\n paretoSignificanceGate,\n runCampaign,\n skillOptOptimizationMethod,\n} from '@tangle-network/agent-eval/campaign'\nexport type {\n CampaignResult,\n CompareOptimizationMethodsOptions,\n DispatchContext,\n ExternalTextOptimizationMethodConfig,\n Gate,\n GepaOptimizationMethodConfig,\n JudgeConfig,\n JudgeDimension,\n JudgeScore,\n LabeledScenarioStore,\n MutableSurface,\n OptimizationMethod,\n OptimizationMethodResult,\n Scenario,\n SkillOptOptimizationMethodConfig,\n SurfaceProposer,\n} from '@tangle-network/agent-eval/campaign'\nexport { selfImprove } from '@tangle-network/agent-eval/contract'\nexport type {\n SelfImproveBudget,\n SelfImproveOptions,\n SelfImproveResult,\n} from '@tangle-network/agent-eval/contract'\n","/**\n * Summarize rater agreement, within-item spread, and surviving-judge coverage.\n * Consumers choose the thresholds that a trustworthy result must satisfy.\n * Agreement does not measure evaluator errors against independent controls.\n * Use agent-eval/meta-eval's auditEvaluator when the consumer needs that separate evidence.\n * Spread stays within each item so differences in task quality cannot mimic rater disagreement.\n */\n\nimport {\n interRaterReliability,\n type DimensionJudgeScore,\n type JudgeVerdict,\n} from '@tangle-network/agent-eval'\n\n/** One item's raters: the per-judge verdicts {@link aggregateJudgeVerdicts}\n * reduces, tagged with the item they scored so spread stays within-item. */\nexport interface TrustItem<D extends string = string> {\n /** Stable item identifier — surfaces in `perItemSpread` and `trustReasons`. */\n itemId: string\n /** The raters' verdicts for THIS item (one per judge call). A failed judge\n * (`perDimension: null`) is dropped before spread/IRR, never folded as 0. */\n verdicts: readonly JudgeVerdict<D>[]\n}\n\n/** Configurable agreement and coverage thresholds for {@link trustVerdicts}. */\nexport interface TrustThresholds {\n /** Minimum corpus inter-rater reliability (Krippendorff-style α). Default 0.2. */\n irrFloor?: number\n /** Maximum per-item rater spread (`max − min` over a single item's surviving\n * raters, across its dimensions). Above this the raters split ON THAT ITEM.\n * Default 0.5. */\n spreadCeiling?: number\n /** Minimum surviving (non-failed) raters required per item. Default 3. */\n minSurvivors?: number\n}\n\n/** Result of the trust gate. `trustworthy` iff every check passed; `trustReasons`\n * is empty iff `trustworthy`. */\nexport interface TrustVerdict {\n /** True iff IRR ≥ floor AND every item's spread ≤ ceiling AND every item has\n * ≥ `minSurvivors` surviving raters. */\n trustworthy: boolean\n /** One entry per FAILED check, each naming its number + the offending value.\n * Empty iff `trustworthy`. */\n trustReasons: string[]\n /** Corpus inter-rater reliability actually measured (the check-1 value). */\n interRaterReliability: number\n /** Per-item spread (`max − min` over surviving raters, max over dimensions),\n * keyed by `itemId`. The check-2 input, surfaced for drill-down. */\n perItemSpread: Record<string, number>\n}\n\nconst DEFAULT_IRR_FLOOR = 0.2\nconst DEFAULT_SPREAD_CEILING = 0.5\nconst DEFAULT_MIN_SURVIVORS = 3\n\n/** Surviving (non-failed) verdicts for an item — those with a real\n * `perDimension` map. A failed judge carries no scores and is excluded from\n * every statistic (it is NOT a zero rater). */\nfunction survivors<D extends string>(item: TrustItem<D>): JudgeVerdict<D>[] {\n return item.verdicts.filter((v) => v.perDimension !== null)\n}\n\n/**\n * Within-item rater spread: for each dimension, `max − min` across the item's\n * surviving raters; the item's spread is the max over its dimensions (the worst\n * dimension the raters split on). Pooled ONLY within this one item — never\n * across items — so a quality gap between items cannot inflate it.\n */\nfunction itemSpread<D extends string>(survivorVerdicts: JudgeVerdict<D>[]): number {\n if (survivorVerdicts.length < 2) return 0\n const dims = new Set<string>()\n for (const v of survivorVerdicts) {\n for (const d of Object.keys(v.perDimension as Record<string, number>)) dims.add(d)\n }\n let worst = 0\n for (const d of dims) {\n let min = Infinity\n let max = -Infinity\n for (const v of survivorVerdicts) {\n const score = (v.perDimension as Record<string, number>)[d]\n if (score === undefined) continue\n if (score < min) min = score\n if (score > max) max = score\n }\n if (max > -Infinity && max - min > worst) worst = max - min\n }\n return worst\n}\n\n/**\n * Check an ensemble against its configured agreement and coverage thresholds. Pure: no LLM, no I/O, no clock, no random —\n * the same `items` + `thresholds` always yield the same verdict.\n *\n * Sibling to {@link aggregateJudgeVerdicts}: that reduces ONE item's raters to a\n * composite; this summarizes agreement across the supplied items.\n * A passing result does not establish evaluator accuracy or authorize a release.\n *\n * @throws if `items` is empty — an empty corpus has no measurable trust, and a\n * silent `trustworthy: true` over zero evidence is the exact lie the gate\n * exists to refuse.\n */\nexport function trustVerdicts<D extends string>(\n items: readonly TrustItem<D>[],\n thresholds: TrustThresholds = {},\n): TrustVerdict {\n if (items.length === 0) {\n throw new Error('trustVerdicts: items is empty — no evidence to trust')\n }\n const irrFloor = thresholds.irrFloor ?? DEFAULT_IRR_FLOOR\n const spreadCeiling = thresholds.spreadCeiling ?? DEFAULT_SPREAD_CEILING\n const minSurvivors = thresholds.minSurvivors ?? DEFAULT_MIN_SURVIVORS\n\n // Rater-major JudgeScore series for the substrate's IRR. Each item's surviving\n // raters are assigned a stable column index so the same rater across items\n // lines up; per (item, dimension) one JudgeScore per rater, in item-then-\n // dimension order — the layout interRaterReliability chunks back into items.\n const maxRaters = items.reduce((m, it) => Math.max(m, survivors(it).length), 0)\n const raterSeries: DimensionJudgeScore[][] = Array.from({ length: maxRaters }, () => [])\n const perItemSpread: Record<string, number> = {}\n const splitItems: Array<{ itemId: string; spread: number }> = []\n const starvedItems: Array<{ itemId: string; n: number }> = []\n\n for (const item of items) {\n const surv = survivors(item)\n if (surv.length < minSurvivors) starvedItems.push({ itemId: item.itemId, n: surv.length })\n\n const spread = itemSpread(surv)\n perItemSpread[item.itemId] = spread\n if (spread > spreadCeiling) splitItems.push({ itemId: item.itemId, spread })\n\n if (surv.length >= 2) {\n const dims = Array.from(\n new Set(surv.flatMap((v) => Object.keys(v.perDimension as Record<string, number>))),\n ).sort()\n surv.forEach((v, raterIdx) => {\n // raterIdx < surv.length ≤ maxRaters = raterSeries.length, so the column\n // always exists; the ??= keeps the access provably defined for the type.\n const column = (raterSeries[raterIdx] ??= [])\n const pd = v.perDimension as Record<string, number>\n for (const d of dims) {\n const score = pd[d]\n if (score === undefined) continue\n column.push({\n judgeName: v.model,\n dimension: `${item.itemId}::${d}`,\n score,\n reasoning: v.rationale ?? '',\n })\n }\n })\n }\n }\n\n const irr = interRaterReliability(raterSeries)\n\n const trustReasons: string[] = []\n if (irr < irrFloor) {\n trustReasons.push(`(1) IRR ${round(irr)} < ${irrFloor}`)\n }\n for (const { itemId, spread } of splitItems) {\n trustReasons.push(`(2) item ${itemId} spread ${round(spread)} > ${spreadCeiling} — raters split`)\n }\n for (const { itemId, n } of starvedItems) {\n trustReasons.push(`(3) item ${itemId}: ${n} surviving raters < ${minSurvivors}`)\n }\n\n return {\n trustworthy: trustReasons.length === 0,\n trustReasons,\n interRaterReliability: irr,\n perItemSpread,\n }\n}\n\n/** Round to 2 decimals for stable, readable reason strings. */\nfunction round(n: number): number {\n return Math.round(n * 100) / 100\n}\n"],"mappings":";AAwBA;AAAA,EACE;AAAA,OAEK;;;ACnBP;AAAA,EACE;AAAA,OAGK;AAwCP,IAAM,oBAAoB;AAC1B,IAAM,yBAAyB;AAC/B,IAAM,wBAAwB;AAK9B,SAAS,UAA4B,MAAuC;AAC1E,SAAO,KAAK,SAAS,OAAO,CAAC,MAAM,EAAE,iBAAiB,IAAI;AAC5D;AAQA,SAAS,WAA6B,kBAA6C;AACjF,MAAI,iBAAiB,SAAS,EAAG,QAAO;AACxC,QAAM,OAAO,oBAAI,IAAY;AAC7B,aAAW,KAAK,kBAAkB;AAChC,eAAW,KAAK,OAAO,KAAK,EAAE,YAAsC,EAAG,MAAK,IAAI,CAAC;AAAA,EACnF;AACA,MAAI,QAAQ;AACZ,aAAW,KAAK,MAAM;AACpB,QAAI,MAAM;AACV,QAAI,MAAM;AACV,eAAW,KAAK,kBAAkB;AAChC,YAAM,QAAS,EAAE,aAAwC,CAAC;AAC1D,UAAI,UAAU,OAAW;AACzB,UAAI,QAAQ,IAAK,OAAM;AACvB,UAAI,QAAQ,IAAK,OAAM;AAAA,IACzB;AACA,QAAI,MAAM,aAAa,MAAM,MAAM,MAAO,SAAQ,MAAM;AAAA,EAC1D;AACA,SAAO;AACT;AAcO,SAAS,cACd,OACA,aAA8B,CAAC,GACjB;AACd,MAAI,MAAM,WAAW,GAAG;AACtB,UAAM,IAAI,MAAM,2DAAsD;AAAA,EACxE;AACA,QAAM,WAAW,WAAW,YAAY;AACxC,QAAM,gBAAgB,WAAW,iBAAiB;AAClD,QAAM,eAAe,WAAW,gBAAgB;AAMhD,QAAM,YAAY,MAAM,OAAO,CAAC,GAAG,OAAO,KAAK,IAAI,GAAG,UAAU,EAAE,EAAE,MAAM,GAAG,CAAC;AAC9E,QAAM,cAAuC,MAAM,KAAK,EAAE,QAAQ,UAAU,GAAG,MAAM,CAAC,CAAC;AACvF,QAAM,gBAAwC,CAAC;AAC/C,QAAM,aAAwD,CAAC;AAC/D,QAAM,eAAqD,CAAC;AAE5D,aAAW,QAAQ,OAAO;AACxB,UAAM,OAAO,UAAU,IAAI;AAC3B,QAAI,KAAK,SAAS,aAAc,cAAa,KAAK,EAAE,QAAQ,KAAK,QAAQ,GAAG,KAAK,OAAO,CAAC;AAEzF,UAAM,SAAS,WAAW,IAAI;AAC9B,kBAAc,KAAK,MAAM,IAAI;AAC7B,QAAI,SAAS,cAAe,YAAW,KAAK,EAAE,QAAQ,KAAK,QAAQ,OAAO,CAAC;AAE3E,QAAI,KAAK,UAAU,GAAG;AACpB,YAAM,OAAO,MAAM;AAAA,QACjB,IAAI,IAAI,KAAK,QAAQ,CAAC,MAAM,OAAO,KAAK,EAAE,YAAsC,CAAC,CAAC;AAAA,MACpF,EAAE,KAAK;AACP,WAAK,QAAQ,CAAC,GAAG,aAAa;AAG5B,cAAM,SAAU,YAAY,QAAQ,MAAM,CAAC;AAC3C,cAAM,KAAK,EAAE;AACb,mBAAW,KAAK,MAAM;AACpB,gBAAM,QAAQ,GAAG,CAAC;AAClB,cAAI,UAAU,OAAW;AACzB,iBAAO,KAAK;AAAA,YACV,WAAW,EAAE;AAAA,YACb,WAAW,GAAG,KAAK,MAAM,KAAK,CAAC;AAAA,YAC/B;AAAA,YACA,WAAW,EAAE,aAAa;AAAA,UAC5B,CAAC;AAAA,QACH;AAAA,MACF,CAAC;AAAA,IACH;AAAA,EACF;AAEA,QAAM,MAAM,sBAAsB,WAAW;AAE7C,QAAM,eAAyB,CAAC;AAChC,MAAI,MAAM,UAAU;AAClB,iBAAa,KAAK,WAAW,MAAM,GAAG,CAAC,MAAM,QAAQ,EAAE;AAAA,EACzD;AACA,aAAW,EAAE,QAAQ,OAAO,KAAK,YAAY;AAC3C,iBAAa,KAAK,YAAY,MAAM,WAAW,MAAM,MAAM,CAAC,MAAM,aAAa,sBAAiB;AAAA,EAClG;AACA,aAAW,EAAE,QAAQ,EAAE,KAAK,cAAc;AACxC,iBAAa,KAAK,YAAY,MAAM,KAAK,CAAC,uBAAuB,YAAY,EAAE;AAAA,EACjF;AAEA,SAAO;AAAA,IACL,aAAa,aAAa,WAAW;AAAA,IACrC;AAAA,IACA,uBAAuB;AAAA,IACvB;AAAA,EACF;AACF;AAGA,SAAS,MAAM,GAAmB;AAChC,SAAO,KAAK,MAAM,IAAI,GAAG,IAAI;AAC/B;;;AD1DA,SAAS,0BAAAA,+BAA8B;AAMvC;AAAA,EACE;AAAA,EACA;AAAA,EACA;AAAA,EACA;AAAA,EACA;AAAA,EACA;AAAA,EACA;AAAA,OACK;AAmBP,SAAS,mBAAmB;AA9ErB,SAAS,mBACd,KACmC;AACnC,QAAM,OAAO,IAAI,aAAa;AAC9B,MAAI,OAAO,GAAG;AACZ,UAAM,IAAI,MAAM,mDAAmD,IAAI,GAAG;AAAA,EAC5E;AACA,MAAI,IAAI,OAAO,WAAW,GAAG;AAC3B,UAAM,IAAI,MAAM,qCAAqC;AAAA,EACvD;AACA,SAAO;AAAA,IACL,MAAM,IAAI;AAAA,IACV,GAAI,IAAI,iBAAiB,SAAY,CAAC,IAAI,EAAE,cAAc,IAAI,aAAa;AAAA,IAC3E,YAAY,IAAI,OAAO,IAAI,CAAC,SAAS,EAAE,KAAK,aAAa,IAAI,WAAW,GAAG,KAAK,IAAI,EAAE;AAAA,IACtF,MAAM,MAAM,OAA4B;AACtC,YAAM,UAAU,MAAM,QAAQ;AAAA,QAC5B,MAAM,KAAK,EAAE,QAAQ,KAAK,GAAG,CAAC,GAAG,QAAQ,IAAI,SAAS,EAAE,GAAG,OAAO,IAAI,CAAC,CAAC;AAAA,MAC1E;AACA,YAAM,WAA8B,QAAQ;AAAA,QAAI,CAAC,GAAG,QAClD,EAAE,WAAW,cACT,EAAE,QACF,EAAE,OAAO,GAAG,IAAI,IAAI,OAAO,GAAG,IAAI,cAAc,MAAM,WAAW,OAAO,EAAE,MAAM,EAAE;AAAA,MACxF;AAEA,YAAM,MAAM,uBAAuB,UAAU,IAAI,QAAQ,IAAI,OAAO;AACpE,aAAO,EAAE,WAAW,IAAI,WAAW,YAAY,IAAI,cAAc,OAAO,IAAI,UAAU;AAAA,IACxF;AAAA,EACF;AACF;","names":["aggregateJudgeVerdicts"]}
|
|
@@ -1,28 +1,9 @@
|
|
|
1
1
|
/**
|
|
2
|
-
*
|
|
3
|
-
*
|
|
4
|
-
*
|
|
5
|
-
*
|
|
6
|
-
*
|
|
7
|
-
* believed").
|
|
8
|
-
*
|
|
9
|
-
* Three checks, each fail-loud and named in `trustReasons`:
|
|
10
|
-
* (1) inter-rater reliability over the corpus ≥ `irrFloor` — raters that
|
|
11
|
-
* disagree no better than chance carry no signal to optimize against.
|
|
12
|
-
* (2) per-item rater spread ≤ `spreadCeiling` — for EACH item, raters must
|
|
13
|
-
* converge on THAT item.
|
|
14
|
-
* (3) surviving raters per item ≥ `minSurvivors` — a mean over one or two
|
|
15
|
-
* raters is an anecdote, not an ensemble.
|
|
16
|
-
*
|
|
17
|
-
* CRITICAL metric semantics — per-item spread is rater disagreement about the
|
|
18
|
-
* SAME item: `max(score) − min(score)` across the raters that scored THAT item
|
|
19
|
-
* (max over its dimensions), never pooled across different items or across the
|
|
20
|
-
* baseline/candidate sides. Pooling reads a genuine quality gap BETWEEN items as
|
|
21
|
-
* "the raters split" and so trips the gate exactly when the finding is largest —
|
|
22
|
-
* the failure mode the after-gate exists to prevent. The corpus IRR (check 1)
|
|
23
|
-
* leans on the substrate's `interRaterReliability`, whose expected-disagreement
|
|
24
|
-
* denominator already pools across items, so genuine item-to-item variation
|
|
25
|
-
* RAISES reliability rather than lowering it.
|
|
2
|
+
* Summarize rater agreement, within-item spread, and surviving-judge coverage.
|
|
3
|
+
* Consumers choose the thresholds that a trustworthy result must satisfy.
|
|
4
|
+
* Agreement does not measure evaluator errors against independent controls.
|
|
5
|
+
* Use agent-eval/meta-eval's auditEvaluator when the consumer needs that separate evidence.
|
|
6
|
+
* Spread stays within each item so differences in task quality cannot mimic rater disagreement.
|
|
26
7
|
*/
|
|
27
8
|
import { type JudgeVerdict } from '@tangle-network/agent-eval';
|
|
28
9
|
/** One item's raters: the per-judge verdicts {@link aggregateJudgeVerdicts}
|
|
@@ -34,11 +15,9 @@ export interface TrustItem<D extends string = string> {
|
|
|
34
15
|
* (`perDimension: null`) is dropped before spread/IRR, never folded as 0. */
|
|
35
16
|
verdicts: readonly JudgeVerdict<D>[];
|
|
36
17
|
}
|
|
37
|
-
/**
|
|
38
|
-
* conservative after-gate bar. */
|
|
18
|
+
/** Configurable agreement and coverage thresholds for {@link trustVerdicts}. */
|
|
39
19
|
export interface TrustThresholds {
|
|
40
|
-
/** Minimum corpus inter-rater reliability (Krippendorff-style α).
|
|
41
|
-
* the raters agree no better than chance. Default 0.2. */
|
|
20
|
+
/** Minimum corpus inter-rater reliability (Krippendorff-style α). Default 0.2. */
|
|
42
21
|
irrFloor?: number;
|
|
43
22
|
/** Maximum per-item rater spread (`max − min` over a single item's surviving
|
|
44
23
|
* raters, across its dimensions). Above this the raters split ON THAT ITEM.
|
|
@@ -63,14 +42,12 @@ export interface TrustVerdict {
|
|
|
63
42
|
perItemSpread: Record<string, number>;
|
|
64
43
|
}
|
|
65
44
|
/**
|
|
66
|
-
*
|
|
67
|
-
* believe a lift computed from them. Pure: no LLM, no I/O, no clock, no random —
|
|
45
|
+
* Check an ensemble against its configured agreement and coverage thresholds. Pure: no LLM, no I/O, no clock, no random —
|
|
68
46
|
* the same `items` + `thresholds` always yield the same verdict.
|
|
69
47
|
*
|
|
70
48
|
* Sibling to {@link aggregateJudgeVerdicts}: that reduces ONE item's raters to a
|
|
71
|
-
* composite; this
|
|
72
|
-
*
|
|
73
|
-
* reporting any lift over their scores.
|
|
49
|
+
* composite; this summarizes agreement across the supplied items.
|
|
50
|
+
* A passing result does not establish evaluator accuracy or authorize a release.
|
|
74
51
|
*
|
|
75
52
|
* @throws if `items` is empty — an empty corpus has no measurable trust, and a
|
|
76
53
|
* silent `trustworthy: true` over zero evidence is the exact lie the gate
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@tangle-network/agent-app",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.48.2",
|
|
4
4
|
"packageManager": "pnpm@11.24.0",
|
|
5
5
|
"description": "Build agent applications with typed chat, tools, sandboxes, integrations, billing, and evaluation.",
|
|
6
6
|
"keywords": [
|
|
@@ -535,15 +535,15 @@
|
|
|
535
535
|
"@storybook/react-vite": "^10.5.10",
|
|
536
536
|
"@tailwindcss/postcss": "^4.3.3",
|
|
537
537
|
"@tangle-network/agent-docs": "0.2.1",
|
|
538
|
-
"@tangle-network/agent-eval": "0.
|
|
538
|
+
"@tangle-network/agent-eval": "0.182.0",
|
|
539
539
|
"@tangle-network/agent-gateway": "0.10.0",
|
|
540
|
-
"@tangle-network/agent-integrations": "0.
|
|
541
|
-
"@tangle-network/agent-interface": "2.6.
|
|
542
|
-
"@tangle-network/agent-knowledge": "
|
|
543
|
-
"@tangle-network/agent-profile-materialize": "0.
|
|
544
|
-
"@tangle-network/agent-runtime": "0.
|
|
540
|
+
"@tangle-network/agent-integrations": "0.54.0",
|
|
541
|
+
"@tangle-network/agent-interface": "2.6.1",
|
|
542
|
+
"@tangle-network/agent-knowledge": "17.0.2",
|
|
543
|
+
"@tangle-network/agent-profile-materialize": "0.20.1",
|
|
544
|
+
"@tangle-network/agent-runtime": "0.231.1",
|
|
545
545
|
"@tangle-network/brand": "1.5.0",
|
|
546
|
-
"@tangle-network/sandbox": "0.
|
|
546
|
+
"@tangle-network/sandbox": "0.40.2",
|
|
547
547
|
"@tangle-network/sandbox-ui": "0.113.3",
|
|
548
548
|
"@tangle-network/ui": "^11.8.0",
|
|
549
549
|
"@testing-library/dom": "^10.4.1",
|
|
@@ -596,14 +596,14 @@
|
|
|
596
596
|
"peerDependencies": {
|
|
597
597
|
"@firecrawl/pdf-inspector-wasm": ">=0.1.3",
|
|
598
598
|
"@huggingface/transformers": ">=3",
|
|
599
|
-
"@tangle-network/agent-eval": ">=0.
|
|
600
|
-
"@tangle-network/agent-integrations": ">=0.53.55 <0.
|
|
601
|
-
"@tangle-network/agent-interface": "^2.6.
|
|
602
|
-
"@tangle-network/agent-knowledge": "^
|
|
603
|
-
"@tangle-network/agent-profile-materialize": ">=0.
|
|
604
|
-
"@tangle-network/agent-runtime": ">=0.
|
|
599
|
+
"@tangle-network/agent-eval": ">=0.182.0 <0.183.0",
|
|
600
|
+
"@tangle-network/agent-integrations": ">=0.53.55 <0.55.0",
|
|
601
|
+
"@tangle-network/agent-interface": "^2.6.1",
|
|
602
|
+
"@tangle-network/agent-knowledge": "^17.0.2",
|
|
603
|
+
"@tangle-network/agent-profile-materialize": ">=0.20.1 <0.21.0",
|
|
604
|
+
"@tangle-network/agent-runtime": ">=0.231.1 <0.232.0",
|
|
605
605
|
"@tangle-network/brand": ">=1.5.0",
|
|
606
|
-
"@tangle-network/sandbox": ">=0.
|
|
606
|
+
"@tangle-network/sandbox": ">=0.40.2 <0.41.0",
|
|
607
607
|
"@tangle-network/sandbox-ui": ">=0.113.3 <0.114.0",
|
|
608
608
|
"@tangle-network/ui": ">=11.6.0 <12.0.0",
|
|
609
609
|
"@tiptap/core": ">=3.28.0 <4.0.0",
|