@agentproto/eval 0.2.11 → 0.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -193,6 +193,43 @@ value >= threshold` (`threshold` defaults to `0.5`).
193
193
  A real adapter wiring `JudgeFn` up to an agent session or the supervisor's
194
194
  judge-gate is a documented follow-up — not built in this package.
195
195
 
196
+ ## Style scorers
197
+
198
+ Five more `eval.*` TOOL contracts, under `src/style/` (see DESIGN.md §7 in the
199
+ `factory` plan for the production bands/gates these back). Two are
200
+ deterministic — bundled into `styleScorersProvider` exactly like
201
+ `evalScorersProvider` — and three are model-backed, each built by a
202
+ `make*Driver(...)` factory that closes over an injected, vendor-neutral
203
+ capability. No LLM SDK, no network dependency, same discipline as
204
+ `eval.llm-judge`.
205
+
206
+ | Tool id | Input | Behavior |
207
+ | ------------------------ | --------------------------------------- | -------- |
208
+ | `eval.text-stats` | `{ text, thresholds? }` | Deterministic. Computes bullet-line ratio, first-person ratio, question rate, and mean sentence length (+ `inBand`) over French text. `passed` is derived from caller-supplied `thresholds` (`maxBulletsRatio` default 0.02, `minFirstPersonRatio` default 0.6, `lengthBand` default `{min: 5, max: 30}` words) — this tool owns no fixed gate. |
209
+ | `eval.lexicon-hit-rate` | `{ text, lexicon (min 1), threshold? }` | Deterministic. Fraction of `lexicon` terms present as whole words in `text`, Unicode-aware (accented terms like `écrire` match correctly; ASCII `\b` would miss them). Text and terms are NFC-normalized before matching. `passed = hitRate >= threshold` (default 0.5). Build a signature lexicon offline with the pure helper `extractLexicon(corpusTexts, { top?, minLen?, background? })` — with `background`, terms are ranked by a log-odds ratio against that corpus instead of raw frequency (a simple frequency-ratio estimator, not the full variance-weighted informative-Dirichlet estimator). |
210
+ | `eval.style-pairwise` | `{ reference, a, b, criteria }` | Model-backed via `makeStylePairwiseDriver(judge: JudgeFn)` — reuses the `eval.llm-judge` `JudgeFn` seam (the raw verdict is validated with `judgeVerdictSchema`; a malformed verdict fails the score instead of propagating `NaN`). `value` encodes preference (1 = `a` wins, 0 = `b` wins, 0.5 = tie). The pure helper `pairwiseWinRate(verdicts)` aggregates `PairwiseVerdict[]` (each carrying a required `item` id) into `{ winRate, kappa, n, nNormal, nSwapped, balanced }`: `winRate` averages the normal-order and swapped-order rates (falling back to whichever order is present, flagged via `balanced: false`, when one is missing entirely); `kappa` is Cohen's kappa between the two most-represented judges, joined by `item` (so a normal+swapped pair on the same item is one observation, not two) — it is `null`, not a default `1`, when there are fewer than two judges or fewer than two items in common, and a `null` must be treated as a gate FAILURE, not skipped. |
211
+ | `eval.style-embedding` | `{ candidate, references[] (min 1) }` | Model-backed via `makeStyleEmbeddingDriver(embed: EmbedFn)`, `EmbedFn = (texts) => Promise<number[][]>`. `value` = cosine similarity of `candidate` to the `references` centroid, clamped to `[0, 1]` via `max(0, cosine)` — NOT remapped via `(cosine + 1) / 2`, which put an uninformative orthogonal candidate at the same 0.5 as the default pass threshold. A candidate equal to the centroid scores 1; orthogonal or opposing candidates score 0. Heterogeneous embedding dimensions (mismatched candidate/reference vectors) fail the score with a rationale instead of silently padding/truncating. Pure helper: `cosineToCentroid(candidate, references)` (throws `EmbeddingDimensionError` on dimension mismatch — callers at the tool boundary must catch it). |
212
+ | `eval.outline-fidelity` | `{ outline, answer }` | Model-backed via `makeOutlineFidelityDriver(judge: JudgeFn)`. `value` = judged outline coverage; `passed = value >= 0.95` — a **fixed** gate, never overridden by the judge's own `passed` (unlike `eval.llm-judge`'s threshold semantics). The raw verdict is validated with `judgeVerdictSchema`; a malformed verdict fails the score. |
213
+
214
+ French text conventions used by `eval.text-stats`: first-person markers are
215
+ `je`, `j'`, `moi`, `mon`/`ma`/`mes`, `me`/`m'`, `nous`, `notre`/`nos`,
216
+ `mien(ne)(s)`; a bullet line starts with `-`, `*`, `•`, or a numbered marker
217
+ like `1.` followed by whitespace (`1.5 million` and `-42 degrés` are prose,
218
+ not bullets). Sentence splitting guards against common French abbreviations
219
+ (`M.`, `Mme`, `Dr`, `etc.`, `cf.`, `p. ex.`) and isolated capital initials
220
+ (`J. Dupont`) so those periods are not treated as sentence ends.
221
+
222
+ ```ts
223
+ import { runTool } from "@agentproto/driver"
224
+ import { textStatsTool, styleScorersProvider } from "@agentproto/eval"
225
+
226
+ const score = await runTool({
227
+ tool: textStatsTool,
228
+ candidates: [styleScorersProvider],
229
+ input: { text: "Je pense que ceci illustre bien mon propos.", thresholds: { minFirstPersonRatio: 0.5 } },
230
+ })
231
+ ```
232
+
196
233
  ## As a CI gate
197
234
 
198
235
  `toVitest` turns a suite into vitest test registrations — one `it(caseId)` per
@@ -230,6 +267,22 @@ toVitest(
230
267
  - `llmJudge` — convenience: build a ready-to-use `ScorerBinding` around a judge
231
268
  - `JudgeFn`, `JudgeVerdict`, `judgeVerdictSchema`, `LlmJudgeInput`,
232
269
  `MakeLlmJudgeDriverOptions`, `LlmJudgeBinding`
270
+ - `textStatsTool` / `textStatsImpl` — plus pure helpers `bulletsRatio`,
271
+ `firstPersonRatio`, `questionRate`, `meanSentenceLength`, `splitSentences`,
272
+ `computeTextStats` (`TextStats`, `LengthBand`)
273
+ - `lexiconHitRateTool` / `lexiconHitRateImpl` — plus `extractLexicon`
274
+ (`ExtractLexiconOptions`)
275
+ - `styleScorersProvider` — the builtin PROVIDER bundling the two
276
+ deterministic style scorers
277
+ - `stylePairwiseTool`, `makeStylePairwiseDriver`, `pairwiseWinRate`
278
+ (`StylePairwiseInput`, `PairwiseWinner`, `PairwiseVerdict`,
279
+ `PairwiseWinRateResult`)
280
+ - `styleEmbeddingTool`, `makeStyleEmbeddingDriver`, `cosineToCentroid`,
281
+ `EmbeddingDimensionError`
282
+ (`StyleEmbeddingInput`, `EmbedFn`, `MakeStyleEmbeddingDriverOptions`)
283
+ - `outlineFidelityTool`, `makeOutlineFidelityDriver` (`OutlineFidelityInput`)
284
+ - `parseVerdict` — shared `style/` helper: validates a raw judge return value
285
+ against `judgeVerdictSchema`, returning `null` on malformed input
233
286
 
234
287
  ## License
235
288
 
package/dist/index.d.ts CHANGED
@@ -409,6 +409,300 @@ type ToVitestOptions<I, A> = RunEvalOptions<I, A>;
409
409
  */
410
410
  declare function toVitest<I, A>(suite: TypedEvalSuite<I, A>, opts: ToVitestOptions<I, A>, hooks: VitestHooks): void;
411
411
 
412
+ /**
413
+ * Split text into non-empty sentences on `.`/`!`/`?` boundaries, guarding
414
+ * against abbreviations (`M.`, `Mme`, `Dr`, `etc.`, `cf.`, `p. ex.`) and
415
+ * isolated capital initials (`J. Dupont`) so those periods don't count as
416
+ * sentence ends.
417
+ */
418
+ declare function splitSentences(text: string): string[];
419
+ /** Fraction of lines that open with a bullet marker (`-`, `*`, `•`, or `1.`) followed by whitespace — a marker glued to the next character (`1.5 million`, `-42 degrés`) is prose, not a bullet. */
420
+ declare function bulletsRatio(text: string): number;
421
+ /** Fraction of sentences carrying a French first-person marker. */
422
+ declare function firstPersonRatio(text: string): number;
423
+ /** Fraction of sentences ending in `?`. */
424
+ declare function questionRate(text: string): number;
425
+ /** Mean number of words per sentence. */
426
+ declare function meanSentenceLength(text: string): number;
427
+ interface TextStats {
428
+ readonly bulletsRatio: number;
429
+ readonly firstPersonRatio: number;
430
+ readonly questionRate: number;
431
+ readonly meanSentenceLength: number;
432
+ readonly inBand: boolean;
433
+ }
434
+ interface LengthBand {
435
+ readonly min: number;
436
+ readonly max: number;
437
+ }
438
+ /** Compute every metric in one pass. */
439
+ declare function computeTextStats(text: string, band?: LengthBand): TextStats;
440
+ declare const textStatsTool: _agentproto_tool.ToolHandle<{
441
+ text: string;
442
+ thresholds?: {
443
+ maxBulletsRatio?: number | undefined;
444
+ minFirstPersonRatio?: number | undefined;
445
+ lengthBand?: {
446
+ min: number;
447
+ max: number;
448
+ } | undefined;
449
+ } | undefined;
450
+ }, {
451
+ value: number;
452
+ passed: boolean;
453
+ label: string;
454
+ rationale?: string | undefined;
455
+ }, _agentproto_tool.ToolContext>;
456
+ declare const textStatsImpl: _agentproto_driver.ToolImplementation<{
457
+ text: string;
458
+ thresholds?: {
459
+ maxBulletsRatio?: number | undefined;
460
+ minFirstPersonRatio?: number | undefined;
461
+ lengthBand?: {
462
+ min: number;
463
+ max: number;
464
+ } | undefined;
465
+ } | undefined;
466
+ }, {
467
+ value: number;
468
+ passed: boolean;
469
+ label: string;
470
+ rationale?: string | undefined;
471
+ }, _agentproto_tool.ToolContext>;
472
+
473
+ interface ExtractLexiconOptions {
474
+ /** Max number of terms to return. Default 20. */
475
+ readonly top?: number;
476
+ /** Minimum term length (characters) to keep. Default 3. */
477
+ readonly minLen?: number;
478
+ /**
479
+ * Optional background corpus. When provided, terms are ranked by an
480
+ * add-one-smoothed log-odds ratio of their rate in `corpusTexts` versus
481
+ * their rate in `background`, surfacing terms disproportionately frequent
482
+ * in the foreground corpus rather than merely frequent overall. This is a
483
+ * simple frequency-ratio estimator, NOT the full variance-weighted
484
+ * informative-Dirichlet log-odds estimator (Monroe et al. 2008) — it has
485
+ * no correction for small counts beyond add-one smoothing, so rare terms
486
+ * in a small background corpus can score unstably high. When omitted,
487
+ * ranking falls back to raw frequency in `corpusTexts`.
488
+ */
489
+ readonly background?: readonly string[];
490
+ }
491
+ /**
492
+ * Pure helper: extract a corpus's signature lexicon — the most frequent
493
+ * non-stopword terms across `corpusTexts` — for use as `eval.lexicon-hit-rate`
494
+ * input. Offline / caller-side; this tool never calls it itself.
495
+ */
496
+ declare function extractLexicon(corpusTexts: readonly string[], opts?: ExtractLexiconOptions): string[];
497
+ declare const lexiconHitRateTool: _agentproto_tool.ToolHandle<{
498
+ text: string;
499
+ lexicon: string[];
500
+ threshold?: number | undefined;
501
+ }, {
502
+ value: number;
503
+ passed: boolean;
504
+ label: string;
505
+ rationale?: string | undefined;
506
+ }, _agentproto_tool.ToolContext>;
507
+ declare const lexiconHitRateImpl: _agentproto_driver.ToolImplementation<{
508
+ text: string;
509
+ lexicon: string[];
510
+ threshold?: number | undefined;
511
+ }, {
512
+ value: number;
513
+ passed: boolean;
514
+ label: string;
515
+ rationale?: string | undefined;
516
+ }, _agentproto_tool.ToolContext>;
517
+
518
+ /**
519
+ * `eval.style-pairwise` — model-backed A/B scorer: which of `a`/`b` reads
520
+ * closer to `reference` under `criteria`. Reuses the existing {@link JudgeFn}
521
+ * seam from judge.ts (no new judge type) — the driver hands the judge
522
+ * `{reference, a, b}` as `output` and reads back a single verdict whose
523
+ * `value` encodes preference: 1 = `a` wins, 0 = `b` wins, 0.5 = tie.
524
+ *
525
+ * `pairwiseWinRate` below is a SEPARATE pure helper: callers run this tool
526
+ * multiple times (varying which side is presented first, and across
527
+ * multiple judges) and feed the resulting {winner} outcomes to it to get an
528
+ * order-debiased win rate plus inter-judge agreement (Cohen's kappa).
529
+ */
530
+ interface StylePairwiseInput {
531
+ readonly reference: string;
532
+ readonly a: string;
533
+ readonly b: string;
534
+ readonly criteria: string;
535
+ }
536
+ declare const stylePairwiseTool: _agentproto_tool.ToolHandle<{
537
+ reference: string;
538
+ a: string;
539
+ b: string;
540
+ criteria: string;
541
+ }, {
542
+ value: number;
543
+ passed: boolean;
544
+ label: string;
545
+ rationale?: string | undefined;
546
+ }, _agentproto_tool.ToolContext>;
547
+ /**
548
+ * Build a DRIVER that implements `eval.style-pairwise` by delegating to
549
+ * `judge` — the same {@link JudgeFn} shape as `eval.llm-judge`, reused rather
550
+ * than threading a bespoke comparison type through this package.
551
+ */
552
+ declare function makeStylePairwiseDriver(judge: JudgeFn): DriverHandle;
553
+ type PairwiseWinner = "a" | "b" | "tie";
554
+ /**
555
+ * One collected pairwise outcome. `order` records which physical position
556
+ * `a` was presented in when this verdict was produced; `winner` is the RAW
557
+ * outcome as reported for that presentation (before un-swapping) — e.g. a
558
+ * judge exhibiting pure position bias always picks the same physical slot,
559
+ * which shows up here as `winner` tracking `order` rather than content.
560
+ * `pairwiseWinRate` un-swaps `winner` back to the logical a/b before
561
+ * aggregating, which is what neutralizes that bias.
562
+ */
563
+ interface PairwiseVerdict {
564
+ /** Identifies which item this verdict scored — required so verdicts from different items are never zipped together. */
565
+ readonly item: string;
566
+ /** Judge identity — distinguishes the (at least two) judges being compared. */
567
+ readonly judge: string;
568
+ /** Physical presentation order this verdict was collected under. */
569
+ readonly order: "normal" | "swapped";
570
+ /** Raw winner as reported under `order` (not yet un-swapped). */
571
+ readonly winner: PairwiseWinner;
572
+ }
573
+ interface PairwiseWinRateResult {
574
+ /** Logical a-win rate (ties count as 0.5), after order de-biasing. Averages the normal-order and swapped-order rates when both are present. */
575
+ readonly winRate: number;
576
+ /**
577
+ * Cohen's kappa agreement between the two most-represented judges, computed
578
+ * over items both judges rated. `null` when there are fewer than two
579
+ * judges, or fewer than two items in common — a `null` here means the
580
+ * agreement gate has no evidence and any threshold check against it (e.g.
581
+ * "kappa >= 0.4") must be treated as a FAILURE, not skipped.
582
+ */
583
+ readonly kappa: number | null;
584
+ /** Number of verdicts folded in. */
585
+ readonly n: number;
586
+ /** Number of verdicts collected under `order: "normal"`. */
587
+ readonly nNormal: number;
588
+ /** Number of verdicts collected under `order: "swapped"`. */
589
+ readonly nSwapped: number;
590
+ /** `false` when one of the two presentation orders has zero verdicts — `winRate` then reflects only the order that is present, and is not order-debiased. */
591
+ readonly balanced: boolean;
592
+ }
593
+ /**
594
+ * Aggregate collected {@link PairwiseVerdict}s into an order-debiased win
595
+ * rate and inter-judge Cohen's kappa.
596
+ *
597
+ * Win rate: each verdict is un-swapped against its `order` (neutralizing
598
+ * position bias), then the normal-order rate and swapped-order rate are
599
+ * averaged — not a flat average over all verdicts — so a judge that always
600
+ * picks the physically-first slot nets out near 0.5 even if one order was
601
+ * sampled more than the other. When only one order was collected at all,
602
+ * `balanced` is `false` and `winRate` falls back to that order's rate alone.
603
+ *
604
+ * Kappa: verdicts are joined by `item` — for each of the two
605
+ * most-represented judges, at most one (first-seen, un-swapped) winner is
606
+ * kept per item, so a normal+swapped pair on the same item is one
607
+ * observation, not two. Kappa is then computed over the items both judges
608
+ * rated; see {@link PairwiseWinRateResult.kappa} for the `null` cases.
609
+ */
610
+ declare function pairwiseWinRate(verdicts: readonly PairwiseVerdict[]): PairwiseWinRateResult;
611
+
612
+ /** Thrown by {@link centroid} / {@link cosineToCentroid} on mismatched embedding dimensions — callers crossing the tool boundary (the driver) must catch this and fail the `Score`, never let it propagate as a thrown error out of the tool contract. */
613
+ declare class EmbeddingDimensionError extends RangeError {
614
+ }
615
+ /**
616
+ * Pure helper: cosine similarity between `candidate` and the centroid of
617
+ * `references`, clamped to `[0, 1]` via `max(0, cosine)` — NOT remapped from
618
+ * `[-1, 1]` via `(cosine + 1) / 2`. That remap put an orthogonal candidate
619
+ * (cosine 0, no signal) at 0.5, the same value as the default `passed`
620
+ * threshold, so an uninformative embedding silently cleared the gate. Under
621
+ * `max(0, cosine)`, orthogonal and opposing candidates both score 0. A
622
+ * candidate equal to the centroid still returns exactly 1.
623
+ *
624
+ * Throws {@link EmbeddingDimensionError} if `candidate` and the references'
625
+ * centroid have different dimensionality — callers at the tool boundary must
626
+ * catch this (see {@link makeStyleEmbeddingDriver}) rather than let it cross
627
+ * into a thrown error from the TOOL contract.
628
+ */
629
+ declare function cosineToCentroid(candidate: readonly number[], references: readonly (readonly number[])[]): number;
630
+ interface StyleEmbeddingInput {
631
+ readonly candidate: string;
632
+ readonly references: readonly string[];
633
+ }
634
+ declare const styleEmbeddingTool: _agentproto_tool.ToolHandle<{
635
+ candidate: string;
636
+ references: string[];
637
+ }, {
638
+ value: number;
639
+ passed: boolean;
640
+ label: string;
641
+ rationale?: string | undefined;
642
+ }, _agentproto_tool.ToolContext>;
643
+ /**
644
+ * The seam a real embedding capability satisfies: given a batch of texts,
645
+ * return one embedding vector per text, same order. Deliberately no model
646
+ * SDK / network types here — vendor-neutral, like {@link JudgeFn}.
647
+ */
648
+ type EmbedFn = (texts: readonly string[]) => Promise<number[][]>;
649
+ interface MakeStyleEmbeddingDriverOptions {
650
+ /** Minimum value to count as passed. Default 0.5. */
651
+ readonly threshold?: number;
652
+ }
653
+ /**
654
+ * Build a DRIVER that implements `eval.style-embedding` by delegating to
655
+ * `embed`. Embeds `[candidate, ...references]` in a single batch call so a
656
+ * remote embedding backend sees one request per score.
657
+ */
658
+ declare function makeStyleEmbeddingDriver(embed: EmbedFn, opts?: MakeStyleEmbeddingDriverOptions): DriverHandle;
659
+
660
+ /**
661
+ * Validate a raw judge return value against {@link judgeVerdictSchema}. An
662
+ * injected `JudgeFn` is caller-supplied and only nominally typed — a
663
+ * misbehaving judge can hand back `NaN`, a missing `value`, or any other
664
+ * malformed shape at runtime. Callers use this to fail the `Score` with a
665
+ * rationale instead of letting a bad verdict propagate as `NaN`/`undefined`
666
+ * downstream.
667
+ */
668
+ declare function parseVerdict(raw: unknown): JudgeVerdict | null;
669
+
670
+ /**
671
+ * `eval.outline-fidelity` — model-backed scorer: does `answer` cover every
672
+ * point in `outline` without adding unsourced facts. Reuses the existing
673
+ * {@link JudgeFn} seam (`output = answer`, `expected = outline`) — no new
674
+ * judge type. Unlike `eval.llm-judge`, `passed` here is a FIXED gate
675
+ * (`value >= 0.95`, per DESIGN.md §7), never overridden by the judge's own
676
+ * `passed`.
677
+ */
678
+ interface OutlineFidelityInput {
679
+ readonly outline: string;
680
+ readonly answer: string;
681
+ }
682
+ declare const outlineFidelityTool: _agentproto_tool.ToolHandle<{
683
+ outline: string;
684
+ answer: string;
685
+ }, {
686
+ value: number;
687
+ passed: boolean;
688
+ label: string;
689
+ rationale?: string | undefined;
690
+ }, _agentproto_tool.ToolContext>;
691
+ /**
692
+ * Build a DRIVER that implements `eval.outline-fidelity` by delegating to
693
+ * `judge`, the same {@link JudgeFn} shape as `eval.llm-judge`.
694
+ */
695
+ declare function makeOutlineFidelityDriver(judge: JudgeFn): DriverHandle;
696
+
697
+ /**
698
+ * Builtin AIP-30 PROVIDER bundling the two deterministic style scorers —
699
+ * the pair DESIGN.md §7 names as the only style scorers safe to sample in
700
+ * prod (free, zero network). The three model-backed style tools have no
701
+ * static provider: each needs an injected judge/embed via its
702
+ * `make*Driver(...)` factory, same pattern as `eval.llm-judge`.
703
+ */
704
+ declare const styleScorersProvider: _agentproto_driver.DriverHandle;
705
+
412
706
  /**
413
707
  * @agentproto/eval — deterministic reference scorers.
414
708
  *
@@ -446,4 +740,4 @@ declare const SPEC_VERSION: "1.0.0-alpha";
446
740
  */
447
741
  declare const evalScorersProvider: _agentproto_driver.DriverHandle;
448
742
 
449
- export { type BoundScorer, type CaseReport, type CaseScore, EVAL_EVENT_SCHEMA, type EvalCase, type EvalEvent, type EvalReport, type EvalSuite, type ExpectApi, type JsonObject, type JsonValue, type JudgeFn, type JudgeVerdict, type LlmJudgeBinding, type LlmJudgeInput, type MakeLlmJudgeDriverOptions, type MinimalJsonSchema, type RunEvalOptions, SPEC_NAME, SPEC_VERSION, type Score, type ScorerBinding, type ScorerInputContext, type ToVitestOptions, type TypedEvalSuite, type VitestHooks, bindScorer, evalScorersProvider, exactMatchImpl, exactMatchTool, jsonSchemaValidImpl, jsonSchemaValidTool, jsonValueSchema, judgeVerdictSchema, latencyBudgetImpl, latencyBudgetTool, llmJudge, llmJudgeTool, makeLlmJudgeDriver, regexMatchImpl, regexMatchTool, runEval, scoreSchema, toVitest };
743
+ export { type BoundScorer, type CaseReport, type CaseScore, EVAL_EVENT_SCHEMA, type EmbedFn, EmbeddingDimensionError, type EvalCase, type EvalEvent, type EvalReport, type EvalSuite, type ExpectApi, type ExtractLexiconOptions, type JsonObject, type JsonValue, type JudgeFn, type JudgeVerdict, type LengthBand, type LlmJudgeBinding, type LlmJudgeInput, type MakeLlmJudgeDriverOptions, type MakeStyleEmbeddingDriverOptions, type MinimalJsonSchema, type OutlineFidelityInput, type PairwiseVerdict, type PairwiseWinRateResult, type PairwiseWinner, type RunEvalOptions, SPEC_NAME, SPEC_VERSION, type Score, type ScorerBinding, type ScorerInputContext, type StyleEmbeddingInput, type StylePairwiseInput, type TextStats, type ToVitestOptions, type TypedEvalSuite, type VitestHooks, bindScorer, bulletsRatio, computeTextStats, cosineToCentroid, evalScorersProvider, exactMatchImpl, exactMatchTool, extractLexicon, firstPersonRatio, jsonSchemaValidImpl, jsonSchemaValidTool, jsonValueSchema, judgeVerdictSchema, latencyBudgetImpl, latencyBudgetTool, lexiconHitRateImpl, lexiconHitRateTool, llmJudge, llmJudgeTool, makeLlmJudgeDriver, makeOutlineFidelityDriver, makeStyleEmbeddingDriver, makeStylePairwiseDriver, meanSentenceLength, outlineFidelityTool, pairwiseWinRate, parseVerdict, questionRate, regexMatchImpl, regexMatchTool, runEval, scoreSchema, splitSentences, styleEmbeddingTool, stylePairwiseTool, styleScorersProvider, textStatsImpl, textStatsTool, toVitest };