@remnic/bench 9.6.19 → 9.6.21
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +55 -1
- package/dist/index.d.ts +107 -40
- package/dist/index.js +1396 -725
- package/package.json +3 -3
package/README.md
CHANGED
|
@@ -70,9 +70,63 @@ remnic bench run --quick memcorrect-v1 --adapter mcp --mcp-demo \
|
|
|
70
70
|
--judge-provider openai --judge-model gpt-5.6
|
|
71
71
|
```
|
|
72
72
|
|
|
73
|
+
The credit-backed Codex CLI path is a distinct provider and model namespace.
|
|
74
|
+
Use `gpt-5.6-luna` for bulk responder and internal work and
|
|
75
|
+
`gpt-5.6-terra` for quality-critical judging. The exact API id `gpt-5.6`
|
|
76
|
+
above is not a CLI alias. Confirm the authenticated catalog before a run with
|
|
77
|
+
`codex debug models`. `gpt-5.6-sol` is opt-in only and is disabled in the
|
|
78
|
+
bounded plan.
|
|
79
|
+
|
|
80
|
+
Each Codex completion is a fresh, non-interactive `codex exec` in a new empty
|
|
81
|
+
temporary workspace. It ignores user configuration and project rules,
|
|
82
|
+
disables hooks, and keeps no session. The sandbox is read-only, approvals are
|
|
83
|
+
denied, and the benchmark prompt instructs the model not to use tools. The
|
|
84
|
+
Build Week plan uses normal service, not fast mode.
|
|
85
|
+
|
|
86
|
+
The Build Week grant has 2,473 Codex credits. Use the account exclusively for
|
|
87
|
+
this one harness process during a bounded run because Codex CLI exposes no
|
|
88
|
+
machine-readable account balance. Bounded mode also requires `codex login
|
|
89
|
+
status` to report ChatGPT authentication. A 473-credit safety reserve leaves
|
|
90
|
+
2,000 usable credits. Configure the atomic completed-turn ledger and guards,
|
|
91
|
+
then measure a quick task before choosing a workload bound:
|
|
92
|
+
|
|
93
|
+
```bash
|
|
94
|
+
export REMNIC_BENCH_CODEX_CREDIT_BUDGET=2473
|
|
95
|
+
export REMNIC_BENCH_CODEX_CREDIT_RESERVE=473
|
|
96
|
+
export REMNIC_BENCH_CODEX_CREDIT_LEDGER="$PWD/.bench-private/codex-credit-ledger.json"
|
|
97
|
+
|
|
98
|
+
remnic bench run --quick longmemeval \
|
|
99
|
+
--runtime-profile real \
|
|
100
|
+
--system-provider codex-cli --system-model gpt-5.6-luna \
|
|
101
|
+
--system-codex-reasoning-effort medium \
|
|
102
|
+
--internal-provider codex-cli --internal-model gpt-5.6-luna \
|
|
103
|
+
--internal-codex-reasoning-effort medium \
|
|
104
|
+
--judge-provider codex-cli --judge-model gpt-5.6-terra \
|
|
105
|
+
--judge-codex-reasoning-effort high
|
|
106
|
+
|
|
107
|
+
remnic bench run longmemeval \
|
|
108
|
+
--runtime-profile real --limit <LEDGER_DERIVED_LIMIT> \
|
|
109
|
+
--system-provider codex-cli --system-model gpt-5.6-luna \
|
|
110
|
+
--system-codex-reasoning-effort medium \
|
|
111
|
+
--internal-provider codex-cli --internal-model gpt-5.6-luna \
|
|
112
|
+
--internal-codex-reasoning-effort medium \
|
|
113
|
+
--judge-provider codex-cli --judge-model gpt-5.6-terra \
|
|
114
|
+
--judge-codex-reasoning-effort high
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
The placeholder is intentional. `--limit` and LoCoMo/MemoryAgentBench's
|
|
118
|
+
`--trial-limit` bound tasks, not token credits. Derive each next batch from
|
|
119
|
+
actual `turn.completed` JSONL usage. Stop dispatching at 2,000 spent; the
|
|
120
|
+
473-credit reserve absorbs only a final in-flight call whose exact cost becomes
|
|
121
|
+
known after completion. Missing exact terminal usage blocks the ledger pending
|
|
122
|
+
manual account reconciliation.
|
|
123
|
+
Rates per one million tokens are Luna: 25 input, 2.5 cached input, 150 output;
|
|
124
|
+
Terra: 62.5 input, 6.25 cached input, 375 output. A bounded result is a trial,
|
|
125
|
+
not a full leaderboard artifact.
|
|
126
|
+
|
|
73
127
|
Codex built and adversarially reviewed the Build Week adapter, Responses
|
|
74
128
|
provider, and report card. The underlying Remnic engine and original benchmark
|
|
75
|
-
harness are prior work. The evidence ledger,
|
|
129
|
+
harness are prior work. The evidence ledger, credit-backed frontier-run
|
|
76
130
|
placeholder, and release status live in the root [`HACKATHON.md`](../../HACKATHON.md).
|
|
77
131
|
|
|
78
132
|
The claimed judge path requires Node.js 22.12+. It is verified from source and
|
package/dist/index.d.ts
CHANGED
|
@@ -1506,6 +1506,36 @@ declare const BENCHMARK_RESULT_SCHEMA: {
|
|
|
1506
1506
|
};
|
|
1507
1507
|
};
|
|
1508
1508
|
|
|
1509
|
+
interface CodexCliNativeUsage {
|
|
1510
|
+
inputTokens: number;
|
|
1511
|
+
cachedInputTokens: number;
|
|
1512
|
+
outputTokens: number;
|
|
1513
|
+
reasoningOutputTokens: number;
|
|
1514
|
+
}
|
|
1515
|
+
interface CodexCreditReceiptScope extends CodexCliNativeUsage {
|
|
1516
|
+
calls: number;
|
|
1517
|
+
credits: number;
|
|
1518
|
+
models: Array<CodexCliNativeUsage & {
|
|
1519
|
+
model: string;
|
|
1520
|
+
calls: number;
|
|
1521
|
+
credits: number;
|
|
1522
|
+
}>;
|
|
1523
|
+
}
|
|
1524
|
+
interface CodexCreditReceipt {
|
|
1525
|
+
schemaVersion: 1;
|
|
1526
|
+
ledgerSha256: string;
|
|
1527
|
+
budgetCredits: number;
|
|
1528
|
+
reserveCredits: number;
|
|
1529
|
+
plannedSpendCeilingCredits: number;
|
|
1530
|
+
totalSpentCredits: number;
|
|
1531
|
+
totalRemainingCredits: number;
|
|
1532
|
+
blocked: boolean;
|
|
1533
|
+
cumulative: CodexCreditReceiptScope;
|
|
1534
|
+
run?: CodexCreditReceiptScope & {
|
|
1535
|
+
id: string;
|
|
1536
|
+
};
|
|
1537
|
+
}
|
|
1538
|
+
|
|
1509
1539
|
declare const BENCHMARK_REPRO_MANIFEST_FILENAME = "MANIFEST.json";
|
|
1510
1540
|
declare const BENCHMARK_REPRO_MANIFEST_SCHEMA_VERSION = 1;
|
|
1511
1541
|
interface BenchmarkReproManifestFile {
|
|
@@ -1591,6 +1621,7 @@ interface BenchmarkReproManifest {
|
|
|
1591
1621
|
}>;
|
|
1592
1622
|
datasets: BenchmarkReproManifestDataset[];
|
|
1593
1623
|
results: BenchmarkReproManifestResult[];
|
|
1624
|
+
codexCredit?: CodexCreditReceipt;
|
|
1594
1625
|
artifactHash: string;
|
|
1595
1626
|
}
|
|
1596
1627
|
interface BuildBenchmarkReproManifestOptions {
|
|
@@ -1878,6 +1909,66 @@ interface ClaudeCliProviderDeps {
|
|
|
1878
1909
|
}
|
|
1879
1910
|
declare function createClaudeCliProvider(config: ClaudeCliProviderConfig, deps?: ClaudeCliProviderDeps): LlmProvider;
|
|
1880
1911
|
|
|
1912
|
+
type StructuredJudgeErrorCode = "api_error" | "rate_limited" | "refusal" | "malformed_response" | "malformed_verdict" | "incomplete_response" | "transport_error" | "aborted";
|
|
1913
|
+
interface StructuredJudgeTelemetry {
|
|
1914
|
+
model: string;
|
|
1915
|
+
rubricVersion: string;
|
|
1916
|
+
inputTokens: number;
|
|
1917
|
+
outputTokens: number;
|
|
1918
|
+
latencyMs: number;
|
|
1919
|
+
errorCode?: StructuredJudgeErrorCode;
|
|
1920
|
+
httpStatus?: number;
|
|
1921
|
+
}
|
|
1922
|
+
interface StructuredJudgeVerdict {
|
|
1923
|
+
score: number;
|
|
1924
|
+
decision: "pass" | "partial" | "fail";
|
|
1925
|
+
reason: string;
|
|
1926
|
+
}
|
|
1927
|
+
type StructuredJudgeVerdictResult = {
|
|
1928
|
+
ok: true;
|
|
1929
|
+
verdict: StructuredJudgeVerdict;
|
|
1930
|
+
telemetry: StructuredJudgeTelemetry;
|
|
1931
|
+
} | {
|
|
1932
|
+
ok: false;
|
|
1933
|
+
error: {
|
|
1934
|
+
code: StructuredJudgeErrorCode;
|
|
1935
|
+
message: string;
|
|
1936
|
+
retryable: boolean;
|
|
1937
|
+
httpStatus?: number;
|
|
1938
|
+
};
|
|
1939
|
+
telemetry: StructuredJudgeTelemetry;
|
|
1940
|
+
};
|
|
1941
|
+
interface StructuredVerdictRequest {
|
|
1942
|
+
rubric: string;
|
|
1943
|
+
rubricVersion: string;
|
|
1944
|
+
input: string;
|
|
1945
|
+
signal?: AbortSignal;
|
|
1946
|
+
maxTokens?: number;
|
|
1947
|
+
}
|
|
1948
|
+
interface AssistantRubricRequest {
|
|
1949
|
+
system: string;
|
|
1950
|
+
user: string;
|
|
1951
|
+
rubricId: string;
|
|
1952
|
+
}
|
|
1953
|
+
interface StructuredJudgeProvider extends LlmProvider {
|
|
1954
|
+
judge(request: StructuredVerdictRequest): Promise<StructuredJudgeVerdictResult>;
|
|
1955
|
+
evaluateAssistantRubric(request: AssistantRubricRequest): Promise<string>;
|
|
1956
|
+
createJudgeError?(failure: Extract<StructuredJudgeVerdictResult, {
|
|
1957
|
+
ok: false;
|
|
1958
|
+
}>): Error;
|
|
1959
|
+
}
|
|
1960
|
+
declare class StructuredJudgeError extends Error {
|
|
1961
|
+
readonly code: StructuredJudgeErrorCode;
|
|
1962
|
+
readonly retryable: boolean;
|
|
1963
|
+
readonly httpStatus?: number;
|
|
1964
|
+
readonly telemetry: StructuredJudgeTelemetry;
|
|
1965
|
+
constructor(failure: Extract<StructuredJudgeVerdictResult, {
|
|
1966
|
+
ok: false;
|
|
1967
|
+
}>);
|
|
1968
|
+
}
|
|
1969
|
+
declare function isStructuredJudgeProvider(provider: LlmProvider): provider is StructuredJudgeProvider;
|
|
1970
|
+
declare function createStructuredBenchJudge(provider: StructuredJudgeProvider, rubricVersion?: string): BenchJudge;
|
|
1971
|
+
|
|
1881
1972
|
interface CodexCliRunRequest {
|
|
1882
1973
|
executable: string;
|
|
1883
1974
|
args: string[];
|
|
@@ -1901,8 +1992,13 @@ interface CodexCliProviderDeps {
|
|
|
1901
1992
|
status: number | null;
|
|
1902
1993
|
stderr: string;
|
|
1903
1994
|
}>;
|
|
1995
|
+
runCodexLoginStatus?: (executable: string, env: NodeJS.ProcessEnv) => Promise<{
|
|
1996
|
+
status: number | null;
|
|
1997
|
+
stdout: string;
|
|
1998
|
+
stderr: string;
|
|
1999
|
+
}>;
|
|
1904
2000
|
}
|
|
1905
|
-
declare function createCodexCliProvider(config: CodexCliProviderConfig, deps?: CodexCliProviderDeps):
|
|
2001
|
+
declare function createCodexCliProvider(config: CodexCliProviderConfig, deps?: CodexCliProviderDeps): StructuredJudgeProvider;
|
|
1906
2002
|
|
|
1907
2003
|
/**
|
|
1908
2004
|
* Result enrichment and JSON writing helpers.
|
|
@@ -2270,47 +2366,15 @@ declare function createOpenAiCompatibleProvider(config: OpenAiCompatibleProvider
|
|
|
2270
2366
|
*/
|
|
2271
2367
|
|
|
2272
2368
|
declare const DEFAULT_OPENAI_RESPONSES_JUDGE_MODEL = "gpt-5.6";
|
|
2273
|
-
type OpenAiResponsesJudgeErrorCode =
|
|
2274
|
-
|
|
2275
|
-
|
|
2276
|
-
|
|
2277
|
-
inputTokens: number;
|
|
2278
|
-
outputTokens: number;
|
|
2279
|
-
latencyMs: number;
|
|
2280
|
-
errorCode?: OpenAiResponsesJudgeErrorCode;
|
|
2281
|
-
httpStatus?: number;
|
|
2282
|
-
}
|
|
2283
|
-
interface OpenAiResponsesVerdict {
|
|
2284
|
-
score: number;
|
|
2285
|
-
decision: "pass" | "partial" | "fail";
|
|
2286
|
-
reason: string;
|
|
2287
|
-
}
|
|
2288
|
-
type OpenAiResponsesVerdictResult = {
|
|
2289
|
-
ok: true;
|
|
2290
|
-
verdict: OpenAiResponsesVerdict;
|
|
2291
|
-
telemetry: OpenAiResponsesJudgeTelemetry;
|
|
2292
|
-
} | {
|
|
2293
|
-
ok: false;
|
|
2294
|
-
error: {
|
|
2295
|
-
code: OpenAiResponsesJudgeErrorCode;
|
|
2296
|
-
message: string;
|
|
2297
|
-
retryable: boolean;
|
|
2298
|
-
httpStatus?: number;
|
|
2299
|
-
};
|
|
2300
|
-
telemetry: OpenAiResponsesJudgeTelemetry;
|
|
2301
|
-
};
|
|
2369
|
+
type OpenAiResponsesJudgeErrorCode = StructuredJudgeErrorCode;
|
|
2370
|
+
type OpenAiResponsesJudgeTelemetry = StructuredJudgeTelemetry;
|
|
2371
|
+
type OpenAiResponsesVerdict = StructuredJudgeVerdict;
|
|
2372
|
+
type OpenAiResponsesVerdictResult = StructuredJudgeVerdictResult;
|
|
2302
2373
|
interface OpenAiResponsesProviderConfig extends Omit<OpenAiCompatibleProviderConfig, "provider" | "model"> {
|
|
2303
2374
|
provider?: "openai";
|
|
2304
2375
|
model?: string;
|
|
2305
2376
|
rubricVersion?: string;
|
|
2306
2377
|
}
|
|
2307
|
-
interface VerdictRequest {
|
|
2308
|
-
rubric: string;
|
|
2309
|
-
rubricVersion: string;
|
|
2310
|
-
input: string;
|
|
2311
|
-
signal?: AbortSignal;
|
|
2312
|
-
maxTokens?: number;
|
|
2313
|
-
}
|
|
2314
2378
|
declare class OpenAiResponsesJudgeError extends Error {
|
|
2315
2379
|
readonly code: OpenAiResponsesJudgeErrorCode;
|
|
2316
2380
|
readonly retryable: boolean;
|
|
@@ -2320,7 +2384,7 @@ declare class OpenAiResponsesJudgeError extends Error {
|
|
|
2320
2384
|
ok: false;
|
|
2321
2385
|
}>);
|
|
2322
2386
|
}
|
|
2323
|
-
declare class OpenAiResponsesProvider implements
|
|
2387
|
+
declare class OpenAiResponsesProvider implements StructuredJudgeProvider {
|
|
2324
2388
|
readonly provider: "openai";
|
|
2325
2389
|
readonly id: string;
|
|
2326
2390
|
readonly name: string;
|
|
@@ -2330,7 +2394,7 @@ declare class OpenAiResponsesProvider implements LlmProvider {
|
|
|
2330
2394
|
private readonly telemetryEvents;
|
|
2331
2395
|
constructor(config?: OpenAiResponsesProviderConfig);
|
|
2332
2396
|
complete(prompt: string, opts?: CompletionOpts): Promise<CompletionResult>;
|
|
2333
|
-
judge(request:
|
|
2397
|
+
judge(request: StructuredVerdictRequest): Promise<OpenAiResponsesVerdictResult>;
|
|
2334
2398
|
evaluateAssistantRubric(request: {
|
|
2335
2399
|
system: string;
|
|
2336
2400
|
user: string;
|
|
@@ -2339,6 +2403,9 @@ declare class OpenAiResponsesProvider implements LlmProvider {
|
|
|
2339
2403
|
getUsage(): TokenUsage;
|
|
2340
2404
|
resetUsage(): void;
|
|
2341
2405
|
getTelemetryEvents(): OpenAiResponsesJudgeTelemetry[];
|
|
2406
|
+
createJudgeError(failure: Extract<OpenAiResponsesVerdictResult, {
|
|
2407
|
+
ok: false;
|
|
2408
|
+
}>): OpenAiResponsesJudgeError;
|
|
2342
2409
|
private parseResponse;
|
|
2343
2410
|
private failure;
|
|
2344
2411
|
private transportFailure;
|
|
@@ -5042,4 +5109,4 @@ declare function checkCodingGraphRegression(report: CodingGraphBenchReport, base
|
|
|
5042
5109
|
*/
|
|
5043
5110
|
declare function buildBaselineFromReport(report: CodingGraphBenchReport, note: string): CodingGraphBaseline;
|
|
5044
5111
|
|
|
5045
|
-
export { AMA_BENCH_DIAGNOSTIC_VARIANTS, ASSISTANT_AGENT_CONFIG_KEY, ASSISTANT_JUDGE_CONFIG_KEY, ASSISTANT_MEETING_PREP_SCENARIOS, ASSISTANT_MEETING_PREP_SMOKE_SCENARIOS, ASSISTANT_MORNING_BRIEF_SCENARIOS, ASSISTANT_MORNING_BRIEF_SMOKE_SCENARIOS, ASSISTANT_NEXT_BEST_ACTION_SCENARIOS, ASSISTANT_NEXT_BEST_ACTION_SMOKE_SCENARIOS, ASSISTANT_RUBRIC_DIMENSIONS, ASSISTANT_RUBRIC_ID_KEY, ASSISTANT_SEEDS_CONFIG_KEY, ASSISTANT_SPOT_CHECK_DIR_KEY, ASSISTANT_SYNTHESIS_SCENARIOS, ASSISTANT_SYNTHESIS_SMOKE_SCENARIOS, type AblationConfigOverrides, type AbstentionRetrievalCase, type AggregateMetrics, type AmaBenchDiagnosticAdapterOptions, type AmaBenchDiagnosticAnswererMode, type AmaBenchDiagnosticBreakdown, type AmaBenchDiagnosticMatrixArtifact, type AmaBenchDiagnosticRecallMode, type AmaBenchDiagnosticRunContext, type AmaBenchDiagnosticTaskEvidence, type AmaBenchDiagnosticTaskRow, type AmaBenchDiagnosticVariant, type AmaBenchDiagnosticVariantSummary, type AnthropicProviderConfig, type AssistantAgent, type AssistantMemoryFact, type AssistantMemoryGraph, type AssistantRubricDimension, type AssistantRubricScores, type AssistantRunnerOptions, type AssistantScenario, type AssistantStance, type AttackRecallOptions, type AttackRetrievalHit, type AttackerMode, BENCHMARK_ARTIFACT_SCHEMA_VERSION, BENCHMARK_INTEGRITY_META_SCHEMA, BENCHMARK_REPRO_MANIFEST_FILENAME, BENCHMARK_REPRO_MANIFEST_SCHEMA_VERSION, BENCHMARK_RESULT_SCHEMA, BENCHMARK_SPLIT_TYPES, type BaselineRow, type BaselineScenario, type BeamDatasetPreview, type BenchConfig, type BenchJudge, type BenchJudgeResult, type BenchMemoryAdapter, type BenchModelSource, type BenchReasoningEffort, type BenchRecallOptions, type BenchRecallSupportAssessment, type BenchRecallSupportRequest, type BenchRecallSupportStatus, type BenchResponder, type BenchResponse, type BenchRuntimeProfile, type BenchTier, type BenchmarkArtifact, type BenchmarkArtifactEnvironment, type BenchmarkArtifactHardware, type BenchmarkArtifactJudgeCalibration, type BenchmarkArtifactPerTaskScore, type BenchmarkArtifactSystem, type BenchmarkArtifactTier, type BenchmarkCategory, type BenchmarkDefinition, type BenchmarkIntegrityMeta, type BenchmarkMeta, type BenchmarkMode, type BenchmarkReport, type BenchmarkReproManifest, type BenchmarkReproManifestDataset, type BenchmarkReproManifestFile, type BenchmarkReproManifestResult, type BenchmarkResult, type BenchmarkSplitType, type BenchmarkStatus, type BenchmarkSuiteResult, type BenchmarkTier, type BootstrapKappaOptions, type BootstrapKappaResult, type BuildBenchmarkArtifactInput, type BuildBenchmarkPublishFeedOptions, type BuildBenchmarkReproManifestOptions, type BuiltInProvider, CALIBRATION_SLICE_SIZE, CANARY_FIXED_RECALL, CANARY_SCORE_FLOOR, DEFAULT_10K_FIXTURE as CODING_GRAPH_10K_FIXTURE, CODING_GRAPH_BENCH_SCHEMA_VERSION, DEFAULT_TOLERANCE_PERCENT as CODING_GRAPH_DEFAULT_TOLERANCE, MIN_ITERATIONS as CODING_GRAPH_MIN_ITERATIONS, DEFAULT_SMOKE_FIXTURE as CODING_GRAPH_SMOKE_FIXTURE, type CalibrationAnswer, type CalibrationVerdictPair, type CanaryAdapterOptions, type CanaryFloorCheck, type ClaudeCliProviderConfig, type CodexCliProviderConfig, type CodingGraphBaseline, type CodingGraphBenchConfig, type CodingGraphBenchReport, type MachineFingerprint as CodingGraphMachineFingerprint, type CodingGraphMetricKey, type RegressionMetricDetail as CodingGraphRegressionDetail, type RegressionMetricKey as CodingGraphRegressionKey, type RegressionGateResult as CodingGraphRegressionResult, type CohenKappaResult, type ComparisonMetricDelta, type ComparisonResult, type CompletionOpts, type CompletionResult, type ConfidenceInterval, type ContaminationCheckResult, type ContaminationEntry, type ContaminationManifest, type CustomBenchmarkScoring, type CustomBenchmarkSpec, type CustomBenchmarkTask, DEFAULT_ABLATION_BENCHMARK, DEFAULT_ABLATION_BOOTSTRAP_SEED, DEFAULT_ASSISTANT_RUBRIC_ID, DEFAULT_BASELINE_SCENARIOS, DEFAULT_JUDGE_BINARIZATION_THRESHOLD, DEFAULT_KAPPA_BOOTSTRAP_SAMPLES, DEFAULT_KAPPA_CONFIDENCE_LEVEL, DEFAULT_OPENAI_RESPONSES_JUDGE_MODEL, type DatasetSource, type DiagnoseLoComoProfileDeltaOptions, type DiscoveredModel, EMPTY_CONTAMINATION_MANIFEST, type EffectSizeInterpretation, type EffectSizeSummary, type ExplainResult, type ExtractedEntity, type ExtractedLink, type ExtractedPage, type ExtractionAttackOptions, type ExtractionAttackResult, type ExtractionAttackTarget, type FixtureGenerator, type FixtureOutput, type FixtureVariant, GENERAL_ANSWER_JUDGE_RUBRIC, type GeneratedFile, type GeneratedRepo, type GoldEntity, type GoldEntityType, type GoldGraph, type GoldLink, type GoldPage, type HarnessRng, INTEGRITY_CIPHER_ALGORITHM, INTEGRITY_HASH_ALGORITHM, INTEGRITY_META_FIELDS, type IngestionBenchAdapter, type IngestionLog, JUDGE_CALIBRATION_KAPPA_THRESHOLD, type JudgeCalibrationIdentities, type JudgeCalibrationResult, type JudgeCategory, type KappaConfidenceInterval, LOCAL_LAB_PROVIDER_KINDS, LOCOMO_DATASET_FILENAMES, LONG_MEM_EVAL_DATASET_FILENAMES, type LeaderboardArtifactWrite, type LettaAdapterConfig, LettaMemCorrectAdapter, type LlmJudge, type LlmProvider, type LoComoCategoryDelta, type LoComoMetricDelta, type LoComoProfileArtifactEvidence, type LoComoProfileDeltaReport, type LoComoTaskRegression, type LoadDatasetOptions, type LoadSealedQrelsOptions, type LoadedDataset, type LoadedJudgeCalibrationState, type LocalLabManifest, type LocalLabManifestNotes, type LocalLabPhase, type LocalLabPhaseDescriptor, type LocalLabPhaseExecute, type LocalLabPhaseName, type LocalLabPhaseOutcome, LocalLabPreflightError, type LocalLabPreflightFailure, type LocalLabPreflightInput, type LocalLabPreflightOptions, type LocalLabPreflightResult, type LocalLabPreflightSuccess, type LocalLabProviderKind, type LocalLabRoleConfig, type LocalLlmProviderConfig, MEMCORRECT_CORRECTION_ACCEPTANCE_RUBRIC, MEMCORRECT_CORRECTION_ACCEPTANCE_RUBRIC_VERSION, MEMCORRECT_STALE_HARM_RUBRIC, MEMCORRECT_STALE_HARM_RUBRIC_VERSION, MEMORY_EVAL_DIMENSIONS, MEMORY_EVAL_PUBLIC_LINE, MIN_CALIBRATION_SOURCE_TASKS, MITIGATED_BASELINE_SCENARIOS, type McpArgumentSemantic, type McpBackendErrorCode, type McpBackendResult, type McpBenchMemoryAdapter, type McpConformanceResult, type McpHttpTransportConfig, type McpListedTool, type McpMemCorrectAdapter, type McpMemoryAdapterOptions, McpMemoryBackendError, type McpMemoryToolMapping, type McpMemoryTransportConfig, type McpStdioTransportConfig, type McpToolCallResult, type McpToolClient, type McpToolMappingEntry, type McpToolMappingValue, type McpToolOperation, type Mem0AdapterConfig, Mem0MemCorrectAdapter, type MemCorrectGeneratorOptions, type MemCorrectJudgeRequest, type MemCorrectJudgeResult, type MemCorrectSystemAdapter, type MemoryEvalCategory, type MemoryEvalDimension, type MemoryEvalDimensionId, type MemoryEvalMetric, type MemoryGraph, type MemoryStats, type MemorySystem, type Message, type MetricAggregate, type MicroMetric, MissingCredentialError, type MitigatedBaselineConfig, type MitigatedTargetConfig, type MultipleChoiceQuestion, OPENAI_RESPONSES_JUDGE_RUBRIC_VERSION, OTHER_NAMESPACE_MEMORIES, type OllamaProviderConfig, type OpenAiCompatibleProviderConfig, OpenAiResponsesJudgeError, type OpenAiResponsesJudgeErrorCode, type OpenAiResponsesJudgeTelemetry, OpenAiResponsesProvider, type OpenAiResponsesProviderConfig, type OpenAiResponsesVerdict, type OpenAiResponsesVerdictResult, PROCEDURAL_REAL_SCENARIOS, PROCEDURAL_REAL_SCENARIOS_SMOKE, PUBLISHED_BENCHMARK_ARTIFACT_IDS, type PersonalizationRetrievalCase, type PreflightDiscoveredModel, type ProceduralAblationArtifact, type ProceduralAblationPerCase, type ProceduralAblationScenario, type ProceduralRealScenario, type ProceduralRealScenarioCategory, type ProviderBaseConfig, type ProviderConfig, type ProviderDiscoveryResult, type ProviderFactoryConfig, type PublishSkipReason, type PublishSkipRecord, type PublishedBenchmarkFeed, type PublishedBenchmarkFeedEntry, type PublishedBenchmarkId, REQUIRED_FRONTMATTER_FIELDS, type RecallMetrics, type RecoveredMemory, type RegressionDetail, type RegressionGateResult$1 as RegressionGateResult, type RemnicAdapterOptions, type ReportCardProvenanceContext, type ResolveBenchRuntimeProfileOptions, type ResolvedBenchRuntimeProfile, type ResolvedLocalLabProfile, type ResolvedLocalLabRole, type ResolvedRunBenchmarkOptions, type RotatedChoices, type RunBenchmarkOptions, type RunJudgeCalibrationOptions, type RunProceduralAblationCliArgs, type RunProceduralAblationOptions, type RunSequentialPhasesOptions, SCHEMA_TIER_FIXTURE, SCHEMA_TIER_SMOKE_FIXTURE, SEALED_PROMPT_REGISTRY, SINGLE_FLAG_ABLATION_MATRIX, SYNTHETIC_MEMORIES, type SanitizedDiagnosticProvider, type SavedBaseline, type SchemaTierCorpus, type SchemaTierFixture, type SchemaTierName, type SchemaTierPage, type SchemaTierPageFrontmatter, type SealedArtifact, type SealedJudgeDecision, type SealedJudgeInput, type SealedQrelsArtifact, type SealedQrelsHandle, type SealedRubric, type SearchResult, type SeededMemory, type SeededRng, type SequentialPhaseHooks, type SingleFlagAblationCell, type SingleFlagAblationId, type SpotCheckLogger, type StatisticalReport, type StructuredJudge, type SyntheticEdge, type SyntheticEmailIngestionAdapterOptions, type SyntheticFileIR, type SyntheticRepoConfig, type SyntheticSymbol, type SyntheticTargetOptions, type TaskResult, type TaskTokenUsage, type TemporalRetrievalCase, type ThirdPartyAdapterConfig, type TierDetail, type TimelineEntry, type TokenUsage, type WallMetric, type WriteBenchmarkArtifactResult, type ZepAdapterConfig, ZepMemCorrectAdapter, addContaminationEntry, aggregateTaskScores, answerBenchmarkQuestion, assertCanaryUnderFloor, assertIntegrityMetaPresent, assertPublishableIntegrity, assertSha256Hex, assistantMeetingPrepDefinition, assistantMorningBriefDefinition, assistantNextBestActionDefinition, assistantSynthesisDefinition, backlinkF1, binarizeJudgeScore, bootstrapCohensKappaConfidenceInterval, bootstrapMeanConfidenceInterval, buildAmaBenchDiagnosticMatrixArtifact, buildAmaBenchDiagnosticVariantSummary, buildAmaBenchLeaderboardRows, buildBaselineFromReport, buildBenchmarkArtifact, buildBenchmarkArtifactFilename, buildBenchmarkPublishFeed, buildBenchmarkReproManifest, buildBenchmarkRunSeeds, buildJudgePayload, buildOracleTrajectoryRecall, buildSchemaTierFixture, buildSchemaTierSmokeFixture, calendarFixture, canonicalJsonStringify, captureMachineFingerprint, chatFixture, checkCodingGraphRegression, checkDatasetContamination, checkRegression, clampScore, cohensD, compareResults, computeCohensKappa, computeSealHash, containsAnswer, createSeededRng$1 as createAdamSeededRng, createAmaBenchDiagnosticAdapter, createAnthropicProvider, createCanaryAdapter, createClaudeCliProvider, createCodexCliProvider, createSeededRng as createCodingGraphSeededRng, createDeterministicSpotCheckLogger, createGatewayResponder, createLightweightAdapter, createLiteLlmProvider, createLocalLlmProvider, createMcpDemoMemCorrectAdapter, createMcpDemoMemoryAdapter, createMcpMemCorrectAdapter, createMcpMemoryAdapter, createMitigatedTarget, createOllamaProvider, createOpenAiCompatibleProvider, createOpenAiResponsesBenchJudge, createOpenAiResponsesProvider, createSeededRandom as createProceduralAblationSeededRandom, createProvider, createProviderBackedAmaBenchRecommendedJudge, createProviderBackedJudge, createProviderBackedResponder, createProviderBackedStructuredJudge, createRemnicAdapter, createResponderFromProvider, createSeededRng$2 as createSeededRng, createSpotCheckFileLogger, createStructuredJudgeFromProvider, createSyntheticEmailIngestionAdapter, createSyntheticTarget, createTimeoutGuardedAdapter, defaultBenchmarkBaselineDir, defaultBenchmarkPublishPath, deleteBenchmarkResults, diagnoseLoComoProfileDelta, discoverAllProviders, discoveryEndpointFor, emailFixture, entityRecall, exactMatch, extractMetrics as extractCodingGraphMetrics, extractMarkdownSectionsByTitle, f1Score, fixtureToAblationScenarios, formatHandoffNote, formatMissingDatasetError, generateReport, generateSyntheticRepo, getAblationCell, getBenchmark, getBenchmarkLowerIsBetter, getMemoryEvalDimension, getRemnicVersion, hashBenchmarkArtifact, hashBytes, hashCanonicalJson, hashString, integrityMetaIsComplete, interpretEffectSize, isAmaBenchUnknownLikeAnswer, isContaminationEntry, isContaminationManifest, isSealedQrelsArtifact, isSha256Hex, judgeMemCorrectCorrectionAcceptance, judgeMemCorrectStaleMemoryHarm, linkMatches, listBenchmarkBaselines, listBenchmarkResults, listBenchmarks, listMemoryEvalBenchmarkIds, listMemoryEvalDimensions, llmJudgeScore, llmJudgeScoreDetailed, loadAblationFixture, loadBaseline, loadBeamDatasetPreview, loadBenchmarkArtifact, loadBenchmarkBaseline, loadBenchmarkReportCardProvenance, loadBenchmarkResult, loadCustomBenchmarkFile, loadJudgeCalibrationState, loadLoCoMo10, loadLocalLabManifest, loadLongMemEvalS, loadSealKeyFromEnv, loadSealedQrels, loadSealedRubric, matchEntity, mergeContaminationManifests, openSeal, orchestrateBenchmarkRuns, pairedDeltaConfidenceInterval, parseBenchmarkArtifact, parseCustomBenchmark, parseLocalLabManifest, parseRubricResponse, parseSealedQrels, pickStableQualifiedName, precisionAtK, preflightLocalLabRole, projectFolderFixture, recallAtK, redactBenchmarkResultSecrets, renderBaselineMarkdown, renderBenchmarkResultExport, renderLoComoProfileDeltaMarkdown, renderMemorySummaryForJudge, renderMemoryViewForAgent, resolveAssistantAgent, resolveAssistantRubricId, resolveAssistantSeeds, resolveAssistantSpotCheckDir, resolveBenchRuntimeProfile, resolveBenchmarkPhaseTimeoutMs, resolveBenchmarkProgressLogging, resolveBenchmarkResultReference, resolveBenchmarkRunCount, resolveLocalLabProfile, resolveLocalLabRole, resolveStructuredJudge, rotateDistractors, rougeL, runAssistantBenchmark, runAssistantMeetingPrepBenchmark, runAssistantMorningBriefBenchmark, runAssistantNextBestActionBenchmark, runAssistantSynthesisBenchmark, runBaseline, runBenchSuite, runBenchmark, runCodingGraphBenchmark, runCustomBenchmarkFile, runExplain, runExtractionAttack, runJudgeCalibration, runMitigatedBaseline, runProceduralAblation, runProceduralAblationCli, runSealedJudge, runSequentialPhases, safeHexEqual, saveBaseline, saveBenchmarkBaseline, schemaCompleteness, sealPayload, selectAmaBenchDiagnosticVariants, selectCalibrationSlice, selectFixtureVariant, serializeBenchmarkArtifact, serializeJsonl, serializeSealedQrels, shuffleTasks, timed, verifyRubricDigest, writeBenchmarkArtifact, writeBenchmarkPublishFeed, writeBenchmarkReproManifest, writeBenchmarkResult, writeJudgeCalibrationState, writeLeaderboardArtifactsForResult, zeroScores };
|
|
5112
|
+
export { AMA_BENCH_DIAGNOSTIC_VARIANTS, ASSISTANT_AGENT_CONFIG_KEY, ASSISTANT_JUDGE_CONFIG_KEY, ASSISTANT_MEETING_PREP_SCENARIOS, ASSISTANT_MEETING_PREP_SMOKE_SCENARIOS, ASSISTANT_MORNING_BRIEF_SCENARIOS, ASSISTANT_MORNING_BRIEF_SMOKE_SCENARIOS, ASSISTANT_NEXT_BEST_ACTION_SCENARIOS, ASSISTANT_NEXT_BEST_ACTION_SMOKE_SCENARIOS, ASSISTANT_RUBRIC_DIMENSIONS, ASSISTANT_RUBRIC_ID_KEY, ASSISTANT_SEEDS_CONFIG_KEY, ASSISTANT_SPOT_CHECK_DIR_KEY, ASSISTANT_SYNTHESIS_SCENARIOS, ASSISTANT_SYNTHESIS_SMOKE_SCENARIOS, type AblationConfigOverrides, type AbstentionRetrievalCase, type AggregateMetrics, type AmaBenchDiagnosticAdapterOptions, type AmaBenchDiagnosticAnswererMode, type AmaBenchDiagnosticBreakdown, type AmaBenchDiagnosticMatrixArtifact, type AmaBenchDiagnosticRecallMode, type AmaBenchDiagnosticRunContext, type AmaBenchDiagnosticTaskEvidence, type AmaBenchDiagnosticTaskRow, type AmaBenchDiagnosticVariant, type AmaBenchDiagnosticVariantSummary, type AnthropicProviderConfig, type AssistantAgent, type AssistantMemoryFact, type AssistantMemoryGraph, type AssistantRubricDimension, type AssistantRubricRequest, type AssistantRubricScores, type AssistantRunnerOptions, type AssistantScenario, type AssistantStance, type AttackRecallOptions, type AttackRetrievalHit, type AttackerMode, BENCHMARK_ARTIFACT_SCHEMA_VERSION, BENCHMARK_INTEGRITY_META_SCHEMA, BENCHMARK_REPRO_MANIFEST_FILENAME, BENCHMARK_REPRO_MANIFEST_SCHEMA_VERSION, BENCHMARK_RESULT_SCHEMA, BENCHMARK_SPLIT_TYPES, type BaselineRow, type BaselineScenario, type BeamDatasetPreview, type BenchConfig, type BenchJudge, type BenchJudgeResult, type BenchMemoryAdapter, type BenchModelSource, type BenchReasoningEffort, type BenchRecallOptions, type BenchRecallSupportAssessment, type BenchRecallSupportRequest, type BenchRecallSupportStatus, type BenchResponder, type BenchResponse, type BenchRuntimeProfile, type BenchTier, type BenchmarkArtifact, type BenchmarkArtifactEnvironment, type BenchmarkArtifactHardware, type BenchmarkArtifactJudgeCalibration, type BenchmarkArtifactPerTaskScore, type BenchmarkArtifactSystem, type BenchmarkArtifactTier, type BenchmarkCategory, type BenchmarkDefinition, type BenchmarkIntegrityMeta, type BenchmarkMeta, type BenchmarkMode, type BenchmarkReport, type BenchmarkReproManifest, type BenchmarkReproManifestDataset, type BenchmarkReproManifestFile, type BenchmarkReproManifestResult, type BenchmarkResult, type BenchmarkSplitType, type BenchmarkStatus, type BenchmarkSuiteResult, type BenchmarkTier, type BootstrapKappaOptions, type BootstrapKappaResult, type BuildBenchmarkArtifactInput, type BuildBenchmarkPublishFeedOptions, type BuildBenchmarkReproManifestOptions, type BuiltInProvider, CALIBRATION_SLICE_SIZE, CANARY_FIXED_RECALL, CANARY_SCORE_FLOOR, DEFAULT_10K_FIXTURE as CODING_GRAPH_10K_FIXTURE, CODING_GRAPH_BENCH_SCHEMA_VERSION, DEFAULT_TOLERANCE_PERCENT as CODING_GRAPH_DEFAULT_TOLERANCE, MIN_ITERATIONS as CODING_GRAPH_MIN_ITERATIONS, DEFAULT_SMOKE_FIXTURE as CODING_GRAPH_SMOKE_FIXTURE, type CalibrationAnswer, type CalibrationVerdictPair, type CanaryAdapterOptions, type CanaryFloorCheck, type ClaudeCliProviderConfig, type CodexCliProviderConfig, type CodingGraphBaseline, type CodingGraphBenchConfig, type CodingGraphBenchReport, type MachineFingerprint as CodingGraphMachineFingerprint, type CodingGraphMetricKey, type RegressionMetricDetail as CodingGraphRegressionDetail, type RegressionMetricKey as CodingGraphRegressionKey, type RegressionGateResult as CodingGraphRegressionResult, type CohenKappaResult, type ComparisonMetricDelta, type ComparisonResult, type CompletionOpts, type CompletionResult, type ConfidenceInterval, type ContaminationCheckResult, type ContaminationEntry, type ContaminationManifest, type CustomBenchmarkScoring, type CustomBenchmarkSpec, type CustomBenchmarkTask, DEFAULT_ABLATION_BENCHMARK, DEFAULT_ABLATION_BOOTSTRAP_SEED, DEFAULT_ASSISTANT_RUBRIC_ID, DEFAULT_BASELINE_SCENARIOS, DEFAULT_JUDGE_BINARIZATION_THRESHOLD, DEFAULT_KAPPA_BOOTSTRAP_SAMPLES, DEFAULT_KAPPA_CONFIDENCE_LEVEL, DEFAULT_OPENAI_RESPONSES_JUDGE_MODEL, type DatasetSource, type DiagnoseLoComoProfileDeltaOptions, type DiscoveredModel, EMPTY_CONTAMINATION_MANIFEST, type EffectSizeInterpretation, type EffectSizeSummary, type ExplainResult, type ExtractedEntity, type ExtractedLink, type ExtractedPage, type ExtractionAttackOptions, type ExtractionAttackResult, type ExtractionAttackTarget, type FixtureGenerator, type FixtureOutput, type FixtureVariant, GENERAL_ANSWER_JUDGE_RUBRIC, type GeneratedFile, type GeneratedRepo, type GoldEntity, type GoldEntityType, type GoldGraph, type GoldLink, type GoldPage, type HarnessRng, INTEGRITY_CIPHER_ALGORITHM, INTEGRITY_HASH_ALGORITHM, INTEGRITY_META_FIELDS, type IngestionBenchAdapter, type IngestionLog, JUDGE_CALIBRATION_KAPPA_THRESHOLD, type JudgeCalibrationIdentities, type JudgeCalibrationResult, type JudgeCategory, type KappaConfidenceInterval, LOCAL_LAB_PROVIDER_KINDS, LOCOMO_DATASET_FILENAMES, LONG_MEM_EVAL_DATASET_FILENAMES, type LeaderboardArtifactWrite, type LettaAdapterConfig, LettaMemCorrectAdapter, type LlmJudge, type LlmProvider, type LoComoCategoryDelta, type LoComoMetricDelta, type LoComoProfileArtifactEvidence, type LoComoProfileDeltaReport, type LoComoTaskRegression, type LoadDatasetOptions, type LoadSealedQrelsOptions, type LoadedDataset, type LoadedJudgeCalibrationState, type LocalLabManifest, type LocalLabManifestNotes, type LocalLabPhase, type LocalLabPhaseDescriptor, type LocalLabPhaseExecute, type LocalLabPhaseName, type LocalLabPhaseOutcome, LocalLabPreflightError, type LocalLabPreflightFailure, type LocalLabPreflightInput, type LocalLabPreflightOptions, type LocalLabPreflightResult, type LocalLabPreflightSuccess, type LocalLabProviderKind, type LocalLabRoleConfig, type LocalLlmProviderConfig, MEMCORRECT_CORRECTION_ACCEPTANCE_RUBRIC, MEMCORRECT_CORRECTION_ACCEPTANCE_RUBRIC_VERSION, MEMCORRECT_STALE_HARM_RUBRIC, MEMCORRECT_STALE_HARM_RUBRIC_VERSION, MEMORY_EVAL_DIMENSIONS, MEMORY_EVAL_PUBLIC_LINE, MIN_CALIBRATION_SOURCE_TASKS, MITIGATED_BASELINE_SCENARIOS, type McpArgumentSemantic, type McpBackendErrorCode, type McpBackendResult, type McpBenchMemoryAdapter, type McpConformanceResult, type McpHttpTransportConfig, type McpListedTool, type McpMemCorrectAdapter, type McpMemoryAdapterOptions, McpMemoryBackendError, type McpMemoryToolMapping, type McpMemoryTransportConfig, type McpStdioTransportConfig, type McpToolCallResult, type McpToolClient, type McpToolMappingEntry, type McpToolMappingValue, type McpToolOperation, type Mem0AdapterConfig, Mem0MemCorrectAdapter, type MemCorrectGeneratorOptions, type MemCorrectJudgeRequest, type MemCorrectJudgeResult, type MemCorrectSystemAdapter, type MemoryEvalCategory, type MemoryEvalDimension, type MemoryEvalDimensionId, type MemoryEvalMetric, type MemoryGraph, type MemoryStats, type MemorySystem, type Message, type MetricAggregate, type MicroMetric, MissingCredentialError, type MitigatedBaselineConfig, type MitigatedTargetConfig, type MultipleChoiceQuestion, OPENAI_RESPONSES_JUDGE_RUBRIC_VERSION, OTHER_NAMESPACE_MEMORIES, type OllamaProviderConfig, type OpenAiCompatibleProviderConfig, OpenAiResponsesJudgeError, type OpenAiResponsesJudgeErrorCode, type OpenAiResponsesJudgeTelemetry, OpenAiResponsesProvider, type OpenAiResponsesProviderConfig, type OpenAiResponsesVerdict, type OpenAiResponsesVerdictResult, PROCEDURAL_REAL_SCENARIOS, PROCEDURAL_REAL_SCENARIOS_SMOKE, PUBLISHED_BENCHMARK_ARTIFACT_IDS, type PersonalizationRetrievalCase, type PreflightDiscoveredModel, type ProceduralAblationArtifact, type ProceduralAblationPerCase, type ProceduralAblationScenario, type ProceduralRealScenario, type ProceduralRealScenarioCategory, type ProviderBaseConfig, type ProviderConfig, type ProviderDiscoveryResult, type ProviderFactoryConfig, type PublishSkipReason, type PublishSkipRecord, type PublishedBenchmarkFeed, type PublishedBenchmarkFeedEntry, type PublishedBenchmarkId, REQUIRED_FRONTMATTER_FIELDS, type RecallMetrics, type RecoveredMemory, type RegressionDetail, type RegressionGateResult$1 as RegressionGateResult, type RemnicAdapterOptions, type ReportCardProvenanceContext, type ResolveBenchRuntimeProfileOptions, type ResolvedBenchRuntimeProfile, type ResolvedLocalLabProfile, type ResolvedLocalLabRole, type ResolvedRunBenchmarkOptions, type RotatedChoices, type RunBenchmarkOptions, type RunJudgeCalibrationOptions, type RunProceduralAblationCliArgs, type RunProceduralAblationOptions, type RunSequentialPhasesOptions, SCHEMA_TIER_FIXTURE, SCHEMA_TIER_SMOKE_FIXTURE, SEALED_PROMPT_REGISTRY, SINGLE_FLAG_ABLATION_MATRIX, SYNTHETIC_MEMORIES, type SanitizedDiagnosticProvider, type SavedBaseline, type SchemaTierCorpus, type SchemaTierFixture, type SchemaTierName, type SchemaTierPage, type SchemaTierPageFrontmatter, type SealedArtifact, type SealedJudgeDecision, type SealedJudgeInput, type SealedQrelsArtifact, type SealedQrelsHandle, type SealedRubric, type SearchResult, type SeededMemory, type SeededRng, type SequentialPhaseHooks, type SingleFlagAblationCell, type SingleFlagAblationId, type SpotCheckLogger, type StatisticalReport, type StructuredJudge, StructuredJudgeError, type StructuredJudgeErrorCode, type StructuredJudgeProvider, type StructuredJudgeTelemetry, type StructuredJudgeVerdict, type StructuredJudgeVerdictResult, type StructuredVerdictRequest, type SyntheticEdge, type SyntheticEmailIngestionAdapterOptions, type SyntheticFileIR, type SyntheticRepoConfig, type SyntheticSymbol, type SyntheticTargetOptions, type TaskResult, type TaskTokenUsage, type TemporalRetrievalCase, type ThirdPartyAdapterConfig, type TierDetail, type TimelineEntry, type TokenUsage, type WallMetric, type WriteBenchmarkArtifactResult, type ZepAdapterConfig, ZepMemCorrectAdapter, addContaminationEntry, aggregateTaskScores, answerBenchmarkQuestion, assertCanaryUnderFloor, assertIntegrityMetaPresent, assertPublishableIntegrity, assertSha256Hex, assistantMeetingPrepDefinition, assistantMorningBriefDefinition, assistantNextBestActionDefinition, assistantSynthesisDefinition, backlinkF1, binarizeJudgeScore, bootstrapCohensKappaConfidenceInterval, bootstrapMeanConfidenceInterval, buildAmaBenchDiagnosticMatrixArtifact, buildAmaBenchDiagnosticVariantSummary, buildAmaBenchLeaderboardRows, buildBaselineFromReport, buildBenchmarkArtifact, buildBenchmarkArtifactFilename, buildBenchmarkPublishFeed, buildBenchmarkReproManifest, buildBenchmarkRunSeeds, buildJudgePayload, buildOracleTrajectoryRecall, buildSchemaTierFixture, buildSchemaTierSmokeFixture, calendarFixture, canonicalJsonStringify, captureMachineFingerprint, chatFixture, checkCodingGraphRegression, checkDatasetContamination, checkRegression, clampScore, cohensD, compareResults, computeCohensKappa, computeSealHash, containsAnswer, createSeededRng$1 as createAdamSeededRng, createAmaBenchDiagnosticAdapter, createAnthropicProvider, createCanaryAdapter, createClaudeCliProvider, createCodexCliProvider, createSeededRng as createCodingGraphSeededRng, createDeterministicSpotCheckLogger, createGatewayResponder, createLightweightAdapter, createLiteLlmProvider, createLocalLlmProvider, createMcpDemoMemCorrectAdapter, createMcpDemoMemoryAdapter, createMcpMemCorrectAdapter, createMcpMemoryAdapter, createMitigatedTarget, createOllamaProvider, createOpenAiCompatibleProvider, createOpenAiResponsesBenchJudge, createOpenAiResponsesProvider, createSeededRandom as createProceduralAblationSeededRandom, createProvider, createProviderBackedAmaBenchRecommendedJudge, createProviderBackedJudge, createProviderBackedResponder, createProviderBackedStructuredJudge, createRemnicAdapter, createResponderFromProvider, createSeededRng$2 as createSeededRng, createSpotCheckFileLogger, createStructuredBenchJudge, createStructuredJudgeFromProvider, createSyntheticEmailIngestionAdapter, createSyntheticTarget, createTimeoutGuardedAdapter, defaultBenchmarkBaselineDir, defaultBenchmarkPublishPath, deleteBenchmarkResults, diagnoseLoComoProfileDelta, discoverAllProviders, discoveryEndpointFor, emailFixture, entityRecall, exactMatch, extractMetrics as extractCodingGraphMetrics, extractMarkdownSectionsByTitle, f1Score, fixtureToAblationScenarios, formatHandoffNote, formatMissingDatasetError, generateReport, generateSyntheticRepo, getAblationCell, getBenchmark, getBenchmarkLowerIsBetter, getMemoryEvalDimension, getRemnicVersion, hashBenchmarkArtifact, hashBytes, hashCanonicalJson, hashString, integrityMetaIsComplete, interpretEffectSize, isAmaBenchUnknownLikeAnswer, isContaminationEntry, isContaminationManifest, isSealedQrelsArtifact, isSha256Hex, isStructuredJudgeProvider, judgeMemCorrectCorrectionAcceptance, judgeMemCorrectStaleMemoryHarm, linkMatches, listBenchmarkBaselines, listBenchmarkResults, listBenchmarks, listMemoryEvalBenchmarkIds, listMemoryEvalDimensions, llmJudgeScore, llmJudgeScoreDetailed, loadAblationFixture, loadBaseline, loadBeamDatasetPreview, loadBenchmarkArtifact, loadBenchmarkBaseline, loadBenchmarkReportCardProvenance, loadBenchmarkResult, loadCustomBenchmarkFile, loadJudgeCalibrationState, loadLoCoMo10, loadLocalLabManifest, loadLongMemEvalS, loadSealKeyFromEnv, loadSealedQrels, loadSealedRubric, matchEntity, mergeContaminationManifests, openSeal, orchestrateBenchmarkRuns, pairedDeltaConfidenceInterval, parseBenchmarkArtifact, parseCustomBenchmark, parseLocalLabManifest, parseRubricResponse, parseSealedQrels, pickStableQualifiedName, precisionAtK, preflightLocalLabRole, projectFolderFixture, recallAtK, redactBenchmarkResultSecrets, renderBaselineMarkdown, renderBenchmarkResultExport, renderLoComoProfileDeltaMarkdown, renderMemorySummaryForJudge, renderMemoryViewForAgent, resolveAssistantAgent, resolveAssistantRubricId, resolveAssistantSeeds, resolveAssistantSpotCheckDir, resolveBenchRuntimeProfile, resolveBenchmarkPhaseTimeoutMs, resolveBenchmarkProgressLogging, resolveBenchmarkResultReference, resolveBenchmarkRunCount, resolveLocalLabProfile, resolveLocalLabRole, resolveStructuredJudge, rotateDistractors, rougeL, runAssistantBenchmark, runAssistantMeetingPrepBenchmark, runAssistantMorningBriefBenchmark, runAssistantNextBestActionBenchmark, runAssistantSynthesisBenchmark, runBaseline, runBenchSuite, runBenchmark, runCodingGraphBenchmark, runCustomBenchmarkFile, runExplain, runExtractionAttack, runJudgeCalibration, runMitigatedBaseline, runProceduralAblation, runProceduralAblationCli, runSealedJudge, runSequentialPhases, safeHexEqual, saveBaseline, saveBenchmarkBaseline, schemaCompleteness, sealPayload, selectAmaBenchDiagnosticVariants, selectCalibrationSlice, selectFixtureVariant, serializeBenchmarkArtifact, serializeJsonl, serializeSealedQrels, shuffleTasks, timed, verifyRubricDigest, writeBenchmarkArtifact, writeBenchmarkPublishFeed, writeBenchmarkReproManifest, writeBenchmarkResult, writeJudgeCalibrationState, writeLeaderboardArtifactsForResult, zeroScores };
|