@tangle-network/agent-knowledge 13.0.1 → 14.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +52 -157
- package/CHANGELOG.md +33 -0
- package/README.md +8 -0
- package/dist/benchmarks/index.js +1 -1
- package/dist/{benchmarks-Qk94Gnj4.js → benchmarks-BbPJdmHe.js} +24 -24
- package/dist/{benchmarks-Qk94Gnj4.js.map → benchmarks-BbPJdmHe.js.map} +1 -1
- package/dist/cli.js +1 -1
- package/dist/index.d.ts +7 -5
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +55 -78
- package/dist/index.js.map +1 -1
- package/dist/{inspect-xAkP4sxi.js → inspect-DCpUoPVb.js} +14 -2
- package/dist/{inspect-xAkP4sxi.js.map → inspect-DCpUoPVb.js.map} +1 -1
- package/dist/memory/index.js +2 -2
- package/dist/{memory-DFSo2iLi.js → memory-BxwOtvd7.js} +2 -2
- package/dist/{memory-DFSo2iLi.js.map → memory-BxwOtvd7.js.map} +1 -1
- package/docs/architecture.md +4 -0
- package/package.json +3 -3
- package/skills/build-with-agent-knowledge/SKILL.md +9 -10
package/AGENTS.md
CHANGED
|
@@ -1,157 +1,52 @@
|
|
|
1
|
-
# Agent Knowledge
|
|
2
|
-
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
-
|
|
24
|
-
|
|
25
|
-
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
agent-knowledge inspect --json
|
|
54
|
-
agent-knowledge explain knowledge/concepts/example.md --json
|
|
55
|
-
agent-knowledge lint --json
|
|
56
|
-
agent-knowledge validate --strict --json
|
|
57
|
-
agent-knowledge viz --json
|
|
58
|
-
```
|
|
59
|
-
|
|
60
|
-
## Write Proposal Format
|
|
61
|
-
|
|
62
|
-
Agents should stage generated edits as FILE blocks:
|
|
63
|
-
|
|
64
|
-
```txt
|
|
65
|
-
---FILE: knowledge/concepts/example.md---
|
|
66
|
-
---
|
|
67
|
-
id: example
|
|
68
|
-
title: Example
|
|
69
|
-
sources:
|
|
70
|
-
- src_abc123
|
|
71
|
-
---
|
|
72
|
-
# Example
|
|
73
|
-
|
|
74
|
-
Sourced knowledge with links to [[Related Page]].
|
|
75
|
-
---END FILE---
|
|
76
|
-
```
|
|
77
|
-
|
|
78
|
-
The parser rejects absolute paths, `..`, control characters, and writes outside the pages directory (`knowledge/` by default).
|
|
79
|
-
Pass the same `pagesDirectory` to `applyKnowledgeWriteBlocks` that the reader uses; the file transaction enforces the same bound.
|
|
80
|
-
|
|
81
|
-
## Tool Boundary
|
|
82
|
-
|
|
83
|
-
Expose knowledge to an agent with `createKnowledgeTools({ stores, runId, retrieverVersion })`.
|
|
84
|
-
It returns `ToolDefinition[]` from `@tangle-network/agent-interface`.
|
|
85
|
-
This package owns every handler; a runtime transports the definitions and the calls and runs no knowledge loop.
|
|
86
|
-
The application supplies the roots, the lineage authority, and the effect authorization.
|
|
87
|
-
|
|
88
|
-
## Eval Boundary
|
|
89
|
-
|
|
90
|
-
Use a complete `OptimizationMethod` from `@tangle-network/agent-eval` with `runRetrievalImprovementLoop()`, `runRagOptimization()`, `optimizeKnowledgeBasePolicy()`, or `runAgentMemoryImprovement()`.
|
|
91
|
-
The method owns candidate search and resume compatibility.
|
|
92
|
-
This package owns serialized knowledge candidates, real KB or memory adapters, isolated data partitions, and safe activation.
|
|
93
|
-
|
|
94
|
-
Use `knowledgeReleaseReport()` before promotion. It folds the candidate and baseline `RunRecord[]` (plus optional traces and the gate decision) into `agent-eval` release confidence evidence.
|
|
95
|
-
|
|
96
|
-
## Integration Boundaries
|
|
97
|
-
|
|
98
|
-
- Use `KbStore` for storage. Applications may provide any durable backend that implements it.
|
|
99
|
-
- Use `new FileSystemKbStore({ root })` when opening a knowledge-base root; it keeps every record under `<root>/.agent-knowledge/` — index, event log, and per-run claim ledgers.
|
|
100
|
-
The string constructor retains the published direct-directory behavior.
|
|
101
|
-
A legacy string that names `<root>/.agent-knowledge` is canonicalized to the same root lock, so the two constructor forms cannot race on one file under different locks.
|
|
102
|
-
The store is the single writer of `index.json`; `writeKnowledgeIndex` goes through it.
|
|
103
|
-
Do not add a second writer for a record this store owns.
|
|
104
|
-
- Research state is durable state. A driver that accumulates belief across rounds takes a store and a `ledgerId` (`createPersistentResearchDrivingDriver`) so corroboration counts, contradiction edges, and open questions survive the process. `runVerifiedResearchLoop` durably announces a fold before its synchronous question generation, calls `driver.checkpoint()` before publishing the round event, and reconstructs an interrupted fold on resume.
|
|
105
|
-
- `TrackedClaim` remains the live Set-based driver API; `ResearchClaimRecord` is its sorted-array durable form. Convert at the persistence boundary rather than changing the published live shape.
|
|
106
|
-
- More than one writer per ledger means `mergeClaimLedger(id, merge)`, never `putClaimLedger`. `putClaimLedger` writes the whole record, so two writers accumulating into one ledger each write what they built from a stale read and the later write erases the earlier writer's claims. `mergeClaimLedger` holds the store's lock across read, merge, and write; `mergeClaimLedgers` is the combining rule and is commutative, associative, and idempotent, so replay and arrival order cannot change the result.
|
|
107
|
-
- One lock per store root, and a consumer with its own lock wrapper joins it rather than building a second one.
|
|
108
|
-
`withKnowledgeMutation` is reentrant per async context, `isKnowledgeMutationHeld(root)` reports whether this context already holds the root, and `runInKnowledgeMutationScope(root, hold, body)` enters the scope on a lock the caller took by its own path.
|
|
109
|
-
Inside either, every lock-taking function in this package runs inline instead of blocking against a lock the caller already holds.
|
|
110
|
-
A second lock over the same root is a second writer, whatever lockfile it uses.
|
|
111
|
-
- Grade a claim's re-executed check with `gradeFor`, and a whole pass with `gradeClaims`.
|
|
112
|
-
The verdict lattice is calibrated: a check that carries its own expected value, one killed at its deadline, and one whose output says it never reached its input are all refusals or environment verdicts, never a refutation of the claim.
|
|
113
|
-
`gradeClaims` adds the one judgment a single claim cannot make — it flags a later claim that repeats an earlier claim's check, expectation and title at the same verdict, so one verification counted N times is visible rather than silent.
|
|
114
|
-
- Use `writeFileDurable` / `writeJsonDurableWithinRoot` from the entrypoint for any file that must survive a crash. They are atomic, fsynced, and symlink-safe; a hand-rolled `writeFile` is none of those.
|
|
115
|
-
- Use `KnowledgeDiscoveryDispatcher` for research workers. Applications should connect it to their own runtime.
|
|
116
|
-
- Do not bypass `lint` or `validate` before using generated knowledge in an agent.
|
|
117
|
-
|
|
118
|
-
## Pluggable Sources + Freshness + Changes
|
|
119
|
-
|
|
120
|
-
Agents that need to stay current against external authorities should compose:
|
|
121
|
-
|
|
122
|
-
- `createCornellLiiSource({ selectors })`: US Code and Wex from law.cornell.edu.
|
|
123
|
-
- `createIrsPublicationsSource({ publications, revenueProcedures })`: IRS index and named publications.
|
|
124
|
-
- `createStateSosSource({ state, baseUrl, entities })`: generic state SOS adapter.
|
|
125
|
-
|
|
126
|
-
Every fetch returns `KnowledgeFragment[]` with `provenance.verifiable` indicating whether the configured URL returned an acceptable response and the expected content was extracted.
|
|
127
|
-
This flag does not authenticate the publisher or cryptographically prove the content.
|
|
128
|
-
Refuse to cite fragments with `verifiable: false`.
|
|
129
|
-
|
|
130
|
-
Track per-tenant freshness with `createFileSystemFreshnessStore({ root })` and re-fetch only when `stale({ workspaceId, sourceId, ttlMs })` returns true.
|
|
131
|
-
|
|
132
|
-
Diff snapshots with `detectChanges(prev, next)`.
|
|
133
|
-
Each `KnowledgeChange` carries `affectedDimensions`; pass those to your eval scheduler to run only the relevant campaigns again.
|
|
134
|
-
|
|
135
|
-
## Authorship
|
|
136
|
-
|
|
137
|
-
Do not add `Co-Authored-By:` trailers or other AI-attribution lines to commits, PR descriptions, or repository artifacts.
|
|
138
|
-
The author is the human running the session.
|
|
139
|
-
|
|
140
|
-
## Comment & doc discipline (no historical narrative)
|
|
141
|
-
|
|
142
|
-
Comments describe **what the code does and why**.
|
|
143
|
-
They must not describe what code used to do, what it replaced, which audit found a bug, or what a prior version looked like.
|
|
144
|
-
History belongs in commit messages and PR descriptions.
|
|
145
|
-
|
|
146
|
-
- Bad: `// replaces the inline retry loop`, `// fix for the silent-zero bug`, `// the 2yr rewrite added this`, `// audit fix`
|
|
147
|
-
- Good: `// value is null when retries exhaust; callers must inspect succeeded`
|
|
148
|
-
|
|
149
|
-
This applies anywhere the repository carries prose.
|
|
150
|
-
|
|
151
|
-
## No fallbacks. Fail loud.
|
|
152
|
-
|
|
153
|
-
Sloppy fallbacks corrupt every signal downstream. No silent zeros, no `?? default` on required fields, no `try/catch { return null }` that erases diagnostic info, no legacy back-compat mode defaulted on for new code.
|
|
154
|
-
|
|
155
|
-
External-boundary calls (LLM, network, FS, subprocess) return *typed outcomes* (`{ succeeded, value, error }`). Callers MUST inspect `succeeded` before using `value`. Named, opted-in fallback rotations (`policy.fallbackModels: [...]`) are fine; deep `?? "kimi"` helpers are not.
|
|
156
|
-
|
|
157
|
-
Full doctrine: `~/dotfiles/claude/AGENTS.md` → "No fallbacks. Fail loud."
|
|
1
|
+
# Agent Knowledge
|
|
2
|
+
|
|
3
|
+
## Read for the task
|
|
4
|
+
|
|
5
|
+
- For package usage, CLI commands, write proposals, or evaluation adapters, read [README.md](README.md).
|
|
6
|
+
- Before changing storage, locks, research persistence, or package ownership, read [architecture](docs/architecture.md).
|
|
7
|
+
- For run isolation and promotion, read [run-scoped citations](docs/run-scoped-citations.md).
|
|
8
|
+
- For retrieval evidence, read [knowledge-use receipts](docs/knowledge-use-receipts.md).
|
|
9
|
+
- Resolve current exports and signatures from [src/index.ts](src/index.ts), the installed declarations, and nearby tests.
|
|
10
|
+
|
|
11
|
+
## Ownership
|
|
12
|
+
|
|
13
|
+
Imports flow from `agent-runtime` to this package, then to `agent-eval` and `agent-interface`.
|
|
14
|
+
This package must not import `agent-runtime`; agent execution enters through callbacks.
|
|
15
|
+
`agent-eval` must not import this package.
|
|
16
|
+
Keep knowledge-domain storage, sources, retrieval, freshness, and candidates here.
|
|
17
|
+
Portable contracts belong to `agent-interface`; shared evaluation and improvement machinery belongs to `agent-eval`.
|
|
18
|
+
Applications own tenant roots, lineage authority, effect authorization, runtime wiring, and promotion decisions.
|
|
19
|
+
|
|
20
|
+
## Evidence and mutation
|
|
21
|
+
|
|
22
|
+
- Register sources before citing them; preserve raw evidence unchanged.
|
|
23
|
+
- Validate write proposals with intake before changing any bytes.
|
|
24
|
+
Writers and readers must use the same pages directory.
|
|
25
|
+
- Build retrieval briefs before execution and mint receipts from their actual results.
|
|
26
|
+
- Reindex page changes and propagate invalidation to citers.
|
|
27
|
+
- Write into run scope; promote through the maintained promotion API after lint and validation pass.
|
|
28
|
+
Missing sources block use and promotion.
|
|
29
|
+
- Source `verifiable` means a fetch and extraction passed configured checks.
|
|
30
|
+
It does not authenticate the publisher.
|
|
31
|
+
Refuse citations to unverifiable fragments.
|
|
32
|
+
|
|
33
|
+
## Storage and outcomes
|
|
34
|
+
|
|
35
|
+
Use `KbStore` as the record owner and durable filesystem helpers for crash-sensitive writes.
|
|
36
|
+
Every writer to one store root shares its mutation lock, including consumer lock wrappers.
|
|
37
|
+
Use the reentrant mutation scope when a caller already holds that root lock.
|
|
38
|
+
Concurrent claim-ledger writers use `mergeClaimLedger`; a whole-record replacement can erase another writer's evidence.
|
|
39
|
+
Keep the live Set-based claim API separate from its serialized records.
|
|
40
|
+
Checkpoint research state before publishing a round event so interrupted work can resume without fabricated progress.
|
|
41
|
+
|
|
42
|
+
Use the maintained claim graders.
|
|
43
|
+
Self-supplied expectations, timeouts, and checks that never reach their input do not refute a claim.
|
|
44
|
+
Repeated copies of one check are not independent evidence.
|
|
45
|
+
External failures remain typed outcomes; callers inspect success before using values.
|
|
46
|
+
Do not replace missing evidence with zero or an implicit fallback.
|
|
47
|
+
|
|
48
|
+
## Validation
|
|
49
|
+
|
|
50
|
+
Run checks relevant to changed behavior from [package.json](package.json).
|
|
51
|
+
For public package or skill changes, include the corresponding package and skill checks.
|
|
52
|
+
Keep development cases separate from final comparisons, and activate only the exact candidate that passed the decision.
|
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,38 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 14.0.1 — 2026-09-07
|
|
4
|
+
|
|
5
|
+
### Fixed
|
|
6
|
+
|
|
7
|
+
- `gradeFor` no longer refutes a claim whose promised observation is in the output because the check exited nonzero.
|
|
8
|
+
A refusal control that exits nonzero and prints the refusal it promised now grades `verified`.
|
|
9
|
+
An unreached-input signature in the output still grades `unrunnable`, and an expectation that itself names a signature keeps the existing rule.
|
|
10
|
+
|
|
11
|
+
## 14.0.0 — 2026-09-05
|
|
12
|
+
|
|
13
|
+
### Changed
|
|
14
|
+
|
|
15
|
+
- `ResearchDriver` gains optional `isComplete()`.
|
|
16
|
+
The research loop requires this result and storage readiness before it reports ready.
|
|
17
|
+
An unfinished driver receives steering rounds even when storage requirements pass.
|
|
18
|
+
Drivers without this method retain their storage readiness behavior.
|
|
19
|
+
The exported interface shape requires a major release under the package compatibility check.
|
|
20
|
+
- Requires `agent-eval` `>=0.174.0 <0.175.0` and tests against `0.174.0`.
|
|
21
|
+
This cohort uses Eval's corrected complete-method result and cost accounting contracts.
|
|
22
|
+
- Default knowledge evaluator version `2` averages only measured dimensions.
|
|
23
|
+
It omits `answer_quality` without answer evaluation, `promotion_decision` without a decision, and `blocking_readiness` without blocking requirements.
|
|
24
|
+
Consumers must handle absent dimension keys and compare scores using their evaluator version.
|
|
25
|
+
Structural-only results state that no task outcome evaluation occurred.
|
|
26
|
+
`candidate-ready` still leaves the candidate detached from the live knowledge base.
|
|
27
|
+
|
|
28
|
+
### Fixed
|
|
29
|
+
|
|
30
|
+
- Knowledge diagnosis runs before acquisition and updates.
|
|
31
|
+
One lifecycle carries findings, acquisition, and update results into final answer checks and the promotion decision.
|
|
32
|
+
Diagnosis alone does not consume final evaluation cases.
|
|
33
|
+
Disabling a required phase fails before candidate work starts.
|
|
34
|
+
Final measurement still uses frozen candidate bytes and cannot feed another adaptive update.
|
|
35
|
+
|
|
3
36
|
## 13.0.1 — 2026-09-01
|
|
4
37
|
|
|
5
38
|
### Changed
|
package/README.md
CHANGED
|
@@ -309,10 +309,18 @@ Reusing a run ID with a different implementation reference fails before cached w
|
|
|
309
309
|
Different run IDs create separate candidate workspaces, so workers can explore in parallel.
|
|
310
310
|
Promotion checks the original base hash and rejects a stale candidate instead of replacing newer work.
|
|
311
311
|
Candidate retries use `evaluateDevelopment` when provided, otherwise they use deterministic validation, readiness, and KB quality checks.
|
|
312
|
+
Diagnosis runs before acquisition and updates, using development data only.
|
|
313
|
+
Its findings and update results remain available to final answer checks and the promotion decision.
|
|
312
314
|
Development evaluation must use only train or selection data.
|
|
313
315
|
The configured `evaluate` callback and final RAG phases run once, on the first candidate that passes those development checks.
|
|
314
316
|
A failed final evaluation ends the run instead of selecting another candidate against final data.
|
|
315
317
|
|
|
318
|
+
The default evaluator reports only measured dimensions and averages those dimensions with equal weight.
|
|
319
|
+
It omits `answer_quality` without answer evaluation, `promotion_decision` without a promotion decision, and `blocking_readiness` without blocking readiness requirements.
|
|
320
|
+
Default evaluator version `2` records this weighting.
|
|
321
|
+
A candidate can pass structural checks without any task outcome evaluation; the metric notes state this limit.
|
|
322
|
+
`candidate-ready` means the configured checks passed and the candidate remains detached from the live knowledge base.
|
|
323
|
+
|
|
316
324
|
Candidate promotion currently requires Linux because it relies on Linux directory descriptors for exact file identity.
|
|
317
325
|
|
|
318
326
|
## Evaluate and improve RAG
|
package/dist/benchmarks/index.js
CHANGED
|
@@ -1,2 +1,2 @@
|
|
|
1
|
-
import { F as buildIndustryRagBenchmarkSmokeCases, H as
|
|
1
|
+
import { F as buildIndustryRagBenchmarkSmokeCases, H as createNoopMemoryBenchmarkAdapter, I as respondToIndustryMemoryBenchmarkSmokeCase, L as respondToIndustryRagBenchmarkSmokeCase, M as INDUSTRY_RAG_BENCHMARKS, N as buildFirstPartyMemoryLifecycleBenchmarkCases, P as buildIndustryMemoryBenchmarkSmokeCases, R as isKnowledgeMemoryBenchmarkCase, V as createInMemoryBenchmarkAdapter, a as buildKnowledgeBenchmarkScenarios, c as runKnowledgeBenchmarkSuite, d as summarizeKnowledgeBenchmarkCampaign, i as runMemoryAdapterBenchmark, j as INDUSTRY_MEMORY_BENCHMARKS, l as scoreKnowledgeBenchmarkArtifact, n as parseKnowledgeBenchmarkJsonl, o as knowledgeBenchmarkJudge, r as parseKnowledgeBenchmarkQrels, s as renderKnowledgeBenchmarkReportMarkdown, t as buildRetrievalBenchmarkCasesFromQrels, u as scoreMemoryBenchmarkArtifact } from "../benchmarks-BbPJdmHe.js";
|
|
2
2
|
export { INDUSTRY_MEMORY_BENCHMARKS, INDUSTRY_RAG_BENCHMARKS, buildFirstPartyMemoryLifecycleBenchmarkCases, buildIndustryMemoryBenchmarkSmokeCases, buildIndustryRagBenchmarkSmokeCases, buildKnowledgeBenchmarkScenarios, buildRetrievalBenchmarkCasesFromQrels, createInMemoryBenchmarkAdapter, createNoopMemoryBenchmarkAdapter, isKnowledgeMemoryBenchmarkCase, knowledgeBenchmarkJudge, parseKnowledgeBenchmarkJsonl, parseKnowledgeBenchmarkQrels, renderKnowledgeBenchmarkReportMarkdown, respondToIndustryMemoryBenchmarkSmokeCase, respondToIndustryRagBenchmarkSmokeCase, runKnowledgeBenchmarkSuite, runMemoryAdapterBenchmark, scoreKnowledgeBenchmarkArtifact, scoreMemoryBenchmarkArtifact, summarizeKnowledgeBenchmarkCampaign };
|
|
@@ -4,6 +4,28 @@ import { randomUUID } from "node:crypto";
|
|
|
4
4
|
import { join } from "node:path";
|
|
5
5
|
import { canonicalJson } from "@tangle-network/agent-eval";
|
|
6
6
|
import { acquireSingleRunLock, createRunCostLedger, fsCampaignStorage, resolveRunDir, runCampaign } from "@tangle-network/agent-eval/campaign";
|
|
7
|
+
//#region src/statistics.ts
|
|
8
|
+
/**
|
|
9
|
+
* The mean of a set of measurements.
|
|
10
|
+
*
|
|
11
|
+
* A measurement that is not a finite number is not a measurement: it is a
|
|
12
|
+
* scorer that divided by zero or an adapter that answered with nothing usable.
|
|
13
|
+
* Averaging it in would spread one broken probe across the whole aggregate,
|
|
14
|
+
* and carrying it through would make the aggregate itself unusable — a
|
|
15
|
+
* non-finite `scoreMean` fails every numeric comparison in the ranking
|
|
16
|
+
* comparator, so the candidate carrying it falls through to the identifier
|
|
17
|
+
* tie-break and places by name rather than by what it scored. Such values are
|
|
18
|
+
* left out, and the mean reports what was actually measured.
|
|
19
|
+
*
|
|
20
|
+
* An empty set answers `0`. That is a reported score, not an absence, and it
|
|
21
|
+
* is the behavior every caller in this package already depends on.
|
|
22
|
+
*/
|
|
23
|
+
function mean(values) {
|
|
24
|
+
const finite = values.filter(Number.isFinite);
|
|
25
|
+
if (finite.length === 0) return 0;
|
|
26
|
+
return finite.reduce((sum, value) => sum + value, 0) / finite.length;
|
|
27
|
+
}
|
|
28
|
+
//#endregion
|
|
7
29
|
//#region src/retrieval-eval.ts
|
|
8
30
|
function retrievalConfigSurface(config) {
|
|
9
31
|
return canonicalJson(config);
|
|
@@ -449,28 +471,6 @@ function rankCandidates(rows) {
|
|
|
449
471
|
}));
|
|
450
472
|
}
|
|
451
473
|
//#endregion
|
|
452
|
-
//#region src/statistics.ts
|
|
453
|
-
/**
|
|
454
|
-
* The mean of a set of measurements.
|
|
455
|
-
*
|
|
456
|
-
* A measurement that is not a finite number is not a measurement: it is a
|
|
457
|
-
* scorer that divided by zero or an adapter that answered with nothing usable.
|
|
458
|
-
* Averaging it in would spread one broken probe across the whole aggregate,
|
|
459
|
-
* and carrying it through would make the aggregate itself unusable — a
|
|
460
|
-
* non-finite `scoreMean` fails every numeric comparison in the ranking
|
|
461
|
-
* comparator, so the candidate carrying it falls through to the identifier
|
|
462
|
-
* tie-break and places by name rather than by what it scored. Such values are
|
|
463
|
-
* left out, and the mean reports what was actually measured.
|
|
464
|
-
*
|
|
465
|
-
* An empty set answers `0`. That is a reported score, not an absence, and it
|
|
466
|
-
* is the behavior every caller in this package already depends on.
|
|
467
|
-
*/
|
|
468
|
-
function mean(values) {
|
|
469
|
-
const finite = values.filter(Number.isFinite);
|
|
470
|
-
if (finite.length === 0) return 0;
|
|
471
|
-
return finite.reduce((sum, value) => sum + value, 0) / finite.length;
|
|
472
|
-
}
|
|
473
|
-
//#endregion
|
|
474
474
|
//#region src/benchmarks/utils.ts
|
|
475
475
|
function unique(values) {
|
|
476
476
|
return [...new Set(values.filter(Boolean))];
|
|
@@ -2819,6 +2819,6 @@ function defaultDocumentTarget(documentId, targetKind) {
|
|
|
2819
2819
|
}
|
|
2820
2820
|
}
|
|
2821
2821
|
//#endregion
|
|
2822
|
-
export { reserveRecoveryAttempts as A,
|
|
2822
|
+
export { reserveRecoveryAttempts as A, rankCandidates as B, sleepForMemoryRecovery as C, hasSettledPaidCall as D, assertNoInterruptedPaidCalls as E, buildIndustryRagBenchmarkSmokeCases as F, buildRetrievalEvalDispatch as G, createNoopMemoryBenchmarkAdapter as H, respondToIndustryMemoryBenchmarkSmokeCase as I, retrievalConfigSurface as J, partitionRetrievalScenarios as K, respondToIndustryRagBenchmarkSmokeCase as L, INDUSTRY_RAG_BENCHMARKS as M, buildFirstPartyMemoryLifecycleBenchmarkCases as N, readActiveAttemptJournal as O, buildIndustryMemoryBenchmarkSmokeCases as P, isKnowledgeMemoryBenchmarkCase as R, runBoundedMemoryLifecycle as S, appendDurableJournalEvent as T, memoryHitToSourceRecord as U, createInMemoryBenchmarkAdapter as V, memoryWriteResultToSourceRecord as W, scoreRetrievalArtifact as X, retrievalRecallJudge as Y, mean as Z, MEMORY_OPERATION_CANCELLATION_TIMEOUT_MS as _, buildKnowledgeBenchmarkScenarios as a, memoryRecoveryDelayMs as b, runKnowledgeBenchmarkSuite as c, summarizeKnowledgeBenchmarkCampaign as d, acquireAgentMemoryRunLease as f, MEMORY_CAMPAIGN_DISPATCH_SHUTDOWN_TIMEOUT_MS as g, DEFAULT_MEMORY_CLEANUP_TIMEOUT_MS as h, runMemoryAdapterBenchmark as i, INDUSTRY_MEMORY_BENCHMARKS as j, reconcileInterruptedMemoryPaidCalls as k, scoreKnowledgeBenchmarkArtifact as l, AgentMemoryLifecycleUnsafeError as m, parseKnowledgeBenchmarkJsonl as n, knowledgeBenchmarkJudge as o, AgentMemoryLifecycleTimeoutError as p, retrievalConfigFromSurface as q, parseKnowledgeBenchmarkQrels as r, renderKnowledgeBenchmarkReportMarkdown as s, buildRetrievalBenchmarkCasesFromQrels as t, scoreMemoryBenchmarkArtifact as u, createBoundedMemoryAdapter as v, appendAttemptJournalEvent as w, resolveMemoryCleanupTimeoutMs as x, createMemoryExecutionPool as y, normalizeUsd as z };
|
|
2823
2823
|
|
|
2824
|
-
//# sourceMappingURL=benchmarks-
|
|
2824
|
+
//# sourceMappingURL=benchmarks-BbPJdmHe.js.map
|