@tangle-network/agent-knowledge 10.6.0 → 10.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +55 -0
- package/dist/benchmarks/index.js +1 -1
- package/dist/{benchmarks-B6fCb6AD.js → benchmarks-Qk94Gnj4.js} +48 -10
- package/dist/benchmarks-Qk94Gnj4.js.map +1 -0
- package/dist/cli.js +1 -1
- package/dist/index-DZeFm-BP.d.ts.map +1 -1
- package/dist/index.d.ts +44 -2
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +43 -37
- package/dist/index.js.map +1 -1
- package/dist/{inspect-DALsvG10.js → inspect-yYqLQuM4.js} +86 -13
- package/dist/inspect-yYqLQuM4.js.map +1 -0
- package/dist/memory/index.js +2 -2
- package/dist/{memory-BRGsy2QN.js → memory-DFSo2iLi.js} +3 -12
- package/dist/memory-DFSo2iLi.js.map +1 -0
- package/package.json +7 -6
- package/dist/benchmarks-B6fCb6AD.js.map +0 -1
- package/dist/inspect-DALsvG10.js.map +0 -1
- package/dist/memory-BRGsy2QN.js.map +0 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,60 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 10.8.0 — 2026-08-22
|
|
4
|
+
|
|
5
|
+
### Fixed
|
|
6
|
+
|
|
7
|
+
- A knowledge root that canonicalizes to a different path than the caller holds now opens its file transactions instead of failing with `knowledge transaction directory escaped its root`.
|
|
8
|
+
`withTransactionRoot` compared the transaction root against the knowledge root with `resolve` on both sides.
|
|
9
|
+
`resolve` is lexical and cannot read a symbolic link, while the transaction root arrives already canonical from `withSafeDirectory`, so the two described one directory with two strings and the containment test rejected a directory that was inside the root.
|
|
10
|
+
On macOS this happened on every run that placed a knowledge base under `os.tmpdir()`, because `/var/folders/...` is a link to `/private/var/folders/...`.
|
|
11
|
+
It also happened on any platform when a caller passed a root reached through a symbolic link.
|
|
12
|
+
- `loadRunState` and `assertStateIdentity` compared a persisted root against a supplied root the same way, and reported a mismatch between two spellings of one directory.
|
|
13
|
+
|
|
14
|
+
### Added
|
|
15
|
+
|
|
16
|
+
- `relativeWithinRoot(root, candidate)`, `canonicalRelativeWithinRoot(root, candidate)`, and `canonicalPathsEqual(left, right)`.
|
|
17
|
+
`src/durable-fs.ts` now owns every comparison between two filesystem paths.
|
|
18
|
+
`canonicalRelativeWithinRoot` and `canonicalPathsEqual` canonicalize both sides first, and neither path has to exist: the deepest existing ancestor is canonicalized and the remaining segments are appended, so a directory that is about to be created is measured against the same root as one that already is.
|
|
19
|
+
Canonicalizing both sides also tightens the boundary, because a candidate that leaves the root through a symbolic link is rejected where a lexical comparison admits it.
|
|
20
|
+
- `pnpm run check:path-containment`, which fails a path comparison written inline instead of through those owners.
|
|
21
|
+
It runs inside `verify:package`, so both workflows enforce it.
|
|
22
|
+
A lexical comparison looks correct on Linux, where `/tmp` is a real directory and the two forms coincide, so this class of defect cannot be caught by running the suite on the machine that gates the merge.
|
|
23
|
+
|
|
24
|
+
## 10.7.1 — 2026-08-22
|
|
25
|
+
|
|
26
|
+
### Changed
|
|
27
|
+
|
|
28
|
+
- Accept Eval `>=0.170.0 <0.171.0`, replacing Eval `>=0.163.2 <0.164.0`, and Interface `^1.6.0`, replacing Interface `^1.4.0`.
|
|
29
|
+
The old Eval range admitted exactly one published version while Eval `latest` was 0.170.0, so a consumer installing the cohort got an unmet peer on this package and could not complete the install.
|
|
30
|
+
- **A consumer must move Eval and Interface with this package.**
|
|
31
|
+
Interface moves because Eval 0.170.0 depends on agent-core 0.9.5, which requires Interface `^1.5.0`.
|
|
32
|
+
Interface 1.4.0 leaves a second physical copy of the contract package in the tree, and one contract package must resolve to one copy.
|
|
33
|
+
Move Eval to 0.170.x and Interface to 1.6.x in the same change.
|
|
34
|
+
- The new Eval ceiling is measured, not assumed.
|
|
35
|
+
This package imports 79 distinct symbols from Eval across five entry points: `.` (37), `/campaign` (35), `/rl` (4), `/experiment` (2), and `/analyst` (1), plus one dynamic `import()` of `/campaign`.
|
|
36
|
+
A compiler probe over every symbol gives the same result against 0.163.2 and against 0.170.0: 78 resolve, and `JsonValue` resolves in neither, because two test files import it from `/campaign` where it has never been exported.
|
|
37
|
+
Eval 0.170.0 removes six `/analyst` exports, which are `createJudgeAdapter`, `createRunCriticAdapter`, `createVerifierAdapter`, and their three option types.
|
|
38
|
+
This package imports none of the six.
|
|
39
|
+
All 15 symbols this package imports from Interface resolve at 1.6.0 exactly as they do at 1.4.0.
|
|
40
|
+
- A pre-1.0 dependency earns a single-minor range, so `<0.171.0` is the boundary the evidence covers.
|
|
41
|
+
An Eval minor above 0.170.x must be verified before the range admits it.
|
|
42
|
+
|
|
43
|
+
## 10.7.0 — 2026-08-21
|
|
44
|
+
|
|
45
|
+
### Added
|
|
46
|
+
|
|
47
|
+
- `createKnowledgeControlLoopAdapter` now returns `stopPolicies.stateFingerprint`. `observe` hands the caller's `AbortSignal` to `act` through the loop state, and Eval's control runtime fingerprints that state with RFC 8785 canonical JSON, which refuses a class instance. Eval's contract names this exact case: a caller whose state is not plain JSON supplies its own fingerprint. The adapter fingerprints the observable knowledge and leaves the signal out, so a caller spreading `...adapter` needs to know none of this.
|
|
48
|
+
|
|
49
|
+
### Fixed
|
|
50
|
+
|
|
51
|
+
- Nothing this package writes encodes an absent optional as `undefined` any more. `createKnowledgeEvent` (`actor`, `target`, `metadata`), the research-loop step (`notes`, `applied`, `readiness`, `metadata`), its event metadata (`written`), and readiness requirement metadata (`validUntil`, `lastVerifiedAt`) all omit the key instead. A key present with no value and an absent key are the same JSON, so the canonical encoder refuses to guess between them — and every control-loop run through `createKnowledgeControlLoopAdapter` aborted at step 0 because of it. `createKnowledgeEvent` is the single owner of every event this package emits, so the fix there covers all of them.
|
|
52
|
+
|
|
53
|
+
### Changed
|
|
54
|
+
|
|
55
|
+
- Accept Eval `>=0.163.2 <0.164.0` and Interface `^1.4.0`, replacing Eval `>=0.149.0 <0.150.0` and Interface `^1.1.0`. Eval 0.150.0 through 0.163.2 removed the paid model transports, deleted 86 unused exports, and changed the raw-finding codec, the GEPA bridge output, and the GEPA engine seed; this package builds and tests against 0.163.2, so the peer range now states that.
|
|
56
|
+
- **A consumer must move Eval and Interface with this package.** Installing 10.7.0 beside Eval 0.149.x is an unmet peer, not a warning to ignore: the GEPA bridge output shape differs between the two. Move Eval to 0.163.2 and Interface to 1.4.0 in the same change.
|
|
57
|
+
|
|
3
58
|
## 10.6.0 — 2026-08-21
|
|
4
59
|
|
|
5
60
|
### Added
|
package/dist/benchmarks/index.js
CHANGED
|
@@ -1,2 +1,2 @@
|
|
|
1
|
-
import {
|
|
1
|
+
import { F as buildIndustryRagBenchmarkSmokeCases, H as createInMemoryBenchmarkAdapter, I as respondToIndustryMemoryBenchmarkSmokeCase, L as respondToIndustryRagBenchmarkSmokeCase, M as INDUSTRY_RAG_BENCHMARKS, N as buildFirstPartyMemoryLifecycleBenchmarkCases, P as buildIndustryMemoryBenchmarkSmokeCases, R as isKnowledgeMemoryBenchmarkCase, U as createNoopMemoryBenchmarkAdapter, a as buildKnowledgeBenchmarkScenarios, c as runKnowledgeBenchmarkSuite, d as summarizeKnowledgeBenchmarkCampaign, i as runMemoryAdapterBenchmark, j as INDUSTRY_MEMORY_BENCHMARKS, l as scoreKnowledgeBenchmarkArtifact, n as parseKnowledgeBenchmarkJsonl, o as knowledgeBenchmarkJudge, r as parseKnowledgeBenchmarkQrels, s as renderKnowledgeBenchmarkReportMarkdown, t as buildRetrievalBenchmarkCasesFromQrels, u as scoreMemoryBenchmarkArtifact } from "../benchmarks-Qk94Gnj4.js";
|
|
2
2
|
export { INDUSTRY_MEMORY_BENCHMARKS, INDUSTRY_RAG_BENCHMARKS, buildFirstPartyMemoryLifecycleBenchmarkCases, buildIndustryMemoryBenchmarkSmokeCases, buildIndustryRagBenchmarkSmokeCases, buildKnowledgeBenchmarkScenarios, buildRetrievalBenchmarkCasesFromQrels, createInMemoryBenchmarkAdapter, createNoopMemoryBenchmarkAdapter, isKnowledgeMemoryBenchmarkCase, knowledgeBenchmarkJudge, parseKnowledgeBenchmarkJsonl, parseKnowledgeBenchmarkQrels, renderKnowledgeBenchmarkReportMarkdown, respondToIndustryMemoryBenchmarkSmokeCase, respondToIndustryRagBenchmarkSmokeCase, runKnowledgeBenchmarkSuite, runMemoryAdapterBenchmark, scoreKnowledgeBenchmarkArtifact, scoreMemoryBenchmarkArtifact, summarizeKnowledgeBenchmarkCampaign };
|
|
@@ -422,12 +422,56 @@ function renderMemoryHits(hits) {
|
|
|
422
422
|
}).join("\n\n");
|
|
423
423
|
}
|
|
424
424
|
//#endregion
|
|
425
|
-
//#region src/
|
|
425
|
+
//#region src/candidate-ranking.ts
|
|
426
|
+
/**
|
|
427
|
+
* Round a US-dollar amount to the precision a cost ledger records.
|
|
428
|
+
*
|
|
429
|
+
* Costs are accumulated by addition, so two runs that spent the same amount
|
|
430
|
+
* can differ in the last bits of the float. Ranking breaks a tie on cost, and
|
|
431
|
+
* without a common precision that float noise decides the order.
|
|
432
|
+
*/
|
|
433
|
+
function normalizeUsd(value) {
|
|
434
|
+
return Number(value.toFixed(12));
|
|
435
|
+
}
|
|
436
|
+
/**
|
|
437
|
+
* Order candidates best-first and stamp each row with the rank it earned.
|
|
438
|
+
*
|
|
439
|
+
* A candidate with a failed cell measured less than it claims to have
|
|
440
|
+
* measured, so it places below every candidate that completed. Among those,
|
|
441
|
+
* the higher mean score wins, then the higher pass rate, then the lower cost.
|
|
442
|
+
* The candidate id breaks the final tie, so ranking the same results twice
|
|
443
|
+
* gives the same order.
|
|
444
|
+
*/
|
|
445
|
+
function rankCandidates(rows) {
|
|
446
|
+
return [...rows].sort((a, b) => Number(a.cellsFailed > 0) - Number(b.cellsFailed > 0) || b.scoreMean - a.scoreMean || b.passRate - a.passRate || a.totalCostUsd - b.totalCostUsd || a.candidateId.localeCompare(b.candidateId)).map((row, index) => ({
|
|
447
|
+
...row,
|
|
448
|
+
rank: index + 1
|
|
449
|
+
}));
|
|
450
|
+
}
|
|
451
|
+
//#endregion
|
|
452
|
+
//#region src/statistics.ts
|
|
453
|
+
/**
|
|
454
|
+
* The mean of a set of measurements.
|
|
455
|
+
*
|
|
456
|
+
* A measurement that is not a finite number is not a measurement: it is a
|
|
457
|
+
* scorer that divided by zero or an adapter that answered with nothing usable.
|
|
458
|
+
* Averaging it in would spread one broken probe across the whole aggregate,
|
|
459
|
+
* and carrying it through would make the aggregate itself unusable — a
|
|
460
|
+
* non-finite `scoreMean` fails every numeric comparison in the ranking
|
|
461
|
+
* comparator, so the candidate carrying it falls through to the identifier
|
|
462
|
+
* tie-break and places by name rather than by what it scored. Such values are
|
|
463
|
+
* left out, and the mean reports what was actually measured.
|
|
464
|
+
*
|
|
465
|
+
* An empty set answers `0`. That is a reported score, not an absence, and it
|
|
466
|
+
* is the behavior every caller in this package already depends on.
|
|
467
|
+
*/
|
|
426
468
|
function mean(values) {
|
|
427
469
|
const finite = values.filter(Number.isFinite);
|
|
428
470
|
if (finite.length === 0) return 0;
|
|
429
471
|
return finite.reduce((sum, value) => sum + value, 0) / finite.length;
|
|
430
472
|
}
|
|
473
|
+
//#endregion
|
|
474
|
+
//#region src/benchmarks/utils.ts
|
|
431
475
|
function unique(values) {
|
|
432
476
|
return [...new Set(values.filter(Boolean))];
|
|
433
477
|
}
|
|
@@ -435,9 +479,6 @@ function formatNumber(value) {
|
|
|
435
479
|
if (!Number.isFinite(value)) return "0";
|
|
436
480
|
return value.toFixed(value === 0 || Math.abs(value) >= 10 ? 0 : 3);
|
|
437
481
|
}
|
|
438
|
-
function normalizeUsd(value) {
|
|
439
|
-
return Number(value.toFixed(12));
|
|
440
|
-
}
|
|
441
482
|
function compactObject(value) {
|
|
442
483
|
if (Array.isArray(value)) return value.map(compactObject);
|
|
443
484
|
if (!value || typeof value !== "object") return value;
|
|
@@ -2591,15 +2632,12 @@ async function runOwnedMemoryAdapterBenchmark(options, storage, runDir, lease, c
|
|
|
2591
2632
|
await lease.assertOwned();
|
|
2592
2633
|
}
|
|
2593
2634
|
const costByCandidate = memoryAdapterBenchmarkCostByCandidate(costLedger, runDir, [...options.candidates, ...options.recoveryCandidates ?? []]);
|
|
2594
|
-
const ranked = rows.map((row) => {
|
|
2635
|
+
const ranked = rankCandidates(rows.map((row) => {
|
|
2595
2636
|
const totalCostUsd = normalizeUsd(costByCandidate.get(row.candidateId) ?? 0);
|
|
2596
2637
|
return {
|
|
2597
2638
|
...row,
|
|
2598
2639
|
totalCostUsd
|
|
2599
2640
|
};
|
|
2600
|
-
}).sort((a, b) => Number(a.cellsFailed > 0) - Number(b.cellsFailed > 0) || b.scoreMean - a.scoreMean || b.passRate - a.passRate || a.totalCostUsd - b.totalCostUsd || a.candidateId.localeCompare(b.candidateId)).map((row, index) => ({
|
|
2601
|
-
...row,
|
|
2602
|
-
rank: index + 1
|
|
2603
2641
|
}));
|
|
2604
2642
|
const rankingJsonPath = join(runDir, "memory-adapter-ranking.json");
|
|
2605
2643
|
const rankingMarkdownPath = join(runDir, "memory-adapter-ranking.md");
|
|
@@ -2781,6 +2819,6 @@ function defaultDocumentTarget(documentId, targetKind) {
|
|
|
2781
2819
|
}
|
|
2782
2820
|
}
|
|
2783
2821
|
//#endregion
|
|
2784
|
-
export { reserveRecoveryAttempts as A,
|
|
2822
|
+
export { reserveRecoveryAttempts as A, normalizeUsd as B, sleepForMemoryRecovery as C, hasSettledPaidCall as D, assertNoInterruptedPaidCalls as E, buildIndustryRagBenchmarkSmokeCases as F, memoryWriteResultToSourceRecord as G, createInMemoryBenchmarkAdapter as H, respondToIndustryMemoryBenchmarkSmokeCase as I, retrievalConfigFromSurface as J, buildRetrievalEvalDispatch as K, respondToIndustryRagBenchmarkSmokeCase as L, INDUSTRY_RAG_BENCHMARKS as M, buildFirstPartyMemoryLifecycleBenchmarkCases as N, readActiveAttemptJournal as O, buildIndustryMemoryBenchmarkSmokeCases as P, isKnowledgeMemoryBenchmarkCase as R, runBoundedMemoryLifecycle as S, appendDurableJournalEvent as T, createNoopMemoryBenchmarkAdapter as U, rankCandidates as V, memoryHitToSourceRecord as W, retrievalRecallJudge as X, retrievalConfigSurface as Y, scoreRetrievalArtifact as Z, MEMORY_OPERATION_CANCELLATION_TIMEOUT_MS as _, buildKnowledgeBenchmarkScenarios as a, memoryRecoveryDelayMs as b, runKnowledgeBenchmarkSuite as c, summarizeKnowledgeBenchmarkCampaign as d, acquireAgentMemoryRunLease as f, MEMORY_CAMPAIGN_DISPATCH_SHUTDOWN_TIMEOUT_MS as g, DEFAULT_MEMORY_CLEANUP_TIMEOUT_MS as h, runMemoryAdapterBenchmark as i, INDUSTRY_MEMORY_BENCHMARKS as j, reconcileInterruptedMemoryPaidCalls as k, scoreKnowledgeBenchmarkArtifact as l, AgentMemoryLifecycleUnsafeError as m, parseKnowledgeBenchmarkJsonl as n, knowledgeBenchmarkJudge as o, AgentMemoryLifecycleTimeoutError as p, partitionRetrievalScenarios as q, parseKnowledgeBenchmarkQrels as r, renderKnowledgeBenchmarkReportMarkdown as s, buildRetrievalBenchmarkCasesFromQrels as t, scoreMemoryBenchmarkArtifact as u, createBoundedMemoryAdapter as v, appendAttemptJournalEvent as w, resolveMemoryCleanupTimeoutMs as x, createMemoryExecutionPool as y, mean as z };
|
|
2785
2823
|
|
|
2786
|
-
//# sourceMappingURL=benchmarks-
|
|
2824
|
+
//# sourceMappingURL=benchmarks-Qk94Gnj4.js.map
|