agent-inspect 6.29.3 → 6.29.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,16 @@
1
1
  # Changelog
2
2
 
3
+ ## 6.29.4
4
+
5
+ ### Patch Changes
6
+
7
+ - d7a3801: Trust and package integrity for 6.29.4:
8
+
9
+ - Advertise honest MCP `tools/list` annotations (readOnlyHint / openWorldHint / destructiveHint) without treating hints as authorization.
10
+ - Ship an exact LICENSE copy in every public fixed-group tarball and enforce it in pack-smoke / docs gates.
11
+ - Keep Evidence demo generator versions synced through `version:packages` (`demo:generate` + verify).
12
+ - Document run comparability and transport retry (429) evaluation guidance; AgentInspect evaluates retries, it does not execute them.
13
+
3
14
  ## 6.29.3
4
15
 
5
16
  ### Patch Changes
package/README.md CHANGED
@@ -212,7 +212,7 @@ The root package is enough for custom capture, the CLI, checks, and Evidence wor
212
212
 
213
213
  ## Status and documentation
214
214
 
215
- **Current published baseline:** **6.29.3** · persisted schema `1.0` · Node.js `>=20` · MIT.
215
+ **Current published baseline:** **6.29.4** · persisted schema `1.0` · Node.js `>=20` · MIT.
216
216
 
217
217
  Legacy v0.1 and v0.2 traces remain readable. Check the npm badge and [changelog](CHANGELOG.md) for the current published version.
218
218
 
@@ -10,6 +10,7 @@ This document states what AgentInspect **does not** provide today. It complement
10
10
  - **No automatic universal instrumentation** of every framework: integration is explicit (manual traces, log ingest, optional adapters).
11
11
  - **No retry execution engine:** TraceContract `retry` rules (including additive `retry.operations[]` recovery oracles in 6.27) evaluate attempt identity and evidence only. AgentInspect does not retry tools, remediates nothing, and does not treat a client `idempotencyKey` as proof of exactly-once writes. Write-tool timeout/unknown completion is unevaluable without authoritative idempotency evidence; write-retry is not safe by default.
12
12
  - **Transport facts are source-boundary only:** HTTP status, server-advertised retry delay, selected delay, and their sources must be captured at the transport/fetch boundary. An MCP wrapper cannot recover facts already collapsed into an SDK error message.
13
+ - **Run comparability is optional guidance:** Ordinary trajectory checks do not fail when cohort/prompt/tool commitments are absent. See [RUN-COMPARABILITY.md](./RUN-COMPARABILITY.md). A machine-readable comparable-run profile is not a default gate in this line.
13
14
  - **Omitted-payload digests** (`createOmittedPayloadCommitment`) are bounded (1 MiB preflight). Digests prove omitted bytes existed; they are not redaction or authorization.
14
15
 
15
16
  ## Correlation metadata (v1.3.0)
@@ -0,0 +1,54 @@
1
+ # Run comparability (guidance)
2
+
3
+ AgentInspect trajectory checks are deterministic **given the evidence they receive**. That does not mean two agent runs are automatically comparable. A changed path may reflect nondeterministic tools, sampling, prompt drift, or missing capture—not only an application regression.
4
+
5
+ Ordinary `check` / TraceContract evaluation does **not** fail when comparability provenance is absent.
6
+
7
+ ## Classification
8
+
9
+ | Class | Meaning |
10
+ | --- | --- |
11
+ | `comparable` | Enough shared identity and commitments to interpret path differences as meaningful |
12
+ | `partially comparable` | Shared cohort/test case, but missing prompt/tools/model commitments |
13
+ | `not comparable` | Identity conflicts or insufficient linkage |
14
+ | `unknown` | No comparability metadata |
15
+
16
+ ## Optional provenance fields
17
+
18
+ Attach as run or step metadata when you need honest before/after interpretation:
19
+
20
+ - `cohortId`, `testCaseId`, `attempt`
21
+ - `promptTemplateVersion` (or a content hash commitment)
22
+ - per-LLM-step **input commitment** (hash; prefer not to store raw prompts)
23
+ - upstream **tool-output commitment** (hash)
24
+ - `availableToolsCommitment`
25
+ - `requestedModel` / `resolvedModel`
26
+ - SDK/framework version
27
+ - sampling parameters (temperature, top_p, …)
28
+ - system fingerprint / finish reason when the provider exposes them
29
+ - contract digest
30
+ - outcome evaluator method/version
31
+
32
+ ## Privacy
33
+
34
+ - Hashes show equality, not content.
35
+ - Low-entropy values may still be guessed.
36
+ - Caller-owned HMAC is allowed; AgentInspect does not manage keys.
37
+ - Cohort labels alone are insufficient.
38
+ - Temperature `0` is not proof of determinism.
39
+ - A digest does not prove complete capture.
40
+
41
+ ## Retry and transport identity
42
+
43
+ Do not infer retries by matching tool names across steps. Prefer explicit:
44
+
45
+ - `operationId`, `attemptId`, `attemptNumber`, `retryOf`
46
+ - HTTP status, server-advertised delay, selected delay, delay source
47
+
48
+ The MCP client wrapper records each call as a separate operation unless the application supplies that identity. AgentInspect **evaluates** retry evidence; it does not execute retries.
49
+
50
+ See recipe `examples/recipes/comparable-cohort-runs/` and `examples/recipes/mcp-transport-retry-429/`.
51
+
52
+ ## Future profile
53
+
54
+ A machine-readable comparable-run profile remains gated for `6.30.0` and requires external validation before it can become a default gate.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "agent-inspect",
3
- "version": "6.29.3",
3
+ "version": "6.29.4",
4
4
  "license": "MIT",
5
5
  "type": "module",
6
6
  "description": "Local evidence debugger and trajectory-test toolkit for TypeScript AI agents — execution trees, TraceContract checks, Evidence v2, and read-only MCP",
@@ -152,6 +152,7 @@
152
152
  "docs/ADAPTERS.md",
153
153
  "docs/ADAPTER-CONFORMANCE.md",
154
154
  "docs/COMPARE.md",
155
+ "docs/RUN-COMPARABILITY.md",
155
156
  "docs/LOG-TO-TREE-QUICKSTART.md",
156
157
  "docs/SCREENSHOTS.md",
157
158
  "docs/SAFE-TRACE-SHARING.md",
@@ -235,7 +236,7 @@
235
236
  "test:all": "pnpm run typecheck && pnpm run linked-versions:check && pnpm run build && pnpm run test && pnpm run size",
236
237
  "prepublish:checks": "pnpm run typecheck && pnpm run test && pnpm run test:coverage && pnpm run build && pnpm run fixtures:check && pnpm run recipes:check && pnpm run size && pnpm run linked-versions:check && pnpm run repo:health && pnpm run pack:smoke",
237
238
  "pack:dry-run": "pnpm run build && npm pack --dry-run",
238
- "pack:smoke": "pnpm run build && node scripts/package-smoke.mjs && node scripts/packed-openai-agents-e2e.mjs && node scripts/packed-ai-sdk-e2e.mjs && node scripts/packed-mcp-e2e.mjs && node scripts/packed-quickstart-e2e.mjs && node scripts/packed-semantic-loop-e2e.mjs && node scripts/packed-swarm-loop-e2e.mjs && node scripts/evidence-ci-golden-paths.mjs",
239
+ "pack:smoke": "pnpm run package-licenses:check && pnpm run build && node scripts/package-smoke.mjs && node scripts/packed-openai-agents-e2e.mjs && node scripts/packed-ai-sdk-e2e.mjs && node scripts/packed-mcp-e2e.mjs && node scripts/packed-quickstart-e2e.mjs && node scripts/packed-semantic-loop-e2e.mjs && node scripts/packed-swarm-loop-e2e.mjs && node scripts/evidence-ci-golden-paths.mjs",
239
240
  "linked-versions:check": "node scripts/check-linked-versions.mjs",
240
241
  "docs:commands": "node scripts/validate-doc-commands.mjs",
241
242
  "docs:links": "node scripts/validate-doc-links.mjs",
@@ -244,7 +245,7 @@
244
245
  "ai-assets:check": "node scripts/validate-ai-assets.mjs",
245
246
  "package-readmes:check": "node scripts/validate-package-readmes.mjs",
246
247
  "actions:check": "node scripts/check-github-actions-pins.mjs",
247
- "docs:check": "pnpm run docs:commands && pnpm run docs:links && pnpm run public-truth:check && pnpm run ai-assets:check && pnpm run package-readmes:check && pnpm run repo:health && pnpm run actions:check && pnpm run build:cli && pnpm run demo:verify",
248
+ "docs:check": "pnpm run docs:commands && pnpm run docs:links && pnpm run public-truth:check && pnpm run ai-assets:check && pnpm run package-readmes:check && pnpm run package-licenses:check && pnpm run build:cli && pnpm run demo-version-sync:check && pnpm run repo:health && pnpm run actions:check && pnpm run demo:verify",
248
249
  "build:cli": "pnpm exec tsup --config tsup.cli.config.ts",
249
250
  "repo:health": "node scripts/validate-repo-health.mjs",
250
251
  "demo:generate": "node scripts/demo-generate.mjs",
@@ -256,10 +257,13 @@
256
257
  "perf:baseline": "node scripts/performance-baseline.mjs",
257
258
  "examples:check": "pnpm install && pnpm --filter agent-inspect-example-01-basic run start",
258
259
  "changeset": "changeset",
259
- "version:packages": "changeset version && pnpm public-truth:sync && pnpm public-truth:check",
260
+ "version:packages": "changeset version && pnpm public-truth:sync && pnpm build && pnpm demo:generate && pnpm public-truth:check && pnpm demo:verify && pnpm package-licenses:check",
260
261
  "release": "changeset publish",
261
262
  "website:dev": "pnpm --filter @agent-inspect/website dev",
262
263
  "website:build": "pnpm --filter @agent-inspect/website build",
263
- "website:typecheck": "pnpm --filter @agent-inspect/website typecheck"
264
+ "website:typecheck": "pnpm --filter @agent-inspect/website typecheck",
265
+ "package-licenses:check": "node scripts/validate-package-licenses.mjs",
266
+ "package-licenses:sync": "node scripts/sync-package-licenses.mjs",
267
+ "demo-version-sync:check": "node scripts/check-demo-version-sync.mjs"
264
268
  }
265
269
  }
@@ -7206,8 +7206,9 @@ ${nav}
7206
7206
  <p>Artifact status: <span class="badge ${statusClass(input.assessmentStatus)}">${escapeHtml(input.assessmentStatus)}</span>
7207
7207
  ${input.sourceStatus !== void 0 ? ` \xB7 Source status: <span class="badge ${statusClass(input.sourceStatus)}">${escapeHtml(input.sourceStatus)}</span>` : ""}</p>
7208
7208
  <p>Profile: <code>${escapeHtml(input.redactionProfile)}</code> \xB7 Verification: <code>${escapeHtml(input.verificationPolicy)}</code></p>
7209
- ${input.contractBinding !== void 0 ? `<p>Contract binding: <code>${escapeHtml(input.contractBinding.status)}</code>${input.contractBinding.ruleCount !== void 0 ? ` \xB7 Rules: <code>${escapeHtml(String(input.contractBinding.ruleCount))}</code>` : ""}${input.contractBinding.sha256 !== void 0 ? ` \xB7 Digest: <code>${escapeHtml(input.contractBinding.sha256.slice(0, 16))}\u2026</code>` : ""}</p>${input.contractBinding.note !== void 0 && input.contractBinding.note.trim() !== "" ? `<p class="muted">${escapeHtml(input.contractBinding.note)}</p>` : ""}` : ""}
7210
- <p>Generator: <code>${escapeHtml(input.generatorName)}@${escapeHtml(input.generatorVersion)}</code>
7209
+ ${input.contractBinding !== void 0 ? ` <p>Contract binding: <code>${escapeHtml(input.contractBinding.status)}</code>${input.contractBinding.ruleCount !== void 0 ? ` \xB7 Rules: <code>${escapeHtml(String(input.contractBinding.ruleCount))}</code>` : ""}${input.contractBinding.sha256 !== void 0 ? ` \xB7 Digest: <code>${escapeHtml(input.contractBinding.sha256.slice(0, 16))}\u2026</code>` : ""}</p>${input.contractBinding.note !== void 0 && input.contractBinding.note.trim() !== "" ? `
7210
+ <p class="muted">${escapeHtml(input.contractBinding.note)}</p>` : ""}
7211
+ ` : ""} <p>Generator: <code>${escapeHtml(input.generatorName)}@${escapeHtml(input.generatorVersion)}</code>
7211
7212
  ${input.createdAt ? ` \xB7 Created: <code>${escapeHtml(input.createdAt)}</code>` : ""}</p>
7212
7213
  <h3>Runs</h3>
7213
7214
  <ul class="runs">${runList}</ul>
@@ -12488,5 +12489,5 @@ function renderGateReport(result, options = {}) {
12488
12489
  }
12489
12490
 
12490
12491
  export { COHORT_METRIC_IDS, DEFAULT_SUITE_ARTIFACTS_DIR, EVIDENCE_FORMAT_VERSION, EVIDENCE_HTML_FILENAME, EVIDENCE_MANIFEST_FILENAME, Redactor, TraceDirectory, TraceReadError, TreeBuilder, aggregateBundleSafeStatus, aggregateSessionCheckResults, analyzeCohort, applyProfileMetadataCaps, assertBundlePathContained, assertEvidenceRelativePath, buildActivitySummary, buildBundleMetadata, buildBundleSummaryMarkdown, buildEvidenceCausalFailureViewHtml, buildEvidenceCiPackage, buildEvidenceCircuitViewHtml, buildEvidenceContractsViewHtml, buildEvidenceDiffViewHtml, buildEvidenceHtmlShell, buildEvidenceManifest, buildEvidenceOutcomesViewHtml, buildEvidenceProvenanceViewHtml, buildEvidenceSafetyViewHtml, buildEvidenceTimelineViewHtml, buildEvidenceToolsLlmViewHtml, buildEvidenceTreeViewHtml, buildLocalExplanation, buildPlaceholderArtifact, buildRunSummary, buildRunTimeline, buildRunWhatSummary, buildSessionIndex, buildTraceStats, buildZipArchive, bundleFailsOnSafety, bundleRunAssetRelativePath, collectTraceSchemaVersions, compactAttributes, createBaselineRegressionRule, createLlmUsageRule, createMaxStepDurationRule, createObservedOutcomeRule, createRequireCompletedRule, createRunDepthRule, createRunDurationRule, createRunStatusRule, createSafetyOversizedAttributeRule, createSafetyRawContentRule, createSafetyRedactionRule, createSafetySecretPatternRule, createStallDetectionRule, createStructureCycleRule, createStructureOrphanRule, createStructureParallelWidthRule, createStructureRelationshipRule, createToolUsageRule, defaultBundleOutputPath, defaultSuiteConfigTemplate, diffRuns, diffTraceEvents, enrichSessionRunRecord, escapeHtml, escapeMarkdown, extractMetadata, extractOutcomesFromTraceEvents, filterMetasBySessionScope, filterTraces, flattenTree, formatDuration2 as formatDuration, formatStepLabel, formatTimestamp, gateHasThresholds, getIndent, getTraceFilePath, inferEvidenceFileRole, isAgentInspectTrace, isPersistedInspectEvent, loadSessionRunRecords, loadSuiteConfig, loadTraceMetadataList, manualTraceEventsToComparableRun, nanoid, normalizeBundleOutputPath, openTrace, parseCohortMetricList, parseDuration, parseDurationFilter, parseGateList, parseTraceJsonl, persistedInspectEventsToTraceEvents, renderActivitySummaryHuman, renderCohortReport, renderErrorLine, renderGateReport, renderObservedOutcomesHtml, renderObservedOutcomesMarkdown, renderRunDiff, renderRunWhat, renderStepLine, renderSuiteReport, renderTimeline, renderTraceStats, resolveBundleRunIds, resolveRedactionProfile, resolveSuiteTemplate, resolveTraceDir, runGate, runSuite, runTraceChecks, safeString, sanitizeBundleRunId, searchTraces, serializeEvidenceManifest, sha256Hex, stableJson, summarizeObservedOutcomes, summarizeSemanticParity, traceEventToPersistedInspectEvent, truncateName, truncateStringForProfile, validateEvent, validateSuiteConfig, verifyEvidenceDirectory, zeroKinds };
12491
- //# sourceMappingURL=chunk-TBTRH7KY.mjs.map
12492
- //# sourceMappingURL=chunk-TBTRH7KY.mjs.map
12492
+ //# sourceMappingURL=chunk-FJRUJSST.mjs.map
12493
+ //# sourceMappingURL=chunk-FJRUJSST.mjs.map