agent-inspect 6.30.0 → 6.31.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +12 -0
- package/README.md +1 -1
- package/docs/COMPARE.md +1 -1
- package/docs/FIRST-TRACE-IN-5-MINUTES.md +21 -17
- package/docs/GETTING-STARTED.md +1 -1
- package/docs/TRACE-CONTRACTS.md +25 -0
- package/package.json +4 -2
- package/packages/cli/dist/{chunk-BDRJENLM.mjs → chunk-SJ6FRNOE.mjs} +201 -49
- package/packages/cli/dist/chunk-SJ6FRNOE.mjs.map +1 -0
- package/packages/cli/dist/index.cjs +288 -50
- package/packages/cli/dist/index.cjs.map +1 -1
- package/packages/cli/dist/index.mjs +91 -5
- package/packages/cli/dist/index.mjs.map +1 -1
- package/packages/cli/dist/{src-ZUGAUXY6.mjs → src-YNEJOPRA.mjs} +3 -3
- package/packages/cli/dist/{src-ZUGAUXY6.mjs.map → src-YNEJOPRA.mjs.map} +1 -1
- package/packages/core/dist/advanced.cjs +39 -1
- package/packages/core/dist/advanced.cjs.map +1 -1
- package/packages/core/dist/advanced.d.cts +2 -2
- package/packages/core/dist/advanced.d.ts +2 -2
- package/packages/core/dist/advanced.mjs +2 -2
- package/packages/core/dist/checks.cjs +207 -47
- package/packages/core/dist/checks.cjs.map +1 -1
- package/packages/core/dist/checks.d.cts +2 -2
- package/packages/core/dist/checks.d.ts +2 -2
- package/packages/core/dist/checks.mjs +1 -1
- package/packages/core/dist/{chunk-OOX25GTH.mjs → chunk-CBPWLHNQ.mjs} +209 -50
- package/packages/core/dist/{chunk-OOX25GTH.mjs.map → chunk-CBPWLHNQ.mjs.map} +1 -1
- package/packages/core/dist/{index-7hG_qggp.d.ts → index-CgOAGuM5.d.ts} +111 -1
- package/packages/core/dist/{index-CQJUNPPe.d.cts → index-EQihlEei.d.cts} +111 -1
- package/packages/cli/dist/chunk-BDRJENLM.mjs.map +0 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,17 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 6.31.1
|
|
4
|
+
|
|
5
|
+
### Patch Changes
|
|
6
|
+
|
|
7
|
+
- 11a0159: Strict `allowedStatuses` validation (typos like `succes` no longer map to permissive `error`); portable `prepublishOnly` skip runner for Trusted Publish; LangChain `handleChainStart` parentage matches CallbackManager runtime arg order.
|
|
8
|
+
|
|
9
|
+
## 6.31.0
|
|
10
|
+
|
|
11
|
+
### Minor Changes
|
|
12
|
+
|
|
13
|
+
- 11a8b90: Add TraceContract `steps.orderRelations` for typed TOOL↔LLM step ordering (additive; TOOL-only `requiredOrder` / `orderRules` unchanged).
|
|
14
|
+
|
|
3
15
|
## 6.30.0
|
|
4
16
|
|
|
5
17
|
### Minor Changes
|
package/README.md
CHANGED
|
@@ -212,7 +212,7 @@ The root package is enough for custom capture, the CLI, checks, and Evidence wor
|
|
|
212
212
|
|
|
213
213
|
## Status and documentation
|
|
214
214
|
|
|
215
|
-
**Current published baseline:** **6.
|
|
215
|
+
**Current published baseline:** **6.31.1** · persisted schema `1.0` · Node.js `>=20` · MIT.
|
|
216
216
|
|
|
217
217
|
Legacy v0.1 and v0.2 traces remain readable. Check the npm badge and [changelog](CHANGELOG.md) for the current published version.
|
|
218
218
|
|
package/docs/COMPARE.md
CHANGED
|
@@ -120,7 +120,7 @@ AgentInspect avoids SDK/collector setup for local debugging:
|
|
|
120
120
|
|
|
121
121
|
## v3.5 positioning (local inner loop)
|
|
122
122
|
|
|
123
|
-
AgentInspect v3.5 is the **adoption release**. The npm map is one core package (`agent-inspect`: APIs + CLI) plus optional packages that stay out of the root dependency graph: framework adapters (`ai-sdk`, `openai-agents`, `langchain`, `mcp`, `adapter-sdk`), CI and quality gates (`vitest`, `jest`, `eval`, `guardrails`, `circuit`, `harness`), and inspection surfaces (`viewer`, `tui`, `mcp-server`, `redact`). See the [package map](
|
|
123
|
+
AgentInspect v3.5 is the **adoption release**. The npm map is one core package (`agent-inspect`: APIs + CLI) plus optional packages that stay out of the root dependency graph: framework adapters (`ai-sdk`, `openai-agents`, `langchain`, `mcp`, `adapter-sdk`), CI and quality gates (`vitest`, `jest`, `eval`, `guardrails`, `circuit`, `harness`), and inspection surfaces (`viewer`, `tui`, `mcp-server`, `redact`). See the [package map](https://github.com/rajudandigam/agent-inspect#package-map).
|
|
124
124
|
|
|
125
125
|
Use it when:
|
|
126
126
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# First trace in 5 minutes
|
|
2
2
|
|
|
3
|
-
Goal: install → one trace → one check → one share-safe bundle.
|
|
3
|
+
Goal: install → one trace → one trajectory check → one share-safe Evidence bundle.
|
|
4
4
|
|
|
5
5
|
**Docs site:** [https://agentinspect.vercel.app/docs/getting-started/](https://agentinspect.vercel.app/docs/getting-started/)
|
|
6
6
|
|
|
@@ -14,12 +14,15 @@ npx agent-inspect list --dir .agent-inspect
|
|
|
14
14
|
Copy a `<run-id>` from `list`, then:
|
|
15
15
|
|
|
16
16
|
```bash
|
|
17
|
-
npx agent-inspect
|
|
18
|
-
npx agent-inspect check <run-id> --dir .agent-inspect
|
|
19
|
-
npx agent-inspect bundle <run-id> --dir .agent-inspect --profile share
|
|
17
|
+
npx agent-inspect view <run-id> --dir .agent-inspect --summary
|
|
18
|
+
npx agent-inspect check <run-id> --dir .agent-inspect --preset trajectory
|
|
20
19
|
npx agent-inspect verify-safe <run-id> --dir .agent-inspect
|
|
20
|
+
npx agent-inspect bundle <run-id> --dir .agent-inspect --profile share --out ./evidence
|
|
21
|
+
npx agent-inspect bundle verify ./evidence
|
|
21
22
|
```
|
|
22
23
|
|
|
24
|
+
`init` scaffolds files into the current directory (`agent-inspect.config.ts`, `.agent-inspect/`, and `examples/agent-inspect-demo.mjs`). It does **not** install dependencies or write a trace by itself. Use `--preset trajectory` for structural CI gates; `--fail-on-observation` belongs only in examples that record explicit OUTCOME events.
|
|
25
|
+
|
|
23
26
|
## Minutes 0–1: Install
|
|
24
27
|
|
|
25
28
|
```bash
|
|
@@ -28,9 +31,6 @@ npm install agent-inspect
|
|
|
28
31
|
npx agent-inspect init --yes
|
|
29
32
|
```
|
|
30
33
|
|
|
31
|
-
Creates `agent-inspect.config.ts`, `.agent-inspect/`, and `examples/agent-inspect-demo.mjs`.
|
|
32
|
-
`init` scaffolds files; it does **not** write a trace by itself.
|
|
33
|
-
|
|
34
34
|
Framework users: pick the correct capture path first — [CHOOSE-YOUR-CAPTURE-PATH.md](./CHOOSE-YOUR-CAPTURE-PATH.md).
|
|
35
35
|
|
|
36
36
|
## Minutes 1–2: Run
|
|
@@ -39,33 +39,36 @@ Framework users: pick the correct capture path first — [CHOOSE-YOUR-CAPTURE-PA
|
|
|
39
39
|
node examples/agent-inspect-demo.mjs
|
|
40
40
|
```
|
|
41
41
|
|
|
42
|
-
No API keys. Deterministic local trace
|
|
42
|
+
No API keys. Deterministic local trace under `.agent-inspect/`.
|
|
43
43
|
|
|
44
44
|
## Minutes 2–3: Inspect
|
|
45
45
|
|
|
46
46
|
```bash
|
|
47
47
|
npx agent-inspect list --dir .agent-inspect
|
|
48
|
-
npx agent-inspect view <run-id> --dir .agent-inspect
|
|
49
|
-
npx agent-inspect report <run-id> --dir .agent-inspect
|
|
48
|
+
npx agent-inspect view <run-id> --dir .agent-inspect --summary
|
|
50
49
|
```
|
|
51
50
|
|
|
51
|
+
Replace `<run-id>` with the ID printed by `list` (do not guess “latest file”).
|
|
52
|
+
|
|
52
53
|
## Minutes 3–4: Check
|
|
53
54
|
|
|
54
55
|
```bash
|
|
55
|
-
npx agent-inspect check <run-id> --dir .agent-inspect
|
|
56
|
+
npx agent-inspect check <run-id> --dir .agent-inspect --preset trajectory
|
|
56
57
|
```
|
|
57
58
|
|
|
59
|
+
Expected exit code `0` for the keyless demo. Add `--required-tool <name>` only when that tool is part of the real expected workflow.
|
|
60
|
+
|
|
58
61
|
## Minutes 4–5: Share-safe artifact
|
|
59
62
|
|
|
60
63
|
```bash
|
|
61
|
-
npx agent-inspect bundle <run-id> --dir .agent-inspect --profile share
|
|
62
64
|
npx agent-inspect verify-safe <run-id> --dir .agent-inspect
|
|
63
|
-
npx agent-inspect bundle
|
|
65
|
+
npx agent-inspect bundle <run-id> --dir .agent-inspect --profile share --out ./evidence
|
|
66
|
+
npx agent-inspect bundle verify ./evidence
|
|
64
67
|
```
|
|
65
68
|
|
|
66
|
-
|
|
69
|
+
`--out ./evidence` writes the Evidence v2 package to a known path; `bundle verify ./evidence` checks that exact directory. Do not use `--allow-unsafe` to force a share.
|
|
67
70
|
|
|
68
|
-
Optional file redaction:
|
|
71
|
+
Optional file redaction before bundling:
|
|
69
72
|
|
|
70
73
|
```bash
|
|
71
74
|
npx agent-inspect redact <run-id> --dir .agent-inspect --profile share -o redacted.jsonl
|
|
@@ -75,8 +78,9 @@ npx agent-inspect redact <run-id> --dir .agent-inspect --profile share -o redact
|
|
|
75
78
|
|
|
76
79
|
| If you use… | Go to |
|
|
77
80
|
| ----------- | ----- |
|
|
78
|
-
|
|
|
79
|
-
|
|
|
81
|
+
| Full install + instrumentation guide | [GETTING-STARTED.md](./GETTING-STARTED.md) (site: [/docs/getting-started/guide](/docs/getting-started/guide)) |
|
|
82
|
+
| Broken agent demo (same answer, wrong path) | [broken-agent-debugging starter](../examples/starters/broken-agent-debugging/README.md) |
|
|
83
|
+
| Coding-agent MCP loop | [CODING-AGENT-LOOP.md](./CODING-AGENT-LOOP.md) |
|
|
80
84
|
| Contracts / CI gates | [TRACE-CONTRACTS.md](./TRACE-CONTRACTS.md) · [SUITES-COHORTS-GATES.md](./SUITES-COHORTS-GATES.md) |
|
|
81
85
|
| AI SDK | [AI SDK adoption](./AI-SDK-ADOPTION.md) |
|
|
82
86
|
| OpenAI Agents | [OpenAI Agents local](./OPENAI-AGENTS-LOCAL.md) |
|
package/docs/GETTING-STARTED.md
CHANGED
|
@@ -14,7 +14,7 @@ pnpm add agent-inspect
|
|
|
14
14
|
|
|
15
15
|
### Quick bootstrap (v3.1+)
|
|
16
16
|
|
|
17
|
-
|
|
17
|
+
Prefer the five-minute path first: [FIRST-TRACE-IN-5-MINUTES.md](./FIRST-TRACE-IN-5-MINUTES.md) (site: [/docs/getting-started](/docs/getting-started)). This page is the longer install and instrumentation guide (site: [/docs/getting-started/guide](/docs/getting-started/guide)).
|
|
18
18
|
|
|
19
19
|
```bash
|
|
20
20
|
npx agent-inspect init --yes
|
package/docs/TRACE-CONTRACTS.md
CHANGED
|
@@ -12,6 +12,7 @@ Contracts compile to deterministic check rules for common cases:
|
|
|
12
12
|
- tool required / forbidden / allowed / maxCalls / order (`requiredTools` / `forbiddenTools` aliases)
|
|
13
13
|
- selectable `requiredOrderMode` (`first-occurrence` | `happens-before` | `all-occurrences`)
|
|
14
14
|
- additive `tools.orderRules` with per-rule occurrence modes
|
|
15
|
+
- typed cross-kind `steps.orderRelations` (TOOL ↔ LLM endpoints; additive in 6.31)
|
|
15
16
|
- bounded `tools.arguments` JSON Pointer checks (`exists` | `type` | `equals` | `oneOf`)
|
|
16
17
|
- `controls` declared-versus-enforced invariants
|
|
17
18
|
- `retry` / side-effect safety using explicit attempt identity (additive `retry.operations[]` recovery oracles in 6.27)
|
|
@@ -61,6 +62,30 @@ For overlapping first calls, omitted / `first-occurrence` warns while `happens-b
|
|
|
61
62
|
|
|
62
63
|
Immediate or positional `all-pairs` matching is not implemented.
|
|
63
64
|
|
|
65
|
+
## `steps.orderRelations` (shipped — experimental, 6.31)
|
|
66
|
+
|
|
67
|
+
Typed before/after relations with explicit `kind: "TOOL" | "LLM"` endpoints. Use this when a TOOL must precede an LLM (or vice versa). It does **not** change TOOL-only `tools.requiredOrder` / `orderRules` semantics.
|
|
68
|
+
|
|
69
|
+
```ts
|
|
70
|
+
defineTraceContract({
|
|
71
|
+
steps: {
|
|
72
|
+
orderRelations: [
|
|
73
|
+
{
|
|
74
|
+
before: { kind: "TOOL", name: "retrieve_policy" },
|
|
75
|
+
after: { kind: "LLM", name: "generate_answer" },
|
|
76
|
+
// mode?: "first-occurrence" | "happens-before" | "all-occurrences"
|
|
77
|
+
// requireEndpoints?: boolean // default true
|
|
78
|
+
},
|
|
79
|
+
],
|
|
80
|
+
},
|
|
81
|
+
});
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
- Default `mode` is `first-occurrence` (same three modes as tool order).
|
|
85
|
+
- Default `requireEndpoints: true` — missing kind+name fails; a same display name under the other kind does **not** satisfy the endpoint.
|
|
86
|
+
- LLM names match finished LLM events after stripping common prefixes (`llm:`, `generation:`, …).
|
|
87
|
+
- Low-level `createStepOrderingRule` is compositional when `requireEndpoints` is omitted/false.
|
|
88
|
+
|
|
64
89
|
### Experimental Vitest / Jest matchers (shipped)
|
|
65
90
|
|
|
66
91
|
| Package | Export | Matchers |
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "agent-inspect",
|
|
3
|
-
"version": "6.
|
|
3
|
+
"version": "6.31.1",
|
|
4
4
|
"license": "MIT",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"description": "Local evidence debugger and trajectory-test toolkit for TypeScript AI agents — execution trees, TraceContract checks, Evidence v2, and read-only MCP",
|
|
@@ -234,7 +234,7 @@
|
|
|
234
234
|
"test:coverage": "vitest run --coverage",
|
|
235
235
|
"size": "size-limit --config size-limit.config.mjs",
|
|
236
236
|
"test:all": "pnpm run typecheck && pnpm run linked-versions:check && pnpm run build && pnpm run test && pnpm run size",
|
|
237
|
-
"prepublish:checks": "pnpm run typecheck && pnpm run
|
|
237
|
+
"prepublish:checks": "pnpm run typecheck && pnpm run build && pnpm run test && pnpm run test:coverage && pnpm run fixtures:check && pnpm run recipes:check && pnpm run size && pnpm run linked-versions:check && pnpm run repo:health && pnpm run pack:smoke",
|
|
238
238
|
"pack:dry-run": "pnpm run build && npm pack --dry-run",
|
|
239
239
|
"pack:smoke": "pnpm run package-licenses:check && pnpm run build && node scripts/package-smoke.mjs && node scripts/packed-openai-agents-e2e.mjs && node scripts/packed-ai-sdk-e2e.mjs && node scripts/packed-mcp-e2e.mjs && node scripts/packed-quickstart-e2e.mjs && node scripts/packed-semantic-loop-e2e.mjs && node scripts/packed-swarm-loop-e2e.mjs && node scripts/evidence-ci-golden-paths.mjs",
|
|
240
240
|
"linked-versions:check": "node scripts/check-linked-versions.mjs",
|
|
@@ -262,6 +262,8 @@
|
|
|
262
262
|
"website:dev": "pnpm --filter @agent-inspect/website dev",
|
|
263
263
|
"website:build": "pnpm --filter @agent-inspect/website build",
|
|
264
264
|
"website:typecheck": "pnpm --filter @agent-inspect/website typecheck",
|
|
265
|
+
"website:check": "pnpm --filter @agent-inspect/website check",
|
|
266
|
+
"website:crawl": "pnpm --filter @agent-inspect/website crawl",
|
|
265
267
|
"package-licenses:check": "node scripts/validate-package-licenses.mjs",
|
|
266
268
|
"package-licenses:sync": "node scripts/sync-package-licenses.mjs",
|
|
267
269
|
"demo-version-sync:check": "node scripts/check-demo-version-sync.mjs"
|
|
@@ -6351,11 +6351,40 @@ function contractFailFinding(ruleId, message, evidence, expected, actual) {
|
|
|
6351
6351
|
evidence: [...evidence]
|
|
6352
6352
|
};
|
|
6353
6353
|
}
|
|
6354
|
+
var ALLOWED_STATUS_ALIASES = {
|
|
6355
|
+
ok: "ok",
|
|
6356
|
+
error: "error",
|
|
6357
|
+
running: "running",
|
|
6358
|
+
success: "ok",
|
|
6359
|
+
failed: "error"
|
|
6360
|
+
};
|
|
6361
|
+
function parseAllowedStatus(status) {
|
|
6362
|
+
return ALLOWED_STATUS_ALIASES[status];
|
|
6363
|
+
}
|
|
6354
6364
|
function normalizeStatus(status) {
|
|
6355
|
-
|
|
6356
|
-
if (
|
|
6357
|
-
|
|
6358
|
-
|
|
6365
|
+
const parsed = parseAllowedStatus(status);
|
|
6366
|
+
if (parsed === void 0) {
|
|
6367
|
+
throw new TypeError(
|
|
6368
|
+
`Unknown run status ${JSON.stringify(status)}. Allowed: ok, error, running (aliases: success, failed).`
|
|
6369
|
+
);
|
|
6370
|
+
}
|
|
6371
|
+
return parsed;
|
|
6372
|
+
}
|
|
6373
|
+
function validateAllowedStatusesShape(allowedStatuses, pathPrefix = "run.allowedStatuses") {
|
|
6374
|
+
if (!allowedStatuses?.length) return [];
|
|
6375
|
+
const out = [];
|
|
6376
|
+
for (let i = 0; i < allowedStatuses.length; i++) {
|
|
6377
|
+
const status = allowedStatuses[i];
|
|
6378
|
+
if (parseAllowedStatus(status) === void 0) {
|
|
6379
|
+
out.push({
|
|
6380
|
+
code: "contract.run.allowedStatuses.unknown",
|
|
6381
|
+
severity: "error",
|
|
6382
|
+
message: `${pathPrefix}[${i}] has unknown status ${JSON.stringify(status)}. Allowed: ok, error, running (aliases: success, failed).`,
|
|
6383
|
+
path: `${pathPrefix}[${i}]`
|
|
6384
|
+
});
|
|
6385
|
+
}
|
|
6386
|
+
}
|
|
6387
|
+
return out;
|
|
6359
6388
|
}
|
|
6360
6389
|
function cloneBody(body) {
|
|
6361
6390
|
return {
|
|
@@ -6368,6 +6397,17 @@ function cloneBody(body) {
|
|
|
6368
6397
|
}
|
|
6369
6398
|
} : {},
|
|
6370
6399
|
...body.llm ? { llm: { ...body.llm } } : {},
|
|
6400
|
+
...body.steps ? {
|
|
6401
|
+
steps: {
|
|
6402
|
+
...body.steps.orderRelations ? {
|
|
6403
|
+
orderRelations: body.steps.orderRelations.map((item) => ({
|
|
6404
|
+
...item,
|
|
6405
|
+
before: { ...item.before },
|
|
6406
|
+
after: { ...item.after }
|
|
6407
|
+
}))
|
|
6408
|
+
} : {}
|
|
6409
|
+
}
|
|
6410
|
+
} : {},
|
|
6371
6411
|
...body.observations ? {
|
|
6372
6412
|
observations: {
|
|
6373
6413
|
...body.observations,
|
|
@@ -6407,7 +6447,7 @@ function cloneBody(body) {
|
|
|
6407
6447
|
};
|
|
6408
6448
|
}
|
|
6409
6449
|
function bodyHasRules(body) {
|
|
6410
|
-
return body.run !== void 0 || body.tools !== void 0 || body.llm !== void 0 || body.observations !== void 0 || body.controls !== void 0 || body.retry !== void 0;
|
|
6450
|
+
return body.run !== void 0 || body.tools !== void 0 || body.llm !== void 0 || body.steps !== void 0 || body.observations !== void 0 || body.controls !== void 0 || body.retry !== void 0;
|
|
6411
6451
|
}
|
|
6412
6452
|
var MAX_EVIDENCE_EVENT_IDS = 16;
|
|
6413
6453
|
var METHOD_VOCABULARY = new Set(OBSERVED_OUTCOME_METHODS);
|
|
@@ -6750,6 +6790,18 @@ function contractToRules(contract) {
|
|
|
6750
6790
|
})
|
|
6751
6791
|
);
|
|
6752
6792
|
}
|
|
6793
|
+
const orderRelations = contract.steps?.orderRelations ?? [];
|
|
6794
|
+
for (const [index, relation] of orderRelations.entries()) {
|
|
6795
|
+
rules.push(
|
|
6796
|
+
createStepOrderingRule({
|
|
6797
|
+
before: { kind: relation.before.kind, name: relation.before.name },
|
|
6798
|
+
after: { kind: relation.after.kind, name: relation.after.name },
|
|
6799
|
+
id: `contract.step.orderRelation.${index}`,
|
|
6800
|
+
mode: relation.mode ?? "first-occurrence",
|
|
6801
|
+
requireEndpoints: relation.requireEndpoints !== false
|
|
6802
|
+
})
|
|
6803
|
+
);
|
|
6804
|
+
}
|
|
6753
6805
|
if (contract.observations) {
|
|
6754
6806
|
const required = contract.observations.required ?? [];
|
|
6755
6807
|
if (required.length > 0) {
|
|
@@ -6999,6 +7051,7 @@ function evaluateBody(input, body, options = {}) {
|
|
|
6999
7051
|
}
|
|
7000
7052
|
function defineTraceContract(input) {
|
|
7001
7053
|
const shapeErrors = [
|
|
7054
|
+
...validateAllowedStatusesShape(input.run?.allowedStatuses),
|
|
7002
7055
|
...validateAlternativesShape(input.alternatives),
|
|
7003
7056
|
...validateScopeShape(input.scope),
|
|
7004
7057
|
...validateProvenanceShape(input.observations)
|
|
@@ -7029,6 +7082,7 @@ function defineTraceContract(input) {
|
|
|
7029
7082
|
}
|
|
7030
7083
|
function evaluateTraceContract(input, contract, options = {}) {
|
|
7031
7084
|
const shapeErrors = [
|
|
7085
|
+
...validateAllowedStatusesShape(contract.run?.allowedStatuses),
|
|
7032
7086
|
...validateAlternativesShape(contract.alternatives),
|
|
7033
7087
|
...validateScopeShape(contract.scope),
|
|
7034
7088
|
...validateProvenanceShape(contract.observations)
|
|
@@ -7536,6 +7590,31 @@ function failFinding(ruleId, message, evidence, expected, actual, meta) {
|
|
|
7536
7590
|
function toolName(event) {
|
|
7537
7591
|
return resolveCanonicalToolName(event);
|
|
7538
7592
|
}
|
|
7593
|
+
function stepEndpointName(event, kind) {
|
|
7594
|
+
if (kind === "TOOL") return resolveCanonicalToolName(event);
|
|
7595
|
+
for (const prefix of ["llm:", "generation:", "transcription:", "speech:"]) {
|
|
7596
|
+
if (event.name.startsWith(prefix)) return event.name.slice(prefix.length);
|
|
7597
|
+
}
|
|
7598
|
+
return event.name;
|
|
7599
|
+
}
|
|
7600
|
+
function formatStepEndpoint(endpoint) {
|
|
7601
|
+
return `${endpoint.kind} ${endpoint.name}`;
|
|
7602
|
+
}
|
|
7603
|
+
function otherKindsForStepName(context, expectedKind, name) {
|
|
7604
|
+
return [
|
|
7605
|
+
...new Set(
|
|
7606
|
+
semanticEvents(context).filter((event) => {
|
|
7607
|
+
const kind = typeof event.kind === "string" ? event.kind.toUpperCase() : "";
|
|
7608
|
+
if (kind === expectedKind) return false;
|
|
7609
|
+
if (kind === "TOOL") return resolveCanonicalToolName(event) === name;
|
|
7610
|
+
if (kind === "LLM") return stepEndpointName(event, "LLM") === name;
|
|
7611
|
+
return event.name === name || event.name === `tool:${name}` || event.name === `llm:${name}` || event.name.endsWith(`:${name}`);
|
|
7612
|
+
}).map(
|
|
7613
|
+
(event) => typeof event.kind === "string" ? event.kind.toUpperCase() : "UNKNOWN"
|
|
7614
|
+
)
|
|
7615
|
+
)
|
|
7616
|
+
].sort();
|
|
7617
|
+
}
|
|
7539
7618
|
function semanticEvents(context) {
|
|
7540
7619
|
return context.logicalEvents ?? context.events;
|
|
7541
7620
|
}
|
|
@@ -7562,14 +7641,6 @@ function finishedEvents(context, kind) {
|
|
|
7562
7641
|
function toolInvocationEvents(context) {
|
|
7563
7642
|
return semanticEvents(context).filter((event) => event.kind === "TOOL");
|
|
7564
7643
|
}
|
|
7565
|
-
function firstIndexByName(events) {
|
|
7566
|
-
const index = /* @__PURE__ */ new Map();
|
|
7567
|
-
for (const [position, event] of events.entries()) {
|
|
7568
|
-
const name = toolName(event);
|
|
7569
|
-
if (!index.has(name)) index.set(name, position);
|
|
7570
|
-
}
|
|
7571
|
-
return index;
|
|
7572
|
-
}
|
|
7573
7644
|
function isRecord9(value) {
|
|
7574
7645
|
return typeof value === "object" && value !== null && !Array.isArray(value);
|
|
7575
7646
|
}
|
|
@@ -8040,18 +8111,85 @@ function createToolUsageRule(options) {
|
|
|
8040
8111
|
}
|
|
8041
8112
|
};
|
|
8042
8113
|
}
|
|
8043
|
-
function
|
|
8044
|
-
const
|
|
8114
|
+
function createStepOrderingRule(options) {
|
|
8115
|
+
const bothTools = options.before.kind === "TOOL" && options.after.kind === "TOOL";
|
|
8116
|
+
const ruleId = options.id ?? (bothTools ? "tool.order" : "step.order");
|
|
8045
8117
|
const mode = options.mode ?? "first-occurrence";
|
|
8118
|
+
const requireEndpoints = options.requireEndpoints === true;
|
|
8119
|
+
const beforeLabel = bothTools ? options.before.name : formatStepEndpoint(options.before);
|
|
8120
|
+
const afterLabel = bothTools ? options.after.name : formatStepEndpoint(options.after);
|
|
8121
|
+
const subject = bothTools ? "Tool" : "Step";
|
|
8046
8122
|
return {
|
|
8047
8123
|
id: ruleId,
|
|
8048
|
-
category: "tool",
|
|
8124
|
+
category: bothTools ? "tool" : "structure",
|
|
8049
8125
|
defaultSeverity: "error",
|
|
8050
8126
|
evaluate(context) {
|
|
8051
|
-
const
|
|
8052
|
-
|
|
8053
|
-
|
|
8054
|
-
const
|
|
8127
|
+
const beforeEvents = finishedEvents(context, options.before.kind).filter(
|
|
8128
|
+
(event) => stepEndpointName(event, options.before.kind) === options.before.name
|
|
8129
|
+
);
|
|
8130
|
+
const afterEvents = finishedEvents(context, options.after.kind).filter(
|
|
8131
|
+
(event) => stepEndpointName(event, options.after.kind) === options.after.name
|
|
8132
|
+
);
|
|
8133
|
+
if (beforeEvents.length === 0 || afterEvents.length === 0) {
|
|
8134
|
+
if (!requireEndpoints) return [];
|
|
8135
|
+
const findings = [];
|
|
8136
|
+
if (beforeEvents.length === 0) {
|
|
8137
|
+
const otherKinds = otherKindsForStepName(
|
|
8138
|
+
context,
|
|
8139
|
+
options.before.kind,
|
|
8140
|
+
options.before.name
|
|
8141
|
+
);
|
|
8142
|
+
const hint = otherKinds.length > 0 ? ` Required name appears under non-${options.before.kind} kind(s): ${otherKinds.join(", ")}.` : "";
|
|
8143
|
+
findings.push(
|
|
8144
|
+
failFinding(
|
|
8145
|
+
ruleId,
|
|
8146
|
+
`Required ${formatStepEndpoint(options.before)} did not appear.${hint}`,
|
|
8147
|
+
runEvidence(context.selectedRun),
|
|
8148
|
+
options.before,
|
|
8149
|
+
{
|
|
8150
|
+
code: bothTools ? "tool.order.missing-endpoint" : "step.order.missing-endpoint",
|
|
8151
|
+
endpoint: "before",
|
|
8152
|
+
otherKinds
|
|
8153
|
+
}
|
|
8154
|
+
)
|
|
8155
|
+
);
|
|
8156
|
+
}
|
|
8157
|
+
if (afterEvents.length === 0) {
|
|
8158
|
+
const otherKinds = otherKindsForStepName(
|
|
8159
|
+
context,
|
|
8160
|
+
options.after.kind,
|
|
8161
|
+
options.after.name
|
|
8162
|
+
);
|
|
8163
|
+
const hint = otherKinds.length > 0 ? ` Required name appears under non-${options.after.kind} kind(s): ${otherKinds.join(", ")}.` : "";
|
|
8164
|
+
findings.push(
|
|
8165
|
+
failFinding(
|
|
8166
|
+
ruleId,
|
|
8167
|
+
`Required ${formatStepEndpoint(options.after)} did not appear.${hint}`,
|
|
8168
|
+
runEvidence(context.selectedRun),
|
|
8169
|
+
options.after,
|
|
8170
|
+
{
|
|
8171
|
+
code: bothTools ? "tool.order.missing-endpoint" : "step.order.missing-endpoint",
|
|
8172
|
+
endpoint: "after",
|
|
8173
|
+
otherKinds
|
|
8174
|
+
}
|
|
8175
|
+
)
|
|
8176
|
+
);
|
|
8177
|
+
}
|
|
8178
|
+
return findings;
|
|
8179
|
+
}
|
|
8180
|
+
const encounter = finishedEvents(context).filter(
|
|
8181
|
+
(event) => event.kind === options.before.kind && stepEndpointName(event, options.before.kind) === options.before.name || event.kind === options.after.kind && stepEndpointName(event, options.after.kind) === options.after.name
|
|
8182
|
+
);
|
|
8183
|
+
let beforeIndex;
|
|
8184
|
+
let afterIndex;
|
|
8185
|
+
for (const [position, event] of encounter.entries()) {
|
|
8186
|
+
if (beforeIndex === void 0 && event.kind === options.before.kind && stepEndpointName(event, options.before.kind) === options.before.name) {
|
|
8187
|
+
beforeIndex = position;
|
|
8188
|
+
}
|
|
8189
|
+
if (afterIndex === void 0 && event.kind === options.after.kind && stepEndpointName(event, options.after.kind) === options.after.name) {
|
|
8190
|
+
afterIndex = position;
|
|
8191
|
+
}
|
|
8192
|
+
}
|
|
8055
8193
|
if (beforeIndex === void 0 || afterIndex === void 0) {
|
|
8056
8194
|
return [];
|
|
8057
8195
|
}
|
|
@@ -8059,32 +8197,37 @@ function createToolOrderingRule(options) {
|
|
|
8059
8197
|
return [
|
|
8060
8198
|
failFinding(
|
|
8061
8199
|
ruleId,
|
|
8062
|
-
|
|
8063
|
-
[
|
|
8064
|
-
|
|
8065
|
-
|
|
8200
|
+
`${subject} ${beforeLabel} must appear before ${afterLabel}.`,
|
|
8201
|
+
[
|
|
8202
|
+
eventEvidence3(encounter[beforeIndex]),
|
|
8203
|
+
eventEvidence3(encounter[afterIndex])
|
|
8204
|
+
],
|
|
8205
|
+
bothTools ? { before: options.before.name, after: options.after.name } : { before: options.before, after: options.after },
|
|
8206
|
+
bothTools ? finishedEvents(context, "TOOL").map(toolName) : encounter.map(
|
|
8207
|
+
(event) => event.kind === "TOOL" || event.kind === "LLM" ? `${event.kind}:${stepEndpointName(event, event.kind)}` : event.name
|
|
8208
|
+
)
|
|
8066
8209
|
)
|
|
8067
8210
|
];
|
|
8068
8211
|
}
|
|
8069
|
-
const beforeEvent =
|
|
8070
|
-
const afterEvent =
|
|
8212
|
+
const beforeEvent = beforeEvents[0];
|
|
8213
|
+
const afterEvent = afterEvents[0];
|
|
8071
8214
|
const beforeEnd = eventEndMs(beforeEvent);
|
|
8072
8215
|
const afterStart = eventStartMs(afterEvent);
|
|
8216
|
+
const endpointExpected = bothTools ? { before: options.before.name, after: options.after.name } : { before: options.before, after: options.after };
|
|
8073
8217
|
if (mode === "happens-before") {
|
|
8074
8218
|
const expected = {
|
|
8075
|
-
|
|
8076
|
-
after: options.after,
|
|
8219
|
+
...endpointExpected,
|
|
8077
8220
|
mode,
|
|
8078
8221
|
relation: "before.end <= after.start"
|
|
8079
8222
|
};
|
|
8080
|
-
if (options.before === options.after) {
|
|
8223
|
+
if (options.before.kind === options.after.kind && options.before.name === options.after.name) {
|
|
8081
8224
|
return [
|
|
8082
8225
|
failFinding(
|
|
8083
8226
|
ruleId,
|
|
8084
|
-
|
|
8227
|
+
`${subject} ${beforeLabel} cannot causally happen before itself.`,
|
|
8085
8228
|
[eventEvidence3(beforeEvent)],
|
|
8086
8229
|
expected,
|
|
8087
|
-
{ code: "tool.order.same-tool" }
|
|
8230
|
+
{ code: bothTools ? "tool.order.same-tool" : "step.order.same-step" }
|
|
8088
8231
|
)
|
|
8089
8232
|
];
|
|
8090
8233
|
}
|
|
@@ -8092,11 +8235,11 @@ function createToolOrderingRule(options) {
|
|
|
8092
8235
|
return [
|
|
8093
8236
|
failFinding(
|
|
8094
8237
|
ruleId,
|
|
8095
|
-
|
|
8238
|
+
`${subject} order ${beforeLabel} before ${afterLabel} could not establish causal timing.`,
|
|
8096
8239
|
[eventEvidence3(beforeEvent), eventEvidence3(afterEvent)],
|
|
8097
8240
|
expected,
|
|
8098
8241
|
{
|
|
8099
|
-
code: "tool.order.interval-unresolved",
|
|
8242
|
+
code: bothTools ? "tool.order.interval-unresolved" : "step.order.interval-unresolved",
|
|
8100
8243
|
beforeEndResolved: beforeEnd !== void 0,
|
|
8101
8244
|
afterStartResolved: afterStart !== void 0
|
|
8102
8245
|
}
|
|
@@ -8107,7 +8250,7 @@ function createToolOrderingRule(options) {
|
|
|
8107
8250
|
return [
|
|
8108
8251
|
failFinding(
|
|
8109
8252
|
ruleId,
|
|
8110
|
-
|
|
8253
|
+
`${subject} ${beforeLabel} must finish before ${afterLabel} starts.`,
|
|
8111
8254
|
[eventEvidence3(beforeEvent), eventEvidence3(afterEvent)],
|
|
8112
8255
|
expected,
|
|
8113
8256
|
{
|
|
@@ -8120,22 +8263,19 @@ function createToolOrderingRule(options) {
|
|
|
8120
8263
|
return [];
|
|
8121
8264
|
}
|
|
8122
8265
|
if (mode === "all-occurrences") {
|
|
8123
|
-
const beforeEvents = tools.filter((event) => toolName(event) === options.before);
|
|
8124
|
-
const afterEvents = tools.filter((event) => toolName(event) === options.after);
|
|
8125
8266
|
const expected = {
|
|
8126
|
-
|
|
8127
|
-
after: options.after,
|
|
8267
|
+
...endpointExpected,
|
|
8128
8268
|
mode,
|
|
8129
8269
|
relation: "max(before.end) <= min(after.start)"
|
|
8130
8270
|
};
|
|
8131
|
-
if (options.before === options.after) {
|
|
8271
|
+
if (options.before.kind === options.after.kind && options.before.name === options.after.name) {
|
|
8132
8272
|
return [
|
|
8133
8273
|
failFinding(
|
|
8134
8274
|
ruleId,
|
|
8135
|
-
|
|
8275
|
+
`${subject} ${beforeLabel} cannot causally happen before itself.`,
|
|
8136
8276
|
[eventEvidence3(beforeEvent)],
|
|
8137
8277
|
expected,
|
|
8138
|
-
{ code: "tool.order.same-tool" }
|
|
8278
|
+
{ code: bothTools ? "tool.order.same-tool" : "step.order.same-step" }
|
|
8139
8279
|
)
|
|
8140
8280
|
];
|
|
8141
8281
|
}
|
|
@@ -8167,11 +8307,11 @@ function createToolOrderingRule(options) {
|
|
|
8167
8307
|
return [
|
|
8168
8308
|
failFinding(
|
|
8169
8309
|
ruleId,
|
|
8170
|
-
|
|
8310
|
+
`${subject} order ${beforeLabel} before ${afterLabel} could not establish all causal intervals.`,
|
|
8171
8311
|
[firstMissingBefore, firstMissingAfter].filter((event) => event !== void 0).map((event) => eventEvidence3(event)),
|
|
8172
8312
|
expected,
|
|
8173
8313
|
{
|
|
8174
|
-
code: "tool.order.interval-unresolved",
|
|
8314
|
+
code: bothTools ? "tool.order.interval-unresolved" : "step.order.interval-unresolved",
|
|
8175
8315
|
missingBeforeEnds,
|
|
8176
8316
|
missingAfterStarts
|
|
8177
8317
|
}
|
|
@@ -8182,7 +8322,7 @@ function createToolOrderingRule(options) {
|
|
|
8182
8322
|
return [
|
|
8183
8323
|
failFinding(
|
|
8184
8324
|
ruleId,
|
|
8185
|
-
`Every ${options.before} tool call must finish before any ${options.after} tool call starts.`,
|
|
8325
|
+
bothTools ? `Every ${options.before.name} tool call must finish before any ${options.after.name} tool call starts.` : `Every ${beforeLabel} must finish before any ${afterLabel} starts.`,
|
|
8186
8326
|
[eventEvidence3(latestBefore.event), eventEvidence3(earliestAfter.event)],
|
|
8187
8327
|
expected,
|
|
8188
8328
|
{
|
|
@@ -8200,10 +8340,13 @@ function createToolOrderingRule(options) {
|
|
|
8200
8340
|
ruleId,
|
|
8201
8341
|
severity: "warning",
|
|
8202
8342
|
status: "warning",
|
|
8203
|
-
message:
|
|
8204
|
-
expected: {
|
|
8343
|
+
message: `${subject} order ${beforeLabel} before ${afterLabel} satisfied start/encounter order, but intervals overlap so causal completion before the next ${bothTools ? "tool" : "step"} was not established.`,
|
|
8344
|
+
expected: {
|
|
8345
|
+
...endpointExpected,
|
|
8346
|
+
mode: "first-occurrence"
|
|
8347
|
+
},
|
|
8205
8348
|
actual: {
|
|
8206
|
-
code: "tool.order.overlap",
|
|
8349
|
+
code: bothTools ? "tool.order.overlap" : "step.order.overlap",
|
|
8207
8350
|
beforeEndedAt: beforeEvent.endedAt,
|
|
8208
8351
|
afterStartedAt: afterEvent.startedAt ?? afterEvent.timestamp
|
|
8209
8352
|
},
|
|
@@ -8215,6 +8358,15 @@ function createToolOrderingRule(options) {
|
|
|
8215
8358
|
}
|
|
8216
8359
|
};
|
|
8217
8360
|
}
|
|
8361
|
+
function createToolOrderingRule(options) {
|
|
8362
|
+
return createStepOrderingRule({
|
|
8363
|
+
before: { kind: "TOOL", name: options.before },
|
|
8364
|
+
after: { kind: "TOOL", name: options.after },
|
|
8365
|
+
...options.id !== void 0 ? { id: options.id } : {},
|
|
8366
|
+
...options.mode !== void 0 ? { mode: options.mode } : {},
|
|
8367
|
+
requireEndpoints: false
|
|
8368
|
+
});
|
|
8369
|
+
}
|
|
8218
8370
|
function createLlmUsageRule(options) {
|
|
8219
8371
|
return {
|
|
8220
8372
|
id: "llm.usage",
|
|
@@ -14733,5 +14885,5 @@ function renderGateReport(result, options = {}) {
|
|
|
14733
14885
|
}
|
|
14734
14886
|
|
|
14735
14887
|
export { ATTRIBUTION_CONFIDENCES, COHORT_METRIC_IDS, DEFAULT_SUITE_ARTIFACTS_DIR, EVIDENCE_FORMAT_VERSION, EVIDENCE_HTML_FILENAME, EVIDENCE_MANIFEST_FILENAME, Redactor, TraceDirectory, TraceReadError, TreeBuilder, aggregateBundleSafeStatus, aggregateSessionCheckResults, analyzeCohort, applyProfileMetadataCaps, assertBundlePathContained, assertEvidenceRelativePath, buildActivitySummary, buildBundleMetadata, buildBundleSummaryMarkdown, buildEvidenceCausalFailureViewHtml, buildEvidenceCiPackage, buildEvidenceCircuitViewHtml, buildEvidenceContractsViewHtml, buildEvidenceDiffViewHtml, buildEvidenceHtmlShell, buildEvidenceManifest, buildEvidenceOutcomesViewHtml, buildEvidenceProvenanceViewHtml, buildEvidenceSafetyViewHtml, buildEvidenceTimelineViewHtml, buildEvidenceToolsLlmViewHtml, buildEvidenceTreeViewHtml, buildLocalExplanation, buildPlaceholderArtifact, buildRunSummary, buildRunTimeline, buildRunWhatSummary, buildSessionIndex, buildTraceStats, buildZipArchive, bundleFailsOnSafety, bundleRunAssetRelativePath, collectTraceSchemaVersions, compactAttributes, createBaselineRegressionRule, createLlmUsageRule, createMaxStepDurationRule, createObservedOutcomeRule, createRequireCompletedRule, createRunDepthRule, createRunDurationRule, createRunStatusRule, createSafetyOversizedAttributeRule, createSafetyRawContentRule, createSafetyRedactionRule, createSafetySecretPatternRule, createStallDetectionRule, createStructureCycleRule, createStructureOrphanRule, createStructureParallelWidthRule, createStructureRelationshipRule, createToolUsageRule, defaultBundleOutputPath, defaultSuiteConfigTemplate, defineTraceContract, diffRuns, diffTraceEvents, enrichSessionRunRecord, escapeHtml, escapeMarkdown, evaluateTraceContract, extractMetadata, extractOutcomesFromTraceEvents, filterMetasBySessionScope, filterTraces, flattenTree, formatDuration2 as formatDuration, formatStepLabel, formatTimestamp, gateHasThresholds, getIndent, getTraceFilePath, inferEvidenceFileRole, isAgentInspectTrace, isAttributionConfidence, isPersistedInspectEvent, loadSessionRunRecords, loadSuiteConfig, loadTraceMetadataList, manualTraceEventsToComparableRun, nanoid, normalizeBundleOutputPath, openTrace, parseCohortMetricList, parseDuration, parseDurationFilter, parseGateList, parseTraceJsonl, persistedInspectEventsToTraceEvents, projectLogicalEvents, renderActivitySummaryHuman, renderCohortReport, renderErrorLine, renderGateReport, renderObservedOutcomesHtml, renderObservedOutcomesMarkdown, renderRunDiff, renderRunWhat, renderStepLine, renderSuiteReport, renderTimeline, renderTraceStats, resolveBundleRunIds, resolveRedactionProfile, resolveSuiteTemplate, resolveTraceDir, runGate, runSuite, runTraceChecks, safeString, sanitizeBundleRunId, searchTraces, serializeEvidenceManifest, sha256Hex, stableJson, summarizeObservedOutcomes, summarizeSemanticParity, traceEventToPersistedInspectEvent, truncateName, truncateStringForProfile, validateEvent, validateSuiteConfig, verifyEvidenceDirectory, zeroKinds };
|
|
14736
|
-
//# sourceMappingURL=chunk-
|
|
14737
|
-
//# sourceMappingURL=chunk-
|
|
14888
|
+
//# sourceMappingURL=chunk-SJ6FRNOE.mjs.map
|
|
14889
|
+
//# sourceMappingURL=chunk-SJ6FRNOE.mjs.map
|