failproofai 1.0.7-beta.1 → 1.0.7-beta.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.next/standalone/.next/BUILD_ID +1 -1
- package/.next/standalone/.next/build-manifest.json +3 -3
- package/.next/standalone/.next/prerender-manifest.json +3 -3
- package/.next/standalone/.next/required-server-files.json +1 -1
- package/.next/standalone/.next/server/app/_global-error/page/server-reference-manifest.json +1 -1
- package/.next/standalone/.next/server/app/_global-error/page.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/_global-error/page_client-reference-manifest.js +1 -1
- package/.next/standalone/.next/server/app/_global-error.html +1 -1
- package/.next/standalone/.next/server/app/_global-error.rsc +7 -7
- package/.next/standalone/.next/server/app/_global-error.segments/__PAGE__.segment.rsc +6 -6
- package/.next/standalone/.next/server/app/_global-error.segments/_full.segment.rsc +7 -7
- package/.next/standalone/.next/server/app/_global-error.segments/_tree.segment.rsc +1 -1
- package/.next/standalone/.next/server/app/_not-found/page/server-reference-manifest.json +1 -1
- package/.next/standalone/.next/server/app/_not-found/page.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/_not-found/page_client-reference-manifest.js +1 -1
- package/.next/standalone/.next/server/app/_not-found.html +1 -1
- package/.next/standalone/.next/server/app/_not-found.rsc +15 -15
- package/.next/standalone/.next/server/app/_not-found.segments/_full.segment.rsc +15 -15
- package/.next/standalone/.next/server/app/_not-found.segments/_not-found/__PAGE__.segment.rsc +14 -14
- package/.next/standalone/.next/server/app/_not-found.segments/_tree.segment.rsc +2 -2
- package/.next/standalone/.next/server/app/api/audit/invite/route.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/api/audit/run/route.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/api/audit/status/route.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/api/auth/login-request/route.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/api/auth/login-verify/route.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/api/auth/logout/route.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/api/auth/status/route.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/api/download/[project]/[session]/route.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/audit/page/server-reference-manifest.json +2 -2
- package/.next/standalone/.next/server/app/audit/page.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/audit/page_client-reference-manifest.js +1 -1
- package/.next/standalone/.next/server/app/index.html +1 -1
- package/.next/standalone/.next/server/app/index.rsc +15 -15
- package/.next/standalone/.next/server/app/index.segments/__PAGE__.segment.rsc +14 -14
- package/.next/standalone/.next/server/app/index.segments/_full.segment.rsc +15 -15
- package/.next/standalone/.next/server/app/index.segments/_tree.segment.rsc +2 -2
- package/.next/standalone/.next/server/app/page/server-reference-manifest.json +1 -1
- package/.next/standalone/.next/server/app/page.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/page_client-reference-manifest.js +1 -1
- package/.next/standalone/.next/server/app/policies/page/server-reference-manifest.json +14 -14
- package/.next/standalone/.next/server/app/policies/page.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/policies/page_client-reference-manifest.js +1 -1
- package/.next/standalone/.next/server/app/project/[name]/page/server-reference-manifest.json +1 -1
- package/.next/standalone/.next/server/app/project/[name]/page.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/project/[name]/page_client-reference-manifest.js +1 -1
- package/.next/standalone/.next/server/app/project/[name]/session/[sessionId]/page/react-loadable-manifest.json +2 -2
- package/.next/standalone/.next/server/app/project/[name]/session/[sessionId]/page/server-reference-manifest.json +2 -2
- package/.next/standalone/.next/server/app/project/[name]/session/[sessionId]/page.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/project/[name]/session/[sessionId]/page_client-reference-manifest.js +1 -1
- package/.next/standalone/.next/server/app/projects/page/server-reference-manifest.json +1 -1
- package/.next/standalone/.next/server/app/projects/page.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/projects/page_client-reference-manifest.js +1 -1
- package/.next/standalone/.next/server/app/settings/page/server-reference-manifest.json +7 -7
- package/.next/standalone/.next/server/app/settings/page.js +2 -2
- package/.next/standalone/.next/server/app/settings/page.js.nft.json +1 -1
- package/.next/standalone/.next/server/app/settings/page_client-reference-manifest.js +1 -1
- package/.next/standalone/.next/server/chunks/[root-of-the-server]__0l3yhx4._.js +2 -2
- package/.next/standalone/.next/server/chunks/[root-of-the-server]__0o07qi9._.js +1 -1
- package/.next/standalone/.next/server/chunks/_0tovk6q._.js +1 -1
- package/.next/standalone/.next/server/chunks/_0trp3yc._.js +1 -1
- package/.next/standalone/.next/server/chunks/node_modules_posthog-node_dist_entrypoints_index_node_mjs_09z9-p7._.js +1 -1
- package/.next/standalone/.next/server/chunks/package_json_[json]_cjs_1nxcc4v._.js +1 -1
- package/.next/standalone/.next/server/chunks/src_hooks_18qtd42._.js +1 -1
- package/.next/standalone/.next/server/chunks/ssr/[root-of-the-server]__04usis8._.js +2 -2
- package/.next/standalone/.next/server/chunks/ssr/[root-of-the-server]__056wjo4._.js +2 -2
- package/.next/standalone/.next/server/chunks/ssr/{[root-of-the-server]__0yxwl6j._.js → [root-of-the-server]__0l44ual._.js} +2 -2
- package/.next/standalone/.next/server/chunks/ssr/[root-of-the-server]__0n0xg95._.js +2 -2
- package/.next/standalone/.next/server/chunks/ssr/[root-of-the-server]__0rwtwpm._.js +2 -2
- package/.next/standalone/.next/server/chunks/ssr/[root-of-the-server]__0s_yomn._.js +2 -2
- package/.next/standalone/.next/server/chunks/ssr/[root-of-the-server]__11mayhe._.js +2 -2
- package/.next/standalone/.next/server/chunks/ssr/[root-of-the-server]__1pprgri._.js +2 -2
- package/.next/standalone/.next/server/chunks/ssr/[root-of-the-server]__1q4p5b8._.js +2 -2
- package/.next/standalone/.next/server/chunks/ssr/_042cgl1._.js +1 -1
- package/.next/standalone/.next/server/chunks/ssr/_08x1r5t._.js +1 -1
- package/.next/standalone/.next/server/chunks/ssr/_0bqoto4._.js +1 -1
- package/.next/standalone/.next/server/chunks/ssr/{_214wgrp._.js → _1mel6y1._.js} +2 -2
- package/.next/standalone/.next/server/chunks/ssr/{_0bn2oo8._.js → _1v-jvrv._.js} +1 -1
- package/.next/standalone/.next/server/chunks/ssr/_1zopuov._.js +1 -1
- package/.next/standalone/.next/server/chunks/ssr/_next-internal_server_app_policies_page_actions_1sp2-yo.js +2 -2
- package/.next/standalone/.next/server/chunks/ssr/app_actions_get-scheduled-audit_ts_0ei9sni._.js +1 -1
- package/.next/standalone/.next/server/chunks/ssr/app_audit__components_audit-dashboard_tsx_0p9ud47._.js +1 -1
- package/.next/standalone/.next/server/chunks/ssr/app_global-error_tsx_1kp6l3x._.js +1 -1
- package/.next/standalone/.next/server/chunks/ssr/app_policies_hooks-client_tsx_19dqvpc._.js +1 -1
- package/.next/standalone/.next/server/chunks/ssr/app_settings_02tf1h4._.js +1 -1
- package/.next/standalone/.next/server/chunks/ssr/src_hooks_15t8kqj._.js +1 -1
- package/.next/standalone/.next/server/chunks/ssr/src_hooks_1fm2w5z._.js +1 -1
- package/.next/standalone/.next/server/chunks/ssr/src_hooks_1j0zy3v._.js +1 -1
- package/.next/standalone/.next/server/chunks/ssr/src_hooks_builtin-policies_ts_09j2ndl._.js +1 -1
- package/.next/standalone/.next/server/middleware-build-manifest.js +3 -3
- package/.next/standalone/.next/server/pages/404.html +1 -1
- package/.next/standalone/.next/server/pages/500.html +1 -1
- package/.next/standalone/.next/server/server-reference-manifest.js +1 -1
- package/.next/standalone/.next/server/server-reference-manifest.json +22 -22
- package/.next/standalone/.next/static/chunks/043j99m8ykg__.css +2 -0
- package/.next/standalone/.next/static/chunks/0fqd7m_u81mi5.js +1 -0
- package/.next/standalone/.next/static/chunks/129ag2bw93bdh.js +1 -0
- package/.next/standalone/.next/static/chunks/{29fql3nbnfc9q.js → 1qd741hzlmjbo.js} +1 -1
- package/.next/standalone/.next/static/chunks/{0vmd180qfntfb.js → 2_pltstd8-xgs.js} +1 -1
- package/.next/standalone/.next/static/chunks/{1u5zsejmgrir_.js → 2c8j9l6j_b1ci.js} +1 -1
- package/.next/standalone/.next/static/chunks/3-k569wzcli8q.js +1 -0
- package/.next/standalone/.next/static/chunks/{258668t68du6b.js → 3brze37td_wnc.js} +1 -1
- package/.next/standalone/.next/static/chunks/3otmypm6j_xfo.js +6 -0
- package/.next/standalone/.next/static/chunks/{12tvm75t5ffui.js → 3ugmd_7dyn0id.js} +1 -1
- package/.next/standalone/.next/static/chunks/3yxro_r2_o9ad.js +69 -0
- package/.next/standalone/PROBE-FOLLOWUP.md +186 -0
- package/.next/standalone/app/actions/get-jev-config.ts +38 -23
- package/.next/standalone/app/actions/update-jev-config.ts +1 -1
- package/.next/standalone/app/settings/jev-panel.tsx +33 -29
- package/.next/standalone/package.json +10 -10
- package/.next/standalone/sdk/python/skill/SKILL.md +60 -14
- package/.next/standalone/sdk/python/skill/agents/openai.yaml +2 -1
- package/.next/standalone/sdk/python/skill/references/evaluator.md +255 -0
- package/.next/standalone/sdk/python/skill/references/events.md +17 -8
- package/.next/standalone/sdk/python/skill/references/frameworks.md +3 -0
- package/.next/standalone/sdk/python/skill/references/install.md +3 -0
- package/.next/standalone/sdk/python/skill/references/integration.md +6 -2
- package/.next/standalone/sdk/python/skill/references/typescript.md +568 -0
- package/.next/standalone/sdk/typescript/CHANGELOG.md +119 -0
- package/.next/standalone/sdk/typescript/LICENSE +42 -0
- package/.next/standalone/sdk/typescript/README.md +552 -0
- package/.next/standalone/sdk/typescript/eslint.config.mjs +59 -0
- package/.next/standalone/sdk/typescript/examples/research-agent.ts +197 -0
- package/.next/standalone/sdk/typescript/integration/ai.test.ts +920 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-4/agent.ts +337 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-4/package-lock.json +281 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-4/package.json +13 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-4/surfaces.ts +605 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-4/tsconfig.json +12 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-4/tsconfig.surfaces.json +4 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-5/agent.ts +342 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-5/package-lock.json +168 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-5/package.json +13 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-5/surfaces.ts +628 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-5/tsconfig.json +12 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-5/tsconfig.surfaces.json +4 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-6/agent.ts +346 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-6/package-lock.json +156 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-6/package.json +13 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-6/surfaces.ts +651 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-6/tsconfig.json +12 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-6/tsconfig.surfaces.json +4 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-7/agent.ts +350 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-7/package-lock.json +153 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-7/package.json +13 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-7/surfaces.ts +651 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-7/tsconfig.json +12 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/ai-7/tsconfig.surfaces.json +4 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-0.3/agent.ts +623 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-0.3/package-lock.json +463 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-0.3/package.json +15 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-0.3/tsconfig.json +12 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-1/agent-v1.ts +99 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-1/agent.ts +623 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-1/package-lock.json +336 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-1/package.json +15 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-1/tsconfig.json +12 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-dup-core/agent.ts +96 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-dup-core/package-lock.json +579 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-dup-core/package.json +15 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-dup-core/tsconfig.json +12 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-dup-core/vendor/lc-weather-provider/index.cjs +53 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-dup-core/vendor/lc-weather-provider/index.mjs +55 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/langchain-dup-core/vendor/lc-weather-provider/package.json +18 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/llamaindex-0.11/agent.ts +659 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/llamaindex-0.11/package-lock.json +635 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/llamaindex-0.11/package.json +15 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/llamaindex-0.11/tsconfig.json +12 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/llamaindex-0.12/agent.ts +659 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/llamaindex-0.12/package-lock.json +553 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/llamaindex-0.12/package.json +15 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/llamaindex-0.12/tsconfig.json +12 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/mastra-0/agent.ts +877 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/mastra-0/mcp-server.mjs +66 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/mastra-0/package-lock.json +6417 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/mastra-0/package.json +15 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/mastra-0/tsconfig.json +12 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/mastra-1/agent.ts +872 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/mastra-1/mcp-server.mjs +66 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/mastra-1/package-lock.json +2540 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/mastra-1/package.json +15 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/mastra-1/tsconfig.json +12 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/app/actions.ts +18 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/app/api/ai/route.ts +42 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/app/api/edge/route.ts +29 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/app/api/langgraph/route.ts +21 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/app/api/llamaindex/route.ts +11 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/app/api/mastra/route.ts +19 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/app/api/status/route.ts +7 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/app/layout.tsx +9 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/app/page.tsx +15 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/instrumentation.ts +15 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/lib/ai.ts +61 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/lib/langgraph.ts +68 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/lib/llamaindex.ts +98 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/lib/mastra.ts +90 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/next.config.ts +52 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/package-lock.json +4343 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/nextjs/package.json +28 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/runtimes/agent.ts +51 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/runtimes/deno-npm.ts +76 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/runtimes/package-lock.json +484 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/runtimes/package.json +15 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/types/agent.ts +79 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/types/package-lock.json +740 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/types/package.json +11 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/vanilla/agent.ts +197 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/vanilla/package-lock.json +70 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/vanilla/package.json +12 -0
- package/.next/standalone/sdk/typescript/integration/fixtures/vanilla/tsconfig.json +12 -0
- package/.next/standalone/sdk/typescript/integration/global-setup.ts +29 -0
- package/.next/standalone/sdk/typescript/integration/harness.ts +451 -0
- package/.next/standalone/sdk/typescript/integration/langchain.test.ts +682 -0
- package/.next/standalone/sdk/typescript/integration/llamaindex.test.ts +709 -0
- package/.next/standalone/sdk/typescript/integration/mastra-coverage.test.ts +386 -0
- package/.next/standalone/sdk/typescript/integration/mastra.test.ts +311 -0
- package/.next/standalone/sdk/typescript/integration/nextjs.test.ts +341 -0
- package/.next/standalone/sdk/typescript/integration/runtime-parity.ts +180 -0
- package/.next/standalone/sdk/typescript/integration/runtimes.bun.test.ts +15 -0
- package/.next/standalone/sdk/typescript/integration/runtimes.core.test.ts +113 -0
- package/.next/standalone/sdk/typescript/integration/runtimes.deno.test.ts +19 -0
- package/.next/standalone/sdk/typescript/integration/types.test.ts +141 -0
- package/.next/standalone/sdk/typescript/integration/vanilla.test.ts +255 -0
- package/.next/standalone/sdk/typescript/package-lock.json +2640 -0
- package/.next/standalone/sdk/typescript/package.json +401 -0
- package/.next/standalone/sdk/typescript/scripts/finalize-build.mjs +123 -0
- package/.next/standalone/sdk/typescript/scripts/release.mjs +177 -0
- package/.next/standalone/sdk/typescript/src/clock.ts +58 -0
- package/.next/standalone/sdk/typescript/src/context.ts +214 -0
- package/.next/standalone/sdk/typescript/src/edge/adapter.ts +18 -0
- package/.next/standalone/sdk/typescript/src/edge/ai.ts +116 -0
- package/.next/standalone/sdk/typescript/src/edge/index.ts +238 -0
- package/.next/standalone/sdk/typescript/src/edge/langchain.ts +18 -0
- package/.next/standalone/sdk/typescript/src/edge/llamaindex.ts +12 -0
- package/.next/standalone/sdk/typescript/src/edge/mastra.ts +17 -0
- package/.next/standalone/sdk/typescript/src/edge/notice.ts +33 -0
- package/.next/standalone/sdk/typescript/src/environment.ts +75 -0
- package/.next/standalone/sdk/typescript/src/evaluator/authoring.ts +480 -0
- package/.next/standalone/sdk/typescript/src/evaluator/cli.ts +96 -0
- package/.next/standalone/sdk/typescript/src/evaluator/client.ts +421 -0
- package/.next/standalone/sdk/typescript/src/evaluator/expression.ts +1292 -0
- package/.next/standalone/sdk/typescript/src/evaluator/index.ts +144 -0
- package/.next/standalone/sdk/typescript/src/evaluator/protocol.ts +747 -0
- package/.next/standalone/sdk/typescript/src/evaluator/runtime.ts +930 -0
- package/.next/standalone/sdk/typescript/src/evaluator/sandbox-worker.ts +171 -0
- package/.next/standalone/sdk/typescript/src/evaluator/source-limits.ts +27 -0
- package/.next/standalone/sdk/typescript/src/evaluator/source.ts +509 -0
- package/.next/standalone/sdk/typescript/src/events.ts +879 -0
- package/.next/standalone/sdk/typescript/src/exit.ts +117 -0
- package/.next/standalone/sdk/typescript/src/index.ts +186 -0
- package/.next/standalone/sdk/typescript/src/integrations/ai.ts +1566 -0
- package/.next/standalone/sdk/typescript/src/integrations/compat.ts +322 -0
- package/.next/standalone/sdk/typescript/src/integrations/core.ts +1321 -0
- package/.next/standalone/sdk/typescript/src/integrations/index.ts +355 -0
- package/.next/standalone/sdk/typescript/src/integrations/langchain.ts +2340 -0
- package/.next/standalone/sdk/typescript/src/integrations/llamaindex.ts +2111 -0
- package/.next/standalone/sdk/typescript/src/integrations/mastra.ts +1802 -0
- package/.next/standalone/sdk/typescript/src/logger.ts +98 -0
- package/.next/standalone/sdk/typescript/src/next.ts +115 -0
- package/.next/standalone/sdk/typescript/src/node-require.ts +446 -0
- package/.next/standalone/sdk/typescript/src/redact.ts +305 -0
- package/.next/standalone/sdk/typescript/src/resolver.ts +120 -0
- package/.next/standalone/sdk/typescript/src/runtime.ts +29 -0
- package/.next/standalone/sdk/typescript/src/schema.ts +410 -0
- package/.next/standalone/sdk/typescript/src/scopes.ts +701 -0
- package/.next/standalone/sdk/typescript/src/shared.ts +33 -0
- package/.next/standalone/sdk/typescript/src/version.ts +5 -0
- package/.next/standalone/sdk/typescript/src/writer.ts +934 -0
- package/.next/standalone/sdk/typescript/test/adapters.test.ts +397 -0
- package/.next/standalone/sdk/typescript/test/ai.test.ts +1076 -0
- package/.next/standalone/sdk/typescript/test/copies.test.ts +204 -0
- package/.next/standalone/sdk/typescript/test/edge.test.ts +183 -0
- package/.next/standalone/sdk/typescript/test/evaluator-client.test.ts +234 -0
- package/.next/standalone/sdk/typescript/test/evaluator-protocol.test.ts +225 -0
- package/.next/standalone/sdk/typescript/test/events.test.ts +193 -0
- package/.next/standalone/sdk/typescript/test/expression.test.ts +181 -0
- package/.next/standalone/sdk/typescript/test/global-setup.ts +26 -0
- package/.next/standalone/sdk/typescript/test/helpers.ts +130 -0
- package/.next/standalone/sdk/typescript/test/integrations.test.ts +369 -0
- package/.next/standalone/sdk/typescript/test/langchain-copies.test.ts +204 -0
- package/.next/standalone/sdk/typescript/test/langchain.test.ts +999 -0
- package/.next/standalone/sdk/typescript/test/llamaindex.test.ts +1760 -0
- package/.next/standalone/sdk/typescript/test/mastra-coverage.test.ts +501 -0
- package/.next/standalone/sdk/typescript/test/mastra-lifecycle.test.ts +479 -0
- package/.next/standalone/sdk/typescript/test/mastra.test.ts +285 -0
- package/.next/standalone/sdk/typescript/test/next.test.ts +109 -0
- package/.next/standalone/sdk/typescript/test/packaging.test.ts +312 -0
- package/.next/standalone/sdk/typescript/test/redaction.test.ts +171 -0
- package/.next/standalone/sdk/typescript/test/runtimes.test.ts +101 -0
- package/.next/standalone/sdk/typescript/test/sandbox.test.ts +189 -0
- package/.next/standalone/sdk/typescript/test/scopes.test.ts +271 -0
- package/.next/standalone/sdk/typescript/test/setup.ts +19 -0
- package/.next/standalone/sdk/typescript/test/skill-snippets.test.ts +73 -0
- package/.next/standalone/sdk/typescript/test/spool-contract.test.ts +124 -0
- package/.next/standalone/sdk/typescript/test/tracker-bounds.test.ts +191 -0
- package/.next/standalone/sdk/typescript/test/wire-format.test.ts +214 -0
- package/.next/standalone/sdk/typescript/test/writer.test.ts +407 -0
- package/.next/standalone/sdk/typescript/tsconfig.build.json +15 -0
- package/.next/standalone/sdk/typescript/tsconfig.cjs.json +19 -0
- package/.next/standalone/sdk/typescript/tsconfig.json +28 -0
- package/.next/standalone/sdk/typescript/vitest.config.ts +33 -0
- package/.next/standalone/sdk/typescript/vitest.integration.config.ts +23 -0
- package/.next/standalone/server.js +1 -1
- package/README.md +2 -2
- package/dist/cli.mjs +208 -27
- package/dist/worker.mjs +120 -14
- package/package.json +10 -10
- package/scripts/build-policy-pack.mjs +14 -0
- package/src/audit/features.ts +3 -2
- package/src/hooks/builtin-policies.ts +156 -5
- package/src/hooks/jev-cli.ts +57 -1
- package/src/hooks/manager.ts +1 -1
- package/src/hooks/pack-cli.ts +150 -18
- package/src/hooks/pack-store.ts +1 -1
- package/src/hooks/policy-catalog.ts +148 -9
- package/src/hooks/types.ts +1 -1
- package/.next/standalone/.next/static/chunks/0cd-_8-c-m1ea.js +0 -6
- package/.next/standalone/.next/static/chunks/0qrbdkv9qmvli.js +0 -69
- package/.next/standalone/.next/static/chunks/0uldbut9y2-e8.js +0 -1
- package/.next/standalone/.next/static/chunks/285spx855h_3r.css +0 -2
- package/.next/standalone/.next/static/chunks/2vkvu9-opa_1z.js +0 -1
- package/.next/standalone/.next/static/chunks/37lhv7wa3ywt6.js +0 -1
- /package/.next/standalone/.next/static/{T8YAYM9h_64fKv705vug7 → gbEOjBgZAxF2UIUwZVHNu}/_buildManifest.js +0 -0
- /package/.next/standalone/.next/static/{T8YAYM9h_64fKv705vug7 → gbEOjBgZAxF2UIUwZVHNu}/_clientMiddlewareManifest.js +0 -0
- /package/.next/standalone/.next/static/{T8YAYM9h_64fKv705vug7 → gbEOjBgZAxF2UIUwZVHNu}/_ssgManifest.js +0 -0
|
@@ -9,13 +9,20 @@
|
|
|
9
9
|
* valid `~/.failproofai/jev.json` exists; without it the hook path is byte for
|
|
10
10
|
* byte what it was before. So this panel has two jobs, in this order:
|
|
11
11
|
*
|
|
12
|
-
* 1. say whether it is on,
|
|
13
|
-
* of the enabled policy set it is allowed to clear, and — once it
|
|
14
|
-
* — how often it fell back to the regex engine. A panel that only
|
|
15
|
-
* input would leave "is this thing working" unanswerable from the
|
|
12
|
+
* 1. say whether it is on, in which mode, against which provider and model,
|
|
13
|
+
* how much of the enabled policy set it is allowed to clear, and — once it
|
|
14
|
+
* has run — how often it fell back to the regex engine. A panel that only
|
|
15
|
+
* took input would leave "is this thing working" unanswerable from the
|
|
16
16
|
* dashboard.
|
|
17
17
|
* 2. take the endpoint and the token.
|
|
18
18
|
*
|
|
19
|
+
* Each fact is stated ONCE. The status block says nothing the form below it
|
|
20
|
+
* already holds: the endpoint and the account id are form values, and printing
|
|
21
|
+
* them above the fields put a per-account URL — and the id inside it — twice on
|
|
22
|
+
* one screen. The config file's path and permission bits are gone for a
|
|
23
|
+
* different reason: nobody repairs a 0644 from a browser, and the loader's own
|
|
24
|
+
* refusal already names the file when it matters.
|
|
25
|
+
*
|
|
19
26
|
* It is a full-width cell in the same hairline console as the scheduled-audit
|
|
20
27
|
* panel, using the same tokens, the same `.btn-press` action and the same
|
|
21
28
|
* `toast()` on success. No new colour and no new control that did not already
|
|
@@ -24,11 +31,12 @@
|
|
|
24
31
|
* ## The token is write-only
|
|
25
32
|
*
|
|
26
33
|
* The field is always blank on load, whatever is stored. `JevSettingsView`
|
|
27
|
-
* carries
|
|
28
|
-
*
|
|
29
|
-
* leaves the machine's filesystem
|
|
30
|
-
* key where it is — the stored one, or
|
|
31
|
-
* it from there — which the server
|
|
34
|
+
* carries presence and nothing else, so the panel can say "configured" — or
|
|
35
|
+
* that the key is read from the environment — and no more; the value itself
|
|
36
|
+
* never leaves the machine's filesystem, and neither does a fragment of it.
|
|
37
|
+
* Leaving the field empty on save KEEPS the key where it is — the stored one, or
|
|
38
|
+
* the environment's for a config that takes it from there — which the server
|
|
39
|
+
* decides, not this component.
|
|
32
40
|
*
|
|
33
41
|
* The stored MODEL is shown the same way when it does not look like a model id,
|
|
34
42
|
* for the same reason: it is the one routing field someone can paste a key into.
|
|
@@ -258,9 +266,11 @@ export default function JevPanel({ initial }: { initial: JevSettingsView | null
|
|
|
258
266
|
: stored
|
|
259
267
|
? stored.source === "env"
|
|
260
268
|
? "configured, read from the environment — leave blank to keep it"
|
|
261
|
-
:
|
|
262
|
-
|
|
263
|
-
|
|
269
|
+
: // Presence and the keep rule, which is everything this field's reader
|
|
270
|
+
// has to decide. Naming the last four characters of the stored key
|
|
271
|
+
// here put that fragment on the page a second time, next to the input
|
|
272
|
+
// it would be re-typed into.
|
|
273
|
+
"configured — leave blank to keep it"
|
|
264
274
|
: // The file names the environment as the key's source and the variable is
|
|
265
275
|
// unset in this process. Leaving the field blank keeps that arrangement,
|
|
266
276
|
// which is not the same thing as "there is no key" — the endpoint and the
|
|
@@ -292,38 +302,32 @@ export default function JevPanel({ initial }: { initial: JevSettingsView | null
|
|
|
292
302
|
</div>
|
|
293
303
|
</div>
|
|
294
304
|
|
|
305
|
+
{/* Where requests GO is not in this block, deliberately. The editable
|
|
306
|
+
field below holds it, and for Cloudflare that URL is
|
|
307
|
+
`/accounts/<id>/ai/run` — so a read-only row above the form printed the
|
|
308
|
+
same per-account address, and the account id inside it, a second time
|
|
309
|
+
on one screen. The path of the config file and its mode were here too;
|
|
310
|
+
nobody acts on either from a browser. */}
|
|
295
311
|
{view && configured && (
|
|
296
312
|
<dl className="set-how-list">
|
|
297
|
-
{view.endpoint && (
|
|
298
|
-
<div className="set-how-row">
|
|
299
|
-
<dt className="set-how-label">endpoint</dt>
|
|
300
|
-
<dd className="set-how-body">{view.endpoint}</dd>
|
|
301
|
-
</div>
|
|
302
|
-
)}
|
|
303
313
|
<div className="set-how-row">
|
|
304
314
|
<dt className="set-how-label">model</dt>
|
|
305
315
|
<dd className="set-how-body">{fmtModel(view.model)}</dd>
|
|
306
316
|
</div>
|
|
317
|
+
{/* Presence, and where it came from. Not four characters of it: that
|
|
318
|
+
is a recognisable fragment of a live credential on a page with no
|
|
319
|
+
authentication, and it tells the reader nothing they cannot get by
|
|
320
|
+
re-pasting the key. The server no longer sends it either. */}
|
|
307
321
|
<div className="set-how-row">
|
|
308
322
|
<dt className="set-how-label">token</dt>
|
|
309
323
|
<dd className="set-how-body">
|
|
310
324
|
{stored
|
|
311
325
|
? stored.source === "env"
|
|
312
326
|
? "from FAILPROOFAI_JEV_API_KEY"
|
|
313
|
-
:
|
|
314
|
-
? `configured, ending ${stored.hint}`
|
|
315
|
-
: "configured"
|
|
327
|
+
: "configured"
|
|
316
328
|
: "none stored"}
|
|
317
329
|
</dd>
|
|
318
330
|
</div>
|
|
319
|
-
{view.permissions && (
|
|
320
|
-
<div className="set-how-row">
|
|
321
|
-
<dt className="set-how-label">file</dt>
|
|
322
|
-
<dd className="set-how-body">
|
|
323
|
-
{`${view.path} · ${view.permissions}`}
|
|
324
|
-
</dd>
|
|
325
|
-
</div>
|
|
326
|
-
)}
|
|
327
331
|
{view.reviewable && (
|
|
328
332
|
<div className="set-how-row">
|
|
329
333
|
<dt className="set-how-label">reviewable</dt>
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "failproofai",
|
|
3
|
-
"version": "1.0.7-beta.
|
|
4
|
-
"description": "Observability and enforcement for AI agent harnesses.
|
|
3
|
+
"version": "1.0.7-beta.2",
|
|
4
|
+
"description": "Observability and enforcement for AI agent harnesses. 40 built-in policies hooked into 12 of them — Claude Code, Codex, Cursor, Hermes, OpenClaw and more — blocking the tool call before it runs. Local dashboard included, no account needed.",
|
|
5
5
|
"bin": {
|
|
6
6
|
"failproofai": "./dist/cli.mjs",
|
|
7
7
|
"failproofaid": "./bin/failproofaid-shim.mjs"
|
|
@@ -122,15 +122,15 @@
|
|
|
122
122
|
"browserslist": "4.28.8"
|
|
123
123
|
},
|
|
124
124
|
"optionalDependencies": {
|
|
125
|
-
"@failproofai/failproofaid-linux-x64": "1.0.7-beta.
|
|
126
|
-
"@failproofai/failproofaid-linux-arm64": "1.0.7-beta.
|
|
127
|
-
"@failproofai/failproofaid-darwin-x64": "1.0.7-beta.
|
|
128
|
-
"@failproofai/failproofaid-darwin-arm64": "1.0.7-beta.
|
|
125
|
+
"@failproofai/failproofaid-linux-x64": "1.0.7-beta.2",
|
|
126
|
+
"@failproofai/failproofaid-linux-arm64": "1.0.7-beta.2",
|
|
127
|
+
"@failproofai/failproofaid-darwin-x64": "1.0.7-beta.2",
|
|
128
|
+
"@failproofai/failproofaid-darwin-arm64": "1.0.7-beta.2"
|
|
129
129
|
},
|
|
130
130
|
"failproofaidBinaries": {
|
|
131
|
-
"linux-x64": "
|
|
132
|
-
"linux-arm64": "
|
|
133
|
-
"darwin-x64": "
|
|
134
|
-
"darwin-arm64": "
|
|
131
|
+
"linux-x64": "5362ab6a4cfb23d1eee598aac67aedd9e43351e62724ee1c6b40c4fc8021c08f",
|
|
132
|
+
"linux-arm64": "f14a2548b41e14ad32b7fdd03a27ffe7160da1046e422e1f651099371b0d27e5",
|
|
133
|
+
"darwin-x64": "14189cade04f17cf5278bd8061a5ec0c2a2381aa21944ae0d05e8c35b98c2816",
|
|
134
|
+
"darwin-arm64": "4ba9f5c321894d953407efa3404f0deec07dccfd7db7bc2511bb35226816b583"
|
|
135
135
|
}
|
|
136
136
|
}
|
|
@@ -1,19 +1,28 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: failproofai-sdk
|
|
3
3
|
description: |-
|
|
4
|
-
|
|
4
|
+
Make a custom AI agent — Python or TypeScript/JavaScript, on a framework or hand-built — report what it did to Failproof AI, and run your own evaluator worker (the "eval pod") that scores those runs. Reach for it on vague phrasing too: "add observability to my agent", "why isn't my agent showing up?", "run an LLM judge on our own infra".
|
|
5
5
|
|
|
6
6
|
Trigger when the user wants to:
|
|
7
|
-
• plan an integration — which points in
|
|
8
|
-
•
|
|
9
|
-
• verify
|
|
7
|
+
• plan an integration — which points in the agent loop to record;
|
|
8
|
+
• instrument — add `failproofai-sdk` (Python) or `@failproofai/sdk` (Node, Bun, Deno, Next.js): turn on an adapter (LangChain/LangGraph, CrewAI, LlamaIndex, Pydantic AI, Vercel AI SDK, Mastra) or wire a hand-built loop;
|
|
9
|
+
• verify — confirm events are written, or debug an integration that produces nothing;
|
|
10
|
+
• evaluate — write, deploy or debug an Evaluator worker in Python or TypeScript.
|
|
10
11
|
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
NOT for reading telemetry that already landed or operating a deployment (that's `fp-cloud-cli`), or building the evaluator service that scores runs (that's `agenteye-evaluator`).
|
|
12
|
+
NOT for reading telemetry or scores that already landed (that's `fp-cloud-cli`), or deciding what is worth evaluating (that's `failproofai-eval-brainstorm`).
|
|
14
13
|
---
|
|
15
14
|
|
|
16
|
-
# Failproof AI Python
|
|
15
|
+
# Failproof AI SDK — Python and TypeScript
|
|
16
|
+
|
|
17
|
+
Two packages, one pipe: `failproofai-sdk` (Python, imported as `failproofai_sdk`)
|
|
18
|
+
and `@failproofai/sdk` (TypeScript/JavaScript). They write the same 15 events in
|
|
19
|
+
the same wire format into the same spool directory. Everything in this file is
|
|
20
|
+
language-neutral unless it says otherwise, and code blocks are Python.
|
|
21
|
+
**For a TypeScript or JavaScript agent, read `references/typescript.md` alongside
|
|
22
|
+
it**: it has the camelCase names, the adapters, bundler and Next.js setup, the
|
|
23
|
+
no-framework wiring, shutdown and verification. Both packages also ship the
|
|
24
|
+
**evaluator worker** that scores finished sessions on your own infrastructure (§7,
|
|
25
|
+
`references/evaluator.md`).
|
|
17
26
|
|
|
18
27
|
The SDK records what your agent did, from inside your agent. You call it at points
|
|
19
28
|
you choose; it appends structured events to local `.jsonl` files. A separate
|
|
@@ -34,10 +43,14 @@ and it is verifiable on a laptop with no server, no API key, and no network.
|
|
|
34
43
|
The API is small — 15 event methods, all keyword-only. The hard parts are
|
|
35
44
|
**deciding where to call them** and **knowing which silences are bugs**, because
|
|
36
45
|
this SDK does not raise when you get it wrong. Sections 1-3 are the plan, 4 is the
|
|
37
|
-
code, 5-6 are the proof.
|
|
46
|
+
code, 5-6 are the proof, 7 is scoring the runs.
|
|
38
47
|
|
|
39
48
|
## 1. Install it
|
|
40
49
|
|
|
50
|
+
TypeScript/JavaScript: `npm install @failproofai/sdk` — zero dependencies, Node ≥
|
|
51
|
+
20.9, Bun or Deno, ESM and CommonJS. The npm name has no lookalike trap; the rest of
|
|
52
|
+
this section is Python. See `references/typescript.md`.
|
|
53
|
+
|
|
41
54
|
```bash
|
|
42
55
|
pip install failproofai-sdk # or: uv add failproofai-sdk
|
|
43
56
|
```
|
|
@@ -136,6 +149,10 @@ Full field-by-field catalog: `references/events.md`.
|
|
|
136
149
|
|
|
137
150
|
Work with these; none of them raise, so none of them show up in testing.
|
|
138
151
|
|
|
152
|
+
(TypeScript: the same contract with camelCase spellings, the same `"dev"` default,
|
|
153
|
+
the same reserved names and the same `duration_ms` rule, but a different shutdown
|
|
154
|
+
recipe — `references/typescript.md` → *The contract, in TypeScript* and *Shutdown*.)
|
|
155
|
+
|
|
139
156
|
- **There IS an ambient session, and it is the ergonomic path.** `session()`,
|
|
140
157
|
`agent()` and `tool_call()` bind identity on contextvars, so `session_id` and
|
|
141
158
|
`agent_id` are optional on all 15 event methods — omitted, they resolve from
|
|
@@ -277,7 +294,10 @@ Threading `session_id` and `agent_id` through every call site by hand is the thi
|
|
|
277
294
|
that makes integrations ugly and abandoned. Don't. Bind identity once per run and
|
|
278
295
|
let the call sites read it.
|
|
279
296
|
|
|
280
|
-
`references/frameworks.md` covers the four adapters
|
|
297
|
+
`references/frameworks.md` covers the four Python adapters; `references/typescript.md`
|
|
298
|
+
covers the four TypeScript ones (LangChain.js/LangGraph.js, Vercel AI SDK, Mastra,
|
|
299
|
+
LlamaIndex.TS), Next.js, bundlers, and the three-edit-site wiring for a hand-built
|
|
300
|
+
TypeScript loop. `references/integration.md` has the hand-written wrapper — one small
|
|
281
301
|
module, correct under `asyncio` and threads, adaptable to any codebase — plus
|
|
282
302
|
worked shapes for a tool dispatcher, an LLM client wrapper, and framework-specific
|
|
283
303
|
callback layers. Read it before writing your own; the naive version (a module
|
|
@@ -289,6 +309,9 @@ has a request context or a trace id, bind to that instead of inventing one.
|
|
|
289
309
|
|
|
290
310
|
## 5. Verify — watch the files
|
|
291
311
|
|
|
312
|
+
(TypeScript: same directory, same checklist; the commands are in
|
|
313
|
+
`references/typescript.md` → *Verify*.)
|
|
314
|
+
|
|
292
315
|
**This is the whole point of the file boundary: you can prove the integration
|
|
293
316
|
without a server.** Run the agent and look.
|
|
294
317
|
|
|
@@ -319,7 +342,7 @@ cat ~/.failproofai/custom-agents/events/*.jsonl | python -m json.tool --json-lin
|
|
|
319
342
|
|
|
320
343
|
Then check, in this order — the first failure explains everything downstream:
|
|
321
344
|
|
|
322
|
-
1. **Any files at all — or do they stop mid-run?** Look at stderr for
|
|
345
|
+
1. **Any files at all — or do they stop mid-run?** (Python) Look at stderr for
|
|
323
346
|
`Exception in thread failproofai-sdk-flush`. **This is the first thing to check and
|
|
324
347
|
the worst thing to miss**: one non-JSON-serializable value killed the writer,
|
|
325
348
|
and everything after it — including the at-exit flush — is gone (§3). The tell
|
|
@@ -342,8 +365,8 @@ Then check, in this order — the first failure explains everything downstream:
|
|
|
342
365
|
even when identity is a module global, and mixing only appears once two runs
|
|
343
366
|
overlap — which is production, not your laptop (§4).
|
|
344
367
|
7. **Do `tool_use` and `tool_result` share a `tool_call_id`?** Unpaired means no
|
|
345
|
-
duration. Also confirm
|
|
346
|
-
the wrong two events and reports a confident wrong duration (§3).
|
|
368
|
+
duration. Also confirm no id repeats *within a session* for the same kind — a
|
|
369
|
+
collision pairs the wrong two events and reports a confident wrong duration (§3).
|
|
347
370
|
|
|
348
371
|
A test-mode loop that costs nothing:
|
|
349
372
|
|
|
@@ -388,6 +411,10 @@ Compare the two paths first: print the directory your agent is actually writing
|
|
|
388
411
|
(`python -c "import failproofai_sdk._resolver as r; print(r.get_base_dir())"` in the
|
|
389
412
|
agent's own environment, with the agent's own env vars) and check the collector is
|
|
390
413
|
running and pointed at the same one. A `.jsonl` count that only grows is the tell.
|
|
414
|
+
The reverse is healthy: with `failproofaid` running, the directory empties seconds
|
|
415
|
+
after each flush, because the daemon ships each file and deletes it — so an empty
|
|
416
|
+
real spool proves nothing either way. Verify content with the throwaway
|
|
417
|
+
`FAILPROOFAI_HOME` loop in §5, and arrival with `fp-cloud-cli`.
|
|
391
418
|
|
|
392
419
|
Confirming events arrived on the *platform* is deliberately not this skill's job —
|
|
393
420
|
that is the `fp-cloud-cli` skill, from a **separate environment** (§1). Collector
|
|
@@ -396,4 +423,23 @@ setup and deployment are your platform's own documentation.
|
|
|
396
423
|
If the files look right (§5) and the collector is running against the same
|
|
397
424
|
directory, the integration is done.
|
|
398
425
|
|
|
399
|
-
|
|
426
|
+
## 7. Score the runs — your own evaluator worker
|
|
427
|
+
|
|
428
|
+
Sessions that land can be scored. Hosted evaluations are written in the dashboard
|
|
429
|
+
(**Analyze → eval authoring**) and run on Failproof AI's managed evaluator: that is
|
|
430
|
+
the default. When an evaluation needs your own model keys, packages, secrets,
|
|
431
|
+
private network or heavy compute, run it in **your own worker** — the eval pod.
|
|
432
|
+
|
|
433
|
+
It ships in the same packages: `failproofai_sdk.evaluator` and
|
|
434
|
+
`@failproofai/sdk/evaluator`. You declare an `Evaluator`, register versioned
|
|
435
|
+
evaluations with an optional `when` condition, and start it with an
|
|
436
|
+
`evaluations:run` key in `FAILPROOFAI_EVALUATOR_TOKEN`. It only calls out over
|
|
437
|
+
HTTPS — claim finished sessions, score them, submit — so a pod needs egress and a
|
|
438
|
+
secret, and no ingress. Its results carry the **customer** tag.
|
|
439
|
+
|
|
440
|
+
Two things to settle before writing one: the agents must already produce finished
|
|
441
|
+
sessions (§2 — nothing scores a run with no `agent_end`), and bump an evaluation's
|
|
442
|
+
`version` whenever its logic changes. Everything else — the API in both languages,
|
|
443
|
+
env vars, a Dockerfile, SIGTERM drain vs. the pod's grace period, scaling, and a
|
|
444
|
+
debugging order for a worker that scores nothing — is in `references/evaluator.md`.
|
|
445
|
+
|
|
@@ -0,0 +1,255 @@
|
|
|
1
|
+
# Your own evaluator worker — the eval pod
|
|
2
|
+
|
|
3
|
+
An evaluation scores a **finished** session. There are two places one can run:
|
|
4
|
+
|
|
5
|
+
| | Hosted | Your own worker (this page) |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| Written | in the dashboard, **Analyze → eval authoring** | in Python or TypeScript, with the SDK you already install |
|
|
8
|
+
| Runs on | Failproof AI's managed evaluator, sandboxed | your infrastructure: a container, a pod, a VM |
|
|
9
|
+
| Use it for | deterministic checks and model-backed checks Failproof AI hosts | LLM judges on your own keys, packages, secrets, your private network, models you host, heavy processing |
|
|
10
|
+
|
|
11
|
+
Default to hosted. Reach for a worker when the evaluation needs one of the
|
|
12
|
+
right-hand column's things. Deciding *what* is worth evaluating is the
|
|
13
|
+
`failproofai-eval-brainstorm` skill. This page is about building and running the
|
|
14
|
+
worker.
|
|
15
|
+
|
|
16
|
+
**How it works.** The worker only makes outbound HTTPS calls, and nothing connects
|
|
17
|
+
in to it, so there is no port, no Service and no ingress:
|
|
18
|
+
|
|
19
|
+
1. It registers its catalogue of evaluations.
|
|
20
|
+
2. It claims finished sessions and runs each applicable evaluation.
|
|
21
|
+
3. It submits the results under a heartbeat.
|
|
22
|
+
|
|
23
|
+
Its results appear on the evaluations page tagged **customer**. Hosted results are
|
|
24
|
+
tagged **managed**.
|
|
25
|
+
|
|
26
|
+
**Nothing scores a session that has no `agent_start`/`agent_end`.** The worker
|
|
27
|
+
evaluates sessions, and sessions exist only through those two events (`SKILL.md`
|
|
28
|
+
§2). Get instrumentation landing first.
|
|
29
|
+
|
|
30
|
+
## Ships in the SDK
|
|
31
|
+
|
|
32
|
+
| | Python | TypeScript |
|
|
33
|
+
|---|---|---|
|
|
34
|
+
| Install | `pip install failproofai-sdk` | `npm install @failproofai/sdk` |
|
|
35
|
+
| Import | `failproofai_sdk.evaluator` | `@failproofai/sdk/evaluator` |
|
|
36
|
+
| Run | `python evaluator.py` (with `app.run_from_env()`) or `python -m failproofai_sdk.evaluator evaluator:app` | `npx failproofai-evaluator ./evals.js` (export named `app`; `./evals.js#other` for another) or `app.runFromEnv()` |
|
|
37
|
+
|
|
38
|
+
Importing the tracing SDK does not load the evaluator. In TypeScript the reverse
|
|
39
|
+
holds too; in Python, importing `failproofai_sdk.evaluator` also imports the tracing
|
|
40
|
+
package and starts its flush thread (harmless, but it is there).
|
|
41
|
+
|
|
42
|
+
## Write the evaluations
|
|
43
|
+
|
|
44
|
+
Python:
|
|
45
|
+
|
|
46
|
+
```python
|
|
47
|
+
from failproofai_sdk.evaluator import ConditionResult, EvalResult, Evaluator, Metric, Score
|
|
48
|
+
|
|
49
|
+
app = Evaluator(name="support-evals", version="2026.09.1")
|
|
50
|
+
|
|
51
|
+
|
|
52
|
+
@app.eval(
|
|
53
|
+
"tool_efficiency",
|
|
54
|
+
version="1.0.0",
|
|
55
|
+
labels=["tools", "deterministic"],
|
|
56
|
+
when=lambda s: ConditionResult(s.count("tool_use") > 0, "no_tool_calls"),
|
|
57
|
+
)
|
|
58
|
+
def tool_efficiency(session):
|
|
59
|
+
calls = session.events_of_type("tool_use")
|
|
60
|
+
distinct = {e.payload.get("tool_name") for e in calls if e.payload.get("tool_name")}
|
|
61
|
+
value = len(distinct) / len(calls)
|
|
62
|
+
return EvalResult(
|
|
63
|
+
score=Score(value, passed=value >= 0.7),
|
|
64
|
+
metrics={"tool_call_count": Metric(len(calls), unit="events")},
|
|
65
|
+
reasoning=f"{len(distinct)} distinct tools across {len(calls)} calls",
|
|
66
|
+
)
|
|
67
|
+
|
|
68
|
+
|
|
69
|
+
@app.eval(
|
|
70
|
+
"answer_relevance",
|
|
71
|
+
version="judge-v1",
|
|
72
|
+
labels=["llm_judge"],
|
|
73
|
+
when=lambda s: ConditionResult(s.count("model_response") > 0, "no_model_response"),
|
|
74
|
+
timeout_seconds=30,
|
|
75
|
+
)
|
|
76
|
+
async def answer_relevance(session):
|
|
77
|
+
answer = session.events_of_type("model_response")[-1].payload.get("content")
|
|
78
|
+
value, why = await ask_judge(answer) # your model call: a 0-1 score and why
|
|
79
|
+
return EvalResult(score=Score(value, passed=value >= 0.7), reasoning=why)
|
|
80
|
+
|
|
81
|
+
|
|
82
|
+
if __name__ == "__main__":
|
|
83
|
+
app.run_from_env()
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
TypeScript (same API, camelCase, options objects):
|
|
87
|
+
|
|
88
|
+
```ts
|
|
89
|
+
import { ConditionResult, EvalResult, Evaluator, Metric, Score } from "@failproofai/sdk/evaluator";
|
|
90
|
+
|
|
91
|
+
export const app = new Evaluator({ name: "support-evals", version: "2026.09.1" });
|
|
92
|
+
|
|
93
|
+
app.eval(
|
|
94
|
+
"tool_efficiency",
|
|
95
|
+
{
|
|
96
|
+
version: "1.0.0",
|
|
97
|
+
labels: ["tools", "deterministic"],
|
|
98
|
+
when: (s) => new ConditionResult(s.count("tool_use") > 0, "no_tool_calls"),
|
|
99
|
+
},
|
|
100
|
+
(session) => {
|
|
101
|
+
const calls = session.eventsOfType("tool_use");
|
|
102
|
+
const distinct = new Set(calls.map((e) => e.payload.tool_name).filter(Boolean));
|
|
103
|
+
const value = distinct.size / calls.length;
|
|
104
|
+
return new EvalResult({
|
|
105
|
+
score: new Score(value, { passed: value >= 0.7 }),
|
|
106
|
+
metrics: { tool_call_count: new Metric(calls.length, { unit: "events" }) },
|
|
107
|
+
reasoning: `${distinct.size} distinct tools across ${calls.length} calls`,
|
|
108
|
+
});
|
|
109
|
+
},
|
|
110
|
+
);
|
|
111
|
+
|
|
112
|
+
app.eval(
|
|
113
|
+
"answer_relevance",
|
|
114
|
+
{
|
|
115
|
+
version: "judge-v1",
|
|
116
|
+
labels: ["llm_judge"],
|
|
117
|
+
when: (s) => new ConditionResult(s.count("model_response") > 0, "no_model_response"),
|
|
118
|
+
timeoutSeconds: 30,
|
|
119
|
+
},
|
|
120
|
+
async (session) => {
|
|
121
|
+
const answer = session.eventsOfType("model_response").at(-1)?.payload.content;
|
|
122
|
+
const { value, why } = await askJudge(answer); // your model call
|
|
123
|
+
return new EvalResult({ score: new Score(value, { passed: value >= 0.7 }), reasoning: why });
|
|
124
|
+
},
|
|
125
|
+
);
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
The evaluator module can be an ES module or CommonJS (`.js`, `.mjs`, `.cjs`, plain
|
|
129
|
+
`tsc` output). For a `.ts` file, either compile it and point `failproofai-evaluator`
|
|
130
|
+
at the `.js`, or call `await app.runFromEnv()` at the bottom and start it with
|
|
131
|
+
`npx tsx evals.ts`.
|
|
132
|
+
|
|
133
|
+
### The rules that bite
|
|
134
|
+
|
|
135
|
+
- **The key is what results chart under.** It must match `^[a-z][a-z0-9_]*$`.
|
|
136
|
+
**Bump `version` whenever the logic changes.** Each result keeps the version that produced it, so a chart shows
|
|
137
|
+
exactly when new logic took over. Keys must be unique, and one worker holds at
|
|
138
|
+
most 100 evaluations.
|
|
139
|
+
- **`when` is how you scope.** Return `ConditionResult(False, "<reason_code>")` to
|
|
140
|
+
skip a session, and the reason is recorded. A judge with no condition costs a
|
|
141
|
+
model call on every session.
|
|
142
|
+
- **Result shapes:**
|
|
143
|
+
- `result_kind` / `resultKind` is `"score"` by default.
|
|
144
|
+
- For a `"metric"` or `"assertion"` evaluation, name one `metrics` or
|
|
145
|
+
`assertions` entry after the key: that entry is the result.
|
|
146
|
+
- Every `EvalResult` carries at least one and at most 25 scores, metrics or
|
|
147
|
+
assertions.
|
|
148
|
+
- `Score` values are 0–1, and anything else throws.
|
|
149
|
+
- **Read payload keys off a real session.** Keys like `tool_name`, `content` and
|
|
150
|
+
`response` are whatever the agents emitted. The adapters add framework fields
|
|
151
|
+
under `fw_*`.
|
|
152
|
+
- **Evaluations must yield.**
|
|
153
|
+
- Every evaluation is bounded by `timeout_seconds` / `timeoutSeconds`, and by
|
|
154
|
+
300 s when you set none.
|
|
155
|
+
- **Python:** a synchronous evaluation that overruns cannot be interrupted. Its
|
|
156
|
+
thread runs on, and a permanently blocked one leaks a thread per session. Write
|
|
157
|
+
judges and network calls as `async def`.
|
|
158
|
+
- **TypeScript:** a synchronous loop blocks the only thread, so no timeout can
|
|
159
|
+
fire. Write evaluations `async`.
|
|
160
|
+
- **The session object:**
|
|
161
|
+
- Python: `session_id`, `agent_id`, `environment`, `started_at`, `ended_at`,
|
|
162
|
+
`events`, `count(type)`, `events_of_type(type)`.
|
|
163
|
+
- TypeScript: `sessionId`, `agentId`, `environment`, `startedAt`, `endedAt`,
|
|
164
|
+
`events`, `count(type)`, `eventsOfType(type)`.
|
|
165
|
+
- Each event has `id`, `ts`, `event_type` (TS `eventType`) and `payload`.
|
|
166
|
+
|
|
167
|
+
## Run it
|
|
168
|
+
|
|
169
|
+
Create a key with the **`evaluations:run`** permission under **Administration →
|
|
170
|
+
Keys**. Inject it from your secret store; never paste it into a command line or an
|
|
171
|
+
image.
|
|
172
|
+
|
|
173
|
+
| Variable | Default | |
|
|
174
|
+
|---|---|---|
|
|
175
|
+
| `FAILPROOFAI_EVALUATOR_URL` | required | `https://app.befailproof.ai` for Cloud, or your instance. HTTPS unless loopback |
|
|
176
|
+
| `FAILPROOFAI_EVALUATOR_TOKEN` | required | the `evaluations:run` key |
|
|
177
|
+
| `FAILPROOFAI_EVALUATOR_WORKER_ID` | `<hostname>-<pid>` | names this worker |
|
|
178
|
+
| `FAILPROOFAI_EVALUATOR_CONCURRENCY` | `1` | sessions scored at once by this process |
|
|
179
|
+
| `FAILPROOFAI_EVALUATOR_REQUEST_TIMEOUT_SECONDS` | `30` | per request to Failproof AI |
|
|
180
|
+
| `FAILPROOFAI_EVALUATOR_DRAIN_TIMEOUT_SECONDS` | `60` | how long a stopping worker waits for runs in flight |
|
|
181
|
+
| `FAILPROOFAI_EVALUATOR_ALLOW_INSECURE_HTTP` | `false` | plain HTTP to a non-loopback URL. **Cleartext token and transcripts.** Isolated dev networks only |
|
|
182
|
+
| `FAILPROOFAI_EVALUATOR_MODULE` | — | the module when the CLI is given none: `module:attr` (Python), `path#export` (TS) |
|
|
183
|
+
|
|
184
|
+
Local smoke run against Cloud:
|
|
185
|
+
|
|
186
|
+
```bash
|
|
187
|
+
export FAILPROOFAI_EVALUATOR_URL=https://app.befailproof.ai
|
|
188
|
+
export FAILPROOFAI_EVALUATOR_TOKEN="$(your-secret-store read failproofai/evaluator)"
|
|
189
|
+
python evaluator.py # or: npx failproofai-evaluator ./evals.js
|
|
190
|
+
```
|
|
191
|
+
|
|
192
|
+
Then finish a session from an instrumented agent and watch for a **customer**
|
|
193
|
+
result on it. Evaluation runs forward: a worker started now scores sessions that
|
|
194
|
+
finish while it runs. The dashboard's "score sessions you already have" backfill
|
|
195
|
+
covers hosted evaluations only; there is no backfill for a worker's own evaluations.
|
|
196
|
+
|
|
197
|
+
## Deploy it as a pod
|
|
198
|
+
|
|
199
|
+
Treat it as a long-running, outbound-only worker:
|
|
200
|
+
|
|
201
|
+
```dockerfile
|
|
202
|
+
# Python
|
|
203
|
+
FROM python:3.12-slim
|
|
204
|
+
RUN pip install --no-cache-dir failproofai-sdk # plus whatever your judges need
|
|
205
|
+
COPY evaluator.py /app/evaluator.py
|
|
206
|
+
CMD ["python", "/app/evaluator.py"]
|
|
207
|
+
```
|
|
208
|
+
|
|
209
|
+
```dockerfile
|
|
210
|
+
# TypeScript (compiled to dist/evals.js)
|
|
211
|
+
FROM node:22-slim
|
|
212
|
+
WORKDIR /app
|
|
213
|
+
COPY package*.json ./
|
|
214
|
+
RUN npm ci --omit=dev
|
|
215
|
+
COPY dist/ ./dist/
|
|
216
|
+
CMD ["npx", "failproofai-evaluator", "./dist/evals.js"]
|
|
217
|
+
```
|
|
218
|
+
|
|
219
|
+
- **Secrets:**
|
|
220
|
+
- `FAILPROOFAI_EVALUATOR_TOKEN` comes from a Secret (Kubernetes) or your secret
|
|
221
|
+
manager.
|
|
222
|
+
- The judge's own model key belongs there too.
|
|
223
|
+
- Egress must reach the Failproof AI URL and your model endpoint. No ingress is
|
|
224
|
+
needed.
|
|
225
|
+
- **Shutdown:** both workers stop on `SIGTERM`/`SIGINT` and drain runs in flight
|
|
226
|
+
for up to `FAILPROOFAI_EVALUATOR_DRAIN_TIMEOUT_SECONDS`. Give the orchestrator
|
|
227
|
+
more grace than that. In Kubernetes, set `terminationGracePeriodSeconds` above
|
|
228
|
+
the drain timeout, e.g. 90 for the default 60. Otherwise a rolling deploy
|
|
229
|
+
hard-kills evaluations mid-run.
|
|
230
|
+
- **Scaling:**
|
|
231
|
+
- Raise `FAILPROOFAI_EVALUATOR_CONCURRENCY` for I/O-bound judges.
|
|
232
|
+
- Add replicas for more throughput. Each replica claims its own sessions.
|
|
233
|
+
- Leave `WORKER_ID` unset, or make it unique per replica: the default
|
|
234
|
+
`<hostname>-<pid>` already is.
|
|
235
|
+
- **Releasing new logic.** Ship the image with a bumped eval `version`. Removing an
|
|
236
|
+
evaluation from the worker, or stopping the worker, stops it running. There is
|
|
237
|
+
nothing to disable in the dashboard, and worker-registered evaluations are not
|
|
238
|
+
listed on the eval authoring page.
|
|
239
|
+
|
|
240
|
+
## Debugging a worker that scores nothing
|
|
241
|
+
|
|
242
|
+
Work through these in order:
|
|
243
|
+
|
|
244
|
+
1. **Does it start?**
|
|
245
|
+
- A missing URL or token fails at startup with the variable's name.
|
|
246
|
+
- An HTTP URL to a non-loopback host is refused unless you opt in.
|
|
247
|
+
2. **Is the key right?** It needs `evaluations:run` in the same organisation as the
|
|
248
|
+
agents.
|
|
249
|
+
3. **Are sessions finishing?** No `agent_end` means no finished session, and
|
|
250
|
+
nothing to claim. Check the instrumentation (`SKILL.md` §5).
|
|
251
|
+
4. **Does `when` skip everything?** The skip reason is recorded per session. A
|
|
252
|
+
condition reading a payload key the agents never emit skips every session.
|
|
253
|
+
5. **Are evaluations timing out?** Check the worker's own log. Make judges `async`
|
|
254
|
+
and give them a `timeout_seconds` / `timeoutSeconds` that covers one model
|
|
255
|
+
call.
|
|
@@ -1,5 +1,11 @@
|
|
|
1
1
|
# Event catalog
|
|
2
2
|
|
|
3
|
+
> TypeScript/JavaScript: the same 15 events on `failproofai.event`, camelCase
|
|
4
|
+
> (`toolUse({ toolName, toolCallId })`) taking one options object; the file written
|
|
5
|
+
> is identical. Name map in `typescript.md`. Code below is Python, and where it
|
|
6
|
+
> says `ValueError` or `TypeError`, TypeScript throws `TypeError` or a plain
|
|
7
|
+
> `Error` (listed in `typescript.md` → *The contract, in TypeScript*).
|
|
8
|
+
|
|
3
9
|
Every method lives on `failproofai_sdk.event`, is **keyword-only**, and returns `None`.
|
|
4
10
|
Nothing here blocks or does I/O — the call queues the event and returns.
|
|
5
11
|
|
|
@@ -93,16 +99,19 @@ its original) oldest-first, and will keep bracketing the wrong pairs. **Always
|
|
|
93
99
|
pass `duration_ms` on the `model_response` too** — that is what keeps the reported
|
|
94
100
|
duration right even when the bracketing is wrong.
|
|
95
101
|
|
|
96
|
-
**Correlation is a
|
|
97
|
-
starts there until their matching end arrives.
|
|
102
|
+
**Correlation is a map keyed by session, kind and your id.** The SDK holds open
|
|
103
|
+
starts there until their matching end arrives. An id only has to be unique *within
|
|
104
|
+
one session, for one kind* (the second bullet). Consequences, in order of how much
|
|
98
105
|
they hurt:
|
|
99
106
|
|
|
100
|
-
- **Per-run counters are unsafe.** `call_1`, `call_2` —
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
107
|
+
- **Per-run counters are unsafe inside a shared session.** `call_1`, `call_2` —
|
|
108
|
+
common in home-grown loops — are fine while every run has its own session. They
|
|
109
|
+
collide when two runs share one: a retried job that reuses its job id as the
|
|
110
|
+
session, or two agents in one session each counting from `call_1`. The failure is
|
|
111
|
+
not the missing duration the docs might lead you to expect; it is a *plausible
|
|
112
|
+
wrong number attributed to the wrong run*, which is worse. Reuse your framework's
|
|
113
|
+
id (Anthropic and OpenAI tool-call ids are globally unique), or a `uuid4`.
|
|
114
|
+
`failproofai_sdk.tool_call()` generates a `uuid4` for you.
|
|
106
115
|
- **`tool_call_id` and `hook_id` no longer collide with each other.** They live in
|
|
107
116
|
separate namespaces, so a `hook_completed(hook_id="x")` cannot pair with a
|
|
108
117
|
pending `tool_use(tool_call_id="x")`. The key is `<kind>:<session_id>:<id>`, so
|
|
@@ -1,5 +1,8 @@
|
|
|
1
1
|
# Framework integrations
|
|
2
2
|
|
|
3
|
+
> TypeScript/JavaScript (LangChain.js/LangGraph.js, Vercel AI SDK, Mastra,
|
|
4
|
+
> LlamaIndex.TS, Next.js): see `typescript.md`. This page is the Python SDK.
|
|
5
|
+
|
|
3
6
|
If the agent runs on LangChain/LangGraph, CrewAI, LlamaIndex or Pydantic AI, you
|
|
4
7
|
do not write the instrumentation — you turn it on. The adapters ship inside the
|
|
5
8
|
SDK wheel and are imported only when you ask for them.
|
|
@@ -1,5 +1,9 @@
|
|
|
1
1
|
# Writing the integration
|
|
2
2
|
|
|
3
|
+
> TypeScript/JavaScript: the same three scopes exist as `session()`, `agent()` and
|
|
4
|
+
> `toolCall()` on `AsyncLocalStorage`, and a hand-built loop is three edit sites —
|
|
5
|
+
> see `typescript.md` → *An agent with no framework*. This page is the Python SDK.
|
|
6
|
+
|
|
3
7
|
Identity is ambient. Bind it once per run with a context manager and every
|
|
4
8
|
`event.*` call inside — including calls in functions that have never heard of
|
|
5
9
|
Failproof AI — lands on the right session and agent.
|
|
@@ -50,8 +54,8 @@ a session is *defined* as something that emitted `agent_start`.
|
|
|
50
54
|
one run rather than splitting it in two.
|
|
51
55
|
- `parent_id` defaults to the enclosing agent from the scope stack. Pass
|
|
52
56
|
`parent_id=None` to force a root span, or a string to override.
|
|
53
|
-
- `tool_call_id` defaults to `uuid4().hex` — unique
|
|
54
|
-
the correlation map
|
|
57
|
+
- `tool_call_id` defaults to `uuid4().hex` — unique everywhere, so it can never
|
|
58
|
+
collide in the correlation map (see `events.md`).
|
|
55
59
|
|
|
56
60
|
## What `agent()` does on the way out
|
|
57
61
|
|