agent-learning-kit 0.1.0__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- agent_learning_kit-0.1.0.dist-info/METADATA +381 -0
- agent_learning_kit-0.1.0.dist-info/RECORD +642 -0
- agent_learning_kit-0.1.0.dist-info/WHEEL +4 -0
- agent_learning_kit-0.1.0.dist-info/entry_points.txt +5 -0
- agent_learning_kit-0.1.0.dist-info/licenses/LICENSE +173 -0
- agent_learning_kit-0.1.0.dist-info/licenses/NOTICE +7 -0
- fi/__init__.py +5 -0
- fi/alk/__init__.py +57 -0
- fi/alk/_facade.py +31 -0
- fi/alk/_module_alias.py +68 -0
- fi/alk/_paths.py +14 -0
- fi/alk/_schema.py +522 -0
- fi/alk/actions.py +727 -0
- fi/alk/bench/__init__.py +517 -0
- fi/alk/bench/_codeexec.py +213 -0
- fi/alk/bench/_coding.py +215 -0
- fi/alk/bench/_docker.py +237 -0
- fi/alk/bench/_grader.py +286 -0
- fi/alk/bench/_pull.py +212 -0
- fi/alk/bench/_voice.py +147 -0
- fi/alk/capabilities.py +627 -0
- fi/alk/cli.py +6396 -0
- fi/alk/config.py +130 -0
- fi/alk/cua_loop.py +562 -0
- fi/alk/evals.py +2351 -0
- fi/alk/extensions.py +163 -0
- fi/alk/harness/ARCHITECTURE.md +231 -0
- fi/alk/harness/DESIGN.md +246 -0
- fi/alk/harness/ENVIRONMENT_CONFORMANCE.md +127 -0
- fi/alk/harness/HOW-IT-WORKS.md +297 -0
- fi/alk/harness/IMPLEMENTATION_AND_VALIDATION_STATUS.md +229 -0
- fi/alk/harness/README.md +417 -0
- fi/alk/harness/__init__.py +77 -0
- fi/alk/harness/__main__.py +3 -0
- fi/alk/harness/amend.py +312 -0
- fi/alk/harness/artifacts.py +319 -0
- fi/alk/harness/authoring_entrypoint.py +189 -0
- fi/alk/harness/authoring_runtime_validation.py +267 -0
- fi/alk/harness/backends/README.md +43 -0
- fi/alk/harness/backends/__init__.py +122 -0
- fi/alk/harness/backends/base.py +241 -0
- fi/alk/harness/backends/claude.py +211 -0
- fi/alk/harness/backends/files.py +182 -0
- fi/alk/harness/backends/vertex_gemini.py +457 -0
- fi/alk/harness/background_noise.py +95 -0
- fi/alk/harness/build.py +385 -0
- fi/alk/harness/bundle.py +593 -0
- fi/alk/harness/bundle_author_v2.py +1831 -0
- fi/alk/harness/bundle_v2.py +719 -0
- fi/alk/harness/call_runner.py +1440 -0
- fi/alk/harness/callback_http_adapter.py +111 -0
- fi/alk/harness/catalogue.py +287 -0
- fi/alk/harness/chat.py +428 -0
- fi/alk/harness/chat_call_runner.py +506 -0
- fi/alk/harness/checks.py +136 -0
- fi/alk/harness/cli.py +1354 -0
- fi/alk/harness/config.py +338 -0
- fi/alk/harness/contract.py +718 -0
- fi/alk/harness/credentials.py +674 -0
- fi/alk/harness/data/persona_vocabulary.json +111 -0
- fi/alk/harness/environment.py +99 -0
- fi/alk/harness/environment_plan.py +168 -0
- fi/alk/harness/events.py +125 -0
- fi/alk/harness/executor.py +304 -0
- fi/alk/harness/folder.py +234 -0
- fi/alk/harness/generated_runtime.py +815 -0
- fi/alk/harness/github.py +72 -0
- fi/alk/harness/hosted_authoring_entrypoint.py +183 -0
- fi/alk/harness/hosted_entrypoint.py +2402 -0
- fi/alk/harness/hosted_scheduler.py +2218 -0
- fi/alk/harness/job.py +426 -0
- fi/alk/harness/judge.py +184 -0
- fi/alk/harness/livekit_source.py +50 -0
- fi/alk/harness/livekit_tool_trace_bootstrap.py +71 -0
- fi/alk/harness/observability.py +208 -0
- fi/alk/harness/outbound.py +3252 -0
- fi/alk/harness/packaging.py +515 -0
- fi/alk/harness/persona_guides.py +157 -0
- fi/alk/harness/platform.py +692 -0
- fi/alk/harness/process_preflight.py +764 -0
- fi/alk/harness/process_runtime.py +5670 -0
- fi/alk/harness/prove.py +425 -0
- fi/alk/harness/provider_import.py +703 -0
- fi/alk/harness/provider_lifecycle.py +392 -0
- fi/alk/harness/provision.py +2896 -0
- fi/alk/harness/reception.py +147 -0
- fi/alk/harness/retell_chat_call_runner.py +373 -0
- fi/alk/harness/run/__init__.py +296 -0
- fi/alk/harness/run/alk.py +184 -0
- fi/alk/harness/run/call.py +162 -0
- fi/alk/harness/run/conversation.py +264 -0
- fi/alk/harness/run/data/voices_by_language_and_gender.json +693 -0
- fi/alk/harness/run/evidence.py +195 -0
- fi/alk/harness/run/grade.py +598 -0
- fi/alk/harness/run/live.py +297 -0
- fi/alk/harness/run/models.py +56 -0
- fi/alk/harness/run/platform_evals.py +227 -0
- fi/alk/harness/run/sdk_voice.py +130 -0
- fi/alk/harness/run/simulation.py +1209 -0
- fi/alk/harness/run/stage.py +91 -0
- fi/alk/harness/run/targets.py +508 -0
- fi/alk/harness/run/tools.py +601 -0
- fi/alk/harness/run/voice.py +340 -0
- fi/alk/harness/runtime.py +172 -0
- fi/alk/harness/sandbox_server.py +2011 -0
- fi/alk/harness/sandbox_worker.py +44 -0
- fi/alk/harness/scenario.py +1048 -0
- fi/alk/harness/scenario_source.py +879 -0
- fi/alk/harness/scenario_tools.py +1143 -0
- fi/alk/harness/scenarios.py +915 -0
- fi/alk/harness/secrets.py +168 -0
- fi/alk/harness/service_catalog.py +97 -0
- fi/alk/harness/session.py +391 -0
- fi/alk/harness/sessions.py +372 -0
- fi/alk/harness/simulator.py +76 -0
- fi/alk/harness/simulator_voice.py +928 -0
- fi/alk/harness/skills/build-environment/SKILL.md +538 -0
- fi/alk/harness/skills/harness.md +131 -0
- fi/alk/harness/skills/kinds/chat.md +48 -0
- fi/alk/harness/skills/kinds/voice-voicemail.md +63 -0
- fi/alk/harness/skills/kinds/voice.md +59 -0
- fi/alk/harness/skills/plan-suite/SKILL.md +103 -0
- fi/alk/harness/skills/provision-environment/SKILL.md +136 -0
- fi/alk/harness/skills/run-scenarios/SKILL.md +112 -0
- fi/alk/harness/skills/understand-agent/SKILL.md +251 -0
- fi/alk/harness/skills/write-scenarios/SKILL.md +606 -0
- fi/alk/harness/skills/write-scenarios/references/refusals.md +28 -0
- fi/alk/harness/skills/write-scenarios/references/world-api.md +92 -0
- fi/alk/harness/source_data_invariants.py +444 -0
- fi/alk/harness/source_tool_evidence.py +79 -0
- fi/alk/harness/sources.py +253 -0
- fi/alk/harness/spend.py +140 -0
- fi/alk/harness/tool_trace_proxy.py +104 -0
- fi/alk/harness/tools.py +1018 -0
- fi/alk/harness/understand.py +169 -0
- fi/alk/harness/voicemail_audio.py +74 -0
- fi/alk/harness/world/__init__.py +33 -0
- fi/alk/harness/world/errors.py +68 -0
- fi/alk/harness/world/expectations.py +91 -0
- fi/alk/harness/world/handle.py +538 -0
- fi/alk/harness/world/kinds.py +196 -0
- fi/alk/harness/world/mutate.py +186 -0
- fi/alk/harness/world/probe.py +413 -0
- fi/alk/harness/world/provision.py +511 -0
- fi/alk/harness/world/provisioned.py +191 -0
- fi/alk/harness/world/runtime.py +616 -0
- fi/alk/harness/world/snapshot.py +288 -0
- fi/alk/harness/world/stores/__init__.py +305 -0
- fi/alk/harness/world/stores/container.py +215 -0
- fi/alk/harness/world/stores/inprocess.py +346 -0
- fi/alk/harness/world/stores/postgres.py +481 -0
- fi/alk/harness/world/stores/prove.py +202 -0
- fi/alk/harness/world/stores/sqlite.py +245 -0
- fi/alk/harness/world/stores/written.py +182 -0
- fi/alk/harness/world/tools.py +1516 -0
- fi/alk/harness/world/workspace.py +144 -0
- fi/alk/image_loop.py +453 -0
- fi/alk/image_perturb.py +241 -0
- fi/alk/improve.py +274 -0
- fi/alk/live/__init__.py +154 -0
- fi/alk/live/_attribution.py +184 -0
- fi/alk/live/_capture.py +264 -0
- fi/alk/live/_codec.py +391 -0
- fi/alk/live/_contract.py +134 -0
- fi/alk/live/_loopback.py +316 -0
- fi/alk/live/_perturb.py +449 -0
- fi/alk/live/_runner.py +386 -0
- fi/alk/live/_stats.py +561 -0
- fi/alk/live/_transcript.py +240 -0
- fi/alk/live/_workers/__init__.py +9 -0
- fi/alk/live/_workers/a2a_worker.py +316 -0
- fi/alk/live/_workers/langgraph_worker.py +217 -0
- fi/alk/live/_workers/livekit_worker.py +207 -0
- fi/alk/live/_workers/mcp_loopback_server.py +46 -0
- fi/alk/live/_workers/mcp_worker.py +158 -0
- fi/alk/live/_workers/pipecat_worker.py +189 -0
- fi/alk/live/a2a_lane.py +138 -0
- fi/alk/live/langgraph_lane.py +339 -0
- fi/alk/live/livekit_lane.py +376 -0
- fi/alk/live/mcp_lane.py +172 -0
- fi/alk/live/pipecat_lane.py +341 -0
- fi/alk/live/voice_redteam.py +494 -0
- fi/alk/loss.py +306 -0
- fi/alk/optimize.py +36260 -0
- fi/alk/practice/__init__.py +51 -0
- fi/alk/practice/_assess.py +103 -0
- fi/alk/practice/_budget.py +81 -0
- fi/alk/practice/_calibrate.py +69 -0
- fi/alk/practice/_capstone.py +86 -0
- fi/alk/practice/_contract.py +91 -0
- fi/alk/practice/_diagnose.py +79 -0
- fi/alk/practice/_drill.py +196 -0
- fi/alk/practice/_experiment.py +720 -0
- fi/alk/practice/_schedule.py +102 -0
- fi/alk/practice/_store.py +194 -0
- fi/alk/practice/_trainer.py +245 -0
- fi/alk/practice/_update.py +125 -0
- fi/alk/redteam.py +2621 -0
- fi/alk/rewardhack.py +237 -0
- fi/alk/simulate.py +10351 -0
- fi/alk/studio/__init__.py +82 -0
- fi/alk/studio/_bias.py +314 -0
- fi/alk/studio/_calibration.py +522 -0
- fi/alk/studio/_coverage.py +262 -0
- fi/alk/studio/_download.py +665 -0
- fi/alk/studio/_fidelity_attack.py +114 -0
- fi/alk/studio/_generate.py +652 -0
- fi/alk/studio/_library.py +370 -0
- fi/alk/studio/_scan.py +134 -0
- fi/alk/studio/_upgrade.py +42 -0
- fi/alk/studio/_vendor.py +172 -0
- fi/alk/suite.py +4200 -0
- fi/alk/tasks.py +828 -0
- fi/alk/telemetry/__init__.py +149 -0
- fi/alk/telemetry/_contract.py +141 -0
- fi/alk/telemetry/_emit.py +182 -0
- fi/alk/telemetry/_ledger.py +296 -0
- fi/alk/telemetry/_queue.py +127 -0
- fi/alk/telemetry/_row.py +294 -0
- fi/alk/telemetry/_run.py +233 -0
- fi/alk/telemetry/_sync.py +193 -0
- fi/alk/telemetry/_url.py +119 -0
- fi/alk/trinity.py +49397 -0
- fi/alk/voice_loop.py +174 -0
- fi/api/__init__.py +1 -0
- fi/api/auth.py +137 -0
- fi/api/types.py +29 -0
- fi/cli/__init__.py +9 -0
- fi/cli/assertions/__init__.py +25 -0
- fi/cli/assertions/conditions.py +76 -0
- fi/cli/assertions/evaluator.py +286 -0
- fi/cli/assertions/exit_codes.py +20 -0
- fi/cli/assertions/parser.py +131 -0
- fi/cli/assertions/reporter.py +194 -0
- fi/cli/commands/__init__.py +9 -0
- fi/cli/commands/config.py +165 -0
- fi/cli/commands/export.py +208 -0
- fi/cli/commands/init.py +112 -0
- fi/cli/commands/list_cmd.py +213 -0
- fi/cli/commands/run.py +486 -0
- fi/cli/commands/validate.py +173 -0
- fi/cli/commands/view.py +424 -0
- fi/cli/config/__init__.py +6 -0
- fi/cli/config/defaults.py +206 -0
- fi/cli/config/loader.py +155 -0
- fi/cli/config/schema.py +174 -0
- fi/cli/main.py +78 -0
- fi/cli/output/__init__.py +6 -0
- fi/cli/output/formatters.py +106 -0
- fi/cli/output/reporters.py +46 -0
- fi/cli/storage/__init__.py +5 -0
- fi/cli/storage/run_history.py +249 -0
- fi/cli/utils/__init__.py +5 -0
- fi/cli/utils/console.py +44 -0
- fi/evals/__init__.py +131 -0
- fi/evals/autoeval/__init__.py +137 -0
- fi/evals/autoeval/analyzer.py +211 -0
- fi/evals/autoeval/config.py +244 -0
- fi/evals/autoeval/export.py +213 -0
- fi/evals/autoeval/interactive.py +283 -0
- fi/evals/autoeval/pipeline.py +625 -0
- fi/evals/autoeval/prompts.py +139 -0
- fi/evals/autoeval/recommender.py +242 -0
- fi/evals/autoeval/rules.py +589 -0
- fi/evals/autoeval/templates.py +299 -0
- fi/evals/autoeval/types.py +232 -0
- fi/evals/core/__init__.py +16 -0
- fi/evals/core/cloud_registry.py +184 -0
- fi/evals/core/engines.py +368 -0
- fi/evals/core/evaluate.py +319 -0
- fi/evals/core/judge_prompt.py +90 -0
- fi/evals/core/prompt_generator.py +83 -0
- fi/evals/core/registry.py +57 -0
- fi/evals/core/result.py +55 -0
- fi/evals/evaluator.py +721 -0
- fi/evals/execution.py +168 -0
- fi/evals/feedback/__init__.py +32 -0
- fi/evals/feedback/calibrator.py +160 -0
- fi/evals/feedback/collector.py +214 -0
- fi/evals/feedback/hooks.py +81 -0
- fi/evals/feedback/retriever.py +128 -0
- fi/evals/feedback/store.py +272 -0
- fi/evals/feedback/types.py +99 -0
- fi/evals/framework/README.md +79 -0
- fi/evals/framework/__init__.py +267 -0
- fi/evals/framework/backends/Dockerfile.eval-runner +33 -0
- fi/evals/framework/backends/__init__.py +99 -0
- fi/evals/framework/backends/_container.py +141 -0
- fi/evals/framework/backends/_utils.py +145 -0
- fi/evals/framework/backends/base.py +223 -0
- fi/evals/framework/backends/celery_backend.py +417 -0
- fi/evals/framework/backends/celery_worker.py +78 -0
- fi/evals/framework/backends/kubernetes_backend.py +665 -0
- fi/evals/framework/backends/ray_backend.py +521 -0
- fi/evals/framework/backends/temporal.py +350 -0
- fi/evals/framework/backends/temporal_worker.py +126 -0
- fi/evals/framework/backends/thread_pool.py +286 -0
- fi/evals/framework/context.py +258 -0
- fi/evals/framework/enrichment.py +306 -0
- fi/evals/framework/evals/__init__.py +68 -0
- fi/evals/framework/evals/agentic.py +399 -0
- fi/evals/framework/evals/builder.py +609 -0
- fi/evals/framework/evals/semantic.py +142 -0
- fi/evals/framework/evaluator.py +647 -0
- fi/evals/framework/evaluators/__init__.py +22 -0
- fi/evals/framework/evaluators/blocking.py +347 -0
- fi/evals/framework/evaluators/non_blocking.py +577 -0
- fi/evals/framework/propagation.py +421 -0
- fi/evals/framework/protocols.py +385 -0
- fi/evals/framework/registry.py +370 -0
- fi/evals/framework/resilience/__init__.py +150 -0
- fi/evals/framework/resilience/circuit_breaker.py +309 -0
- fi/evals/framework/resilience/degradation.py +355 -0
- fi/evals/framework/resilience/health.py +505 -0
- fi/evals/framework/resilience/rate_limiter.py +228 -0
- fi/evals/framework/resilience/retry.py +274 -0
- fi/evals/framework/resilience/types.py +288 -0
- fi/evals/framework/resilience/wrapper.py +433 -0
- fi/evals/framework/types.py +218 -0
- fi/evals/guardrails/README.md +915 -0
- fi/evals/guardrails/__init__.py +96 -0
- fi/evals/guardrails/backends/__init__.py +43 -0
- fi/evals/guardrails/backends/azure.py +361 -0
- fi/evals/guardrails/backends/base.py +88 -0
- fi/evals/guardrails/backends/generic_llm.py +163 -0
- fi/evals/guardrails/backends/granite.py +216 -0
- fi/evals/guardrails/backends/llamaguard.py +221 -0
- fi/evals/guardrails/backends/local_base.py +479 -0
- fi/evals/guardrails/backends/openai.py +365 -0
- fi/evals/guardrails/backends/qwen.py +170 -0
- fi/evals/guardrails/backends/shieldgemma.py +154 -0
- fi/evals/guardrails/backends/turing.py +235 -0
- fi/evals/guardrails/backends/vllm_client.py +321 -0
- fi/evals/guardrails/backends/wildguard.py +188 -0
- fi/evals/guardrails/base.py +888 -0
- fi/evals/guardrails/config.py +221 -0
- fi/evals/guardrails/discovery.py +243 -0
- fi/evals/guardrails/gateway.py +437 -0
- fi/evals/guardrails/registry.py +231 -0
- fi/evals/guardrails/scanners/__init__.py +127 -0
- fi/evals/guardrails/scanners/base.py +191 -0
- fi/evals/guardrails/scanners/code_injection.py +243 -0
- fi/evals/guardrails/scanners/eval_delegate.py +574 -0
- fi/evals/guardrails/scanners/invisible_chars.py +351 -0
- fi/evals/guardrails/scanners/jailbreak.py +412 -0
- fi/evals/guardrails/scanners/language.py +288 -0
- fi/evals/guardrails/scanners/pipeline.py +260 -0
- fi/evals/guardrails/scanners/regex.py +311 -0
- fi/evals/guardrails/scanners/secrets.py +274 -0
- fi/evals/guardrails/scanners/topics.py +649 -0
- fi/evals/guardrails/scanners/urls.py +341 -0
- fi/evals/guardrails/types.py +96 -0
- fi/evals/llm/__init__.py +3 -0
- fi/evals/llm/base_llm_provider.py +35 -0
- fi/evals/llm/providers/litellm.py +70 -0
- fi/evals/local/__init__.py +90 -0
- fi/evals/local/evaluator.py +690 -0
- fi/evals/local/execution_mode.py +121 -0
- fi/evals/local/llm.py +489 -0
- fi/evals/local/metrics/__init__.py +19 -0
- fi/evals/local/registry.py +360 -0
- fi/evals/manager.py +1018 -0
- fi/evals/manager_types.py +362 -0
- fi/evals/metrics/__init__.py +185 -0
- fi/evals/metrics/agents/__init__.py +74 -0
- fi/evals/metrics/agents/metrics.py +693 -0
- fi/evals/metrics/agents/report.py +36463 -0
- fi/evals/metrics/agents/types.py +160 -0
- fi/evals/metrics/base_llm_metric.py +111 -0
- fi/evals/metrics/base_metric.py +138 -0
- fi/evals/metrics/code_security/__init__.py +305 -0
- fi/evals/metrics/code_security/analyzer.py +985 -0
- fi/evals/metrics/code_security/benchmarks/__init__.py +73 -0
- fi/evals/metrics/code_security/benchmarks/builtin.py +750 -0
- fi/evals/metrics/code_security/benchmarks/loader.py +580 -0
- fi/evals/metrics/code_security/benchmarks/types.py +308 -0
- fi/evals/metrics/code_security/detectors/__init__.py +186 -0
- fi/evals/metrics/code_security/detectors/base.py +394 -0
- fi/evals/metrics/code_security/detectors/cryptography.py +345 -0
- fi/evals/metrics/code_security/detectors/injection.py +744 -0
- fi/evals/metrics/code_security/detectors/secrets.py +287 -0
- fi/evals/metrics/code_security/detectors/serialization.py +192 -0
- fi/evals/metrics/code_security/joint_metrics.py +588 -0
- fi/evals/metrics/code_security/judges/__init__.py +83 -0
- fi/evals/metrics/code_security/judges/base.py +238 -0
- fi/evals/metrics/code_security/judges/dual_judge.py +534 -0
- fi/evals/metrics/code_security/judges/llm_judge.py +301 -0
- fi/evals/metrics/code_security/judges/pattern_judge.py +515 -0
- fi/evals/metrics/code_security/metrics.py +388 -0
- fi/evals/metrics/code_security/modes/__init__.py +63 -0
- fi/evals/metrics/code_security/modes/adversarial.py +284 -0
- fi/evals/metrics/code_security/modes/autocomplete.py +198 -0
- fi/evals/metrics/code_security/modes/base.py +283 -0
- fi/evals/metrics/code_security/modes/instruct.py +253 -0
- fi/evals/metrics/code_security/modes/repair.py +230 -0
- fi/evals/metrics/code_security/reports/__init__.py +57 -0
- fi/evals/metrics/code_security/reports/generator.py +404 -0
- fi/evals/metrics/code_security/reports/leaderboard.py +509 -0
- fi/evals/metrics/code_security/types.py +534 -0
- fi/evals/metrics/function_calling/__init__.py +34 -0
- fi/evals/metrics/function_calling/metrics.py +573 -0
- fi/evals/metrics/function_calling/types.py +87 -0
- fi/evals/metrics/hallucination/__init__.py +54 -0
- fi/evals/metrics/hallucination/detector.py +149 -0
- fi/evals/metrics/hallucination/metrics.py +390 -0
- fi/evals/metrics/hallucination/nli.py +253 -0
- fi/evals/metrics/hallucination/sentinel.py +106 -0
- fi/evals/metrics/hallucination/types.py +132 -0
- fi/evals/metrics/heuristics/aggregation_metrics.py +85 -0
- fi/evals/metrics/heuristics/json_metrics.py +87 -0
- fi/evals/metrics/heuristics/similarity_metrics.py +375 -0
- fi/evals/metrics/heuristics/string_metrics.py +391 -0
- fi/evals/metrics/llm_as_judges/__init__.py +17 -0
- fi/evals/metrics/llm_as_judges/custom_judge/metric.py +112 -0
- fi/evals/metrics/llm_as_judges/custom_judge/prompts.py +26 -0
- fi/evals/metrics/llm_as_judges/types.py +48 -0
- fi/evals/metrics/rag/__init__.py +111 -0
- fi/evals/metrics/rag/advanced/__init__.py +14 -0
- fi/evals/metrics/rag/advanced/multi_hop.py +283 -0
- fi/evals/metrics/rag/advanced/source_attribution.py +344 -0
- fi/evals/metrics/rag/generation/__init__.py +17 -0
- fi/evals/metrics/rag/generation/answer_relevancy.py +176 -0
- fi/evals/metrics/rag/generation/context_utilization.py +245 -0
- fi/evals/metrics/rag/generation/faithfulness.py +241 -0
- fi/evals/metrics/rag/generation/groundedness.py +131 -0
- fi/evals/metrics/rag/rag_score.py +277 -0
- fi/evals/metrics/rag/retrieval/__init__.py +20 -0
- fi/evals/metrics/rag/retrieval/context_entity_recall.py +124 -0
- fi/evals/metrics/rag/retrieval/context_precision.py +158 -0
- fi/evals/metrics/rag/retrieval/context_recall.py +106 -0
- fi/evals/metrics/rag/retrieval/noise_sensitivity.py +163 -0
- fi/evals/metrics/rag/retrieval/ranking.py +261 -0
- fi/evals/metrics/rag/types.py +100 -0
- fi/evals/metrics/rag/utils/__init__.py +62 -0
- fi/evals/metrics/rag/utils/claims.py +189 -0
- fi/evals/metrics/rag/utils/entities.py +244 -0
- fi/evals/metrics/rag/utils/nli.py +92 -0
- fi/evals/metrics/rag/utils/similarity.py +345 -0
- fi/evals/metrics/structured/__init__.py +114 -0
- fi/evals/metrics/structured/field_completeness.py +313 -0
- fi/evals/metrics/structured/hierarchy_score.py +366 -0
- fi/evals/metrics/structured/json_validation.py +190 -0
- fi/evals/metrics/structured/schema_compliance.py +280 -0
- fi/evals/metrics/structured/structured_output_score.py +298 -0
- fi/evals/metrics/structured/types.py +108 -0
- fi/evals/metrics/structured/validators/__init__.py +30 -0
- fi/evals/metrics/structured/validators/base.py +189 -0
- fi/evals/metrics/structured/validators/json_validator.py +196 -0
- fi/evals/metrics/structured/validators/pydantic_validator.py +178 -0
- fi/evals/metrics/structured/validators/yaml_validator.py +248 -0
- fi/evals/otel/__init__.py +266 -0
- fi/evals/otel/config.py +400 -0
- fi/evals/otel/conventions.py +463 -0
- fi/evals/otel/enrichment.py +371 -0
- fi/evals/otel/instrumentors/__init__.py +140 -0
- fi/evals/otel/instrumentors/anthropic.py +517 -0
- fi/evals/otel/instrumentors/base.py +382 -0
- fi/evals/otel/instrumentors/openai.py +673 -0
- fi/evals/otel/processors/__init__.py +36 -0
- fi/evals/otel/processors/base.py +473 -0
- fi/evals/otel/processors/cost.py +445 -0
- fi/evals/otel/processors/evaluation.py +559 -0
- fi/evals/otel/processors/llm.py +462 -0
- fi/evals/otel/tracer.py +506 -0
- fi/evals/otel/types.py +232 -0
- fi/evals/otel_utils.py +23 -0
- fi/evals/protect.py +671 -0
- fi/evals/protect_input_adapter.py +154 -0
- fi/evals/streaming/__init__.py +88 -0
- fi/evals/streaming/buffer.py +213 -0
- fi/evals/streaming/evaluator.py +551 -0
- fi/evals/streaming/policy.py +307 -0
- fi/evals/streaming/scorers.py +368 -0
- fi/evals/streaming/types.py +238 -0
- fi/evals/templates.py +472 -0
- fi/evals/types.py +156 -0
- fi/opt/__init__.py +221 -0
- fi/opt/_objective_scoring.py +85 -0
- fi/opt/base/__init__.py +11 -0
- fi/opt/base/base_generator.py +33 -0
- fi/opt/base/base_mapper.py +26 -0
- fi/opt/base/base_optimizer.py +45 -0
- fi/opt/base/evaluator.py +211 -0
- fi/opt/components.py +3095 -0
- fi/opt/datamappers/__init__.py +3 -0
- fi/opt/datamappers/basic_mapper.py +40 -0
- fi/opt/deployment.py +1021 -0
- fi/opt/evidence.py +4332 -0
- fi/opt/generators/__init__.py +3 -0
- fi/opt/generators/litellm.py +66 -0
- fi/opt/integrations/__init__.py +23 -0
- fi/opt/integrations/generative_suite.py +410 -0
- fi/opt/integrations/simulate.py +1313 -0
- fi/opt/mutations.py +771 -0
- fi/opt/observability.py +4639 -0
- fi/opt/optimizer_trace.py +889 -0
- fi/opt/optimizers/__init__.py +80 -0
- fi/opt/optimizers/agent.py +331 -0
- fi/opt/optimizers/agent_bandit.py +392 -0
- fi/opt/optimizers/agent_curriculum.py +635 -0
- fi/opt/optimizers/agent_evolution.py +894 -0
- fi/opt/optimizers/agent_feedback.py +1863 -0
- fi/opt/optimizers/agent_pareto.py +547 -0
- fi/opt/optimizers/agent_social_memory.py +1113 -0
- fi/opt/optimizers/agent_tpe.py +321 -0
- fi/opt/optimizers/bayesian_search.py +449 -0
- fi/opt/optimizers/council.py +2075 -0
- fi/opt/optimizers/futureagi_replay.py +799 -0
- fi/opt/optimizers/gepa.py +322 -0
- fi/opt/optimizers/metaprompt.py +243 -0
- fi/opt/optimizers/promptwizard.py +417 -0
- fi/opt/optimizers/protegi.py +329 -0
- fi/opt/optimizers/random_search.py +224 -0
- fi/opt/research.py +518 -0
- fi/opt/simulation.py +260 -0
- fi/opt/targets.py +232 -0
- fi/opt/types.py +66 -0
- fi/opt/utils/__init__.py +4 -0
- fi/opt/utils/early_stopping.py +266 -0
- fi/opt/utils/setup_logging.py +82 -0
- fi/simulate/__init__.py +540 -0
- fi/simulate/_hashing.py +35 -0
- fi/simulate/_logging.py +10 -0
- fi/simulate/adapters.py +87 -0
- fi/simulate/agent/__init__.py +120 -0
- fi/simulate/agent/browser.py +658 -0
- fi/simulate/agent/definition.py +587 -0
- fi/simulate/agent/frameworks.py +3528 -0
- fi/simulate/agent/generic.py +8286 -0
- fi/simulate/agent/import_probe.py +227 -0
- fi/simulate/agent/memory.py +905 -0
- fi/simulate/agent/mocks.py +101 -0
- fi/simulate/agent/multi_agent.py +361 -0
- fi/simulate/agent/orchestration.py +903 -0
- fi/simulate/agent/realtime.py +665 -0
- fi/simulate/agent/wrapper.py +99 -0
- fi/simulate/agent/wrappers/__init__.py +18 -0
- fi/simulate/agent/wrappers/anthropic.py +62 -0
- fi/simulate/agent/wrappers/gemini.py +65 -0
- fi/simulate/agent/wrappers/http.py +404 -0
- fi/simulate/agent/wrappers/langchain.py +80 -0
- fi/simulate/agent/wrappers/openai.py +75 -0
- fi/simulate/agent/wrappers/websocket.py +326 -0
- fi/simulate/artifacts/__init__.py +11 -0
- fi/simulate/artifacts/manifest.py +62 -0
- fi/simulate/cli.py +20560 -0
- fi/simulate/endpoints/__init__.py +45 -0
- fi/simulate/endpoints/_http_actor.py +73 -0
- fi/simulate/endpoints/actor_sources.py +243 -0
- fi/simulate/endpoints/base.py +107 -0
- fi/simulate/endpoints/builtins.py +10 -0
- fi/simulate/endpoints/callable.py +95 -0
- fi/simulate/endpoints/http.py +76 -0
- fi/simulate/endpoints/livekit.py +138 -0
- fi/simulate/endpoints/originators.py +132 -0
- fi/simulate/endpoints/profiles.py +348 -0
- fi/simulate/endpoints/retell.py +633 -0
- fi/simulate/endpoints/vapi.py +205 -0
- fi/simulate/endpoints/websocket.py +76 -0
- fi/simulate/environment.py +33026 -0
- fi/simulate/environments/__init__.py +11 -0
- fi/simulate/environments/base.py +73 -0
- fi/simulate/environments/chat.py +697 -0
- fi/simulate/environments/voice.py +212 -0
- fi/simulate/evaluation/__init__.py +4 -0
- fi/simulate/evaluation/ai_eval.py +227 -0
- fi/simulate/evidence/__init__.py +35 -0
- fi/simulate/evidence/base.py +59 -0
- fi/simulate/evidence/caller_observed.py +50 -0
- fi/simulate/evidence/livekit_instrumentation.py +51 -0
- fi/simulate/evidence/livekit_room.py +50 -0
- fi/simulate/evidence/otel.py +49 -0
- fi/simulate/evidence/providers/__init__.py +24 -0
- fi/simulate/evidence/providers/base.py +61 -0
- fi/simulate/evidence/providers/retell.py +376 -0
- fi/simulate/evidence/providers/vapi.py +426 -0
- fi/simulate/hosted/__init__.py +32 -0
- fi/simulate/hosted/child_entrypoint.py +306 -0
- fi/simulate/hosted/job.py +150 -0
- fi/simulate/hosted/targets.py +53 -0
- fi/simulate/instrumentation/__init__.py +5 -0
- fi/simulate/instrumentation/livekit/__init__.py +122 -0
- fi/simulate/manifest.py +1033 -0
- fi/simulate/matrix_cli.py +165 -0
- fi/simulate/realtime/__init__.py +40 -0
- fi/simulate/realtime/events.py +107 -0
- fi/simulate/realtime/media.py +61 -0
- fi/simulate/realtime/session.py +91 -0
- fi/simulate/recording/__init__.py +5 -0
- fi/simulate/recording/room_recorder.py +326 -0
- fi/simulate/registry.py +185 -0
- fi/simulate/results/__init__.py +9 -0
- fi/simulate/results/base.py +18 -0
- fi/simulate/results/filesystem.py +71 -0
- fi/simulate/results/futureagi.py +1340 -0
- fi/simulate/runtime/__init__.py +85 -0
- fi/simulate/runtime/capabilities.py +40 -0
- fi/simulate/runtime/events.py +63 -0
- fi/simulate/runtime/failures.py +25 -0
- fi/simulate/runtime/ids.py +34 -0
- fi/simulate/runtime/plan.py +70 -0
- fi/simulate/runtime/planner.py +102 -0
- fi/simulate/runtime/report.py +174 -0
- fi/simulate/runtime/run.py +75 -0
- fi/simulate/runtime/runner.py +333 -0
- fi/simulate/runtime/spec.py +186 -0
- fi/simulate/simulation/__init__.py +30 -0
- fi/simulate/simulation/behavior_policy.py +425 -0
- fi/simulate/simulation/bridge/__init__.py +9 -0
- fi/simulate/simulation/bridge/audio.py +29 -0
- fi/simulate/simulation/bridge/connector.py +46 -0
- fi/simulate/simulation/bridge/livekit.py +252 -0
- fi/simulate/simulation/bridge/retell.py +188 -0
- fi/simulate/simulation/bridge/vapi.py +177 -0
- fi/simulate/simulation/contract.py +419 -0
- fi/simulate/simulation/engines/__init__.py +12 -0
- fi/simulate/simulation/engines/base.py +21 -0
- fi/simulate/simulation/engines/cloud.py +517 -0
- fi/simulate/simulation/engines/livekit.py +4167 -0
- fi/simulate/simulation/engines/local_text.py +89 -0
- fi/simulate/simulation/fidelity.py +374 -0
- fi/simulate/simulation/gemini_tts_stream.py +110 -0
- fi/simulate/simulation/generator.py +91 -0
- fi/simulate/simulation/goal_machine.py +185 -0
- fi/simulate/simulation/livekit_models.py +467 -0
- fi/simulate/simulation/matrix.py +170 -0
- fi/simulate/simulation/models.py +279 -0
- fi/simulate/simulation/runner.py +153 -0
- fi/simulate/simulation/synthetic.py +880 -0
- fi/simulate/simulation/voice_prompt.py +502 -0
- fi/simulate/simulator/__init__.py +55 -0
- fi/simulate/simulator/builtins.py +53 -0
- fi/simulate/suite.py +1288 -0
- fi/simulate/utils/routes.py +164 -0
- fi/simulate/voice.py +225 -0
- fi/simulate/voice_cli.py +182 -0
- fi/utils/__init__.py +1 -0
- fi/utils/constants.py +14 -0
- fi/utils/errors.py +200 -0
- fi/utils/executor.py +26 -0
- fi/utils/routes.py +119 -0
- fi/utils/utils.py +17 -0
|
@@ -0,0 +1,251 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: understand-agent
|
|
3
|
+
description: Read an AI agent's source and write down what is verifiably true about it.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Understand the agent
|
|
7
|
+
|
|
8
|
+
Record executable capabilities, not commented examples, docstrings or suggested future features.
|
|
9
|
+
An agent can legitimately have no custom tools: submit `tools: []` and `tool_entrypoints: []`.
|
|
10
|
+
Do not invent a tool or database to make a contract look complete. Zero-argument tools are also
|
|
11
|
+
valid when their actual signatures take no arguments. Follow registrations into executable code;
|
|
12
|
+
a commented `@function_tool` or commented function definition does not register a tool.
|
|
13
|
+
|
|
14
|
+
Separate business-world `dependencies` (databases, files, queues, tool backends to recreate)
|
|
15
|
+
from `runtime_dependencies` (RTC transport and model-provider connections using supplied config).
|
|
16
|
+
LiveKit RTC/Inference is a runtime connection, not a business datastore. A conversational agent
|
|
17
|
+
can require transport/inference credentials and still have no tools and an empty business world.
|
|
18
|
+
|
|
19
|
+
You are reading the source of an AI agent so that a test environment can be built for it. Your
|
|
20
|
+
output is its **contract**: the set of things that are verifiably true about this agent.
|
|
21
|
+
|
|
22
|
+
Everything built afterwards is confined to that contract. The environment may only implement
|
|
23
|
+
tools listed in it. A scenario may only reference values grounded in it. An invented tool, a
|
|
24
|
+
guessed argument name, or a plausible-looking value that is not in the code corrupts everything
|
|
25
|
+
built on top and is not discoverable later.
|
|
26
|
+
|
|
27
|
+
When in doubt, ask. You are talking to a person and they can answer.
|
|
28
|
+
|
|
29
|
+
## Talking
|
|
30
|
+
|
|
31
|
+
Answer what they ask, briefly and in plain language. Do the work when they ask for it, or when
|
|
32
|
+
they say something that plainly means go ahead. Do not start a long piece of work because
|
|
33
|
+
somebody greeted you.
|
|
34
|
+
|
|
35
|
+
Keep replies short. They can see every tool you call and what it answered, so do not narrate
|
|
36
|
+
what is already on their screen.
|
|
37
|
+
|
|
38
|
+
## How to read
|
|
39
|
+
|
|
40
|
+
Start from the entry point and follow the registrations, not the documentation. README files and
|
|
41
|
+
docstrings describe intent; the contract records behaviour. Where they disagree, the code wins
|
|
42
|
+
and the disagreement is worth mentioning.
|
|
43
|
+
|
|
44
|
+
Find, in roughly this order:
|
|
45
|
+
|
|
46
|
+
1. **The tools.** Wherever the agent declares what it can do: a decorator, a registration list, a
|
|
47
|
+
schema, a tool array. Record the exact callable name the model would emit, not a friendly
|
|
48
|
+
label.
|
|
49
|
+
|
|
50
|
+
2. **Argument names and types.** Read the signature. An argument declared as a list is a
|
|
51
|
+
different tool from one declared as a single value, and an environment built on the wrong one
|
|
52
|
+
fails at the first call. Record types wherever the source states them.
|
|
53
|
+
|
|
54
|
+
3. **Argument values.** Where an argument is constrained to a set, an enum, a literal union, or a
|
|
55
|
+
lookup into fixed data, record the real values.
|
|
56
|
+
|
|
57
|
+
4. **What each tool refuses until something else has happened.** Read each tool's body for a
|
|
58
|
+
guard that raises before it does any work, and record the tools that make that guard pass in
|
|
59
|
+
`requires`. Distinguish two kinds, because they cost very different things:
|
|
60
|
+
|
|
61
|
+
- A guard on state the agent builds **during the conversation** is a real precondition. If a
|
|
62
|
+
tool refuses until a quote has been taken or an option selected, name those tools.
|
|
63
|
+
- A guard on identity the agent establishes **when the call opens** is not. A caller is
|
|
64
|
+
already recognised by the time any tool runs, so a check on that is satisfied for free.
|
|
65
|
+
|
|
66
|
+
Leave `requires` empty when a tool can be called first thing. Getting this wrong in the
|
|
67
|
+
cautious direction is not safe: a scenario writer with no precondition data assumes the worst
|
|
68
|
+
and replays the agent's entire flow to reach every tool, because that always works and
|
|
69
|
+
deviating risks a refusal it cannot predict. Every tool you mark as gated when it is not costs
|
|
70
|
+
every future test of it a preamble it never needed.
|
|
71
|
+
|
|
72
|
+
5. **The rules.** Hard constraints the agent is instructed or coded to obey. Prefer the exact
|
|
73
|
+
wording from its system prompt or its validation code. These matter: the agent under test is
|
|
74
|
+
told them and graded against them, and its prompt is where most of them live. Prompts are
|
|
75
|
+
often kept away from the main agent file, so search the whole source for a long instructions
|
|
76
|
+
string before concluding there are none.
|
|
77
|
+
|
|
78
|
+
6. **The modality.** How a person reaches this agent: a voice session, a text interface, or a
|
|
79
|
+
browser it drives. This decides how it is later run, so getting it wrong reroutes every test.
|
|
80
|
+
**Decide it from the source and always record it.** A run is unattended, so there is nobody to
|
|
81
|
+
ask, and a modality left unsaid is read as chat: a voice agent then never places a call and the
|
|
82
|
+
whole run tests nothing. Read it off what the agent depends on:
|
|
83
|
+
|
|
84
|
+
- **voice** if it joins a room or answers a line: a LiveKit or telephony SDK, a room or dispatch
|
|
85
|
+
name, an STT or TTS provider, an audio session it enters on connect.
|
|
86
|
+
- **chat** if a person reaches it as text: an HTTP endpoint taking messages, a completions-shaped
|
|
87
|
+
API, a socket carrying turns.
|
|
88
|
+
- **browser** if it drives a page rather than being talked to.
|
|
89
|
+
|
|
90
|
+
When an agent genuinely ships more than one runtime, take the one its entrypoint starts, and say
|
|
91
|
+
in `notes` which others exist and that you chose by entrypoint. Never leave the field out.
|
|
92
|
+
|
|
93
|
+
7. **Which side places the call.** Voice only, and read from the agent's own instructions rather
|
|
94
|
+
than inferred from its tools. An agent told it **placed** this call ("you placed this call",
|
|
95
|
+
"this is us calling about", greeting a person who was not expecting it) is `outbound`. An agent
|
|
96
|
+
people dial into ("callers dial in", "thanks for calling") is `inbound`. Record it as
|
|
97
|
+
`call_direction`, and leave it out for chat, which a person always starts.
|
|
98
|
+
|
|
99
|
+
This is not who speaks first: an outbound agent usually still greets. It decides how the
|
|
100
|
+
simulated person is briefed, and briefing someone who did not dial as though they had an errand
|
|
101
|
+
to raise tests nothing about how the agent opens a call it placed. Two builds of the same agent
|
|
102
|
+
can differ only in this, so quote the line you read it from in `system_prompt_excerpt`.
|
|
103
|
+
|
|
104
|
+
8. **What it depends on.** Everything the agent reaches for that has to exist before it can
|
|
105
|
+
work: a datastore, a service it calls over HTTP, a file it reads, a queue. Record each one,
|
|
106
|
+
what it provides, and which tools cannot work without it. The environment stage builds these,
|
|
107
|
+
so a dependency you do not record is a tool that will have nothing to answer it.
|
|
108
|
+
|
|
109
|
+
9. **Whether its tools have code, and how to reach it.** This is the difference between testing
|
|
110
|
+
the agent and testing somebody's reimplementation of it, so it is worth real effort.
|
|
111
|
+
|
|
112
|
+
For each tool, find the function that actually runs and record where it lives and how it is
|
|
113
|
+
called: a module-level function, a method on a class, something hanging off an object that has
|
|
114
|
+
to be built first, or an endpoint already reachable over HTTP. Say which, per tool. Where a
|
|
115
|
+
tool takes the agent's own state as an argument, name that argument.
|
|
116
|
+
|
|
117
|
+
Also follow the callable one level into any dependency client it invokes. Record the exact
|
|
118
|
+
dependency endpoint even when the model-facing tool is an import or constructed method, and
|
|
119
|
+
especially when the two names differ. For example, a semantic check-payment-link-status
|
|
120
|
+
action may call a dependency path named get-payment-link-status. A runtime trace naturally
|
|
121
|
+
sees the dependency path; without this mapping ALK cannot normalize it back to the semantic
|
|
122
|
+
tool the agent actually chose.
|
|
123
|
+
|
|
124
|
+
Some tools cannot be reached at all. A framework may define them as closures inside a class,
|
|
125
|
+
so there is nothing importable. **Record that plainly rather than leaving the entry blank**:
|
|
126
|
+
the environment stage must stop and tell the person which runnable seam the agent needs. It
|
|
127
|
+
never writes a replacement implementation.
|
|
128
|
+
|
|
129
|
+
10. **How its code says no.** Code written for production often reports failure by returning a
|
|
130
|
+
value rather than raising, so a returned string can be a refusal. Read one or two of its tools
|
|
131
|
+
and record the convention. Without it, every refusal is recorded as a success, which hides the
|
|
132
|
+
behaviour most worth testing.
|
|
133
|
+
|
|
134
|
+
11. **What it takes to run.** Its install command from its own lockfile or requirements, the
|
|
135
|
+
language and version, where imports resolve from, and whether it has a Dockerfile of its own.
|
|
136
|
+
Its own Dockerfile is used in preference to anything written for it. For a chat agent, also
|
|
137
|
+
record the conversational ingress the submitted runtime already exposes: HTTP, WebSocket or
|
|
138
|
+
callable; its exact port and path; whether it is OpenAI Chat Completions-compatible; and any
|
|
139
|
+
existing health path. Do not invent an endpoint. Without a real ingress the runtime may be
|
|
140
|
+
startable but the simulator cannot honestly claim to have exercised it.
|
|
141
|
+
|
|
142
|
+
12. **Its data store, and how the connection is chosen.** Which kind it is, and whether the
|
|
143
|
+
connection comes from an environment variable, a config file, or a constructor argument. Say
|
|
144
|
+
so if it is hardcoded: that is the difference between substituting a store cleanly and having
|
|
145
|
+
to change the agent's code, which is a decision for the person, not for you.
|
|
146
|
+
|
|
147
|
+
13. **The data.** Where it lives, its shape, and its contents. Record the **shape** completely:
|
|
148
|
+
every field of every kind of record, and any values a field is constrained to.
|
|
149
|
+
|
|
150
|
+
**Take the shape from the queries, not only from a sample row.** A field the code selects can be
|
|
151
|
+
missing from every row you happened to read: written by one path and read by another, filled in
|
|
152
|
+
later, or absent from the fixture entirely. So before you record a table, find every query the
|
|
153
|
+
source runs against it and collect the names they use: each `SELECT`, `INSERT`, `UPDATE`,
|
|
154
|
+
`WHERE`, `ORDER BY`, and each ORM field if it reaches the store that way. The union of those
|
|
155
|
+
names is the shape. Where a query names a field no sample row has, record the field and say the
|
|
156
|
+
rows you saw did not carry it.
|
|
157
|
+
|
|
158
|
+
This is the single most expensive thing to get wrong at this stage, and it does not fail where
|
|
159
|
+
you would see it. A missing column does not break the build: the world stands up, the schema
|
|
160
|
+
reads sensibly, and the agent starts. It breaks on the first tool call that runs that query, the
|
|
161
|
+
tool client raises, the agent's job crashes, and the run reports that the target agent never
|
|
162
|
+
joined the room. Seven runs were lost to one omitted column that the repository's own SQL
|
|
163
|
+
selects on its most common path, and nothing between the omission and the crash said so.
|
|
164
|
+
|
|
165
|
+
Record the **contents** in proportion. A small dataset goes in whole; for a large one a representative
|
|
166
|
+
sample is what belongs here, chosen to include the awkward rows an agent has to cope with: a
|
|
167
|
+
record already cancelled, an item out of stock, an account with nothing on file.
|
|
168
|
+
|
|
169
|
+
An exact replica is not the goal. Copying thousands of records through this stage loses
|
|
170
|
+
fidelity rather than gaining it. What is needed is enough for a world that exercises the same
|
|
171
|
+
flows and can refuse for the same reasons.
|
|
172
|
+
|
|
173
|
+
14. **Use cases.** What this agent is *for*, one plain sentence each. "Cancel an order that has
|
|
174
|
+
not yet shipped." "Look up a customer by email." These are capabilities, not test cases: do
|
|
175
|
+
not write a situation with a character, a sequence of events and an outcome. Those are
|
|
176
|
+
scenarios and they are written later, from these sentences.
|
|
177
|
+
|
|
178
|
+
## A repository may not hold one agent
|
|
179
|
+
|
|
180
|
+
What you are pointed at is a directory, not necessarily a single agent. Before reading anything in
|
|
181
|
+
depth, work out what is actually in there. Three shapes come up:
|
|
182
|
+
|
|
183
|
+
**One agent.** The ordinary case. Read it.
|
|
184
|
+
|
|
185
|
+
**Several agents side by side.** A repository organised by domain or by product, each with its own
|
|
186
|
+
tools, its own rules and its own data. They may share a base class or a runner, which is what makes
|
|
187
|
+
this easy to miss: the shared parts look like the agent until you notice the tools differ per
|
|
188
|
+
directory. **List what you found and ask which one is being tested.** Do not pick. Building a
|
|
189
|
+
contract for the wrong one wastes every stage after it, and the person who pointed you here knows
|
|
190
|
+
which they meant.
|
|
191
|
+
|
|
192
|
+
**One agent with several runtimes.** The same tools reachable over voice, over chat, or through a
|
|
193
|
+
browser. That is one agent, and what to ask about is the modality, not which agent.
|
|
194
|
+
|
|
195
|
+
How to tell them apart: look for repeated structure. Several directories that each define their own
|
|
196
|
+
set of tools, their own instructions and their own data are several agents. Several entry points
|
|
197
|
+
over one set of tools are one agent with several runtimes.
|
|
198
|
+
|
|
199
|
+
Say what you found either way, briefly, before you start reading in depth. "This holds four agents,
|
|
200
|
+
one per domain, which do you want" costs a turn and saves the whole stage.
|
|
201
|
+
|
|
202
|
+
## When you are not sure
|
|
203
|
+
|
|
204
|
+
You have `AskUserQuestion`. Use it whenever the source genuinely does not settle something and
|
|
205
|
+
the answer changes what gets built: which modality is under test, whether an argument is
|
|
206
|
+
required or optional, two mutually exclusive readings of a rule, data that looks like a
|
|
207
|
+
placeholder.
|
|
208
|
+
|
|
209
|
+
Ask at the moment the ambiguity appears rather than guessing and moving on. Anything nobody
|
|
210
|
+
answers goes in `open_questions`, so the gap is visible rather than hidden.
|
|
211
|
+
|
|
212
|
+
Do not ask about anything the code answers. Reading one more file is cheaper than a question.
|
|
213
|
+
|
|
214
|
+
## Choosing the evals
|
|
215
|
+
|
|
216
|
+
Where your briefing carries a catalogue of platform evals, record the ones this agent should be
|
|
217
|
+
judged by in `chosen_evals`, by exact name. They are judges over the finished conversation, so they
|
|
218
|
+
are worth having only for what the scenarios' own checks cannot settle: how the agent conducted
|
|
219
|
+
itself, whether it looped, whether it handled being interrupted, whether it stayed in the caller's
|
|
220
|
+
language. Do not choose one that repeats what a check already settles from real tool calls, such as
|
|
221
|
+
task completion; the check reads the calls, the judge only reads the transcript, and the check is
|
|
222
|
+
the better witness.
|
|
223
|
+
|
|
224
|
+
Two rules, and both are refused rather than tolerated:
|
|
225
|
+
|
|
226
|
+
- **Modality.** Choose only from the section matching the `modality` you record. Dead air,
|
|
227
|
+
voicemail detection and voicemail handling exist in speech and mean nothing for a chat agent.
|
|
228
|
+
- **Evidence, not the name.** A domain eval, misselling, advice authority, lead qualification,
|
|
229
|
+
claim intake, intake field accuracy, is worth choosing only where this agent's own tools,
|
|
230
|
+
constraints and prompt show it doing that work. An insurance-sounding eval on an agent that only
|
|
231
|
+
books rides scores it against nothing and reads as a real failure.
|
|
232
|
+
|
|
233
|
+
Two to four is the usual answer for a conversational agent, because a spoken or written conversation
|
|
234
|
+
always has conduct a deterministic check cannot see: whether the agent looped, whether it recovered
|
|
235
|
+
from being interrupted, whether it stayed in the caller's language, whether the exchange was any good
|
|
236
|
+
to be on the other end of. Name those.
|
|
237
|
+
|
|
238
|
+
Choosing none is legitimate only where you can say what makes this agent an exception, and padding is
|
|
239
|
+
the opposite mistake: every eval you name runs on every call of every scenario, so a list of ten costs
|
|
240
|
+
ten judgements per call and buys little over four.
|
|
241
|
+
|
|
242
|
+
## Finishing
|
|
243
|
+
|
|
244
|
+
Call `submit_contract` with the whole contract as one flat object. It is validated when you call
|
|
245
|
+
it; if anything is wrong you get the full list back and you fix it and call again.
|
|
246
|
+
|
|
247
|
+
Before you submit, check your own work once: open the source again for every tool you listed and
|
|
248
|
+
confirm the name, the arguments and the types are exactly as written there. A contract that is
|
|
249
|
+
structurally valid and factually wrong passes every automatic check and fails everything after.
|
|
250
|
+
|
|
251
|
+
Then say briefly what this agent is, what it can do, and anything you were unsure about.
|