agentv 5.3.0-next.1 → 5.3.2-next.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (75) hide show
  1. package/README.md +70 -60
  2. package/dist/{artifact-writer-JFNIPMKW.js → artifact-writer-KJEOROKQ.js} +5 -5
  3. package/dist/{chunk-6ZZCDZPD.js → chunk-6262OKXM.js} +21774 -21034
  4. package/dist/chunk-6262OKXM.js.map +1 -0
  5. package/dist/chunk-AQ5BIAXF.js +604 -0
  6. package/dist/chunk-AQ5BIAXF.js.map +1 -0
  7. package/dist/chunk-BV5VQLI2.js +2 -0
  8. package/dist/{chunk-V52ATPTT.js → chunk-JGSRUJZQ.js} +186 -32
  9. package/dist/chunk-JGSRUJZQ.js.map +1 -0
  10. package/dist/chunk-MMDLXYBX.js +721 -0
  11. package/dist/chunk-MMDLXYBX.js.map +1 -0
  12. package/dist/{chunk-FVR4RQFK.js → chunk-QSK3PC44.js} +1131 -580
  13. package/dist/chunk-QSK3PC44.js.map +1 -0
  14. package/dist/chunk-TEEXVJWM.js +2 -0
  15. package/dist/{chunk-BHKQHG26.js → chunk-WIAJ7JAV.js} +186 -64
  16. package/dist/chunk-WIAJ7JAV.js.map +1 -0
  17. package/dist/cli.d.ts +1 -0
  18. package/dist/cli.js +17517 -10
  19. package/dist/cli.js.map +1 -1
  20. package/dist/config.d.ts +2 -0
  21. package/dist/{ts-eval-loader-DQDYRULE-V3377FOL.js → config.js} +7 -7
  22. package/dist/config.js.map +1 -0
  23. package/dist/contracts-DsZmZLl8.d.ts +742 -0
  24. package/dist/contracts.d.ts +2 -0
  25. package/dist/contracts.js +28 -0
  26. package/dist/contracts.js.map +1 -0
  27. package/dist/dashboard/assets/{index-DNgf3qJ2.js → index-BfqOlLOF.js} +1 -1
  28. package/dist/dashboard/assets/index-Bh8lpGce.css +1 -0
  29. package/dist/dashboard/assets/index-oPR3ywQb.js +121 -0
  30. package/dist/dashboard/index.html +2 -2
  31. package/dist/{dist-6Z7U473R.js → dist-A3SGR7TW.js} +30 -18
  32. package/dist/dist-A3SGR7TW.js.map +1 -0
  33. package/dist/index.d.ts +4 -0
  34. package/dist/index.js +137 -18
  35. package/dist/{interactive-RUY3OCBI.js → interactive-DL7N2C7K.js} +24 -24
  36. package/dist/interactive-DL7N2C7K.js.map +1 -0
  37. package/dist/provider.d.ts +2 -0
  38. package/dist/provider.js +24 -0
  39. package/dist/provider.js.map +1 -0
  40. package/dist/sdk.d.ts +802 -0
  41. package/dist/sdk.js +145 -0
  42. package/dist/sdk.js.map +1 -0
  43. package/dist/skills/agentv-bench/SKILL.md +14 -13
  44. package/dist/skills/agentv-bench/agents/analyzer.md +1 -1
  45. package/dist/skills/agentv-bench/agents/executor.md +1 -1
  46. package/dist/skills/agentv-bench/references/autoresearch.md +9 -9
  47. package/dist/skills/agentv-bench/references/environment-adaptation.md +4 -4
  48. package/dist/skills/agentv-bench/references/eval-yaml-spec.md +30 -47
  49. package/dist/skills/agentv-bench/references/schemas.md +44 -60
  50. package/dist/skills/agentv-bench/references/subagent-pipeline.md +20 -18
  51. package/dist/skills/agentv-eval-migrations/SKILL.md +13 -0
  52. package/dist/skills/agentv-eval-migrations/references/breaking-changes.md +56 -39
  53. package/dist/skills/agentv-eval-writer/SKILL.md +101 -60
  54. package/dist/skills/agentv-eval-writer/references/custom-evaluators.md +15 -10
  55. package/dist/skills/agentv-eval-writer/references/eval.schema.json +4543 -5189
  56. package/dist/skills/agentv-eval-writer/references/python-helpers.md +2 -2
  57. package/dist/skills/agentv-eval-writer/references/rubric-evaluator.md +20 -3
  58. package/dist/templates/.agentv/providers.yaml +42 -0
  59. package/dist/templates/.env.example +2 -2
  60. package/dist/ts-eval-loader-3G5GEAEC-6F52JLPP.js +18 -0
  61. package/dist/ts-eval-loader-3G5GEAEC-6F52JLPP.js.map +1 -0
  62. package/package.json +29 -4
  63. package/dist/chunk-6ZZCDZPD.js.map +0 -1
  64. package/dist/chunk-BHKQHG26.js.map +0 -1
  65. package/dist/chunk-FVR4RQFK.js.map +0 -1
  66. package/dist/chunk-T32NL3E6.js +0 -17973
  67. package/dist/chunk-T32NL3E6.js.map +0 -1
  68. package/dist/chunk-V52ATPTT.js.map +0 -1
  69. package/dist/dashboard/assets/index-D_bokML8.css +0 -1
  70. package/dist/dashboard/assets/index-r_jSJmlw.js +0 -121
  71. package/dist/interactive-RUY3OCBI.js.map +0 -1
  72. package/dist/templates/.agentv/targets.yaml +0 -97
  73. /package/dist/{artifact-writer-JFNIPMKW.js.map → artifact-writer-KJEOROKQ.js.map} +0 -0
  74. /package/dist/{dist-6Z7U473R.js.map → chunk-BV5VQLI2.js.map} +0 -0
  75. /package/dist/{ts-eval-loader-DQDYRULE-V3377FOL.js.map → chunk-TEEXVJWM.js.map} +0 -0
@@ -30,10 +30,12 @@ specific target, set `subagent_mode_allowed: false` in `.agentv/targets.yaml`:
30
30
  ```yaml
31
31
  # .agentv/targets.yaml
32
32
  targets:
33
- - name: my-target
33
+ - id: my-target
34
34
  provider: openai
35
- model: ${{ OPENAI_MODEL }}
36
- api_key: ${{ OPENAI_API_KEY }}
35
+ runtime: host
36
+ config:
37
+ model: "{{ env.OPENAI_MODEL }}"
38
+ api_key: "{{ env.OPENAI_API_KEY }}"
37
39
  subagent_mode_allowed: false # forces CLI invocation instead of executor subagent
38
40
  ```
39
41
 
@@ -43,8 +45,8 @@ even in subagent mode.
43
45
  ## CLI Targets: Single Command
44
46
 
45
47
  For evals with CLI targets, `pipeline run` handles input extraction, target invocation, and
46
- code grading in one step. When `--out` is omitted, the output directory defaults to
47
- `.agentv/results/default/<timestamp>` (same convention as `agentv eval`):
48
+ code grading in one step. When `--out` is omitted, the output directory uses the
49
+ canonical `.agentv/results/<run_id>` convention:
48
50
 
49
51
  ```bash
50
52
  # Extract inputs and invoke all CLI targets in parallel:
@@ -70,7 +72,7 @@ opted out via `subagent_mode_allowed: false` in `.agentv/targets.yaml`), fall ba
70
72
  ### Step 1: Extract inputs
71
73
 
72
74
  ```bash
73
- # Defaults to .agentv/results/default/<timestamp>
75
+ # Defaults to .agentv/results/<run_id>
74
76
  agentv pipeline input evals/repro.eval.yaml
75
77
  ```
76
78
 
@@ -117,7 +119,7 @@ LLM grading → merge and validate).
117
119
  Use individual commands when you need control over each step with CLI targets:
118
120
 
119
121
  ```bash
120
- # Step 1: Extract inputs (defaults to .agentv/results/default/<timestamp>)
122
+ # Step 1: Extract inputs (defaults to .agentv/results/<run_id>)
121
123
  agentv pipeline input evals/repro.eval.yaml
122
124
 
123
125
  # Step 2: run_tests.py invokes CLI targets (or use pipeline run instead)
@@ -127,7 +129,7 @@ agentv pipeline grade <run-dir>
127
129
 
128
130
  # Step 4: Subagent does LLM grading, writes results to llm_grader_results/<name>.json per test
129
131
 
130
- # Step 5: Merge scores (writes index.jsonl with full scores[] for dashboard)
132
+ # Step 5: Merge scores into the canonical run bundle
131
133
  agentv pipeline bench <run-dir>
132
134
 
133
135
  # Step 6: Validate
@@ -152,22 +154,22 @@ The agent reads `llm_graders/<name>.json` for each test, grades the response usi
152
154
 
153
155
  ## Pipeline Bench and Dashboard
154
156
 
155
- `pipeline bench` merges LLM scores into `index.jsonl` with a full `scores[]` array per entry,
156
- matching the CLI-mode schema. The web dashboard (`agentv results serve`) reads this format
157
- directly — no separate conversion script is needed. Run `agentv results validate <run-dir>`
158
- to verify compatibility.
157
+ `pipeline bench` merges LLM scores into the canonical run bundle, including
158
+ `.internal/index.jsonl`, `summary.json`, and per-attempt `grading.json` with
159
+ `component_results`. The web dashboard reads this format directly — no separate
160
+ conversion script is needed. Run `agentv results validate <run-dir>` to verify
161
+ compatibility.
159
162
 
160
163
  ## Output Structure
161
164
 
162
- The path hierarchy mirrors the CLI mode: `<evalset-name>` comes from the `name` field in
163
- the eval.yaml. The target is recorded in `manifest.json` — one run = one target.
165
+ The path hierarchy mirrors CLI mode. The target is recorded in run metadata.
164
166
 
165
167
  ```
166
- .agentv/results/<experiment>/<timestamp>/
168
+ .agentv/results/<run_id>/
167
169
  ├── manifest.json ← eval metadata, target, test_ids
168
- ├── index.jsonl ← per-test scores
169
- ├── summary.json ← aggregate statistics
170
- └── <evalset-name>/ ← eval.yaml "name" field, or eval file basename if absent (same as CLI mode)
170
+ ├── summary.json ← aggregate statistics
171
+ ├── .internal/index.jsonl ← canonical row index
172
+ └── <result-dir>/sample-1/grading.json
171
173
  └── <test-id>/ ← test case id
172
174
  ├── input.json ← test input text + messages
173
175
  ├── invoke.json ← target command or agent instructions
@@ -16,6 +16,19 @@ v4.42.4-era shape, current shape, migration steps, verification commands, and
16
16
  compatibility notes for each major breaking authoring change. Then compare the
17
17
  eval file against the current portable contract:
18
18
 
19
+ - Treat migrations as rule-based rewrites first. Do not rely on a manual
20
+ "looks current" pass when the old field has a direct current replacement.
21
+ - Rewrite authored prompt data into top-level `prompts` plus `tests[].vars`;
22
+ keep `input` only for documented raw-case import compatibility.
23
+ - Move sibling `expected_output` to `vars.expected_output`, then add or keep an
24
+ explicit assertion that consumes it, usually `llm-rubric` with
25
+ `{{ expected_output }}`.
26
+ - Rewrite target identity to `id`, keep `provider` for the backend/adapter kind,
27
+ put provider settings under `config`, and rewrite env references to
28
+ `{{ env.NAME }}`.
29
+ - Remove `use_target`, `eval_cases`, `evalcases`, `providerPromptMap`, and
30
+ `provider_prompt_map` from authored eval/config YAML. Current AgentV rejects
31
+ these fields, so do not preserve them as compatibility aliases.
19
32
  - Keep committed eval YAML portable: prompts, cases, assertions, workspace
20
33
  templates, repos, hooks, env checks, Docker preflight/container config, and
21
34
  `workspace.scope`.
@@ -23,7 +23,7 @@ For a v4.42.4-era eval:
23
23
 
24
24
  1. Move top-level `execution.target` to top-level `target`.
25
25
  2. Move top-level `execution.targets` to top-level `targets`, and rename target
26
- object `name` to `label`.
26
+ object `name` to `id`.
27
27
  3. Rename suite/test/turn `assertions` to `assert`.
28
28
  4. If a test uses `tests[].execution.assertions` or
29
29
  `tests[].execution.evaluators`, rename that nested list to
@@ -56,7 +56,14 @@ For a v4.42.4-era eval:
56
56
  assertions where there is a direct mapping.
57
57
  18. Keep raw cases under `tests` / `tests: file://...`; run full eval suites
58
58
  directly with CLI multi-file selection and tags.
59
- 19. Validate with `bun apps/cli/src/cli.ts validate <eval-file>`.
59
+ 19. Move authored prompt input from top-level or `tests[].input` into
60
+ `prompts` plus `tests[].vars.input`.
61
+ 20. Move authored reference answers from sibling `expected_output` into
62
+ `vars.expected_output`; add or keep an explicit assertion that consumes it.
63
+ 21. Rewrite target env interpolation from `${{ NAME }}` to `{{ env.NAME }}`.
64
+ 22. Remove `use_target`, `eval_cases`, `evalcases`, `providerPromptMap`, and
65
+ `provider_prompt_map`; current AgentV rejects them.
66
+ 23. Validate with `bun apps/cli/src/cli.ts validate <eval-file>`.
60
67
 
61
68
  ## Assertions Renamed To `assert`
62
69
 
@@ -67,7 +74,7 @@ v4.42.4 docs and schema used `assertions` for suite-level and per-test graders:
67
74
  ```yaml
68
75
  assertions:
69
76
  - name: correctness
70
- type: llm-grader
77
+ type: llm-rubric
71
78
  prompt: ./graders/correctness.md
72
79
 
73
80
  tests:
@@ -532,11 +539,13 @@ experiment:
532
539
 
533
540
  ### Current Shape
534
541
 
535
- Current schema accepts top-level `experiment` only as a non-empty string run
536
- grouping label. Current docs also support promptfoo-shaped `tags.experiment`.
542
+ Current schema rejects top-level `experiment` in authored eval YAML. Use the
543
+ promptfoo-shaped `tags.experiment` key when the label belongs in the eval file,
544
+ or pass CLI `--experiment` when the label is run-time context.
537
545
 
538
546
  ```yaml
539
- experiment: with-skills
547
+ tags:
548
+ experiment: with-skills
540
549
  target: codex
541
550
  timeout_seconds: 600
542
551
  evaluate_options:
@@ -545,14 +554,6 @@ evaluate_options:
545
554
  strategy: pass_any
546
555
  ```
547
556
 
548
- or:
549
-
550
- ```yaml
551
- tags:
552
- experiment: with-skills
553
- target: codex
554
- ```
555
-
556
557
  ### Migration Steps
557
558
 
558
559
  - If `experiment` is an object, move runtime fields out:
@@ -561,7 +562,8 @@ target: codex
561
562
  - budget -> `evaluate_options.budget_usd`
562
563
  - timeout -> top-level `timeout_seconds`
563
564
  - threshold -> top-level `threshold`
564
- - Keep only a string label in `experiment`, or use `tags.experiment`.
565
+ - Move string experiment labels to `tags.experiment`, or supply them at run time
566
+ with CLI `--experiment`.
565
567
 
566
568
  ### Verification
567
569
 
@@ -570,6 +572,8 @@ bun apps/cli/src/cli.ts validate path/to/eval.eval.yaml
570
572
  rg -n "^experiment:" path/to/evals
571
573
  ```
572
574
 
575
+ Any match under authored eval YAML should be removed or migrated.
576
+
573
577
  ### Compatibility Notes
574
578
 
575
579
  Do not describe a v4.42.4 eval as if `experiment:` was already the main runtime
@@ -844,54 +848,66 @@ targets:
844
848
  ### Current Shape
845
849
 
846
850
  Current eval YAML uses top-level `target` or `targets`. Target references use
847
- `label`; `id` is reserved for provider/backend identity when needed:
851
+ the target `id`; `provider` names the backend or adapter kind:
848
852
 
849
853
  ```yaml
850
854
  target: azure-base
851
855
 
852
856
  targets:
853
- - label: with-skills
854
- use_target: default
857
+ - id: with-skills
858
+ provider: codex-cli
859
+ runtime: host
860
+ config:
861
+ command: ["codex", "exec", "--json"]
855
862
  hooks:
856
863
  before_each:
857
864
  command: ["setup-plugins.sh", "skills"]
858
865
  ```
859
866
 
860
- Current `.agentv/targets.yaml` uses `label` and nests provider settings under
867
+ Current `.agentv/targets.yaml` uses `id` and nests provider settings under
861
868
  `config`:
862
869
 
863
870
  ```yaml
864
871
  targets:
865
- - label: azure-base
872
+ - id: azure-base
866
873
  provider: azure
874
+ runtime: host
867
875
  config:
868
- endpoint: ${{ AZURE_OPENAI_ENDPOINT }}
869
- api_key: ${{ AZURE_OPENAI_API_KEY }}
870
- model: ${{ AZURE_DEPLOYMENT_NAME }}
876
+ endpoint: "{{ env.AZURE_OPENAI_ENDPOINT }}"
877
+ api_key: "{{ env.AZURE_OPENAI_API_KEY }}"
878
+ model: "{{ env.AZURE_DEPLOYMENT_NAME }}"
871
879
  ```
872
880
 
873
881
  ### Migration Steps
874
882
 
875
883
  - Eval YAML: `execution.target` -> top-level `target`.
876
884
  - Eval YAML: `execution.targets` -> top-level `targets`.
877
- - Eval target object `name` -> `label`.
878
- - If an eval-local target object has provider configuration, include a
879
- `label`; use `extends` to derive from a base target.
880
- - Targets file: rename `name` to `label` and move provider-specific settings
881
- into `config`.
882
- - Keep AgentV target extensions such as `grader_target`, `use_target`,
883
- `fallback_targets`, `workers`, and `batch_requests` as top-level fields on
884
- target objects.
885
+ - Eval target object `name` -> `id`.
886
+ - If an eval-local target object has provider configuration, include a concrete
887
+ `provider`; `id` is AgentV's stable target identity and `provider` names the
888
+ backend or adapter kind.
889
+ - Targets file: rename `name` to `id`, move provider-specific settings into
890
+ `config`, and rewrite environment references from `${{ NAME }}` to
891
+ `{{ env.NAME }}`.
892
+ - Remove `use_target`; current authored target definitions must resolve to
893
+ concrete provider objects.
894
+ - Keep supported AgentV target extensions such as `grader_target`,
895
+ `fallback_targets`, and `workers` as top-level fields on target objects.
896
+ - Remove runner-level request batching from migrated configs. CLI providers are
897
+ invoked once per eval case; throughput batching belongs inside provider
898
+ adapters or CLIs without eval YAML batch configuration.
885
899
 
886
900
  ### Verification
887
901
 
888
902
  ```bash
889
903
  bun apps/cli/src/cli.ts validate path/to/eval.eval.yaml
890
904
  bun apps/cli/src/cli.ts validate .agentv/targets.yaml
891
- rg -n "execution:|name:|endpoint:|api_key:|model:" path/to/evals .agentv/targets.yaml
905
+ rg -n "execution:|name:|use_target:|\\$\\{\\{|endpoint:|api_key:|model:" path/to/evals .agentv/targets.yaml
892
906
  ```
893
907
 
894
- Inspect `name:` matches manually because grader names still exist.
908
+ Inspect `name:`, top-level provider fields, and `${{ ... }}` matches manually
909
+ because grader names, non-target YAML, and historical examples can still use
910
+ similar strings legitimately.
895
911
 
896
912
  ### Compatibility Notes
897
913
 
@@ -944,7 +960,7 @@ assert:
944
960
  - `type: g-eval` -> `type: llm-rubric`.
945
961
  - `type: code-grader`, `code-judge`, `code_grader`, or `code_judge` ->
946
962
  `type: script`.
947
- - `type: llm_judge` or `llm_grader` -> `type: llm-grader`.
963
+ - `type: llm_judge` or `llm_grader` -> `type: llm-rubric`.
948
964
  - Convert multi-word snake_case deterministic types to kebab-case:
949
965
  `is_json` -> `is-json`, `contains_all` -> `contains-all`,
950
966
  `starts_with` -> `starts-with`, and so on.
@@ -973,9 +989,9 @@ v4.42.4 LLM grader docs allowed both:
973
989
 
974
990
  ```yaml
975
991
  assertions:
976
- - type: llm-grader
992
+ - type: llm-rubric
977
993
  prompt: ./graders/correctness.md
978
- - type: llm-grader
994
+ - type: llm-rubric
979
995
  prompt: file://graders/correctness.md
980
996
  ```
981
997
 
@@ -1082,9 +1098,10 @@ rg -n "include:|tests:" path/to/evals
1082
1098
 
1083
1099
  ### Compatibility Notes
1084
1100
 
1085
- `eval_cases` remains a deprecated alias in the current schema, but migrated
1086
- YAML should use `tests`. The current convention is that runnable suites use
1087
- `*.eval.yaml`; reusable raw case files commonly use `*.cases.yaml` or JSONL.
1101
+ `eval_cases` and `evalcases` have been removed from authored eval YAML. Migrate
1102
+ them to `tests` before validating or running the suite. The current convention is
1103
+ that runnable suites use `*.eval.yaml`; reusable raw case files commonly use
1104
+ `*.cases.yaml` or JSONL.
1088
1105
 
1089
1106
  ## Result Artifact Path Changes Are Not Eval YAML Migrations
1090
1107