agentv 5.3.0-next.1 → 5.3.2-next.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +70 -60
- package/dist/{artifact-writer-JFNIPMKW.js → artifact-writer-KJEOROKQ.js} +5 -5
- package/dist/{chunk-6ZZCDZPD.js → chunk-6262OKXM.js} +21774 -21034
- package/dist/chunk-6262OKXM.js.map +1 -0
- package/dist/chunk-AQ5BIAXF.js +604 -0
- package/dist/chunk-AQ5BIAXF.js.map +1 -0
- package/dist/chunk-BV5VQLI2.js +2 -0
- package/dist/{chunk-V52ATPTT.js → chunk-JGSRUJZQ.js} +186 -32
- package/dist/chunk-JGSRUJZQ.js.map +1 -0
- package/dist/chunk-MMDLXYBX.js +721 -0
- package/dist/chunk-MMDLXYBX.js.map +1 -0
- package/dist/{chunk-FVR4RQFK.js → chunk-QSK3PC44.js} +1131 -580
- package/dist/chunk-QSK3PC44.js.map +1 -0
- package/dist/chunk-TEEXVJWM.js +2 -0
- package/dist/{chunk-BHKQHG26.js → chunk-WIAJ7JAV.js} +186 -64
- package/dist/chunk-WIAJ7JAV.js.map +1 -0
- package/dist/cli.d.ts +1 -0
- package/dist/cli.js +17517 -10
- package/dist/cli.js.map +1 -1
- package/dist/config.d.ts +2 -0
- package/dist/{ts-eval-loader-DQDYRULE-V3377FOL.js → config.js} +7 -7
- package/dist/config.js.map +1 -0
- package/dist/contracts-DsZmZLl8.d.ts +742 -0
- package/dist/contracts.d.ts +2 -0
- package/dist/contracts.js +28 -0
- package/dist/contracts.js.map +1 -0
- package/dist/dashboard/assets/{index-DNgf3qJ2.js → index-BfqOlLOF.js} +1 -1
- package/dist/dashboard/assets/index-Bh8lpGce.css +1 -0
- package/dist/dashboard/assets/index-oPR3ywQb.js +121 -0
- package/dist/dashboard/index.html +2 -2
- package/dist/{dist-6Z7U473R.js → dist-A3SGR7TW.js} +30 -18
- package/dist/dist-A3SGR7TW.js.map +1 -0
- package/dist/index.d.ts +4 -0
- package/dist/index.js +137 -18
- package/dist/{interactive-RUY3OCBI.js → interactive-DL7N2C7K.js} +24 -24
- package/dist/interactive-DL7N2C7K.js.map +1 -0
- package/dist/provider.d.ts +2 -0
- package/dist/provider.js +24 -0
- package/dist/provider.js.map +1 -0
- package/dist/sdk.d.ts +802 -0
- package/dist/sdk.js +145 -0
- package/dist/sdk.js.map +1 -0
- package/dist/skills/agentv-bench/SKILL.md +14 -13
- package/dist/skills/agentv-bench/agents/analyzer.md +1 -1
- package/dist/skills/agentv-bench/agents/executor.md +1 -1
- package/dist/skills/agentv-bench/references/autoresearch.md +9 -9
- package/dist/skills/agentv-bench/references/environment-adaptation.md +4 -4
- package/dist/skills/agentv-bench/references/eval-yaml-spec.md +30 -47
- package/dist/skills/agentv-bench/references/schemas.md +44 -60
- package/dist/skills/agentv-bench/references/subagent-pipeline.md +20 -18
- package/dist/skills/agentv-eval-migrations/SKILL.md +13 -0
- package/dist/skills/agentv-eval-migrations/references/breaking-changes.md +56 -39
- package/dist/skills/agentv-eval-writer/SKILL.md +101 -60
- package/dist/skills/agentv-eval-writer/references/custom-evaluators.md +15 -10
- package/dist/skills/agentv-eval-writer/references/eval.schema.json +4543 -5189
- package/dist/skills/agentv-eval-writer/references/python-helpers.md +2 -2
- package/dist/skills/agentv-eval-writer/references/rubric-evaluator.md +20 -3
- package/dist/templates/.agentv/providers.yaml +42 -0
- package/dist/templates/.env.example +2 -2
- package/dist/ts-eval-loader-3G5GEAEC-6F52JLPP.js +18 -0
- package/dist/ts-eval-loader-3G5GEAEC-6F52JLPP.js.map +1 -0
- package/package.json +29 -4
- package/dist/chunk-6ZZCDZPD.js.map +0 -1
- package/dist/chunk-BHKQHG26.js.map +0 -1
- package/dist/chunk-FVR4RQFK.js.map +0 -1
- package/dist/chunk-T32NL3E6.js +0 -17973
- package/dist/chunk-T32NL3E6.js.map +0 -1
- package/dist/chunk-V52ATPTT.js.map +0 -1
- package/dist/dashboard/assets/index-D_bokML8.css +0 -1
- package/dist/dashboard/assets/index-r_jSJmlw.js +0 -121
- package/dist/interactive-RUY3OCBI.js.map +0 -1
- package/dist/templates/.agentv/targets.yaml +0 -97
- /package/dist/{artifact-writer-JFNIPMKW.js.map → artifact-writer-KJEOROKQ.js.map} +0 -0
- /package/dist/{dist-6Z7U473R.js.map → chunk-BV5VQLI2.js.map} +0 -0
- /package/dist/{ts-eval-loader-DQDYRULE-V3377FOL.js.map → chunk-TEEXVJWM.js.map} +0 -0
|
@@ -30,10 +30,12 @@ specific target, set `subagent_mode_allowed: false` in `.agentv/targets.yaml`:
|
|
|
30
30
|
```yaml
|
|
31
31
|
# .agentv/targets.yaml
|
|
32
32
|
targets:
|
|
33
|
-
-
|
|
33
|
+
- id: my-target
|
|
34
34
|
provider: openai
|
|
35
|
-
|
|
36
|
-
|
|
35
|
+
runtime: host
|
|
36
|
+
config:
|
|
37
|
+
model: "{{ env.OPENAI_MODEL }}"
|
|
38
|
+
api_key: "{{ env.OPENAI_API_KEY }}"
|
|
37
39
|
subagent_mode_allowed: false # forces CLI invocation instead of executor subagent
|
|
38
40
|
```
|
|
39
41
|
|
|
@@ -43,8 +45,8 @@ even in subagent mode.
|
|
|
43
45
|
## CLI Targets: Single Command
|
|
44
46
|
|
|
45
47
|
For evals with CLI targets, `pipeline run` handles input extraction, target invocation, and
|
|
46
|
-
code grading in one step. When `--out` is omitted, the output directory
|
|
47
|
-
`.agentv/results
|
|
48
|
+
code grading in one step. When `--out` is omitted, the output directory uses the
|
|
49
|
+
canonical `.agentv/results/<run_id>` convention:
|
|
48
50
|
|
|
49
51
|
```bash
|
|
50
52
|
# Extract inputs and invoke all CLI targets in parallel:
|
|
@@ -70,7 +72,7 @@ opted out via `subagent_mode_allowed: false` in `.agentv/targets.yaml`), fall ba
|
|
|
70
72
|
### Step 1: Extract inputs
|
|
71
73
|
|
|
72
74
|
```bash
|
|
73
|
-
# Defaults to .agentv/results
|
|
75
|
+
# Defaults to .agentv/results/<run_id>
|
|
74
76
|
agentv pipeline input evals/repro.eval.yaml
|
|
75
77
|
```
|
|
76
78
|
|
|
@@ -117,7 +119,7 @@ LLM grading → merge and validate).
|
|
|
117
119
|
Use individual commands when you need control over each step with CLI targets:
|
|
118
120
|
|
|
119
121
|
```bash
|
|
120
|
-
# Step 1: Extract inputs (defaults to .agentv/results
|
|
122
|
+
# Step 1: Extract inputs (defaults to .agentv/results/<run_id>)
|
|
121
123
|
agentv pipeline input evals/repro.eval.yaml
|
|
122
124
|
|
|
123
125
|
# Step 2: run_tests.py invokes CLI targets (or use pipeline run instead)
|
|
@@ -127,7 +129,7 @@ agentv pipeline grade <run-dir>
|
|
|
127
129
|
|
|
128
130
|
# Step 4: Subagent does LLM grading, writes results to llm_grader_results/<name>.json per test
|
|
129
131
|
|
|
130
|
-
# Step 5: Merge scores
|
|
132
|
+
# Step 5: Merge scores into the canonical run bundle
|
|
131
133
|
agentv pipeline bench <run-dir>
|
|
132
134
|
|
|
133
135
|
# Step 6: Validate
|
|
@@ -152,22 +154,22 @@ The agent reads `llm_graders/<name>.json` for each test, grades the response usi
|
|
|
152
154
|
|
|
153
155
|
## Pipeline Bench and Dashboard
|
|
154
156
|
|
|
155
|
-
`pipeline bench` merges LLM scores into
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
to verify
|
|
157
|
+
`pipeline bench` merges LLM scores into the canonical run bundle, including
|
|
158
|
+
`.internal/index.jsonl`, `summary.json`, and per-attempt `grading.json` with
|
|
159
|
+
`component_results`. The web dashboard reads this format directly — no separate
|
|
160
|
+
conversion script is needed. Run `agentv results validate <run-dir>` to verify
|
|
161
|
+
compatibility.
|
|
159
162
|
|
|
160
163
|
## Output Structure
|
|
161
164
|
|
|
162
|
-
The path hierarchy mirrors
|
|
163
|
-
the eval.yaml. The target is recorded in `manifest.json` — one run = one target.
|
|
165
|
+
The path hierarchy mirrors CLI mode. The target is recorded in run metadata.
|
|
164
166
|
|
|
165
167
|
```
|
|
166
|
-
.agentv/results/<
|
|
168
|
+
.agentv/results/<run_id>/
|
|
167
169
|
├── manifest.json ← eval metadata, target, test_ids
|
|
168
|
-
├──
|
|
169
|
-
├──
|
|
170
|
-
└── <
|
|
170
|
+
├── summary.json ← aggregate statistics
|
|
171
|
+
├── .internal/index.jsonl ← canonical row index
|
|
172
|
+
└── <result-dir>/sample-1/grading.json
|
|
171
173
|
└── <test-id>/ ← test case id
|
|
172
174
|
├── input.json ← test input text + messages
|
|
173
175
|
├── invoke.json ← target command or agent instructions
|
|
@@ -16,6 +16,19 @@ v4.42.4-era shape, current shape, migration steps, verification commands, and
|
|
|
16
16
|
compatibility notes for each major breaking authoring change. Then compare the
|
|
17
17
|
eval file against the current portable contract:
|
|
18
18
|
|
|
19
|
+
- Treat migrations as rule-based rewrites first. Do not rely on a manual
|
|
20
|
+
"looks current" pass when the old field has a direct current replacement.
|
|
21
|
+
- Rewrite authored prompt data into top-level `prompts` plus `tests[].vars`;
|
|
22
|
+
keep `input` only for documented raw-case import compatibility.
|
|
23
|
+
- Move sibling `expected_output` to `vars.expected_output`, then add or keep an
|
|
24
|
+
explicit assertion that consumes it, usually `llm-rubric` with
|
|
25
|
+
`{{ expected_output }}`.
|
|
26
|
+
- Rewrite target identity to `id`, keep `provider` for the backend/adapter kind,
|
|
27
|
+
put provider settings under `config`, and rewrite env references to
|
|
28
|
+
`{{ env.NAME }}`.
|
|
29
|
+
- Remove `use_target`, `eval_cases`, `evalcases`, `providerPromptMap`, and
|
|
30
|
+
`provider_prompt_map` from authored eval/config YAML. Current AgentV rejects
|
|
31
|
+
these fields, so do not preserve them as compatibility aliases.
|
|
19
32
|
- Keep committed eval YAML portable: prompts, cases, assertions, workspace
|
|
20
33
|
templates, repos, hooks, env checks, Docker preflight/container config, and
|
|
21
34
|
`workspace.scope`.
|
|
@@ -23,7 +23,7 @@ For a v4.42.4-era eval:
|
|
|
23
23
|
|
|
24
24
|
1. Move top-level `execution.target` to top-level `target`.
|
|
25
25
|
2. Move top-level `execution.targets` to top-level `targets`, and rename target
|
|
26
|
-
object `name` to `
|
|
26
|
+
object `name` to `id`.
|
|
27
27
|
3. Rename suite/test/turn `assertions` to `assert`.
|
|
28
28
|
4. If a test uses `tests[].execution.assertions` or
|
|
29
29
|
`tests[].execution.evaluators`, rename that nested list to
|
|
@@ -56,7 +56,14 @@ For a v4.42.4-era eval:
|
|
|
56
56
|
assertions where there is a direct mapping.
|
|
57
57
|
18. Keep raw cases under `tests` / `tests: file://...`; run full eval suites
|
|
58
58
|
directly with CLI multi-file selection and tags.
|
|
59
|
-
19.
|
|
59
|
+
19. Move authored prompt input from top-level or `tests[].input` into
|
|
60
|
+
`prompts` plus `tests[].vars.input`.
|
|
61
|
+
20. Move authored reference answers from sibling `expected_output` into
|
|
62
|
+
`vars.expected_output`; add or keep an explicit assertion that consumes it.
|
|
63
|
+
21. Rewrite target env interpolation from `${{ NAME }}` to `{{ env.NAME }}`.
|
|
64
|
+
22. Remove `use_target`, `eval_cases`, `evalcases`, `providerPromptMap`, and
|
|
65
|
+
`provider_prompt_map`; current AgentV rejects them.
|
|
66
|
+
23. Validate with `bun apps/cli/src/cli.ts validate <eval-file>`.
|
|
60
67
|
|
|
61
68
|
## Assertions Renamed To `assert`
|
|
62
69
|
|
|
@@ -67,7 +74,7 @@ v4.42.4 docs and schema used `assertions` for suite-level and per-test graders:
|
|
|
67
74
|
```yaml
|
|
68
75
|
assertions:
|
|
69
76
|
- name: correctness
|
|
70
|
-
type: llm-
|
|
77
|
+
type: llm-rubric
|
|
71
78
|
prompt: ./graders/correctness.md
|
|
72
79
|
|
|
73
80
|
tests:
|
|
@@ -532,11 +539,13 @@ experiment:
|
|
|
532
539
|
|
|
533
540
|
### Current Shape
|
|
534
541
|
|
|
535
|
-
Current schema
|
|
536
|
-
|
|
542
|
+
Current schema rejects top-level `experiment` in authored eval YAML. Use the
|
|
543
|
+
promptfoo-shaped `tags.experiment` key when the label belongs in the eval file,
|
|
544
|
+
or pass CLI `--experiment` when the label is run-time context.
|
|
537
545
|
|
|
538
546
|
```yaml
|
|
539
|
-
|
|
547
|
+
tags:
|
|
548
|
+
experiment: with-skills
|
|
540
549
|
target: codex
|
|
541
550
|
timeout_seconds: 600
|
|
542
551
|
evaluate_options:
|
|
@@ -545,14 +554,6 @@ evaluate_options:
|
|
|
545
554
|
strategy: pass_any
|
|
546
555
|
```
|
|
547
556
|
|
|
548
|
-
or:
|
|
549
|
-
|
|
550
|
-
```yaml
|
|
551
|
-
tags:
|
|
552
|
-
experiment: with-skills
|
|
553
|
-
target: codex
|
|
554
|
-
```
|
|
555
|
-
|
|
556
557
|
### Migration Steps
|
|
557
558
|
|
|
558
559
|
- If `experiment` is an object, move runtime fields out:
|
|
@@ -561,7 +562,8 @@ target: codex
|
|
|
561
562
|
- budget -> `evaluate_options.budget_usd`
|
|
562
563
|
- timeout -> top-level `timeout_seconds`
|
|
563
564
|
- threshold -> top-level `threshold`
|
|
564
|
-
-
|
|
565
|
+
- Move string experiment labels to `tags.experiment`, or supply them at run time
|
|
566
|
+
with CLI `--experiment`.
|
|
565
567
|
|
|
566
568
|
### Verification
|
|
567
569
|
|
|
@@ -570,6 +572,8 @@ bun apps/cli/src/cli.ts validate path/to/eval.eval.yaml
|
|
|
570
572
|
rg -n "^experiment:" path/to/evals
|
|
571
573
|
```
|
|
572
574
|
|
|
575
|
+
Any match under authored eval YAML should be removed or migrated.
|
|
576
|
+
|
|
573
577
|
### Compatibility Notes
|
|
574
578
|
|
|
575
579
|
Do not describe a v4.42.4 eval as if `experiment:` was already the main runtime
|
|
@@ -844,54 +848,66 @@ targets:
|
|
|
844
848
|
### Current Shape
|
|
845
849
|
|
|
846
850
|
Current eval YAML uses top-level `target` or `targets`. Target references use
|
|
847
|
-
`
|
|
851
|
+
the target `id`; `provider` names the backend or adapter kind:
|
|
848
852
|
|
|
849
853
|
```yaml
|
|
850
854
|
target: azure-base
|
|
851
855
|
|
|
852
856
|
targets:
|
|
853
|
-
-
|
|
854
|
-
|
|
857
|
+
- id: with-skills
|
|
858
|
+
provider: codex-cli
|
|
859
|
+
runtime: host
|
|
860
|
+
config:
|
|
861
|
+
command: ["codex", "exec", "--json"]
|
|
855
862
|
hooks:
|
|
856
863
|
before_each:
|
|
857
864
|
command: ["setup-plugins.sh", "skills"]
|
|
858
865
|
```
|
|
859
866
|
|
|
860
|
-
Current `.agentv/targets.yaml` uses `
|
|
867
|
+
Current `.agentv/targets.yaml` uses `id` and nests provider settings under
|
|
861
868
|
`config`:
|
|
862
869
|
|
|
863
870
|
```yaml
|
|
864
871
|
targets:
|
|
865
|
-
-
|
|
872
|
+
- id: azure-base
|
|
866
873
|
provider: azure
|
|
874
|
+
runtime: host
|
|
867
875
|
config:
|
|
868
|
-
endpoint:
|
|
869
|
-
api_key:
|
|
870
|
-
model:
|
|
876
|
+
endpoint: "{{ env.AZURE_OPENAI_ENDPOINT }}"
|
|
877
|
+
api_key: "{{ env.AZURE_OPENAI_API_KEY }}"
|
|
878
|
+
model: "{{ env.AZURE_DEPLOYMENT_NAME }}"
|
|
871
879
|
```
|
|
872
880
|
|
|
873
881
|
### Migration Steps
|
|
874
882
|
|
|
875
883
|
- Eval YAML: `execution.target` -> top-level `target`.
|
|
876
884
|
- Eval YAML: `execution.targets` -> top-level `targets`.
|
|
877
|
-
- Eval target object `name` -> `
|
|
878
|
-
- If an eval-local target object has provider configuration, include a
|
|
879
|
-
`
|
|
880
|
-
|
|
881
|
-
|
|
882
|
-
|
|
883
|
-
`
|
|
884
|
-
|
|
885
|
+
- Eval target object `name` -> `id`.
|
|
886
|
+
- If an eval-local target object has provider configuration, include a concrete
|
|
887
|
+
`provider`; `id` is AgentV's stable target identity and `provider` names the
|
|
888
|
+
backend or adapter kind.
|
|
889
|
+
- Targets file: rename `name` to `id`, move provider-specific settings into
|
|
890
|
+
`config`, and rewrite environment references from `${{ NAME }}` to
|
|
891
|
+
`{{ env.NAME }}`.
|
|
892
|
+
- Remove `use_target`; current authored target definitions must resolve to
|
|
893
|
+
concrete provider objects.
|
|
894
|
+
- Keep supported AgentV target extensions such as `grader_target`,
|
|
895
|
+
`fallback_targets`, and `workers` as top-level fields on target objects.
|
|
896
|
+
- Remove runner-level request batching from migrated configs. CLI providers are
|
|
897
|
+
invoked once per eval case; throughput batching belongs inside provider
|
|
898
|
+
adapters or CLIs without eval YAML batch configuration.
|
|
885
899
|
|
|
886
900
|
### Verification
|
|
887
901
|
|
|
888
902
|
```bash
|
|
889
903
|
bun apps/cli/src/cli.ts validate path/to/eval.eval.yaml
|
|
890
904
|
bun apps/cli/src/cli.ts validate .agentv/targets.yaml
|
|
891
|
-
rg -n "execution:|name:|endpoint:|api_key:|model:" path/to/evals .agentv/targets.yaml
|
|
905
|
+
rg -n "execution:|name:|use_target:|\\$\\{\\{|endpoint:|api_key:|model:" path/to/evals .agentv/targets.yaml
|
|
892
906
|
```
|
|
893
907
|
|
|
894
|
-
Inspect `name
|
|
908
|
+
Inspect `name:`, top-level provider fields, and `${{ ... }}` matches manually
|
|
909
|
+
because grader names, non-target YAML, and historical examples can still use
|
|
910
|
+
similar strings legitimately.
|
|
895
911
|
|
|
896
912
|
### Compatibility Notes
|
|
897
913
|
|
|
@@ -944,7 +960,7 @@ assert:
|
|
|
944
960
|
- `type: g-eval` -> `type: llm-rubric`.
|
|
945
961
|
- `type: code-grader`, `code-judge`, `code_grader`, or `code_judge` ->
|
|
946
962
|
`type: script`.
|
|
947
|
-
- `type: llm_judge` or `llm_grader` -> `type: llm-
|
|
963
|
+
- `type: llm_judge` or `llm_grader` -> `type: llm-rubric`.
|
|
948
964
|
- Convert multi-word snake_case deterministic types to kebab-case:
|
|
949
965
|
`is_json` -> `is-json`, `contains_all` -> `contains-all`,
|
|
950
966
|
`starts_with` -> `starts-with`, and so on.
|
|
@@ -973,9 +989,9 @@ v4.42.4 LLM grader docs allowed both:
|
|
|
973
989
|
|
|
974
990
|
```yaml
|
|
975
991
|
assertions:
|
|
976
|
-
- type: llm-
|
|
992
|
+
- type: llm-rubric
|
|
977
993
|
prompt: ./graders/correctness.md
|
|
978
|
-
- type: llm-
|
|
994
|
+
- type: llm-rubric
|
|
979
995
|
prompt: file://graders/correctness.md
|
|
980
996
|
```
|
|
981
997
|
|
|
@@ -1082,9 +1098,10 @@ rg -n "include:|tests:" path/to/evals
|
|
|
1082
1098
|
|
|
1083
1099
|
### Compatibility Notes
|
|
1084
1100
|
|
|
1085
|
-
`eval_cases`
|
|
1086
|
-
|
|
1087
|
-
`*.eval.yaml`; reusable raw case files commonly use
|
|
1101
|
+
`eval_cases` and `evalcases` have been removed from authored eval YAML. Migrate
|
|
1102
|
+
them to `tests` before validating or running the suite. The current convention is
|
|
1103
|
+
that runnable suites use `*.eval.yaml`; reusable raw case files commonly use
|
|
1104
|
+
`*.cases.yaml` or JSONL.
|
|
1088
1105
|
|
|
1089
1106
|
## Result Artifact Path Changes Are Not Eval YAML Migrations
|
|
1090
1107
|
|