gooddata-eval 1.74.1.dev5__tar.gz → 1.75.1.dev1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/AGENTS.md +119 -2
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/PKG-INFO +202 -4
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/README.md +200 -2
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/pyproject.toml +5 -2
- gooddata_eval-1.75.1.dev1/scripts/verify_guardrail_refusal_criteria.py +126 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/cli/agentic_runner.py +18 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/cli/main.py +152 -3
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/__init__.py +14 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/alert_skill.py +2 -2
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/conversation.py +2 -2
- gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/agentic/dashboard_skill.py +742 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/guardrail.py +5 -3
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/metric_skill.py +2 -2
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/visualization.py +2 -2
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/chat/sse_client.py +126 -17
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/config.py +6 -0
- gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/dataset/from_insights.py +855 -0
- gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/evaluators/_guardrail_criteria.py +39 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/general_question.py +2 -4
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/guardrail.py +8 -7
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/search_tool.py +2 -4
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/visualization.py +34 -9
- gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/granularity.py +65 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/models.py +80 -1
- gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/reporting/html_report.py +121 -0
- gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/reporting/report_template.html +452 -0
- gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/scoring.py +393 -0
- gooddata_eval-1.75.1.dev1/tests/conftest.py +45 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_alert_skill.py +2 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_conversation.py +2 -0
- gooddata_eval-1.75.1.dev1/tests/test_agentic_dashboard_skill.py +679 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_guardrail.py +2 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_metric_skill.py +2 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_runner.py +5 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_visualization.py +8 -0
- gooddata_eval-1.75.1.dev1/tests/test_from_insights.py +1194 -0
- gooddata_eval-1.75.1.dev1/tests/test_guardrail_criteria.py +59 -0
- gooddata_eval-1.75.1.dev1/tests/test_html_report.py +133 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_models.py +56 -0
- gooddata_eval-1.75.1.dev1/tests/test_scoring.py +539 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_sse_client.py +83 -4
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_trace_linker.py +1 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_visualization_evaluator.py +42 -3
- gooddata_eval-1.74.1.dev5/src/gooddata_eval/core/scoring.py +0 -221
- gooddata_eval-1.74.1.dev5/tests/conftest.py +0 -22
- gooddata_eval-1.74.1.dev5/tests/test_scoring.py +0 -207
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/.gitignore +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/CLAUDE.md +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/LICENSE.txt +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/Makefile +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/__init__.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/_version.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/cli/__init__.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/__init__.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/_output.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/_catalog.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/_gate.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/_langfuse.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/_trace_linker.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/general_question.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/kda_skill.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/search_tool.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/chat/__init__.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/chat/render.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/connection.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/dataset/__init__.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/dataset/langfuse_source.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/dataset/local.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/__init__.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/_deep_subset.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/_llm_judge.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/_maql.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/_text_utils.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/alert_skill.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/base.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/metric_skill.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/summary.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/__init__.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/_env.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/client.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/experiment.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/observations.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/otlp.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/sink.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/reporting/__init__.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/reporting/console.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/reporting/json_report.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/runner.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/summary/__init__.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/summary/http_client.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/timing.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/workspace.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/__init__.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/_fake_langfuse.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/fixtures/sample_dataset/metric_skill_create.json +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/fixtures/sample_dataset/visualization_revenue.json +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/fixtures/sse_visualization_stream.txt +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_gate.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_general_question.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_kda_skill.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_langfuse_trace.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_observe_experiment.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_run_context.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_search_tool.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_alert_skill_evaluator.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_chat_render.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_cli.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_connection.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_deep_subset.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_fake_langfuse.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_client.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_e2e_fake_server.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_env.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_experiment.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_observations.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_otlp.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_sink.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_source.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_llm_judge.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_local_loader.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_maql_normalize.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_metric_skill_evaluator.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_reporting.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_runner.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_search_tool_evaluator.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_summary_client.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_summary_evaluator.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_text_evaluators.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_timing.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_workspace.py +0 -0
- {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tox.ini +0 -0
|
@@ -9,7 +9,7 @@ experiment. The newest and most actively developed package in the repo.
|
|
|
9
9
|
|
|
10
10
|
## Owns
|
|
11
11
|
|
|
12
|
-
- The `gd-eval` CLI (`
|
|
12
|
+
- The `gd-eval` CLI (`generate`, `run`, `report`, `models`)
|
|
13
13
|
- Dataset loading and the evaluation run loop
|
|
14
14
|
- Per-capability evaluators and their scoring
|
|
15
15
|
- Result reporting, and pushing experiments, scores and trace links to Langfuse
|
|
@@ -27,7 +27,7 @@ experiment. The newest and most actively developed package in the repo.
|
|
|
27
27
|
| `core/agentic/` | multi-turn agentic evaluation per capability, **plus** all Langfuse trace polling and linking (`_langfuse.py`, `_trace_linker.py`) |
|
|
28
28
|
| `core/chat/` | SSE client for the agent's streaming chat endpoint |
|
|
29
29
|
| `core/summary/` | HTTP client for the dedicated dashboard-summary endpoint — a single-shot chat backend, not reporting |
|
|
30
|
-
| `core/dataset/` | dataset format and
|
|
30
|
+
| `core/dataset/` | dataset format, loading, and `from_insights.py` — dataset generation from a workspace's real insights |
|
|
31
31
|
| `core/evaluators/` | single-shot evaluators and their registry |
|
|
32
32
|
| `core/langfuse/` | the whole Langfuse v4 client: `_env` (base URL + credentials), `otlp` (OTLP/JSON encoding), `experiment` (root-span construction, score targets), `observations` (trace reads), `client` (httpx calls), `sink` (single-shot results as experiments) |
|
|
33
33
|
| `core/reporting/` | console and JSON output rendering |
|
|
@@ -61,6 +61,106 @@ its own shape. `test_kind` on the item is what labels the result, not the evalua
|
|
|
61
61
|
which is why `knowledge_question` can reuse `GeneralQuestionEvaluator` verbatim.
|
|
62
62
|
`dashboard_summary` items additionally need `summary_input`.
|
|
63
63
|
|
|
64
|
+
## Running the pipeline
|
|
65
|
+
|
|
66
|
+
Four subcommands, in the order you use them. Everything runs through `uv`; never a bare
|
|
67
|
+
`python`. There is no build step -- `uv run` syncs the environment from `uv.lock` on first
|
|
68
|
+
use, so a fresh clone needs nothing but:
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
uv run --package gooddata-eval gd-eval <subcommand> --help
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
`openai` is an optional extra (`llm-judge`) so the published package stays installable
|
|
75
|
+
without it, but the `dev` dependency group pulls it in, which is why a plain `uv run` here
|
|
76
|
+
has the phrasing step and the LLM judge. Installing `gooddata-eval` from PyPI does not --
|
|
77
|
+
there the extra is explicit, and every `openai` import site is guarded or deferred.
|
|
78
|
+
|
|
79
|
+
Connection is the same for every subcommand that talks to the platform: `--host` +
|
|
80
|
+
`--token`, or `GOODDATA_TOKEN` in the environment, or `--profile <name>` reading
|
|
81
|
+
`~/.gooddata/profiles.yaml`. Precedence is flags > env > profile.
|
|
82
|
+
|
|
83
|
+
### 1. `generate` — build a dataset from a workspace
|
|
84
|
+
|
|
85
|
+
Reverse-engineers `visualization` items out of the charts a workspace already has, so the
|
|
86
|
+
expected output is copied from a live object rather than invented. Needs `OPENAI_API_KEY`
|
|
87
|
+
for the phrasing step, or `--no-phrase` to emit mechanical `Show <title>` questions.
|
|
88
|
+
|
|
89
|
+
```bash
|
|
90
|
+
uv run --package gooddata-eval gd-eval generate \
|
|
91
|
+
--host "$GOODDATA_HOST" --workspace "$WORKSPACE_ID" \
|
|
92
|
+
--dataset-name ecommerce --out ./datasets/ecommerce \
|
|
93
|
+
--snapshot-out /tmp/ws.json \
|
|
94
|
+
--phrase-model gpt-4o --skip-ambiguous
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
Iterate offline instead of re-fetching: `--snapshot-out` writes everything the generator
|
|
98
|
+
read as one JSON file, and `--snapshot-in` replays it with no host, token or network. Add
|
|
99
|
+
`--dry-run` to print the shape counts and each spec's brief without writing anything —
|
|
100
|
+
the fastest way to see what a workspace yields.
|
|
101
|
+
|
|
102
|
+
Quality gates fail the command (exit 1) below `--min-questions` (15) or `--min-filtered` (1). Lower them for a smoke test; do not lower them to ship a dataset.
|
|
103
|
+
`--langfuse-out` additionally writes a Langfuse-importable file, and `--id-prefix` rewrites
|
|
104
|
+
ids on that export only, because Langfuse item ids are unique per project.
|
|
105
|
+
|
|
106
|
+
Insights the AAC spec cannot express without guessing are skipped with a printed reason
|
|
107
|
+
(`SKIP <id>: derived measure (previousPeriodMeasure)`). Read those — they are the
|
|
108
|
+
generator telling you what it refused to invent, not noise.
|
|
109
|
+
|
|
110
|
+
### 2. `run` — evaluate
|
|
111
|
+
|
|
112
|
+
```bash
|
|
113
|
+
uv run --package gooddata-eval gd-eval run \
|
|
114
|
+
--host "$GOODDATA_HOST" --workspace "$WORKSPACE_ID" \
|
|
115
|
+
--dataset ./datasets/ecommerce --kind visualization \
|
|
116
|
+
--model gpt-5.2 --model ProviderName/gpt-4o \
|
|
117
|
+
--runs 3 --gate power --concurrency 4 \
|
|
118
|
+
--json ./results/run.json --html ./results/run.html
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
`--dataset` reads a local folder; `--langfuse-dataset` pulls one by name instead. `--kind`
|
|
122
|
+
only supplies a default for items that do not carry their own `test_kind`. Repeat
|
|
123
|
+
`--model` to compare models in one run. `--runs` with `--gate power` measures stability
|
|
124
|
+
(every run must pass) rather than pass@K. `--langfuse` pushes the run as a scored
|
|
125
|
+
experiment, needing `LANGFUSE_HOST`, `LANGFUSE_PUBLIC_KEY` and `LANGFUSE_SECRET_KEY`.
|
|
126
|
+
|
|
127
|
+
`--concurrency` is capped for you where it matters: kinds that create workspace objects
|
|
128
|
+
run one at a time regardless, see the parallel-safety gotcha below.
|
|
129
|
+
|
|
130
|
+
### 3. `report` — compare runs
|
|
131
|
+
|
|
132
|
+
```bash
|
|
133
|
+
uv run --package gooddata-eval gd-eval report \
|
|
134
|
+
./results/*.json -o ./results/comparison.html --title "luna vs 4o" --redact
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
Several JSON reports become side-by-side columns keyed by file name. `--redact` is the
|
|
138
|
+
customer-safe form: conversation ids, response ids and raw reasoning dropped, model names
|
|
139
|
+
replaced with "Model A", "Model B".
|
|
140
|
+
|
|
141
|
+
### 4. `models` — what the org has configured
|
|
142
|
+
|
|
143
|
+
```bash
|
|
144
|
+
uv run --package gooddata-eval gd-eval models --host "$GOODDATA_HOST"
|
|
145
|
+
```
|
|
146
|
+
|
|
147
|
+
Run this before guessing a `--model` string.
|
|
148
|
+
|
|
149
|
+
### Environment
|
|
150
|
+
|
|
151
|
+
| Variable | Used by |
|
|
152
|
+
|---|---|
|
|
153
|
+
| `GOODDATA_TOKEN` | every platform-facing subcommand |
|
|
154
|
+
| `OPENAI_API_KEY` | `generate` phrasing, and the LLM-as-judge evaluators |
|
|
155
|
+
| `GD_EVAL_JUDGE_MODEL` | judge model, same as `--judge-model` |
|
|
156
|
+
| `GD_EVAL_AGENT_ID` | which agent to drive, same as `--agent-id` |
|
|
157
|
+
| `LANGFUSE_HOST`, `LANGFUSE_PUBLIC_KEY`, `LANGFUSE_SECRET_KEY` | `--langfuse`, `--langfuse-dataset` |
|
|
158
|
+
| `GOODDATA_EVAL_CHAT_*` | SSE retry, backoff and timeout knobs |
|
|
159
|
+
| `GD_EVAL_TIMERS` | same as `--timers` |
|
|
160
|
+
|
|
161
|
+
A gitignored `.env` at the repo root is the normal place for these; load it with
|
|
162
|
+
`set -a && . ./.env && set +a` before the command.
|
|
163
|
+
|
|
64
164
|
## Gotchas
|
|
65
165
|
|
|
66
166
|
**Adding an evaluator is a registry change, not a naming convention.** Single-shot kinds go
|
|
@@ -89,6 +189,23 @@ ingestion has no pass/fail signal and inflates or misattributes per-item latency
|
|
|
89
189
|
(`run_trace_link_inline` is the synchronous alternative). Do not "fix" a slow item by
|
|
90
190
|
making trace scoring synchronous again.
|
|
91
191
|
|
|
192
|
+
**A generated item's `expected_output` is copied, never invented — keep it that way.**
|
|
193
|
+
`core/dataset/from_insights.py` converts each insight with the platform's own
|
|
194
|
+
`declarative_visualization_to_aac()` (from `gooddata-code-convertors`, via `gooddata-sdk`),
|
|
195
|
+
so the mapping is not ours to get wrong. What is ours is deciding what the evaluator cannot
|
|
196
|
+
yet score — derived measures and measure-level filters convert fine and then compare wrong,
|
|
197
|
+
so they are skipped with a printed reason — and stripping the no-op filters AD saves for an
|
|
198
|
+
"All" selection, which would otherwise let a question claim a filter its chart lacks. Teach
|
|
199
|
+
the comparator about a construct and the matching skip can go; do not make one convert by
|
|
200
|
+
hand. Chart type names are the convertor's, which are also the agent's — do not rename them. One
|
|
201
|
+
granularity is patched in `CONVERTOR_GRANULARITY_FIXES`: `week_us` → `WEEK_US` is a convertor
|
|
202
|
+
bug, the platform enum is `WEEK`.
|
|
203
|
+
|
|
204
|
+
**The snapshot is a plain-JSON contract.** `--snapshot-in`/`--snapshot-out` is what makes
|
|
205
|
+
the generator testable offline and iterable without re-fetching, and it is why the
|
|
206
|
+
generator reads the declarative analytics model rather than `sdk.visualizations`. Anything
|
|
207
|
+
that changes the fetch shape invalidates every saved snapshot.
|
|
208
|
+
|
|
92
209
|
**Scoring weights do not sum to 1.** `quality_score` is the fraction of boolean-valued keys
|
|
93
210
|
in `best_detail` that are true, falling back to `pass_at_k` when there are none (text
|
|
94
211
|
evaluators). `value_score` is `0.6 * quality + 0.2 * speed` — the 0.8 total is what the
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: gooddata-eval
|
|
3
|
-
Version: 1.
|
|
3
|
+
Version: 1.75.1.dev1
|
|
4
4
|
Summary: Evaluate the GoodData AI agent against your own questions and models.
|
|
5
5
|
Project-URL: Source, https://github.com/gooddata/gooddata-python-sdk
|
|
6
6
|
Author-email: GoodData <support@gooddata.com>
|
|
@@ -17,7 +17,7 @@ Classifier: Topic :: Scientific/Engineering
|
|
|
17
17
|
Classifier: Topic :: Software Development
|
|
18
18
|
Classifier: Typing :: Typed
|
|
19
19
|
Requires-Python: >=3.10
|
|
20
|
-
Requires-Dist: gooddata-sdk~=1.
|
|
20
|
+
Requires-Dist: gooddata-sdk~=1.75.1.dev1
|
|
21
21
|
Requires-Dist: httpx<1.0,>=0.27
|
|
22
22
|
Requires-Dist: orjson<4.0.0,>=3.9.15
|
|
23
23
|
Requires-Dist: pydantic<3.0,>=2.6
|
|
@@ -44,7 +44,10 @@ Or install `gd-eval` as a standalone tool:
|
|
|
44
44
|
| Command | Description |
|
|
45
45
|
|---|---|
|
|
46
46
|
| `gd-eval run` | Run an evaluation dataset against one or more models. |
|
|
47
|
+
| `gd-eval report` | Render JSON report(s) as one self-contained HTML file. |
|
|
47
48
|
| `gd-eval models` | List LLM providers and models configured in the org. |
|
|
49
|
+
| `gd-eval generate` | Generate a `visualization` dataset from a workspace's existing insights. |
|
|
50
|
+
|
|
48
51
|
|
|
49
52
|
---
|
|
50
53
|
|
|
@@ -173,6 +176,8 @@ interleaves when K > 1, and per-item latencies rise, so they stop being clean si
|
|
|
173
176
|
| Flag | Description |
|
|
174
177
|
|---|---|
|
|
175
178
|
| `--json PATH` | Write a JSON report to this path. Always uses the nested `{models, runs, comparison}` shape even for a single model. |
|
|
179
|
+
| `--html PATH` | Write a self-contained HTML report to this path (same output as `gd-eval report`). |
|
|
180
|
+
| `--redact` | Make the HTML customer-safe. See `gd-eval report`. |
|
|
176
181
|
| `--quiet` | Suppress per-item progress. Per-model result tables and the comparison summary are still printed. |
|
|
177
182
|
| `--preserve-failed` | Keep failed conversations on the server instead of deleting them, so they can be inspected afterwards. Applies to the single-turn chat path; agentic kinds manage their own conversation lifecycle. |
|
|
178
183
|
| `--timers` | Print per-turn `[timer]` diagnostics — GoodData response, judge, and simulated-user seconds as they happen. Off by default: an 18-item `--runs 2` run emits ~72 lines and buries the progress output. The same measurements are always in the JSON report's `latency_breakdown_s`, so this only adds a live view. Also settable via `GD_EVAL_TIMERS=1`. |
|
|
@@ -333,6 +338,59 @@ linking ran. Pass `TAVERN_E2E_SKIP_TRACE_LINK=1` to opt out of linking altogethe
|
|
|
333
338
|
|
|
334
339
|
---
|
|
335
340
|
|
|
341
|
+
## `gd-eval report`
|
|
342
|
+
|
|
343
|
+
Turns JSON report(s) into one HTML file you can actually navigate. No server, no
|
|
344
|
+
credentials, no external assets — it opens over `file://`, attaches to a Jira issue and
|
|
345
|
+
survives a Slack thread.
|
|
346
|
+
|
|
347
|
+
```bash
|
|
348
|
+
# one run
|
|
349
|
+
gd-eval report results.json -o report.html
|
|
350
|
+
|
|
351
|
+
# several runs side by side -- each file becomes its own column
|
|
352
|
+
gd-eval report aug-21.json sep-07.json -o comparison.html --title "H200 regression check"
|
|
353
|
+
|
|
354
|
+
# customer-safe
|
|
355
|
+
gd-eval report results.json -o customer.html --redact
|
|
356
|
+
```
|
|
357
|
+
|
|
358
|
+
| Flag | Description |
|
|
359
|
+
|---|---|
|
|
360
|
+
| `-o, --out PATH` | Where to write the HTML. Required. |
|
|
361
|
+
| `--title TEXT` | Title shown in the report header. |
|
|
362
|
+
| `--redact` | Drop conversation/response ids and raw reasoning, and rename models to `Model A`, `Model B`, … Pass rate, per-item results, questions and latency survive. |
|
|
363
|
+
|
|
364
|
+
The report is a *view* over the JSON — it computes no numbers of its own. It gives you:
|
|
365
|
+
|
|
366
|
+
- **Run cards and a comparison table** — pass rate, quality, latency per run.
|
|
367
|
+
- **An item table** with a pass/fail column per run, so a model or run-over-run
|
|
368
|
+
regression is one glance rather than a hand-assembled spreadsheet.
|
|
369
|
+
- **An expression filter** for cross-cutting questions the fixed filters can't
|
|
370
|
+
anticipate, e.g. `d.filter_ranking_score === false` or
|
|
371
|
+
`d.expected_metric_uris.length > 1 && !d.metrics_correct`. Available variables:
|
|
372
|
+
`d` (the focused run's `detail`), `it` (its item), `i` (the row, `i.per[label]` for any
|
|
373
|
+
run), `q` (question), `kind`.
|
|
374
|
+
- **The conversation**, when the item ran the agentic multi-turn path — every turn in
|
|
375
|
+
order, with the simulated user marked apart from a real question, so you can see
|
|
376
|
+
whether the agent got there or was handed the answer.
|
|
377
|
+
- **A per-item drawer** — checks as pass/fail chips, expected vs actual side by side,
|
|
378
|
+
full reasoning, conversation/response ids.
|
|
379
|
+
- **A latency timeline** from `detail.latency_breakdown`, in execution order, one bar per
|
|
380
|
+
step. Clicking a step expands its full record, joined by `index`: the paragraph a
|
|
381
|
+
reasoning step was summarised from, or a tool call's arguments and result from
|
|
382
|
+
`detail.tool_calls`.
|
|
383
|
+
|
|
384
|
+
`--redact` additionally drops `transcript` and `tool_calls`: the exchange shows that the
|
|
385
|
+
simulated user is primed with the expected output, and a tool result carries
|
|
386
|
+
semantic-layer internals and real query rows. The turn count and the timeline shape
|
|
387
|
+
survive.
|
|
388
|
+
|
|
389
|
+
Passing several files keyed by file name is the whole run-over-run mechanism: no
|
|
390
|
+
database, no run registry, just the JSON files you already have on disk.
|
|
391
|
+
|
|
392
|
+
---
|
|
393
|
+
|
|
336
394
|
## `gd-eval models`
|
|
337
395
|
|
|
338
396
|
List all LLM providers and their models in the org. Marks the active model
|
|
@@ -353,6 +411,143 @@ gd-eval models \
|
|
|
353
411
|
|
|
354
412
|
---
|
|
355
413
|
|
|
414
|
+
## `gd-eval generate`
|
|
415
|
+
|
|
416
|
+
Reverse-engineers a `visualization` dataset out of the charts a customer has already
|
|
417
|
+
built, so you get eval questions without hand-authoring any. Reads the workspace's
|
|
418
|
+
declarative analytics model (read-only), translates each visible insight's buckets,
|
|
419
|
+
sorts and filters into an `expected_output.visualization` spec, then asks an LLM to
|
|
420
|
+
write the analyst question that chart answers. Because `expected_output` is copied from
|
|
421
|
+
a live object rather than authored, every question is grounded in the real LDM by
|
|
422
|
+
construction — the LLM only writes English.
|
|
423
|
+
|
|
424
|
+
**Setup:** host + token (read access to the workspace), and `OPENAI_API_KEY` plus the
|
|
425
|
+
`llm-judge` extra for the phrasing step (`uv add 'gooddata-eval[llm-judge]'`; skip both
|
|
426
|
+
with `--no-phrase`).
|
|
427
|
+
|
|
428
|
+
```bash
|
|
429
|
+
export GOODDATA_TOKEN='your-api-token'
|
|
430
|
+
|
|
431
|
+
# 1. see what a workspace yields before writing anything
|
|
432
|
+
gd-eval generate \
|
|
433
|
+
--host https://your.gooddata.cloud \
|
|
434
|
+
--workspace ecommerce_demo \
|
|
435
|
+
--dataset-name ecommerce \
|
|
436
|
+
--dry-run
|
|
437
|
+
|
|
438
|
+
# 2. generate, phrase, validate, and export
|
|
439
|
+
gd-eval generate \
|
|
440
|
+
--host https://your.gooddata.cloud \
|
|
441
|
+
--workspace ecommerce_demo \
|
|
442
|
+
--dataset-name ecommerce \
|
|
443
|
+
--dashboard dash_1_returns \
|
|
444
|
+
--out ./my-dataset \
|
|
445
|
+
--langfuse-out out/langfuse-dataset.json
|
|
446
|
+
|
|
447
|
+
# 3. run it
|
|
448
|
+
gd-eval run --host … --workspace ecommerce_demo --dataset ./my-dataset --model gpt-5.2
|
|
449
|
+
```
|
|
450
|
+
|
|
451
|
+
`--workspace` is where insights are read from; `--dataset-name` is the `dataset_name`
|
|
452
|
+
written into every item (and the default output folder).
|
|
453
|
+
|
|
454
|
+
| Flag | Effect |
|
|
455
|
+
|---|---|
|
|
456
|
+
| `--dashboard <id>` | restrict to insights on that dashboard (repeatable); default is the whole workspace |
|
|
457
|
+
| `--out <dir>` | output folder (default `./<dataset-name>`); this is what `gd-eval run --dataset` reads |
|
|
458
|
+
| `--snapshot-out` / `--snapshot-in` | save/replay the fetched model — replay needs no host, token, or network |
|
|
459
|
+
| `--langfuse-out <file>` | also write a Langfuse-importable dataset JSON |
|
|
460
|
+
| `--id-prefix` | prefix exported Langfuse item ids (they're unique per *project*, so re-importing an item under its original id is a 409) |
|
|
461
|
+
| `--no-phrase` | skip the LLM; emit mechanical `Show <title>` questions |
|
|
462
|
+
| `--phrase-model` | OpenAI model for phrasing (default `gpt-4o`) |
|
|
463
|
+
| `--no-viz-type` | always blank the expected chart type |
|
|
464
|
+
| `--enrich-ranked <N>` | additionally derive up to N ranked questions (see below); default 0 (off) |
|
|
465
|
+
| `--skip-ambiguous` | drop items naming something the model carries more than once; reported either way |
|
|
466
|
+
| `--min-questions` / `--min-shapes` / `--min-filtered` | quality gate, default 15, 3 and 1 |
|
|
467
|
+
|
|
468
|
+
### Ranked questions (`--enrich-ranked`)
|
|
469
|
+
|
|
470
|
+
Analysts sort in Analytical Designer and save the chart without persisting the sort, so
|
|
471
|
+
`sort_by`/`ranking_filter` coverage is near zero on most real models — the eval can
|
|
472
|
+
punish a spurious ranking but never confirm the agent builds a required one.
|
|
473
|
+
`--enrich-ranked N` fills that gap by *deriving* ranked items from the specs already
|
|
474
|
+
extracted. Adding a limit or a sort to a definition that executes cannot make it
|
|
475
|
+
unanswerable, and "the top 3 X by Y" has exactly one correct spec, so a derived item is
|
|
476
|
+
less ambiguous to grade than the insight it came from.
|
|
477
|
+
|
|
478
|
+
The budget is spent best-grounded first:
|
|
479
|
+
|
|
480
|
+
1. **Insights whose own title promised a ranking their definition never implemented** —
|
|
481
|
+
"Top Returned Reasons" saved with `sorts: []`. The direction comes from the title
|
|
482
|
+
(`highest`/`most`/`largest` vs `lowest`/`least`/`worst`) and the N too when it states
|
|
483
|
+
one; a title naming both ends names neither and is still skipped.
|
|
484
|
+
2. **Ranking filters added to a plain breakdown** — one metric, one non-date dimension,
|
|
485
|
+
no existing sort. N follows the dimension's element count, so a top-5 over six values
|
|
486
|
+
is never emitted.
|
|
487
|
+
3. **Sort-only variants**, which order without limiting.
|
|
488
|
+
|
|
489
|
+
Eligibility is deliberately narrow: two metrics leave "top 3 by what?" unanswered, a
|
|
490
|
+
second dimension leaves the N ambiguous between the pair and within a group, and a date
|
|
491
|
+
dimension turns the result into "top 3 months", which nobody asks. Variants are
|
|
492
|
+
deduplicated by resolved definition — differently-titled insights over one metric and
|
|
493
|
+
dimension would otherwise produce the same question twice — and bases are taken
|
|
494
|
+
round-robin by metric so one popular metric cannot become a third of the corpus.
|
|
495
|
+
|
|
496
|
+
Derived items carry `derived_from` (the insight id) and `derived_basis` (`title` when a
|
|
497
|
+
human's chart title asked for the ranking, `shape` when this generator chose to add
|
|
498
|
+
one), so a pass rate over each can be computed separately.
|
|
499
|
+
|
|
500
|
+
### Items that cannot say what they mean
|
|
501
|
+
|
|
502
|
+
Two classes of question are unwinnable however well the agent behaves, and both are
|
|
503
|
+
reported:
|
|
504
|
+
|
|
505
|
+
- **A name the model carries more than once.** One workspace has six labels all titled
|
|
506
|
+
"Product Title"; a question naming one cannot say which is meant, and a perfect chart
|
|
507
|
+
over the wrong one scores zero. `--skip-ambiguous` drops them; the count and the
|
|
508
|
+
offending names are printed either way.
|
|
509
|
+
- **A date granularity's cyclical twin.** `MONTH` walks consecutive calendar months,
|
|
510
|
+
`MONTH_OF_YEAR` stacks every January together. Date dimensions are therefore briefed
|
|
511
|
+
by what they do ("one point per calendar month over time, not month-of-year") and the
|
|
512
|
+
writer is told to say it in natural words while keeping the date dataset's name —
|
|
513
|
+
never as a label id in prose ("Order Created At - Month").
|
|
514
|
+
|
|
515
|
+
**The question must never contradict its own expected output.** Four rules enforce that:
|
|
516
|
+
|
|
517
|
+
- The writer is briefed on buckets, sorts and filters only — never the insight title,
|
|
518
|
+
and never the chart type. Titles routinely describe intent the definition doesn't
|
|
519
|
+
implement ("Products by Most Items Sold" over `sorts: []`).
|
|
520
|
+
- Every generated question is checked against its spec, and any hit is a hard error:
|
|
521
|
+
ranking words (`top`, `most`, `highest`, …) require a real sort or ranking filter;
|
|
522
|
+
filter words (`only`, `last quarter`, `in 2025`, …) require a real date or attribute
|
|
523
|
+
filter; a breakdown clause requires a non-empty `view_by`/`segment_by` and vice versa;
|
|
524
|
+
a metric may never be broken down by itself; and no template residue (`breakdown
|
|
525
|
+
dimension`, `{…}`) may survive. A violation is fed back once for a rewrite, then
|
|
526
|
+
dropped — and a drop fails the run.
|
|
527
|
+
- The writer's rules are built per insight, so an insight with no `view_by` is never
|
|
528
|
+
asked to name a breakdown at all.
|
|
529
|
+
- `type` is set only when the question actually names a chart form. An insight's
|
|
530
|
+
`visualizationUrl` records what a human clicked, not what the question constrains —
|
|
531
|
+
with one exception: a chart with no breakdown *must* name its form ("as a KPI", "as a
|
|
532
|
+
single number"). Without it the agent reads a bare "Show me Gross Revenue" as a metric
|
|
533
|
+
lookup, activates only its search skill and builds nothing.
|
|
534
|
+
|
|
535
|
+
Everything the writer sees is a display name (`Spend Amount`, `Merchant Name`), never a
|
|
536
|
+
raw URI, so questions read like a person wrote them.
|
|
537
|
+
|
|
538
|
+
**What it won't do.** Insights it can't express without guessing are skipped with a
|
|
539
|
+
printed reason, never approximated: derived (arithmetic/PoP) measures, measure-level
|
|
540
|
+
filters, `uris`-form attribute filters, unmapped chart types, hidden objects, and
|
|
541
|
+
insights whose title promises behaviour their definition lacks (though `--enrich-ranked`
|
|
542
|
+
implements a promised *ranking* rather than discarding it). If too few survive, the
|
|
543
|
+
quality gate fails the run rather than fabricating items to hit the minimum — point at
|
|
544
|
+
more dashboards, or lower `--min-questions`.
|
|
545
|
+
|
|
546
|
+
Every written item is validated as a `DatasetItem` with a scorable AAC visualization
|
|
547
|
+
before the command reports success.
|
|
548
|
+
|
|
549
|
+
---
|
|
550
|
+
|
|
356
551
|
## Dataset format
|
|
357
552
|
|
|
358
553
|
A dataset is a folder of `.json` files, one per question:
|
|
@@ -423,7 +618,8 @@ is the fraction of satisfied criteria.
|
|
|
423
618
|
|
|
424
619
|
### `[llm-judge]` — LLM-as-judge evaluators
|
|
425
620
|
|
|
426
|
-
`general_question` and `guardrail` items are scored by a GPT-4o judge
|
|
621
|
+
`general_question` and `guardrail` items are scored by a GPT-4o judge, and
|
|
622
|
+
`gd-eval generate` uses the same package to write question text.
|
|
427
623
|
Requires the OpenAI package and `OPENAI_API_KEY`:
|
|
428
624
|
|
|
429
625
|
```bash
|
|
@@ -432,13 +628,15 @@ uv add 'gooddata-eval[llm-judge]'
|
|
|
432
628
|
uv tool install 'gooddata-eval[llm-judge]'
|
|
433
629
|
```
|
|
434
630
|
|
|
435
|
-
Without `[llm-judge]`, those items are **skipped
|
|
631
|
+
Without `[llm-judge]`, those items are **skipped** and `gd-eval generate` needs
|
|
632
|
+
`--no-phrase`.
|
|
436
633
|
|
|
437
634
|
## Exit codes
|
|
438
635
|
|
|
439
636
|
| Code | Meaning |
|
|
440
637
|
|---|---|
|
|
441
638
|
| `0` | Run completed. Evaluation failures do **not** cause a non-zero exit. |
|
|
639
|
+
| `1` | `gd-eval generate` only: a quality gate failed, an item was dropped, or a written item failed validation. |
|
|
442
640
|
| `2` | Operational error: bad connection, missing model, unreadable dataset, missing credentials. |
|
|
443
641
|
|
|
444
642
|
## Scores (in JSON report and Langfuse)
|
|
@@ -16,7 +16,10 @@ Or install `gd-eval` as a standalone tool:
|
|
|
16
16
|
| Command | Description |
|
|
17
17
|
|---|---|
|
|
18
18
|
| `gd-eval run` | Run an evaluation dataset against one or more models. |
|
|
19
|
+
| `gd-eval report` | Render JSON report(s) as one self-contained HTML file. |
|
|
19
20
|
| `gd-eval models` | List LLM providers and models configured in the org. |
|
|
21
|
+
| `gd-eval generate` | Generate a `visualization` dataset from a workspace's existing insights. |
|
|
22
|
+
|
|
20
23
|
|
|
21
24
|
---
|
|
22
25
|
|
|
@@ -145,6 +148,8 @@ interleaves when K > 1, and per-item latencies rise, so they stop being clean si
|
|
|
145
148
|
| Flag | Description |
|
|
146
149
|
|---|---|
|
|
147
150
|
| `--json PATH` | Write a JSON report to this path. Always uses the nested `{models, runs, comparison}` shape even for a single model. |
|
|
151
|
+
| `--html PATH` | Write a self-contained HTML report to this path (same output as `gd-eval report`). |
|
|
152
|
+
| `--redact` | Make the HTML customer-safe. See `gd-eval report`. |
|
|
148
153
|
| `--quiet` | Suppress per-item progress. Per-model result tables and the comparison summary are still printed. |
|
|
149
154
|
| `--preserve-failed` | Keep failed conversations on the server instead of deleting them, so they can be inspected afterwards. Applies to the single-turn chat path; agentic kinds manage their own conversation lifecycle. |
|
|
150
155
|
| `--timers` | Print per-turn `[timer]` diagnostics — GoodData response, judge, and simulated-user seconds as they happen. Off by default: an 18-item `--runs 2` run emits ~72 lines and buries the progress output. The same measurements are always in the JSON report's `latency_breakdown_s`, so this only adds a live view. Also settable via `GD_EVAL_TIMERS=1`. |
|
|
@@ -305,6 +310,59 @@ linking ran. Pass `TAVERN_E2E_SKIP_TRACE_LINK=1` to opt out of linking altogethe
|
|
|
305
310
|
|
|
306
311
|
---
|
|
307
312
|
|
|
313
|
+
## `gd-eval report`
|
|
314
|
+
|
|
315
|
+
Turns JSON report(s) into one HTML file you can actually navigate. No server, no
|
|
316
|
+
credentials, no external assets — it opens over `file://`, attaches to a Jira issue and
|
|
317
|
+
survives a Slack thread.
|
|
318
|
+
|
|
319
|
+
```bash
|
|
320
|
+
# one run
|
|
321
|
+
gd-eval report results.json -o report.html
|
|
322
|
+
|
|
323
|
+
# several runs side by side -- each file becomes its own column
|
|
324
|
+
gd-eval report aug-21.json sep-07.json -o comparison.html --title "H200 regression check"
|
|
325
|
+
|
|
326
|
+
# customer-safe
|
|
327
|
+
gd-eval report results.json -o customer.html --redact
|
|
328
|
+
```
|
|
329
|
+
|
|
330
|
+
| Flag | Description |
|
|
331
|
+
|---|---|
|
|
332
|
+
| `-o, --out PATH` | Where to write the HTML. Required. |
|
|
333
|
+
| `--title TEXT` | Title shown in the report header. |
|
|
334
|
+
| `--redact` | Drop conversation/response ids and raw reasoning, and rename models to `Model A`, `Model B`, … Pass rate, per-item results, questions and latency survive. |
|
|
335
|
+
|
|
336
|
+
The report is a *view* over the JSON — it computes no numbers of its own. It gives you:
|
|
337
|
+
|
|
338
|
+
- **Run cards and a comparison table** — pass rate, quality, latency per run.
|
|
339
|
+
- **An item table** with a pass/fail column per run, so a model or run-over-run
|
|
340
|
+
regression is one glance rather than a hand-assembled spreadsheet.
|
|
341
|
+
- **An expression filter** for cross-cutting questions the fixed filters can't
|
|
342
|
+
anticipate, e.g. `d.filter_ranking_score === false` or
|
|
343
|
+
`d.expected_metric_uris.length > 1 && !d.metrics_correct`. Available variables:
|
|
344
|
+
`d` (the focused run's `detail`), `it` (its item), `i` (the row, `i.per[label]` for any
|
|
345
|
+
run), `q` (question), `kind`.
|
|
346
|
+
- **The conversation**, when the item ran the agentic multi-turn path — every turn in
|
|
347
|
+
order, with the simulated user marked apart from a real question, so you can see
|
|
348
|
+
whether the agent got there or was handed the answer.
|
|
349
|
+
- **A per-item drawer** — checks as pass/fail chips, expected vs actual side by side,
|
|
350
|
+
full reasoning, conversation/response ids.
|
|
351
|
+
- **A latency timeline** from `detail.latency_breakdown`, in execution order, one bar per
|
|
352
|
+
step. Clicking a step expands its full record, joined by `index`: the paragraph a
|
|
353
|
+
reasoning step was summarised from, or a tool call's arguments and result from
|
|
354
|
+
`detail.tool_calls`.
|
|
355
|
+
|
|
356
|
+
`--redact` additionally drops `transcript` and `tool_calls`: the exchange shows that the
|
|
357
|
+
simulated user is primed with the expected output, and a tool result carries
|
|
358
|
+
semantic-layer internals and real query rows. The turn count and the timeline shape
|
|
359
|
+
survive.
|
|
360
|
+
|
|
361
|
+
Passing several files keyed by file name is the whole run-over-run mechanism: no
|
|
362
|
+
database, no run registry, just the JSON files you already have on disk.
|
|
363
|
+
|
|
364
|
+
---
|
|
365
|
+
|
|
308
366
|
## `gd-eval models`
|
|
309
367
|
|
|
310
368
|
List all LLM providers and their models in the org. Marks the active model
|
|
@@ -325,6 +383,143 @@ gd-eval models \
|
|
|
325
383
|
|
|
326
384
|
---
|
|
327
385
|
|
|
386
|
+
## `gd-eval generate`
|
|
387
|
+
|
|
388
|
+
Reverse-engineers a `visualization` dataset out of the charts a customer has already
|
|
389
|
+
built, so you get eval questions without hand-authoring any. Reads the workspace's
|
|
390
|
+
declarative analytics model (read-only), translates each visible insight's buckets,
|
|
391
|
+
sorts and filters into an `expected_output.visualization` spec, then asks an LLM to
|
|
392
|
+
write the analyst question that chart answers. Because `expected_output` is copied from
|
|
393
|
+
a live object rather than authored, every question is grounded in the real LDM by
|
|
394
|
+
construction — the LLM only writes English.
|
|
395
|
+
|
|
396
|
+
**Setup:** host + token (read access to the workspace), and `OPENAI_API_KEY` plus the
|
|
397
|
+
`llm-judge` extra for the phrasing step (`uv add 'gooddata-eval[llm-judge]'`; skip both
|
|
398
|
+
with `--no-phrase`).
|
|
399
|
+
|
|
400
|
+
```bash
|
|
401
|
+
export GOODDATA_TOKEN='your-api-token'
|
|
402
|
+
|
|
403
|
+
# 1. see what a workspace yields before writing anything
|
|
404
|
+
gd-eval generate \
|
|
405
|
+
--host https://your.gooddata.cloud \
|
|
406
|
+
--workspace ecommerce_demo \
|
|
407
|
+
--dataset-name ecommerce \
|
|
408
|
+
--dry-run
|
|
409
|
+
|
|
410
|
+
# 2. generate, phrase, validate, and export
|
|
411
|
+
gd-eval generate \
|
|
412
|
+
--host https://your.gooddata.cloud \
|
|
413
|
+
--workspace ecommerce_demo \
|
|
414
|
+
--dataset-name ecommerce \
|
|
415
|
+
--dashboard dash_1_returns \
|
|
416
|
+
--out ./my-dataset \
|
|
417
|
+
--langfuse-out out/langfuse-dataset.json
|
|
418
|
+
|
|
419
|
+
# 3. run it
|
|
420
|
+
gd-eval run --host … --workspace ecommerce_demo --dataset ./my-dataset --model gpt-5.2
|
|
421
|
+
```
|
|
422
|
+
|
|
423
|
+
`--workspace` is where insights are read from; `--dataset-name` is the `dataset_name`
|
|
424
|
+
written into every item (and the default output folder).
|
|
425
|
+
|
|
426
|
+
| Flag | Effect |
|
|
427
|
+
|---|---|
|
|
428
|
+
| `--dashboard <id>` | restrict to insights on that dashboard (repeatable); default is the whole workspace |
|
|
429
|
+
| `--out <dir>` | output folder (default `./<dataset-name>`); this is what `gd-eval run --dataset` reads |
|
|
430
|
+
| `--snapshot-out` / `--snapshot-in` | save/replay the fetched model — replay needs no host, token, or network |
|
|
431
|
+
| `--langfuse-out <file>` | also write a Langfuse-importable dataset JSON |
|
|
432
|
+
| `--id-prefix` | prefix exported Langfuse item ids (they're unique per *project*, so re-importing an item under its original id is a 409) |
|
|
433
|
+
| `--no-phrase` | skip the LLM; emit mechanical `Show <title>` questions |
|
|
434
|
+
| `--phrase-model` | OpenAI model for phrasing (default `gpt-4o`) |
|
|
435
|
+
| `--no-viz-type` | always blank the expected chart type |
|
|
436
|
+
| `--enrich-ranked <N>` | additionally derive up to N ranked questions (see below); default 0 (off) |
|
|
437
|
+
| `--skip-ambiguous` | drop items naming something the model carries more than once; reported either way |
|
|
438
|
+
| `--min-questions` / `--min-shapes` / `--min-filtered` | quality gate, default 15, 3 and 1 |
|
|
439
|
+
|
|
440
|
+
### Ranked questions (`--enrich-ranked`)
|
|
441
|
+
|
|
442
|
+
Analysts sort in Analytical Designer and save the chart without persisting the sort, so
|
|
443
|
+
`sort_by`/`ranking_filter` coverage is near zero on most real models — the eval can
|
|
444
|
+
punish a spurious ranking but never confirm the agent builds a required one.
|
|
445
|
+
`--enrich-ranked N` fills that gap by *deriving* ranked items from the specs already
|
|
446
|
+
extracted. Adding a limit or a sort to a definition that executes cannot make it
|
|
447
|
+
unanswerable, and "the top 3 X by Y" has exactly one correct spec, so a derived item is
|
|
448
|
+
less ambiguous to grade than the insight it came from.
|
|
449
|
+
|
|
450
|
+
The budget is spent best-grounded first:
|
|
451
|
+
|
|
452
|
+
1. **Insights whose own title promised a ranking their definition never implemented** —
|
|
453
|
+
"Top Returned Reasons" saved with `sorts: []`. The direction comes from the title
|
|
454
|
+
(`highest`/`most`/`largest` vs `lowest`/`least`/`worst`) and the N too when it states
|
|
455
|
+
one; a title naming both ends names neither and is still skipped.
|
|
456
|
+
2. **Ranking filters added to a plain breakdown** — one metric, one non-date dimension,
|
|
457
|
+
no existing sort. N follows the dimension's element count, so a top-5 over six values
|
|
458
|
+
is never emitted.
|
|
459
|
+
3. **Sort-only variants**, which order without limiting.
|
|
460
|
+
|
|
461
|
+
Eligibility is deliberately narrow: two metrics leave "top 3 by what?" unanswered, a
|
|
462
|
+
second dimension leaves the N ambiguous between the pair and within a group, and a date
|
|
463
|
+
dimension turns the result into "top 3 months", which nobody asks. Variants are
|
|
464
|
+
deduplicated by resolved definition — differently-titled insights over one metric and
|
|
465
|
+
dimension would otherwise produce the same question twice — and bases are taken
|
|
466
|
+
round-robin by metric so one popular metric cannot become a third of the corpus.
|
|
467
|
+
|
|
468
|
+
Derived items carry `derived_from` (the insight id) and `derived_basis` (`title` when a
|
|
469
|
+
human's chart title asked for the ranking, `shape` when this generator chose to add
|
|
470
|
+
one), so a pass rate over each can be computed separately.
|
|
471
|
+
|
|
472
|
+
### Items that cannot say what they mean
|
|
473
|
+
|
|
474
|
+
Two classes of question are unwinnable however well the agent behaves, and both are
|
|
475
|
+
reported:
|
|
476
|
+
|
|
477
|
+
- **A name the model carries more than once.** One workspace has six labels all titled
|
|
478
|
+
"Product Title"; a question naming one cannot say which is meant, and a perfect chart
|
|
479
|
+
over the wrong one scores zero. `--skip-ambiguous` drops them; the count and the
|
|
480
|
+
offending names are printed either way.
|
|
481
|
+
- **A date granularity's cyclical twin.** `MONTH` walks consecutive calendar months,
|
|
482
|
+
`MONTH_OF_YEAR` stacks every January together. Date dimensions are therefore briefed
|
|
483
|
+
by what they do ("one point per calendar month over time, not month-of-year") and the
|
|
484
|
+
writer is told to say it in natural words while keeping the date dataset's name —
|
|
485
|
+
never as a label id in prose ("Order Created At - Month").
|
|
486
|
+
|
|
487
|
+
**The question must never contradict its own expected output.** Four rules enforce that:
|
|
488
|
+
|
|
489
|
+
- The writer is briefed on buckets, sorts and filters only — never the insight title,
|
|
490
|
+
and never the chart type. Titles routinely describe intent the definition doesn't
|
|
491
|
+
implement ("Products by Most Items Sold" over `sorts: []`).
|
|
492
|
+
- Every generated question is checked against its spec, and any hit is a hard error:
|
|
493
|
+
ranking words (`top`, `most`, `highest`, …) require a real sort or ranking filter;
|
|
494
|
+
filter words (`only`, `last quarter`, `in 2025`, …) require a real date or attribute
|
|
495
|
+
filter; a breakdown clause requires a non-empty `view_by`/`segment_by` and vice versa;
|
|
496
|
+
a metric may never be broken down by itself; and no template residue (`breakdown
|
|
497
|
+
dimension`, `{…}`) may survive. A violation is fed back once for a rewrite, then
|
|
498
|
+
dropped — and a drop fails the run.
|
|
499
|
+
- The writer's rules are built per insight, so an insight with no `view_by` is never
|
|
500
|
+
asked to name a breakdown at all.
|
|
501
|
+
- `type` is set only when the question actually names a chart form. An insight's
|
|
502
|
+
`visualizationUrl` records what a human clicked, not what the question constrains —
|
|
503
|
+
with one exception: a chart with no breakdown *must* name its form ("as a KPI", "as a
|
|
504
|
+
single number"). Without it the agent reads a bare "Show me Gross Revenue" as a metric
|
|
505
|
+
lookup, activates only its search skill and builds nothing.
|
|
506
|
+
|
|
507
|
+
Everything the writer sees is a display name (`Spend Amount`, `Merchant Name`), never a
|
|
508
|
+
raw URI, so questions read like a person wrote them.
|
|
509
|
+
|
|
510
|
+
**What it won't do.** Insights it can't express without guessing are skipped with a
|
|
511
|
+
printed reason, never approximated: derived (arithmetic/PoP) measures, measure-level
|
|
512
|
+
filters, `uris`-form attribute filters, unmapped chart types, hidden objects, and
|
|
513
|
+
insights whose title promises behaviour their definition lacks (though `--enrich-ranked`
|
|
514
|
+
implements a promised *ranking* rather than discarding it). If too few survive, the
|
|
515
|
+
quality gate fails the run rather than fabricating items to hit the minimum — point at
|
|
516
|
+
more dashboards, or lower `--min-questions`.
|
|
517
|
+
|
|
518
|
+
Every written item is validated as a `DatasetItem` with a scorable AAC visualization
|
|
519
|
+
before the command reports success.
|
|
520
|
+
|
|
521
|
+
---
|
|
522
|
+
|
|
328
523
|
## Dataset format
|
|
329
524
|
|
|
330
525
|
A dataset is a folder of `.json` files, one per question:
|
|
@@ -395,7 +590,8 @@ is the fraction of satisfied criteria.
|
|
|
395
590
|
|
|
396
591
|
### `[llm-judge]` — LLM-as-judge evaluators
|
|
397
592
|
|
|
398
|
-
`general_question` and `guardrail` items are scored by a GPT-4o judge
|
|
593
|
+
`general_question` and `guardrail` items are scored by a GPT-4o judge, and
|
|
594
|
+
`gd-eval generate` uses the same package to write question text.
|
|
399
595
|
Requires the OpenAI package and `OPENAI_API_KEY`:
|
|
400
596
|
|
|
401
597
|
```bash
|
|
@@ -404,13 +600,15 @@ uv add 'gooddata-eval[llm-judge]'
|
|
|
404
600
|
uv tool install 'gooddata-eval[llm-judge]'
|
|
405
601
|
```
|
|
406
602
|
|
|
407
|
-
Without `[llm-judge]`, those items are **skipped
|
|
603
|
+
Without `[llm-judge]`, those items are **skipped** and `gd-eval generate` needs
|
|
604
|
+
`--no-phrase`.
|
|
408
605
|
|
|
409
606
|
## Exit codes
|
|
410
607
|
|
|
411
608
|
| Code | Meaning |
|
|
412
609
|
|---|---|
|
|
413
610
|
| `0` | Run completed. Evaluation failures do **not** cause a non-zero exit. |
|
|
611
|
+
| `1` | `gd-eval generate` only: a quality gate failed, an item was dropped, or a written item failed validation. |
|
|
414
612
|
| `2` | Operational error: bad connection, missing model, unreadable dataset, missing credentials. |
|
|
415
613
|
|
|
416
614
|
## Scores (in JSON report and Langfuse)
|