gooddata-eval 1.74.1.dev5__tar.gz → 1.75.1.dev1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (131) hide show
  1. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/AGENTS.md +119 -2
  2. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/PKG-INFO +202 -4
  3. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/README.md +200 -2
  4. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/pyproject.toml +5 -2
  5. gooddata_eval-1.75.1.dev1/scripts/verify_guardrail_refusal_criteria.py +126 -0
  6. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/cli/agentic_runner.py +18 -0
  7. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/cli/main.py +152 -3
  8. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/__init__.py +14 -0
  9. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/alert_skill.py +2 -2
  10. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/conversation.py +2 -2
  11. gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/agentic/dashboard_skill.py +742 -0
  12. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/guardrail.py +5 -3
  13. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/metric_skill.py +2 -2
  14. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/visualization.py +2 -2
  15. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/chat/sse_client.py +126 -17
  16. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/config.py +6 -0
  17. gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/dataset/from_insights.py +855 -0
  18. gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/evaluators/_guardrail_criteria.py +39 -0
  19. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/general_question.py +2 -4
  20. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/guardrail.py +8 -7
  21. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/search_tool.py +2 -4
  22. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/visualization.py +34 -9
  23. gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/granularity.py +65 -0
  24. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/models.py +80 -1
  25. gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/reporting/html_report.py +121 -0
  26. gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/reporting/report_template.html +452 -0
  27. gooddata_eval-1.75.1.dev1/src/gooddata_eval/core/scoring.py +393 -0
  28. gooddata_eval-1.75.1.dev1/tests/conftest.py +45 -0
  29. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_alert_skill.py +2 -0
  30. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_conversation.py +2 -0
  31. gooddata_eval-1.75.1.dev1/tests/test_agentic_dashboard_skill.py +679 -0
  32. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_guardrail.py +2 -0
  33. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_metric_skill.py +2 -0
  34. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_runner.py +5 -0
  35. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_visualization.py +8 -0
  36. gooddata_eval-1.75.1.dev1/tests/test_from_insights.py +1194 -0
  37. gooddata_eval-1.75.1.dev1/tests/test_guardrail_criteria.py +59 -0
  38. gooddata_eval-1.75.1.dev1/tests/test_html_report.py +133 -0
  39. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_models.py +56 -0
  40. gooddata_eval-1.75.1.dev1/tests/test_scoring.py +539 -0
  41. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_sse_client.py +83 -4
  42. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_trace_linker.py +1 -0
  43. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_visualization_evaluator.py +42 -3
  44. gooddata_eval-1.74.1.dev5/src/gooddata_eval/core/scoring.py +0 -221
  45. gooddata_eval-1.74.1.dev5/tests/conftest.py +0 -22
  46. gooddata_eval-1.74.1.dev5/tests/test_scoring.py +0 -207
  47. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/.gitignore +0 -0
  48. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/CLAUDE.md +0 -0
  49. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/LICENSE.txt +0 -0
  50. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/Makefile +0 -0
  51. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/__init__.py +0 -0
  52. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/_version.py +0 -0
  53. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/cli/__init__.py +0 -0
  54. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/__init__.py +0 -0
  55. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/_output.py +0 -0
  56. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/_catalog.py +0 -0
  57. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/_gate.py +0 -0
  58. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/_langfuse.py +0 -0
  59. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/_trace_linker.py +0 -0
  60. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/general_question.py +0 -0
  61. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/kda_skill.py +0 -0
  62. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/agentic/search_tool.py +0 -0
  63. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/chat/__init__.py +0 -0
  64. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/chat/render.py +0 -0
  65. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/connection.py +0 -0
  66. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/dataset/__init__.py +0 -0
  67. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/dataset/langfuse_source.py +0 -0
  68. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/dataset/local.py +0 -0
  69. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/__init__.py +0 -0
  70. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/_deep_subset.py +0 -0
  71. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/_llm_judge.py +0 -0
  72. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/_maql.py +0 -0
  73. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/_text_utils.py +0 -0
  74. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/alert_skill.py +0 -0
  75. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/base.py +0 -0
  76. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/metric_skill.py +0 -0
  77. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/evaluators/summary.py +0 -0
  78. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/__init__.py +0 -0
  79. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/_env.py +0 -0
  80. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/client.py +0 -0
  81. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/experiment.py +0 -0
  82. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/observations.py +0 -0
  83. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/otlp.py +0 -0
  84. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/langfuse/sink.py +0 -0
  85. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/reporting/__init__.py +0 -0
  86. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/reporting/console.py +0 -0
  87. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/reporting/json_report.py +0 -0
  88. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/runner.py +0 -0
  89. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/summary/__init__.py +0 -0
  90. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/summary/http_client.py +0 -0
  91. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/timing.py +0 -0
  92. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/src/gooddata_eval/core/workspace.py +0 -0
  93. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/__init__.py +0 -0
  94. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/_fake_langfuse.py +0 -0
  95. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/fixtures/sample_dataset/metric_skill_create.json +0 -0
  96. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/fixtures/sample_dataset/visualization_revenue.json +0 -0
  97. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/fixtures/sse_visualization_stream.txt +0 -0
  98. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_gate.py +0 -0
  99. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_general_question.py +0 -0
  100. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_kda_skill.py +0 -0
  101. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_langfuse_trace.py +0 -0
  102. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_observe_experiment.py +0 -0
  103. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_run_context.py +0 -0
  104. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_agentic_search_tool.py +0 -0
  105. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_alert_skill_evaluator.py +0 -0
  106. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_chat_render.py +0 -0
  107. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_cli.py +0 -0
  108. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_connection.py +0 -0
  109. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_deep_subset.py +0 -0
  110. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_fake_langfuse.py +0 -0
  111. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_client.py +0 -0
  112. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_e2e_fake_server.py +0 -0
  113. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_env.py +0 -0
  114. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_experiment.py +0 -0
  115. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_observations.py +0 -0
  116. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_otlp.py +0 -0
  117. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_sink.py +0 -0
  118. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_langfuse_source.py +0 -0
  119. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_llm_judge.py +0 -0
  120. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_local_loader.py +0 -0
  121. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_maql_normalize.py +0 -0
  122. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_metric_skill_evaluator.py +0 -0
  123. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_reporting.py +0 -0
  124. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_runner.py +0 -0
  125. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_search_tool_evaluator.py +0 -0
  126. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_summary_client.py +0 -0
  127. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_summary_evaluator.py +0 -0
  128. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_text_evaluators.py +0 -0
  129. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_timing.py +0 -0
  130. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tests/test_workspace.py +0 -0
  131. {gooddata_eval-1.74.1.dev5 → gooddata_eval-1.75.1.dev1}/tox.ini +0 -0
@@ -9,7 +9,7 @@ experiment. The newest and most actively developed package in the repo.
9
9
 
10
10
  ## Owns
11
11
 
12
- - The `gd-eval` CLI (`gd-eval run`, `gd-eval models`)
12
+ - The `gd-eval` CLI (`generate`, `run`, `report`, `models`)
13
13
  - Dataset loading and the evaluation run loop
14
14
  - Per-capability evaluators and their scoring
15
15
  - Result reporting, and pushing experiments, scores and trace links to Langfuse
@@ -27,7 +27,7 @@ experiment. The newest and most actively developed package in the repo.
27
27
  | `core/agentic/` | multi-turn agentic evaluation per capability, **plus** all Langfuse trace polling and linking (`_langfuse.py`, `_trace_linker.py`) |
28
28
  | `core/chat/` | SSE client for the agent's streaming chat endpoint |
29
29
  | `core/summary/` | HTTP client for the dedicated dashboard-summary endpoint — a single-shot chat backend, not reporting |
30
- | `core/dataset/` | dataset format and loading |
30
+ | `core/dataset/` | dataset format, loading, and `from_insights.py` — dataset generation from a workspace's real insights |
31
31
  | `core/evaluators/` | single-shot evaluators and their registry |
32
32
  | `core/langfuse/` | the whole Langfuse v4 client: `_env` (base URL + credentials), `otlp` (OTLP/JSON encoding), `experiment` (root-span construction, score targets), `observations` (trace reads), `client` (httpx calls), `sink` (single-shot results as experiments) |
33
33
  | `core/reporting/` | console and JSON output rendering |
@@ -61,6 +61,106 @@ its own shape. `test_kind` on the item is what labels the result, not the evalua
61
61
  which is why `knowledge_question` can reuse `GeneralQuestionEvaluator` verbatim.
62
62
  `dashboard_summary` items additionally need `summary_input`.
63
63
 
64
+ ## Running the pipeline
65
+
66
+ Four subcommands, in the order you use them. Everything runs through `uv`; never a bare
67
+ `python`. There is no build step -- `uv run` syncs the environment from `uv.lock` on first
68
+ use, so a fresh clone needs nothing but:
69
+
70
+ ```bash
71
+ uv run --package gooddata-eval gd-eval <subcommand> --help
72
+ ```
73
+
74
+ `openai` is an optional extra (`llm-judge`) so the published package stays installable
75
+ without it, but the `dev` dependency group pulls it in, which is why a plain `uv run` here
76
+ has the phrasing step and the LLM judge. Installing `gooddata-eval` from PyPI does not --
77
+ there the extra is explicit, and every `openai` import site is guarded or deferred.
78
+
79
+ Connection is the same for every subcommand that talks to the platform: `--host` +
80
+ `--token`, or `GOODDATA_TOKEN` in the environment, or `--profile <name>` reading
81
+ `~/.gooddata/profiles.yaml`. Precedence is flags > env > profile.
82
+
83
+ ### 1. `generate` — build a dataset from a workspace
84
+
85
+ Reverse-engineers `visualization` items out of the charts a workspace already has, so the
86
+ expected output is copied from a live object rather than invented. Needs `OPENAI_API_KEY`
87
+ for the phrasing step, or `--no-phrase` to emit mechanical `Show <title>` questions.
88
+
89
+ ```bash
90
+ uv run --package gooddata-eval gd-eval generate \
91
+ --host "$GOODDATA_HOST" --workspace "$WORKSPACE_ID" \
92
+ --dataset-name ecommerce --out ./datasets/ecommerce \
93
+ --snapshot-out /tmp/ws.json \
94
+ --phrase-model gpt-4o --skip-ambiguous
95
+ ```
96
+
97
+ Iterate offline instead of re-fetching: `--snapshot-out` writes everything the generator
98
+ read as one JSON file, and `--snapshot-in` replays it with no host, token or network. Add
99
+ `--dry-run` to print the shape counts and each spec's brief without writing anything —
100
+ the fastest way to see what a workspace yields.
101
+
102
+ Quality gates fail the command (exit 1) below `--min-questions` (15) or `--min-filtered` (1). Lower them for a smoke test; do not lower them to ship a dataset.
103
+ `--langfuse-out` additionally writes a Langfuse-importable file, and `--id-prefix` rewrites
104
+ ids on that export only, because Langfuse item ids are unique per project.
105
+
106
+ Insights the AAC spec cannot express without guessing are skipped with a printed reason
107
+ (`SKIP <id>: derived measure (previousPeriodMeasure)`). Read those — they are the
108
+ generator telling you what it refused to invent, not noise.
109
+
110
+ ### 2. `run` — evaluate
111
+
112
+ ```bash
113
+ uv run --package gooddata-eval gd-eval run \
114
+ --host "$GOODDATA_HOST" --workspace "$WORKSPACE_ID" \
115
+ --dataset ./datasets/ecommerce --kind visualization \
116
+ --model gpt-5.2 --model ProviderName/gpt-4o \
117
+ --runs 3 --gate power --concurrency 4 \
118
+ --json ./results/run.json --html ./results/run.html
119
+ ```
120
+
121
+ `--dataset` reads a local folder; `--langfuse-dataset` pulls one by name instead. `--kind`
122
+ only supplies a default for items that do not carry their own `test_kind`. Repeat
123
+ `--model` to compare models in one run. `--runs` with `--gate power` measures stability
124
+ (every run must pass) rather than pass@K. `--langfuse` pushes the run as a scored
125
+ experiment, needing `LANGFUSE_HOST`, `LANGFUSE_PUBLIC_KEY` and `LANGFUSE_SECRET_KEY`.
126
+
127
+ `--concurrency` is capped for you where it matters: kinds that create workspace objects
128
+ run one at a time regardless, see the parallel-safety gotcha below.
129
+
130
+ ### 3. `report` — compare runs
131
+
132
+ ```bash
133
+ uv run --package gooddata-eval gd-eval report \
134
+ ./results/*.json -o ./results/comparison.html --title "luna vs 4o" --redact
135
+ ```
136
+
137
+ Several JSON reports become side-by-side columns keyed by file name. `--redact` is the
138
+ customer-safe form: conversation ids, response ids and raw reasoning dropped, model names
139
+ replaced with "Model A", "Model B".
140
+
141
+ ### 4. `models` — what the org has configured
142
+
143
+ ```bash
144
+ uv run --package gooddata-eval gd-eval models --host "$GOODDATA_HOST"
145
+ ```
146
+
147
+ Run this before guessing a `--model` string.
148
+
149
+ ### Environment
150
+
151
+ | Variable | Used by |
152
+ |---|---|
153
+ | `GOODDATA_TOKEN` | every platform-facing subcommand |
154
+ | `OPENAI_API_KEY` | `generate` phrasing, and the LLM-as-judge evaluators |
155
+ | `GD_EVAL_JUDGE_MODEL` | judge model, same as `--judge-model` |
156
+ | `GD_EVAL_AGENT_ID` | which agent to drive, same as `--agent-id` |
157
+ | `LANGFUSE_HOST`, `LANGFUSE_PUBLIC_KEY`, `LANGFUSE_SECRET_KEY` | `--langfuse`, `--langfuse-dataset` |
158
+ | `GOODDATA_EVAL_CHAT_*` | SSE retry, backoff and timeout knobs |
159
+ | `GD_EVAL_TIMERS` | same as `--timers` |
160
+
161
+ A gitignored `.env` at the repo root is the normal place for these; load it with
162
+ `set -a && . ./.env && set +a` before the command.
163
+
64
164
  ## Gotchas
65
165
 
66
166
  **Adding an evaluator is a registry change, not a naming convention.** Single-shot kinds go
@@ -89,6 +189,23 @@ ingestion has no pass/fail signal and inflates or misattributes per-item latency
89
189
  (`run_trace_link_inline` is the synchronous alternative). Do not "fix" a slow item by
90
190
  making trace scoring synchronous again.
91
191
 
192
+ **A generated item's `expected_output` is copied, never invented — keep it that way.**
193
+ `core/dataset/from_insights.py` converts each insight with the platform's own
194
+ `declarative_visualization_to_aac()` (from `gooddata-code-convertors`, via `gooddata-sdk`),
195
+ so the mapping is not ours to get wrong. What is ours is deciding what the evaluator cannot
196
+ yet score — derived measures and measure-level filters convert fine and then compare wrong,
197
+ so they are skipped with a printed reason — and stripping the no-op filters AD saves for an
198
+ "All" selection, which would otherwise let a question claim a filter its chart lacks. Teach
199
+ the comparator about a construct and the matching skip can go; do not make one convert by
200
+ hand. Chart type names are the convertor's, which are also the agent's — do not rename them. One
201
+ granularity is patched in `CONVERTOR_GRANULARITY_FIXES`: `week_us` → `WEEK_US` is a convertor
202
+ bug, the platform enum is `WEEK`.
203
+
204
+ **The snapshot is a plain-JSON contract.** `--snapshot-in`/`--snapshot-out` is what makes
205
+ the generator testable offline and iterable without re-fetching, and it is why the
206
+ generator reads the declarative analytics model rather than `sdk.visualizations`. Anything
207
+ that changes the fetch shape invalidates every saved snapshot.
208
+
92
209
  **Scoring weights do not sum to 1.** `quality_score` is the fraction of boolean-valued keys
93
210
  in `best_detail` that are true, falling back to `pass_at_k` when there are none (text
94
211
  evaluators). `value_score` is `0.6 * quality + 0.2 * speed` — the 0.8 total is what the
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: gooddata-eval
3
- Version: 1.74.1.dev5
3
+ Version: 1.75.1.dev1
4
4
  Summary: Evaluate the GoodData AI agent against your own questions and models.
5
5
  Project-URL: Source, https://github.com/gooddata/gooddata-python-sdk
6
6
  Author-email: GoodData <support@gooddata.com>
@@ -17,7 +17,7 @@ Classifier: Topic :: Scientific/Engineering
17
17
  Classifier: Topic :: Software Development
18
18
  Classifier: Typing :: Typed
19
19
  Requires-Python: >=3.10
20
- Requires-Dist: gooddata-sdk~=1.74.1.dev5
20
+ Requires-Dist: gooddata-sdk~=1.75.1.dev1
21
21
  Requires-Dist: httpx<1.0,>=0.27
22
22
  Requires-Dist: orjson<4.0.0,>=3.9.15
23
23
  Requires-Dist: pydantic<3.0,>=2.6
@@ -44,7 +44,10 @@ Or install `gd-eval` as a standalone tool:
44
44
  | Command | Description |
45
45
  |---|---|
46
46
  | `gd-eval run` | Run an evaluation dataset against one or more models. |
47
+ | `gd-eval report` | Render JSON report(s) as one self-contained HTML file. |
47
48
  | `gd-eval models` | List LLM providers and models configured in the org. |
49
+ | `gd-eval generate` | Generate a `visualization` dataset from a workspace's existing insights. |
50
+
48
51
 
49
52
  ---
50
53
 
@@ -173,6 +176,8 @@ interleaves when K > 1, and per-item latencies rise, so they stop being clean si
173
176
  | Flag | Description |
174
177
  |---|---|
175
178
  | `--json PATH` | Write a JSON report to this path. Always uses the nested `{models, runs, comparison}` shape even for a single model. |
179
+ | `--html PATH` | Write a self-contained HTML report to this path (same output as `gd-eval report`). |
180
+ | `--redact` | Make the HTML customer-safe. See `gd-eval report`. |
176
181
  | `--quiet` | Suppress per-item progress. Per-model result tables and the comparison summary are still printed. |
177
182
  | `--preserve-failed` | Keep failed conversations on the server instead of deleting them, so they can be inspected afterwards. Applies to the single-turn chat path; agentic kinds manage their own conversation lifecycle. |
178
183
  | `--timers` | Print per-turn `[timer]` diagnostics — GoodData response, judge, and simulated-user seconds as they happen. Off by default: an 18-item `--runs 2` run emits ~72 lines and buries the progress output. The same measurements are always in the JSON report's `latency_breakdown_s`, so this only adds a live view. Also settable via `GD_EVAL_TIMERS=1`. |
@@ -333,6 +338,59 @@ linking ran. Pass `TAVERN_E2E_SKIP_TRACE_LINK=1` to opt out of linking altogethe
333
338
 
334
339
  ---
335
340
 
341
+ ## `gd-eval report`
342
+
343
+ Turns JSON report(s) into one HTML file you can actually navigate. No server, no
344
+ credentials, no external assets — it opens over `file://`, attaches to a Jira issue and
345
+ survives a Slack thread.
346
+
347
+ ```bash
348
+ # one run
349
+ gd-eval report results.json -o report.html
350
+
351
+ # several runs side by side -- each file becomes its own column
352
+ gd-eval report aug-21.json sep-07.json -o comparison.html --title "H200 regression check"
353
+
354
+ # customer-safe
355
+ gd-eval report results.json -o customer.html --redact
356
+ ```
357
+
358
+ | Flag | Description |
359
+ |---|---|
360
+ | `-o, --out PATH` | Where to write the HTML. Required. |
361
+ | `--title TEXT` | Title shown in the report header. |
362
+ | `--redact` | Drop conversation/response ids and raw reasoning, and rename models to `Model A`, `Model B`, … Pass rate, per-item results, questions and latency survive. |
363
+
364
+ The report is a *view* over the JSON — it computes no numbers of its own. It gives you:
365
+
366
+ - **Run cards and a comparison table** — pass rate, quality, latency per run.
367
+ - **An item table** with a pass/fail column per run, so a model or run-over-run
368
+ regression is one glance rather than a hand-assembled spreadsheet.
369
+ - **An expression filter** for cross-cutting questions the fixed filters can't
370
+ anticipate, e.g. `d.filter_ranking_score === false` or
371
+ `d.expected_metric_uris.length > 1 && !d.metrics_correct`. Available variables:
372
+ `d` (the focused run's `detail`), `it` (its item), `i` (the row, `i.per[label]` for any
373
+ run), `q` (question), `kind`.
374
+ - **The conversation**, when the item ran the agentic multi-turn path — every turn in
375
+ order, with the simulated user marked apart from a real question, so you can see
376
+ whether the agent got there or was handed the answer.
377
+ - **A per-item drawer** — checks as pass/fail chips, expected vs actual side by side,
378
+ full reasoning, conversation/response ids.
379
+ - **A latency timeline** from `detail.latency_breakdown`, in execution order, one bar per
380
+ step. Clicking a step expands its full record, joined by `index`: the paragraph a
381
+ reasoning step was summarised from, or a tool call's arguments and result from
382
+ `detail.tool_calls`.
383
+
384
+ `--redact` additionally drops `transcript` and `tool_calls`: the exchange shows that the
385
+ simulated user is primed with the expected output, and a tool result carries
386
+ semantic-layer internals and real query rows. The turn count and the timeline shape
387
+ survive.
388
+
389
+ Passing several files keyed by file name is the whole run-over-run mechanism: no
390
+ database, no run registry, just the JSON files you already have on disk.
391
+
392
+ ---
393
+
336
394
  ## `gd-eval models`
337
395
 
338
396
  List all LLM providers and their models in the org. Marks the active model
@@ -353,6 +411,143 @@ gd-eval models \
353
411
 
354
412
  ---
355
413
 
414
+ ## `gd-eval generate`
415
+
416
+ Reverse-engineers a `visualization` dataset out of the charts a customer has already
417
+ built, so you get eval questions without hand-authoring any. Reads the workspace's
418
+ declarative analytics model (read-only), translates each visible insight's buckets,
419
+ sorts and filters into an `expected_output.visualization` spec, then asks an LLM to
420
+ write the analyst question that chart answers. Because `expected_output` is copied from
421
+ a live object rather than authored, every question is grounded in the real LDM by
422
+ construction — the LLM only writes English.
423
+
424
+ **Setup:** host + token (read access to the workspace), and `OPENAI_API_KEY` plus the
425
+ `llm-judge` extra for the phrasing step (`uv add 'gooddata-eval[llm-judge]'`; skip both
426
+ with `--no-phrase`).
427
+
428
+ ```bash
429
+ export GOODDATA_TOKEN='your-api-token'
430
+
431
+ # 1. see what a workspace yields before writing anything
432
+ gd-eval generate \
433
+ --host https://your.gooddata.cloud \
434
+ --workspace ecommerce_demo \
435
+ --dataset-name ecommerce \
436
+ --dry-run
437
+
438
+ # 2. generate, phrase, validate, and export
439
+ gd-eval generate \
440
+ --host https://your.gooddata.cloud \
441
+ --workspace ecommerce_demo \
442
+ --dataset-name ecommerce \
443
+ --dashboard dash_1_returns \
444
+ --out ./my-dataset \
445
+ --langfuse-out out/langfuse-dataset.json
446
+
447
+ # 3. run it
448
+ gd-eval run --host … --workspace ecommerce_demo --dataset ./my-dataset --model gpt-5.2
449
+ ```
450
+
451
+ `--workspace` is where insights are read from; `--dataset-name` is the `dataset_name`
452
+ written into every item (and the default output folder).
453
+
454
+ | Flag | Effect |
455
+ |---|---|
456
+ | `--dashboard <id>` | restrict to insights on that dashboard (repeatable); default is the whole workspace |
457
+ | `--out <dir>` | output folder (default `./<dataset-name>`); this is what `gd-eval run --dataset` reads |
458
+ | `--snapshot-out` / `--snapshot-in` | save/replay the fetched model — replay needs no host, token, or network |
459
+ | `--langfuse-out <file>` | also write a Langfuse-importable dataset JSON |
460
+ | `--id-prefix` | prefix exported Langfuse item ids (they're unique per *project*, so re-importing an item under its original id is a 409) |
461
+ | `--no-phrase` | skip the LLM; emit mechanical `Show <title>` questions |
462
+ | `--phrase-model` | OpenAI model for phrasing (default `gpt-4o`) |
463
+ | `--no-viz-type` | always blank the expected chart type |
464
+ | `--enrich-ranked <N>` | additionally derive up to N ranked questions (see below); default 0 (off) |
465
+ | `--skip-ambiguous` | drop items naming something the model carries more than once; reported either way |
466
+ | `--min-questions` / `--min-shapes` / `--min-filtered` | quality gate, default 15, 3 and 1 |
467
+
468
+ ### Ranked questions (`--enrich-ranked`)
469
+
470
+ Analysts sort in Analytical Designer and save the chart without persisting the sort, so
471
+ `sort_by`/`ranking_filter` coverage is near zero on most real models — the eval can
472
+ punish a spurious ranking but never confirm the agent builds a required one.
473
+ `--enrich-ranked N` fills that gap by *deriving* ranked items from the specs already
474
+ extracted. Adding a limit or a sort to a definition that executes cannot make it
475
+ unanswerable, and "the top 3 X by Y" has exactly one correct spec, so a derived item is
476
+ less ambiguous to grade than the insight it came from.
477
+
478
+ The budget is spent best-grounded first:
479
+
480
+ 1. **Insights whose own title promised a ranking their definition never implemented** —
481
+ "Top Returned Reasons" saved with `sorts: []`. The direction comes from the title
482
+ (`highest`/`most`/`largest` vs `lowest`/`least`/`worst`) and the N too when it states
483
+ one; a title naming both ends names neither and is still skipped.
484
+ 2. **Ranking filters added to a plain breakdown** — one metric, one non-date dimension,
485
+ no existing sort. N follows the dimension's element count, so a top-5 over six values
486
+ is never emitted.
487
+ 3. **Sort-only variants**, which order without limiting.
488
+
489
+ Eligibility is deliberately narrow: two metrics leave "top 3 by what?" unanswered, a
490
+ second dimension leaves the N ambiguous between the pair and within a group, and a date
491
+ dimension turns the result into "top 3 months", which nobody asks. Variants are
492
+ deduplicated by resolved definition — differently-titled insights over one metric and
493
+ dimension would otherwise produce the same question twice — and bases are taken
494
+ round-robin by metric so one popular metric cannot become a third of the corpus.
495
+
496
+ Derived items carry `derived_from` (the insight id) and `derived_basis` (`title` when a
497
+ human's chart title asked for the ranking, `shape` when this generator chose to add
498
+ one), so a pass rate over each can be computed separately.
499
+
500
+ ### Items that cannot say what they mean
501
+
502
+ Two classes of question are unwinnable however well the agent behaves, and both are
503
+ reported:
504
+
505
+ - **A name the model carries more than once.** One workspace has six labels all titled
506
+ "Product Title"; a question naming one cannot say which is meant, and a perfect chart
507
+ over the wrong one scores zero. `--skip-ambiguous` drops them; the count and the
508
+ offending names are printed either way.
509
+ - **A date granularity's cyclical twin.** `MONTH` walks consecutive calendar months,
510
+ `MONTH_OF_YEAR` stacks every January together. Date dimensions are therefore briefed
511
+ by what they do ("one point per calendar month over time, not month-of-year") and the
512
+ writer is told to say it in natural words while keeping the date dataset's name —
513
+ never as a label id in prose ("Order Created At - Month").
514
+
515
+ **The question must never contradict its own expected output.** Four rules enforce that:
516
+
517
+ - The writer is briefed on buckets, sorts and filters only — never the insight title,
518
+ and never the chart type. Titles routinely describe intent the definition doesn't
519
+ implement ("Products by Most Items Sold" over `sorts: []`).
520
+ - Every generated question is checked against its spec, and any hit is a hard error:
521
+ ranking words (`top`, `most`, `highest`, …) require a real sort or ranking filter;
522
+ filter words (`only`, `last quarter`, `in 2025`, …) require a real date or attribute
523
+ filter; a breakdown clause requires a non-empty `view_by`/`segment_by` and vice versa;
524
+ a metric may never be broken down by itself; and no template residue (`breakdown
525
+ dimension`, `{…}`) may survive. A violation is fed back once for a rewrite, then
526
+ dropped — and a drop fails the run.
527
+ - The writer's rules are built per insight, so an insight with no `view_by` is never
528
+ asked to name a breakdown at all.
529
+ - `type` is set only when the question actually names a chart form. An insight's
530
+ `visualizationUrl` records what a human clicked, not what the question constrains —
531
+ with one exception: a chart with no breakdown *must* name its form ("as a KPI", "as a
532
+ single number"). Without it the agent reads a bare "Show me Gross Revenue" as a metric
533
+ lookup, activates only its search skill and builds nothing.
534
+
535
+ Everything the writer sees is a display name (`Spend Amount`, `Merchant Name`), never a
536
+ raw URI, so questions read like a person wrote them.
537
+
538
+ **What it won't do.** Insights it can't express without guessing are skipped with a
539
+ printed reason, never approximated: derived (arithmetic/PoP) measures, measure-level
540
+ filters, `uris`-form attribute filters, unmapped chart types, hidden objects, and
541
+ insights whose title promises behaviour their definition lacks (though `--enrich-ranked`
542
+ implements a promised *ranking* rather than discarding it). If too few survive, the
543
+ quality gate fails the run rather than fabricating items to hit the minimum — point at
544
+ more dashboards, or lower `--min-questions`.
545
+
546
+ Every written item is validated as a `DatasetItem` with a scorable AAC visualization
547
+ before the command reports success.
548
+
549
+ ---
550
+
356
551
  ## Dataset format
357
552
 
358
553
  A dataset is a folder of `.json` files, one per question:
@@ -423,7 +618,8 @@ is the fraction of satisfied criteria.
423
618
 
424
619
  ### `[llm-judge]` — LLM-as-judge evaluators
425
620
 
426
- `general_question` and `guardrail` items are scored by a GPT-4o judge.
621
+ `general_question` and `guardrail` items are scored by a GPT-4o judge, and
622
+ `gd-eval generate` uses the same package to write question text.
427
623
  Requires the OpenAI package and `OPENAI_API_KEY`:
428
624
 
429
625
  ```bash
@@ -432,13 +628,15 @@ uv add 'gooddata-eval[llm-judge]'
432
628
  uv tool install 'gooddata-eval[llm-judge]'
433
629
  ```
434
630
 
435
- Without `[llm-judge]`, those items are **skipped**.
631
+ Without `[llm-judge]`, those items are **skipped** and `gd-eval generate` needs
632
+ `--no-phrase`.
436
633
 
437
634
  ## Exit codes
438
635
 
439
636
  | Code | Meaning |
440
637
  |---|---|
441
638
  | `0` | Run completed. Evaluation failures do **not** cause a non-zero exit. |
639
+ | `1` | `gd-eval generate` only: a quality gate failed, an item was dropped, or a written item failed validation. |
442
640
  | `2` | Operational error: bad connection, missing model, unreadable dataset, missing credentials. |
443
641
 
444
642
  ## Scores (in JSON report and Langfuse)
@@ -16,7 +16,10 @@ Or install `gd-eval` as a standalone tool:
16
16
  | Command | Description |
17
17
  |---|---|
18
18
  | `gd-eval run` | Run an evaluation dataset against one or more models. |
19
+ | `gd-eval report` | Render JSON report(s) as one self-contained HTML file. |
19
20
  | `gd-eval models` | List LLM providers and models configured in the org. |
21
+ | `gd-eval generate` | Generate a `visualization` dataset from a workspace's existing insights. |
22
+
20
23
 
21
24
  ---
22
25
 
@@ -145,6 +148,8 @@ interleaves when K > 1, and per-item latencies rise, so they stop being clean si
145
148
  | Flag | Description |
146
149
  |---|---|
147
150
  | `--json PATH` | Write a JSON report to this path. Always uses the nested `{models, runs, comparison}` shape even for a single model. |
151
+ | `--html PATH` | Write a self-contained HTML report to this path (same output as `gd-eval report`). |
152
+ | `--redact` | Make the HTML customer-safe. See `gd-eval report`. |
148
153
  | `--quiet` | Suppress per-item progress. Per-model result tables and the comparison summary are still printed. |
149
154
  | `--preserve-failed` | Keep failed conversations on the server instead of deleting them, so they can be inspected afterwards. Applies to the single-turn chat path; agentic kinds manage their own conversation lifecycle. |
150
155
  | `--timers` | Print per-turn `[timer]` diagnostics — GoodData response, judge, and simulated-user seconds as they happen. Off by default: an 18-item `--runs 2` run emits ~72 lines and buries the progress output. The same measurements are always in the JSON report's `latency_breakdown_s`, so this only adds a live view. Also settable via `GD_EVAL_TIMERS=1`. |
@@ -305,6 +310,59 @@ linking ran. Pass `TAVERN_E2E_SKIP_TRACE_LINK=1` to opt out of linking altogethe
305
310
 
306
311
  ---
307
312
 
313
+ ## `gd-eval report`
314
+
315
+ Turns JSON report(s) into one HTML file you can actually navigate. No server, no
316
+ credentials, no external assets — it opens over `file://`, attaches to a Jira issue and
317
+ survives a Slack thread.
318
+
319
+ ```bash
320
+ # one run
321
+ gd-eval report results.json -o report.html
322
+
323
+ # several runs side by side -- each file becomes its own column
324
+ gd-eval report aug-21.json sep-07.json -o comparison.html --title "H200 regression check"
325
+
326
+ # customer-safe
327
+ gd-eval report results.json -o customer.html --redact
328
+ ```
329
+
330
+ | Flag | Description |
331
+ |---|---|
332
+ | `-o, --out PATH` | Where to write the HTML. Required. |
333
+ | `--title TEXT` | Title shown in the report header. |
334
+ | `--redact` | Drop conversation/response ids and raw reasoning, and rename models to `Model A`, `Model B`, … Pass rate, per-item results, questions and latency survive. |
335
+
336
+ The report is a *view* over the JSON — it computes no numbers of its own. It gives you:
337
+
338
+ - **Run cards and a comparison table** — pass rate, quality, latency per run.
339
+ - **An item table** with a pass/fail column per run, so a model or run-over-run
340
+ regression is one glance rather than a hand-assembled spreadsheet.
341
+ - **An expression filter** for cross-cutting questions the fixed filters can't
342
+ anticipate, e.g. `d.filter_ranking_score === false` or
343
+ `d.expected_metric_uris.length > 1 && !d.metrics_correct`. Available variables:
344
+ `d` (the focused run's `detail`), `it` (its item), `i` (the row, `i.per[label]` for any
345
+ run), `q` (question), `kind`.
346
+ - **The conversation**, when the item ran the agentic multi-turn path — every turn in
347
+ order, with the simulated user marked apart from a real question, so you can see
348
+ whether the agent got there or was handed the answer.
349
+ - **A per-item drawer** — checks as pass/fail chips, expected vs actual side by side,
350
+ full reasoning, conversation/response ids.
351
+ - **A latency timeline** from `detail.latency_breakdown`, in execution order, one bar per
352
+ step. Clicking a step expands its full record, joined by `index`: the paragraph a
353
+ reasoning step was summarised from, or a tool call's arguments and result from
354
+ `detail.tool_calls`.
355
+
356
+ `--redact` additionally drops `transcript` and `tool_calls`: the exchange shows that the
357
+ simulated user is primed with the expected output, and a tool result carries
358
+ semantic-layer internals and real query rows. The turn count and the timeline shape
359
+ survive.
360
+
361
+ Passing several files keyed by file name is the whole run-over-run mechanism: no
362
+ database, no run registry, just the JSON files you already have on disk.
363
+
364
+ ---
365
+
308
366
  ## `gd-eval models`
309
367
 
310
368
  List all LLM providers and their models in the org. Marks the active model
@@ -325,6 +383,143 @@ gd-eval models \
325
383
 
326
384
  ---
327
385
 
386
+ ## `gd-eval generate`
387
+
388
+ Reverse-engineers a `visualization` dataset out of the charts a customer has already
389
+ built, so you get eval questions without hand-authoring any. Reads the workspace's
390
+ declarative analytics model (read-only), translates each visible insight's buckets,
391
+ sorts and filters into an `expected_output.visualization` spec, then asks an LLM to
392
+ write the analyst question that chart answers. Because `expected_output` is copied from
393
+ a live object rather than authored, every question is grounded in the real LDM by
394
+ construction — the LLM only writes English.
395
+
396
+ **Setup:** host + token (read access to the workspace), and `OPENAI_API_KEY` plus the
397
+ `llm-judge` extra for the phrasing step (`uv add 'gooddata-eval[llm-judge]'`; skip both
398
+ with `--no-phrase`).
399
+
400
+ ```bash
401
+ export GOODDATA_TOKEN='your-api-token'
402
+
403
+ # 1. see what a workspace yields before writing anything
404
+ gd-eval generate \
405
+ --host https://your.gooddata.cloud \
406
+ --workspace ecommerce_demo \
407
+ --dataset-name ecommerce \
408
+ --dry-run
409
+
410
+ # 2. generate, phrase, validate, and export
411
+ gd-eval generate \
412
+ --host https://your.gooddata.cloud \
413
+ --workspace ecommerce_demo \
414
+ --dataset-name ecommerce \
415
+ --dashboard dash_1_returns \
416
+ --out ./my-dataset \
417
+ --langfuse-out out/langfuse-dataset.json
418
+
419
+ # 3. run it
420
+ gd-eval run --host … --workspace ecommerce_demo --dataset ./my-dataset --model gpt-5.2
421
+ ```
422
+
423
+ `--workspace` is where insights are read from; `--dataset-name` is the `dataset_name`
424
+ written into every item (and the default output folder).
425
+
426
+ | Flag | Effect |
427
+ |---|---|
428
+ | `--dashboard <id>` | restrict to insights on that dashboard (repeatable); default is the whole workspace |
429
+ | `--out <dir>` | output folder (default `./<dataset-name>`); this is what `gd-eval run --dataset` reads |
430
+ | `--snapshot-out` / `--snapshot-in` | save/replay the fetched model — replay needs no host, token, or network |
431
+ | `--langfuse-out <file>` | also write a Langfuse-importable dataset JSON |
432
+ | `--id-prefix` | prefix exported Langfuse item ids (they're unique per *project*, so re-importing an item under its original id is a 409) |
433
+ | `--no-phrase` | skip the LLM; emit mechanical `Show <title>` questions |
434
+ | `--phrase-model` | OpenAI model for phrasing (default `gpt-4o`) |
435
+ | `--no-viz-type` | always blank the expected chart type |
436
+ | `--enrich-ranked <N>` | additionally derive up to N ranked questions (see below); default 0 (off) |
437
+ | `--skip-ambiguous` | drop items naming something the model carries more than once; reported either way |
438
+ | `--min-questions` / `--min-shapes` / `--min-filtered` | quality gate, default 15, 3 and 1 |
439
+
440
+ ### Ranked questions (`--enrich-ranked`)
441
+
442
+ Analysts sort in Analytical Designer and save the chart without persisting the sort, so
443
+ `sort_by`/`ranking_filter` coverage is near zero on most real models — the eval can
444
+ punish a spurious ranking but never confirm the agent builds a required one.
445
+ `--enrich-ranked N` fills that gap by *deriving* ranked items from the specs already
446
+ extracted. Adding a limit or a sort to a definition that executes cannot make it
447
+ unanswerable, and "the top 3 X by Y" has exactly one correct spec, so a derived item is
448
+ less ambiguous to grade than the insight it came from.
449
+
450
+ The budget is spent best-grounded first:
451
+
452
+ 1. **Insights whose own title promised a ranking their definition never implemented** —
453
+ "Top Returned Reasons" saved with `sorts: []`. The direction comes from the title
454
+ (`highest`/`most`/`largest` vs `lowest`/`least`/`worst`) and the N too when it states
455
+ one; a title naming both ends names neither and is still skipped.
456
+ 2. **Ranking filters added to a plain breakdown** — one metric, one non-date dimension,
457
+ no existing sort. N follows the dimension's element count, so a top-5 over six values
458
+ is never emitted.
459
+ 3. **Sort-only variants**, which order without limiting.
460
+
461
+ Eligibility is deliberately narrow: two metrics leave "top 3 by what?" unanswered, a
462
+ second dimension leaves the N ambiguous between the pair and within a group, and a date
463
+ dimension turns the result into "top 3 months", which nobody asks. Variants are
464
+ deduplicated by resolved definition — differently-titled insights over one metric and
465
+ dimension would otherwise produce the same question twice — and bases are taken
466
+ round-robin by metric so one popular metric cannot become a third of the corpus.
467
+
468
+ Derived items carry `derived_from` (the insight id) and `derived_basis` (`title` when a
469
+ human's chart title asked for the ranking, `shape` when this generator chose to add
470
+ one), so a pass rate over each can be computed separately.
471
+
472
+ ### Items that cannot say what they mean
473
+
474
+ Two classes of question are unwinnable however well the agent behaves, and both are
475
+ reported:
476
+
477
+ - **A name the model carries more than once.** One workspace has six labels all titled
478
+ "Product Title"; a question naming one cannot say which is meant, and a perfect chart
479
+ over the wrong one scores zero. `--skip-ambiguous` drops them; the count and the
480
+ offending names are printed either way.
481
+ - **A date granularity's cyclical twin.** `MONTH` walks consecutive calendar months,
482
+ `MONTH_OF_YEAR` stacks every January together. Date dimensions are therefore briefed
483
+ by what they do ("one point per calendar month over time, not month-of-year") and the
484
+ writer is told to say it in natural words while keeping the date dataset's name —
485
+ never as a label id in prose ("Order Created At - Month").
486
+
487
+ **The question must never contradict its own expected output.** Four rules enforce that:
488
+
489
+ - The writer is briefed on buckets, sorts and filters only — never the insight title,
490
+ and never the chart type. Titles routinely describe intent the definition doesn't
491
+ implement ("Products by Most Items Sold" over `sorts: []`).
492
+ - Every generated question is checked against its spec, and any hit is a hard error:
493
+ ranking words (`top`, `most`, `highest`, …) require a real sort or ranking filter;
494
+ filter words (`only`, `last quarter`, `in 2025`, …) require a real date or attribute
495
+ filter; a breakdown clause requires a non-empty `view_by`/`segment_by` and vice versa;
496
+ a metric may never be broken down by itself; and no template residue (`breakdown
497
+ dimension`, `{…}`) may survive. A violation is fed back once for a rewrite, then
498
+ dropped — and a drop fails the run.
499
+ - The writer's rules are built per insight, so an insight with no `view_by` is never
500
+ asked to name a breakdown at all.
501
+ - `type` is set only when the question actually names a chart form. An insight's
502
+ `visualizationUrl` records what a human clicked, not what the question constrains —
503
+ with one exception: a chart with no breakdown *must* name its form ("as a KPI", "as a
504
+ single number"). Without it the agent reads a bare "Show me Gross Revenue" as a metric
505
+ lookup, activates only its search skill and builds nothing.
506
+
507
+ Everything the writer sees is a display name (`Spend Amount`, `Merchant Name`), never a
508
+ raw URI, so questions read like a person wrote them.
509
+
510
+ **What it won't do.** Insights it can't express without guessing are skipped with a
511
+ printed reason, never approximated: derived (arithmetic/PoP) measures, measure-level
512
+ filters, `uris`-form attribute filters, unmapped chart types, hidden objects, and
513
+ insights whose title promises behaviour their definition lacks (though `--enrich-ranked`
514
+ implements a promised *ranking* rather than discarding it). If too few survive, the
515
+ quality gate fails the run rather than fabricating items to hit the minimum — point at
516
+ more dashboards, or lower `--min-questions`.
517
+
518
+ Every written item is validated as a `DatasetItem` with a scorable AAC visualization
519
+ before the command reports success.
520
+
521
+ ---
522
+
328
523
  ## Dataset format
329
524
 
330
525
  A dataset is a folder of `.json` files, one per question:
@@ -395,7 +590,8 @@ is the fraction of satisfied criteria.
395
590
 
396
591
  ### `[llm-judge]` — LLM-as-judge evaluators
397
592
 
398
- `general_question` and `guardrail` items are scored by a GPT-4o judge.
593
+ `general_question` and `guardrail` items are scored by a GPT-4o judge, and
594
+ `gd-eval generate` uses the same package to write question text.
399
595
  Requires the OpenAI package and `OPENAI_API_KEY`:
400
596
 
401
597
  ```bash
@@ -404,13 +600,15 @@ uv add 'gooddata-eval[llm-judge]'
404
600
  uv tool install 'gooddata-eval[llm-judge]'
405
601
  ```
406
602
 
407
- Without `[llm-judge]`, those items are **skipped**.
603
+ Without `[llm-judge]`, those items are **skipped** and `gd-eval generate` needs
604
+ `--no-phrase`.
408
605
 
409
606
  ## Exit codes
410
607
 
411
608
  | Code | Meaning |
412
609
  |---|---|
413
610
  | `0` | Run completed. Evaluation failures do **not** cause a non-zero exit. |
611
+ | `1` | `gd-eval generate` only: a quality gate failed, an item was dropped, or a written item failed validation. |
414
612
  | `2` | Operational error: bad connection, missing model, unreadable dataset, missing credentials. |
415
613
 
416
614
  ## Scores (in JSON report and Langfuse)