tuieval 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- tuieval-0.1.0/.gitignore +13 -0
- tuieval-0.1.0/CHANGELOG.md +12 -0
- tuieval-0.1.0/LICENSE +21 -0
- tuieval-0.1.0/PKG-INFO +156 -0
- tuieval-0.1.0/README.md +130 -0
- tuieval-0.1.0/docs/images/tui-setup.png +0 -0
- tuieval-0.1.0/docs/models.md +120 -0
- tuieval-0.1.0/docs/writing-packs.md +160 -0
- tuieval-0.1.0/packaging/homebrew/README.md +19 -0
- tuieval-0.1.0/packaging/homebrew/tuieval.rb +36 -0
- tuieval-0.1.0/pyproject.toml +47 -0
- tuieval-0.1.0/src/tuieval/__init__.py +2 -0
- tuieval-0.1.0/src/tuieval/__main__.py +5 -0
- tuieval-0.1.0/src/tuieval/cli.py +78 -0
- tuieval-0.1.0/src/tuieval/client.py +196 -0
- tuieval-0.1.0/src/tuieval/compare.py +511 -0
- tuieval-0.1.0/src/tuieval/engine.py +1827 -0
- tuieval-0.1.0/src/tuieval/graders/__init__.py +138 -0
- tuieval-0.1.0/src/tuieval/graders/answer.py +52 -0
- tuieval-0.1.0/src/tuieval/graders/code.py +104 -0
- tuieval-0.1.0/src/tuieval/graders/rag.py +33 -0
- tuieval-0.1.0/src/tuieval/graders/reply.py +44 -0
- tuieval-0.1.0/src/tuieval/graders/tool_call.py +75 -0
- tuieval-0.1.0/src/tuieval/machines.py +293 -0
- tuieval-0.1.0/src/tuieval/packs.py +239 -0
- tuieval-0.1.0/src/tuieval/profiles.py +98 -0
- tuieval-0.1.0/src/tuieval/run_evals.py +588 -0
- tuieval-0.1.0/src/tuieval/scaffold.py +106 -0
- tuieval-0.1.0/src/tuieval/selftest.py +185 -0
- tuieval-0.1.0/src/tuieval/templates/models.toml +137 -0
- tuieval-0.1.0/src/tuieval/templates/packs/answer/pack.toml +17 -0
- tuieval-0.1.0/src/tuieval/templates/packs/answer/system.txt +4 -0
- tuieval-0.1.0/src/tuieval/templates/packs/answer/tests.yaml +67 -0
- tuieval-0.1.0/src/tuieval/templates/packs/code/pack.toml +16 -0
- tuieval-0.1.0/src/tuieval/templates/packs/code/system.txt +3 -0
- tuieval-0.1.0/src/tuieval/templates/packs/code/tests.yaml +197 -0
- tuieval-0.1.0/src/tuieval/templates/packs/rag/pack.toml +16 -0
- tuieval-0.1.0/src/tuieval/templates/packs/rag/system.txt +9 -0
- tuieval-0.1.0/src/tuieval/templates/packs/rag/tests.yaml +99 -0
- tuieval-0.1.0/src/tuieval/templates/packs/reply/pack.toml +15 -0
- tuieval-0.1.0/src/tuieval/templates/packs/reply/system.txt +2 -0
- tuieval-0.1.0/src/tuieval/templates/packs/reply/tests.yaml +67 -0
- tuieval-0.1.0/src/tuieval/templates/packs/tool_call/pack.toml +18 -0
- tuieval-0.1.0/src/tuieval/templates/packs/tool_call/system.txt +2 -0
- tuieval-0.1.0/src/tuieval/templates/packs/tool_call/tests.yaml +59 -0
- tuieval-0.1.0/src/tuieval/templates/packs/tool_call/tools.yaml +32 -0
- tuieval-0.1.0/src/tuieval/tui.py +2558 -0
- tuieval-0.1.0/src/tuieval/tune.py +518 -0
- tuieval-0.1.0/src/tuieval/verdict.py +346 -0
- tuieval-0.1.0/src/tuieval/watch_proxy.py +260 -0
- tuieval-0.1.0/src/tuieval/workspace.py +27 -0
- tuieval-0.1.0/src/tuieval/yamlout.py +21 -0
- tuieval-0.1.0/tests/mock_server.py +121 -0
- tuieval-0.1.0/tests/test_tuieval.py +197 -0
tuieval-0.1.0/.gitignore
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.1.0
|
|
4
|
+
|
|
5
|
+
First public release.
|
|
6
|
+
|
|
7
|
+
- TUI and CLI for evaluating local models on your own eval packs: accuracy, speed and token use, with PASS / FAIL / INCONCLUSIVE verdicts per use case.
|
|
8
|
+
- Workspaces (`tuieval init`, `--workspace`, `TUIEVAL_HOME`) keep your packs, models and results apart from the tool.
|
|
9
|
+
- `tuieval new-pack` starter templates for the built-in graders: answer, rag, reply, tool_call, code.
|
|
10
|
+
- Custom graders from the workspace's `graders/` folder or a pack's own `grader.py`.
|
|
11
|
+
- Servers: llama.cpp (with per-machine fit check and speed tuning), any OpenAI-compatible server, OpenRouter (pinned to one provider endpoint).
|
|
12
|
+
- Speed tuning uses a built-in workload, so it needs no packs.
|
tuieval-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 tuieval contributors
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
tuieval-0.1.0/PKG-INFO
ADDED
|
@@ -0,0 +1,156 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: tuieval
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Evaluate local LLMs on your own eval packs, in the terminal: accuracy, speed and tokens, with PASS/FAIL verdicts.
|
|
5
|
+
Project-URL: Homepage, https://github.com/ashe-wb/tuieval
|
|
6
|
+
Project-URL: Issues, https://github.com/ashe-wb/tuieval/issues
|
|
7
|
+
Author: tuieval contributors
|
|
8
|
+
License-Expression: MIT
|
|
9
|
+
License-File: LICENSE
|
|
10
|
+
Keywords: benchmark,evals,evaluation,llama.cpp,llm,local-llm,tui
|
|
11
|
+
Classifier: Development Status :: 4 - Beta
|
|
12
|
+
Classifier: Environment :: Console :: Curses
|
|
13
|
+
Classifier: Intended Audience :: Developers
|
|
14
|
+
Classifier: Intended Audience :: Science/Research
|
|
15
|
+
Classifier: Operating System :: MacOS
|
|
16
|
+
Classifier: Operating System :: POSIX :: Linux
|
|
17
|
+
Classifier: Programming Language :: Python :: 3
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
20
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
21
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
22
|
+
Requires-Python: >=3.11
|
|
23
|
+
Requires-Dist: pyyaml>=6
|
|
24
|
+
Requires-Dist: textual<9,>=8
|
|
25
|
+
Description-Content-Type: text/markdown
|
|
26
|
+
|
|
27
|
+
# tuieval
|
|
28
|
+
|
|
29
|
+
**Evaluate local LLMs on your own questions, in the terminal.** Compare models (tiny, dense, MoE; llama.cpp GGUFs, LM Studio, Ollama, vLLM, OpenRouter) on what *you* use them for, measuring **accuracy, speed and token use** together, and get a **PASS / FAIL / INCONCLUSIVE** verdict per use case. Grading is automatic; there's no LLM judge.
|
|
30
|
+
|
|
31
|
+
tuieval ships with **no built-in benchmark**. Public benchmarks leak into training data and rarely match your work. You write *eval packs* (folders of your own questions with checkable answers), and tuieval runs them, grades them, and tells you which model is ready for that job on this machine.
|
|
32
|
+
|
|
33
|
+
```bash
|
|
34
|
+
pipx install tuieval # or: pip install tuieval (Python 3.11+)
|
|
35
|
+
tuieval init my-evals && cd my-evals
|
|
36
|
+
tuieval new-pack my-first-pack # a pack of example questions to edit
|
|
37
|
+
tuieval add ~/models/Some-Model-Q4_K_M.gguf
|
|
38
|
+
tuieval # open the TUI
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
Once you've added your own packs and models, the setup screen looks like this (packs on the left, models with a verdict code per use case on the right, the highlighted model's details below):
|
|
42
|
+
|
|
43
|
+

|
|
44
|
+
|
|
45
|
+
## What you get
|
|
46
|
+
|
|
47
|
+
- **A TUI** to pick packs and models, watch reasoning and answers stream live with TTFT, tokens/s, memory and a running score, and browse results.
|
|
48
|
+
- **Verdicts you can act on.** A pack passes only with zero critical failures over enough trials, and an accuracy whose 95% lower bound clears your bar. *INCONCLUSIVE* says what evidence is missing.
|
|
49
|
+
- **Two tiers:** *Screen* a sample of every pack to drop weak models fast, then *Certify* finalists on every test with repeats. Certification reuses the screening answers.
|
|
50
|
+
- **Honest numbers:** every request is a fresh single-turn conversation, with prompt caching off and repeat rounds in different orders. Each answer records what was sent, and the run warns about reused prompts or identical repeats.
|
|
51
|
+
- **Per-machine speed.** A fit check picks the largest context that fits your Mac's GPU memory, `tuieval tune` finds the fastest speed-only server flags (with a guard that rejects flags that change answers), and readiness includes a *Fast enough?* table per machine, measured or projected.
|
|
52
|
+
- **Tests that check themselves.** Every test carries a reference answer and known-wrong answers, and `tuieval selftest` checks the grader accepts the first and rejects the second, and that each gate is reachable at all.
|
|
53
|
+
- **History.** Verdict changes are appended to `results/verdicts.jsonl`, and replaced results are moved to `history/`, never overwritten.
|
|
54
|
+
|
|
55
|
+
## Eval packs
|
|
56
|
+
|
|
57
|
+
A pack is a folder in your workspace's `packs/`:
|
|
58
|
+
|
|
59
|
+
```
|
|
60
|
+
packs/support-bot/
|
|
61
|
+
pack.toml label, use case, grader, gate (what PASS means)
|
|
62
|
+
system.txt the system prompt
|
|
63
|
+
tests.yaml the questions, with expected answers, references and known-wrong answers
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
```yaml
|
|
67
|
+
- id: refund-window
|
|
68
|
+
difficulty: medium
|
|
69
|
+
input: A customer bought shoes 20 days ago and wants a refund. Our policy allows 30 days. Can they get one?
|
|
70
|
+
max_words: 40
|
|
71
|
+
must_include: ["yes|can"]
|
|
72
|
+
reference: "Yes, they're within the 30-day window, so they can get a refund."
|
|
73
|
+
wrong: ["No, the refund window has passed."]
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
Built-in graders: `answer` (a number or word on an `ANSWER:` line, or a correct refusal when the data isn't there), `rag` (grounded answers with citations), `reply` (free-form replies against rules), `tool_call` (the right tool with the right arguments, or rightly none) and `code` (Python run against hidden tests). Your own grader is one Python file in the workspace's `graders/` folder.
|
|
77
|
+
|
|
78
|
+
`tuieval new-pack <name> --grader <grader>` creates a pack with working examples for any grader. **See [docs/writing-packs.md](docs/writing-packs.md)** for every field, gates, critical failures, difficulty labels, custom graders and how to write tests that separate good models from weak ones.
|
|
79
|
+
|
|
80
|
+
## The workspace
|
|
81
|
+
|
|
82
|
+
Your packs, models and results live in a **workspace** folder, apart from the tool:
|
|
83
|
+
|
|
84
|
+
```
|
|
85
|
+
my-evals/
|
|
86
|
+
models.toml your servers and models
|
|
87
|
+
packs/ your eval packs
|
|
88
|
+
graders/ your own graders (optional)
|
|
89
|
+
results/ logs/ tuning/ reports/ presets.toml written by tuieval
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
tuieval uses the current folder, or `--workspace DIR` / `TUIEVAL_HOME`. Keep it in git or a synced folder: results from one machine count on another when the answer-changing settings match, so you can Certify on your fastest machine.
|
|
93
|
+
|
|
94
|
+
## Models
|
|
95
|
+
|
|
96
|
+
```bash
|
|
97
|
+
tuieval add ~/models/Some-Model-Q4_K_M.gguf --tags 9b,dense,q4 # llama.cpp (llama-server on PATH)
|
|
98
|
+
tuieval add ~/models/VL-Q4.gguf --mmproj ~/models/VL-mmproj.gguf # vision
|
|
99
|
+
tuieval add qwen3:8b --server local # an already-running server (set its url)
|
|
100
|
+
tuieval run --tier smoke --only openrouter:qwen/qwen3-32b # any OpenRouter model, no config needed
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
See **[docs/models.md](docs/models.md)** for servers, OpenRouter endpoint pinning, machines, tuning and every option.
|
|
104
|
+
|
|
105
|
+
## The TUI
|
|
106
|
+
|
|
107
|
+
1. **Pick packs and models** (space ticks; type to filter models by label or tag). Each model shows a short code per use case, e.g. `C✓ S?` (✓ pass, ✗ fail, ? inconclusive; grey = from earlier results).
|
|
108
|
+
2. **Pick a tier** (Smoke to check setup, Screen, Certify) and **press `s`**. The line above the buttons shows how many answers that is and roughly how long it will take. For each model, tuieval starts its server (or uses a running one), checks the right model is loaded, runs every selected pack, then stops it.
|
|
109
|
+
3. **Watch the run:** progress with ETA, live score, tok/s, TTFT and memory, the current test with reasoning and answer side by side, recent results with the grader's reason. `k` skips a model, `c` cancels (finished work is kept and resumes next time). Runs started while one is going are queued.
|
|
110
|
+
4. **Press `r` for results:** production readiness, verdict history, scorecard, speed & tokens (★ marks models nothing beats on both accuracy and time), per question, is the difference real?, tests that separate models, by difficulty, failures, and test quality.
|
|
111
|
+
|
|
112
|
+
Other keys: `t` tunes the ticked models' speed flags, `a` adds a model, `m` scans for GGUFs, `p` saves or loads a preset, `x` hides a model.
|
|
113
|
+
|
|
114
|
+
## Command line
|
|
115
|
+
|
|
116
|
+
```bash
|
|
117
|
+
tuieval run # Screen every model on every pack
|
|
118
|
+
tuieval run --tier certify --only a,b --packs support-bot,coding
|
|
119
|
+
tuieval run --dry-run # the plan and server commands
|
|
120
|
+
tuieval verdict # PASS / FAIL / INCONCLUSIVE per model and use case
|
|
121
|
+
tuieval report # the same with evidence, as a markdown file
|
|
122
|
+
tuieval history # every verdict change over time
|
|
123
|
+
tuieval compare --speed --pairwise --failures # scorecards
|
|
124
|
+
tuieval selftest # check every test's reference and wrong answers
|
|
125
|
+
tuieval items # tests that don't separate models or look broken
|
|
126
|
+
tuieval capture logs/live/<file> --pack X # turn a real failure into a new test
|
|
127
|
+
tuieval regrade # re-score stored answers after changing a grader
|
|
128
|
+
tuieval machines # this machine and others, fit and tuning per model
|
|
129
|
+
tuieval tune <model> # fastest speed flags for a model on this machine
|
|
130
|
+
tuieval watch --upstream http://localhost:8080 # show the reasoning of any app using your server
|
|
131
|
+
tuieval help
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
## Reading results
|
|
135
|
+
|
|
136
|
+
- **Start with Production readiness.** A single critical failure already means FAIL, whatever the accuracy: a model that breaks a hard rule once in 300 answers will do it in production. The Failures tab and `tuieval report` list exactly which answers failed.
|
|
137
|
+
- **INCONCLUSIVE is not "nearly passed".** It says what's missing (usually a Certify run, or more trials).
|
|
138
|
+
- **Trust "Is the difference real?" over raw percentages.** Differences of one or two tests are usually noise.
|
|
139
|
+
- **Read `trunc` before accuracy.** A model that runs out of tokens isn't wrong; it's thinking too long for the budget.
|
|
140
|
+
- **Watch TTFT for interactive use.** A model that is 5% more accurate but takes 3 s longer to start answering may be the worse choice.
|
|
141
|
+
|
|
142
|
+
## Safety
|
|
143
|
+
|
|
144
|
+
⚠️ The `code` grader executes model-written code on your machine. Run code packs inside a container or VM with no credentials in the environment.
|
|
145
|
+
|
|
146
|
+
## Development
|
|
147
|
+
|
|
148
|
+
```bash
|
|
149
|
+
git clone https://github.com/ashe-wb/tuieval && cd tuieval
|
|
150
|
+
python -m venv .venv && .venv/bin/pip install -e .
|
|
151
|
+
.venv/bin/python -m unittest discover -s tests # uses a mock server; no model needed
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
## License
|
|
155
|
+
|
|
156
|
+
MIT. See [LICENSE](LICENSE).
|
tuieval-0.1.0/README.md
ADDED
|
@@ -0,0 +1,130 @@
|
|
|
1
|
+
# tuieval
|
|
2
|
+
|
|
3
|
+
**Evaluate local LLMs on your own questions, in the terminal.** Compare models (tiny, dense, MoE; llama.cpp GGUFs, LM Studio, Ollama, vLLM, OpenRouter) on what *you* use them for, measuring **accuracy, speed and token use** together, and get a **PASS / FAIL / INCONCLUSIVE** verdict per use case. Grading is automatic; there's no LLM judge.
|
|
4
|
+
|
|
5
|
+
tuieval ships with **no built-in benchmark**. Public benchmarks leak into training data and rarely match your work. You write *eval packs* (folders of your own questions with checkable answers), and tuieval runs them, grades them, and tells you which model is ready for that job on this machine.
|
|
6
|
+
|
|
7
|
+
```bash
|
|
8
|
+
pipx install tuieval # or: pip install tuieval (Python 3.11+)
|
|
9
|
+
tuieval init my-evals && cd my-evals
|
|
10
|
+
tuieval new-pack my-first-pack # a pack of example questions to edit
|
|
11
|
+
tuieval add ~/models/Some-Model-Q4_K_M.gguf
|
|
12
|
+
tuieval # open the TUI
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
Once you've added your own packs and models, the setup screen looks like this (packs on the left, models with a verdict code per use case on the right, the highlighted model's details below):
|
|
16
|
+
|
|
17
|
+

|
|
18
|
+
|
|
19
|
+
## What you get
|
|
20
|
+
|
|
21
|
+
- **A TUI** to pick packs and models, watch reasoning and answers stream live with TTFT, tokens/s, memory and a running score, and browse results.
|
|
22
|
+
- **Verdicts you can act on.** A pack passes only with zero critical failures over enough trials, and an accuracy whose 95% lower bound clears your bar. *INCONCLUSIVE* says what evidence is missing.
|
|
23
|
+
- **Two tiers:** *Screen* a sample of every pack to drop weak models fast, then *Certify* finalists on every test with repeats. Certification reuses the screening answers.
|
|
24
|
+
- **Honest numbers:** every request is a fresh single-turn conversation, with prompt caching off and repeat rounds in different orders. Each answer records what was sent, and the run warns about reused prompts or identical repeats.
|
|
25
|
+
- **Per-machine speed.** A fit check picks the largest context that fits your Mac's GPU memory, `tuieval tune` finds the fastest speed-only server flags (with a guard that rejects flags that change answers), and readiness includes a *Fast enough?* table per machine, measured or projected.
|
|
26
|
+
- **Tests that check themselves.** Every test carries a reference answer and known-wrong answers, and `tuieval selftest` checks the grader accepts the first and rejects the second, and that each gate is reachable at all.
|
|
27
|
+
- **History.** Verdict changes are appended to `results/verdicts.jsonl`, and replaced results are moved to `history/`, never overwritten.
|
|
28
|
+
|
|
29
|
+
## Eval packs
|
|
30
|
+
|
|
31
|
+
A pack is a folder in your workspace's `packs/`:
|
|
32
|
+
|
|
33
|
+
```
|
|
34
|
+
packs/support-bot/
|
|
35
|
+
pack.toml label, use case, grader, gate (what PASS means)
|
|
36
|
+
system.txt the system prompt
|
|
37
|
+
tests.yaml the questions, with expected answers, references and known-wrong answers
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
```yaml
|
|
41
|
+
- id: refund-window
|
|
42
|
+
difficulty: medium
|
|
43
|
+
input: A customer bought shoes 20 days ago and wants a refund. Our policy allows 30 days. Can they get one?
|
|
44
|
+
max_words: 40
|
|
45
|
+
must_include: ["yes|can"]
|
|
46
|
+
reference: "Yes, they're within the 30-day window, so they can get a refund."
|
|
47
|
+
wrong: ["No, the refund window has passed."]
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
Built-in graders: `answer` (a number or word on an `ANSWER:` line, or a correct refusal when the data isn't there), `rag` (grounded answers with citations), `reply` (free-form replies against rules), `tool_call` (the right tool with the right arguments, or rightly none) and `code` (Python run against hidden tests). Your own grader is one Python file in the workspace's `graders/` folder.
|
|
51
|
+
|
|
52
|
+
`tuieval new-pack <name> --grader <grader>` creates a pack with working examples for any grader. **See [docs/writing-packs.md](docs/writing-packs.md)** for every field, gates, critical failures, difficulty labels, custom graders and how to write tests that separate good models from weak ones.
|
|
53
|
+
|
|
54
|
+
## The workspace
|
|
55
|
+
|
|
56
|
+
Your packs, models and results live in a **workspace** folder, apart from the tool:
|
|
57
|
+
|
|
58
|
+
```
|
|
59
|
+
my-evals/
|
|
60
|
+
models.toml your servers and models
|
|
61
|
+
packs/ your eval packs
|
|
62
|
+
graders/ your own graders (optional)
|
|
63
|
+
results/ logs/ tuning/ reports/ presets.toml written by tuieval
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
tuieval uses the current folder, or `--workspace DIR` / `TUIEVAL_HOME`. Keep it in git or a synced folder: results from one machine count on another when the answer-changing settings match, so you can Certify on your fastest machine.
|
|
67
|
+
|
|
68
|
+
## Models
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
tuieval add ~/models/Some-Model-Q4_K_M.gguf --tags 9b,dense,q4 # llama.cpp (llama-server on PATH)
|
|
72
|
+
tuieval add ~/models/VL-Q4.gguf --mmproj ~/models/VL-mmproj.gguf # vision
|
|
73
|
+
tuieval add qwen3:8b --server local # an already-running server (set its url)
|
|
74
|
+
tuieval run --tier smoke --only openrouter:qwen/qwen3-32b # any OpenRouter model, no config needed
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
See **[docs/models.md](docs/models.md)** for servers, OpenRouter endpoint pinning, machines, tuning and every option.
|
|
78
|
+
|
|
79
|
+
## The TUI
|
|
80
|
+
|
|
81
|
+
1. **Pick packs and models** (space ticks; type to filter models by label or tag). Each model shows a short code per use case, e.g. `C✓ S?` (✓ pass, ✗ fail, ? inconclusive; grey = from earlier results).
|
|
82
|
+
2. **Pick a tier** (Smoke to check setup, Screen, Certify) and **press `s`**. The line above the buttons shows how many answers that is and roughly how long it will take. For each model, tuieval starts its server (or uses a running one), checks the right model is loaded, runs every selected pack, then stops it.
|
|
83
|
+
3. **Watch the run:** progress with ETA, live score, tok/s, TTFT and memory, the current test with reasoning and answer side by side, recent results with the grader's reason. `k` skips a model, `c` cancels (finished work is kept and resumes next time). Runs started while one is going are queued.
|
|
84
|
+
4. **Press `r` for results:** production readiness, verdict history, scorecard, speed & tokens (★ marks models nothing beats on both accuracy and time), per question, is the difference real?, tests that separate models, by difficulty, failures, and test quality.
|
|
85
|
+
|
|
86
|
+
Other keys: `t` tunes the ticked models' speed flags, `a` adds a model, `m` scans for GGUFs, `p` saves or loads a preset, `x` hides a model.
|
|
87
|
+
|
|
88
|
+
## Command line
|
|
89
|
+
|
|
90
|
+
```bash
|
|
91
|
+
tuieval run # Screen every model on every pack
|
|
92
|
+
tuieval run --tier certify --only a,b --packs support-bot,coding
|
|
93
|
+
tuieval run --dry-run # the plan and server commands
|
|
94
|
+
tuieval verdict # PASS / FAIL / INCONCLUSIVE per model and use case
|
|
95
|
+
tuieval report # the same with evidence, as a markdown file
|
|
96
|
+
tuieval history # every verdict change over time
|
|
97
|
+
tuieval compare --speed --pairwise --failures # scorecards
|
|
98
|
+
tuieval selftest # check every test's reference and wrong answers
|
|
99
|
+
tuieval items # tests that don't separate models or look broken
|
|
100
|
+
tuieval capture logs/live/<file> --pack X # turn a real failure into a new test
|
|
101
|
+
tuieval regrade # re-score stored answers after changing a grader
|
|
102
|
+
tuieval machines # this machine and others, fit and tuning per model
|
|
103
|
+
tuieval tune <model> # fastest speed flags for a model on this machine
|
|
104
|
+
tuieval watch --upstream http://localhost:8080 # show the reasoning of any app using your server
|
|
105
|
+
tuieval help
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
## Reading results
|
|
109
|
+
|
|
110
|
+
- **Start with Production readiness.** A single critical failure already means FAIL, whatever the accuracy: a model that breaks a hard rule once in 300 answers will do it in production. The Failures tab and `tuieval report` list exactly which answers failed.
|
|
111
|
+
- **INCONCLUSIVE is not "nearly passed".** It says what's missing (usually a Certify run, or more trials).
|
|
112
|
+
- **Trust "Is the difference real?" over raw percentages.** Differences of one or two tests are usually noise.
|
|
113
|
+
- **Read `trunc` before accuracy.** A model that runs out of tokens isn't wrong; it's thinking too long for the budget.
|
|
114
|
+
- **Watch TTFT for interactive use.** A model that is 5% more accurate but takes 3 s longer to start answering may be the worse choice.
|
|
115
|
+
|
|
116
|
+
## Safety
|
|
117
|
+
|
|
118
|
+
⚠️ The `code` grader executes model-written code on your machine. Run code packs inside a container or VM with no credentials in the environment.
|
|
119
|
+
|
|
120
|
+
## Development
|
|
121
|
+
|
|
122
|
+
```bash
|
|
123
|
+
git clone https://github.com/ashe-wb/tuieval && cd tuieval
|
|
124
|
+
python -m venv .venv && .venv/bin/pip install -e .
|
|
125
|
+
.venv/bin/python -m unittest discover -s tests # uses a mock server; no model needed
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
## License
|
|
129
|
+
|
|
130
|
+
MIT. See [LICENSE](LICENSE).
|
|
Binary file
|
|
@@ -0,0 +1,120 @@
|
|
|
1
|
+
# Models, servers and machines
|
|
2
|
+
|
|
3
|
+
Everything about models lives in your workspace's `models.toml`. `tuieval init` writes one with three servers ready to use:
|
|
4
|
+
|
|
5
|
+
| server | what it is |
|
|
6
|
+
|---|---|
|
|
7
|
+
| `llama` | [llama.cpp](https://github.com/ggml-org/llama.cpp)'s `llama-server`, started by tuieval for each GGUF model, with a per-machine fit check and speed tuning |
|
|
8
|
+
| `local` | any OpenAI-compatible server that's already running (LM Studio, Ollama, vLLM, a remote box): just a `url` |
|
|
9
|
+
| `openrouter` | any model on [OpenRouter](https://openrouter.ai/models), named at run time, pinned to one provider endpoint |
|
|
10
|
+
|
|
11
|
+
## Adding models
|
|
12
|
+
|
|
13
|
+
```bash
|
|
14
|
+
tuieval add ~/models/Some-Model-Q4_K_M.gguf --tags 9b,dense,q4 # GGUF -> llama
|
|
15
|
+
tuieval add ~/models/Some-Model-Q4_K_M.gguf --no-think # same model, thinking off (label …-nothink)
|
|
16
|
+
tuieval add ~/models/VL-Q4.gguf --mmproj ~/models/VL-mmproj.gguf # vision model
|
|
17
|
+
tuieval add qwen3:8b --server local # a model your running server serves
|
|
18
|
+
tuieval scan --add # every new GGUF under model_dirs
|
|
19
|
+
tuieval list # models, packs, result status
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
In the TUI, `a` adds a model and `m` scans your model folders. Each model is one `[[models]]` block:
|
|
23
|
+
|
|
24
|
+
| field | meaning |
|
|
25
|
+
|---|---|
|
|
26
|
+
| `server` | which `[servers.*]` block serves it |
|
|
27
|
+
| `model` | GGUF path (llama) or model id (other servers) |
|
|
28
|
+
| `label` | names `results/<label>/`; default: the file or repo name, lowercased. Never rename a label that has results. |
|
|
29
|
+
| `served_name` | the model name sent with each request (default: the label for llama, the model id elsewhere) |
|
|
30
|
+
| `vision`, `mmproj` | the model reads images; `mmproj` is llama's vision projector and implies `vision` |
|
|
31
|
+
| `thinking = false` | run with thinking off |
|
|
32
|
+
| `tags = [...]` | filter by them in the TUI and with `--tags` |
|
|
33
|
+
| `sampling = {...}` | override `[sampling]` for this model |
|
|
34
|
+
| `server_args = [...]` | extra server flags (treated as answer-changing) |
|
|
35
|
+
| `ctx` (llama) | cap the context; the fit check may lower it further per machine |
|
|
36
|
+
| `max_context` (other servers) | the context the server provides; packs whose prompts need more are skipped |
|
|
37
|
+
| `kv_type` (llama) | e.g. `"q8_0"`: a quantized KV cache, which makes it a different model (give it its own label) |
|
|
38
|
+
| `tools = false` | packs that need tool calling are skipped |
|
|
39
|
+
|
|
40
|
+
Results record the settings each model actually ran with, and the results screen notes any differences between models.
|
|
41
|
+
|
|
42
|
+
## OpenRouter
|
|
43
|
+
|
|
44
|
+
OpenRouter models need no block. Name any model from openrouter.ai/models when you run, e.g. `tuieval run --tier smoke --only openrouter:qwen/qwen3-235b-a22b-2507`. In the TUI, tick **+ openrouter: any model** at the bottom of the model list and pick from the live list (filter by typing; it shows context and price), or press `a` and type `openrouter:<model id>`. A wrong id is refused with the closest matches.
|
|
45
|
+
|
|
46
|
+
- Needs `OPENROUTER_API_KEY` exported in the shell that starts tuieval. A plain `tuieval run` never includes OpenRouter models, since they cost money.
|
|
47
|
+
- Vision, tool calling and context size come from OpenRouter's model list; packs the model can't do are skipped.
|
|
48
|
+
- **One provider per model.** Left alone, OpenRouter spreads requests over providers running different quantizations. `pin_endpoint = true` picks one endpoint per model (closest to the released weights first: bf16 > fp8 > undeclared > lower; then tool and sampling support, uptime and price), sends every request only there with no fallback, keeps it while it's offered, and records it with the results. Answers from any other provider aren't counted.
|
|
49
|
+
- Hosted providers cache shared prompt openings and that can't be switched off. It doesn't change answers, but those answers (⟲ in the run) are left out of TTFT and prompt-speed numbers.
|
|
50
|
+
- To keep providers that store or train on prompts (and could learn your tests) out entirely, add `request = { provider = { data_collection = "deny" } }` under `[servers.openrouter]`.
|
|
51
|
+
|
|
52
|
+
## Every answer is independent
|
|
53
|
+
|
|
54
|
+
Each request is a fresh single-turn conversation: the system prompt and one question, never an earlier answer. Repeats go round by round, each round asking the tests in its own fixed shuffled order, so a question never follows itself. The llama server is told not to reuse earlier prompts (`request = { cache_prompt = false }`). Each answer records the messages sent and any prompt tokens the server reports reusing, and the run warns if that's ever not 0. If a question gets the exact same long answer twice with sampling on, the run flags it as a probable cached response. `tuieval selftest` checks all of this.
|
|
55
|
+
|
|
56
|
+
## Answer-changing vs speed-only flags
|
|
57
|
+
|
|
58
|
+
The best server flags differ per model and per machine, so tuieval splits them by what they change:
|
|
59
|
+
|
|
60
|
+
| Changes **answers**: same everywhere, part of each result's fingerprint | Changes **only speed**: tuned per model per machine |
|
|
61
|
+
|---|---|
|
|
62
|
+
| model file and quant, KV-cache type, `--jinja`, `--reasoning-format`, sampling, mmproj (`cmd` in `models.toml`) | threads, batch sizes, flash attention, cache reuse (`perf` and `[servers.llama.tune]`) |
|
|
63
|
+
|
|
64
|
+
**Quality results travel.** Tuning never invalidates results, and a result from one machine counts on another as long as the answer-changing settings match. So you can Certify on your fastest machine and only tune (and optionally Screen) on the others: sync the workspace folder between them.
|
|
65
|
+
|
|
66
|
+
## Machines
|
|
67
|
+
|
|
68
|
+
- **Machine id** is detected automatically (`m1max-32gb`, `m4pro-24gb`, …; set `EVALS_MACHINE` to rename it). Every result records the machine, the server version and the exact speed flags used. `tuieval machines` lists this machine and every other machine that has run the evals (they record themselves in `tuning/`).
|
|
69
|
+
- **Fit check (Apple Silicon, llama):** before starting a GGUF, tuieval reads its header (layers, KV heads, hybrid attention layers) and picks the largest context that fits this machine's GPU memory, capped at `max_ctx`. A model that can't fit at 8k context is skipped with *doesn't fit on <machine>* instead of swapping. Packs that need more context than fits are skipped. It never quantizes the KV cache on its own, since that changes answers.
|
|
70
|
+
- **Per-machine settings** go under `[machines.<id>]`: `memory_headroom_gb` (GPU memory kept free for macOS, default 4) and `gpu_residency_gb` (see the stall guard).
|
|
71
|
+
|
|
72
|
+
## Tuning
|
|
73
|
+
|
|
74
|
+
`tuieval tune <model>` (or `t` on the setup screen, for the ticked models) finds the fastest speed flags for a model on this machine:
|
|
75
|
+
|
|
76
|
+
1. If `llama-bench` is installed, it sweeps threads, micro-batch and flash attention first (fast, no server starts).
|
|
77
|
+
2. Then it starts the real server with one knob changed at a time and times a fixed **built-in** workload (three short prompts, three medium ones and one ~8k-token prompt), so tuning needs no packs and speeds are comparable between workspaces.
|
|
78
|
+
3. An **output guard** rejects any option that changes greedy answers beyond noise. Candidates that make macOS swap are rejected.
|
|
79
|
+
|
|
80
|
+
Expect 8–15 server starts, about 20–30 minutes for a 27B model, once per model per machine. The result is saved in `tuning/<machine>/<model>.toml` and used by every later run there. Models without a profile run with each knob's first option and show *untuned*. A profile is marked for retuning when the model file or server version changes.
|
|
81
|
+
|
|
82
|
+
The knobs are `[servers.<name>.tune]` in `models.toml`: each knob is a list of options, each option a list of flags. Placeholders: `{p}` P-cores, `{p_minus_2}`, `{all}` all cores, `{gpu_safe_gb}` (the GPU residency limit minus 2 GB; also `_minus_1`, `_minus_2`, `_plus_1`).
|
|
83
|
+
|
|
84
|
+
## Speed verdicts
|
|
85
|
+
|
|
86
|
+
Speed is judged separately, per machine: the same model can be production-grade in quality and still too slow on a smaller machine. Results → Production readiness has a **Fast enough?** table: p90 seconds per answer against each pack's `max_p90_s`, for every machine. It's *measured* where the model ran, and otherwise *projected* from each answer's token counts and that machine's tuned speeds. `tuieval compare --machine <id>` shows the same on the command line.
|
|
87
|
+
|
|
88
|
+
## The stall guard (Apple Silicon)
|
|
89
|
+
|
|
90
|
+
Measured on an M1 Max 32 GB: once the system's GPU allocations pass about half of RAM, the GPU driver evicts and re-maps memory on every GPU job, the server spends 70–95% of its CPU in the kernel while the GPU idles, and servers that submit many small GPU jobs slow to a crawl. For servers with `stall_guard = true`, tuieval samples the server's own vs kernel CPU time and the system's GPU allocation every 5 s during runs and tuning. If more than 70% of its CPU goes to the kernel for a minute, the run stops with an explanation instead of crawling for hours. Finished answers are kept and resume next time. If your machine behaves differently, set `gpu_residency_gb` under `[machines.<id>]`.
|
|
91
|
+
|
|
92
|
+
## Server options
|
|
93
|
+
|
|
94
|
+
| option | meaning |
|
|
95
|
+
|---|---|
|
|
96
|
+
| `cmd` | command that starts the server (no `cmd` = an already-running server at `url`). Placeholders: `{model}`, `{served_name}`, `{port}`, `{mmproj}`, `{ctx}`, `{kv_type}`, `{root}` (the workspace), and the tune placeholders. An argument `env:NAME=value` sets an environment variable instead. |
|
|
97
|
+
| `url` | an already-running server's base URL |
|
|
98
|
+
| `port` | the port `cmd` serves on |
|
|
99
|
+
| `cwd` | folder to start `cmd` in |
|
|
100
|
+
| `model_is_path` | `model` is a file or folder: check it exists (and, for GGUFs, that it fits) before starting |
|
|
101
|
+
| `vision_args` | flags appended for models with `mmproj` |
|
|
102
|
+
| `kv_type`, `max_ctx` | llama defaults: KV-cache type, upper bound on context |
|
|
103
|
+
| `ctx_flag` | the server's context flag, so long packs are skipped if a configured context is too small |
|
|
104
|
+
| `request` | fields added to every request (e.g. `{ cache_prompt = false }`) |
|
|
105
|
+
| `perf` | speed-only flags always applied |
|
|
106
|
+
| `tune` | speed-only knobs for `tuieval tune` |
|
|
107
|
+
| `tune_objective` | `total` (default: workload time) or `decode` (tokens/s after a warm-up pass) |
|
|
108
|
+
| `version_cmd` | prints the server version, recorded with every result |
|
|
109
|
+
| `health` | readiness path for servers without `/v1/models` (must return JSON with `model`) |
|
|
110
|
+
| `before_start` | a command run before starting the server (e.g. to free memory another process holds) |
|
|
111
|
+
| `env` | environment variables for the server |
|
|
112
|
+
| `log_facts` | `{name = "regex"}` read from the server's startup log and recorded with every run |
|
|
113
|
+
| `require_facts` | `{name = "regex"}` the startup log must show, or the run refuses to start (e.g. a setting that must be off for fair evals) |
|
|
114
|
+
| `stall_guard` | watch for GPU-driver stalls (see above) |
|
|
115
|
+
| `outputs_depend_on_machine` | answers depend on this machine's memory settings: results only count on the machine that produced them, and tuned flags are fingerprinted |
|
|
116
|
+
| `any_model`, `label_prefix` | any model the server lists can be named at run time as `<server>:<id>` (OpenRouter) |
|
|
117
|
+
| `pin_endpoint` | pin each model to one provider endpoint (OpenRouter) |
|
|
118
|
+
| `api_key_env` | environment variable holding the API key |
|
|
119
|
+
| `thinking_param` | `reasoning` to send `[sampling] enable_thinking` as OpenRouter's `reasoning.enabled` |
|
|
120
|
+
| `headers` | extra HTTP headers |
|
|
@@ -0,0 +1,160 @@
|
|
|
1
|
+
# Writing eval packs
|
|
2
|
+
|
|
3
|
+
tuieval ships with no packs: you write the questions that matter for what you use models for. A pack is a folder in your workspace's `packs/` with a `pack.toml` and one or more test files. Every folder there with a `pack.toml` shows up in the TUI. Delete or rename a folder (prefix it with `_` to hide it) and it's gone.
|
|
4
|
+
|
|
5
|
+
The quickest start is a starter template with example tests that already pass `tuieval selftest`:
|
|
6
|
+
|
|
7
|
+
```bash
|
|
8
|
+
tuieval new-pack support-bot --grader reply # answer | rag | reply | tool_call | code
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
Edit `packs/support-bot/tests.yaml`, replace the examples with your own questions, and run `tuieval selftest support-bot` after every change.
|
|
12
|
+
|
|
13
|
+
To try a different set of questions, copy a pack (`cp -r packs/support-bot packs/support-bot-v2`), edit the copy, and pick whichever you want in the TUI.
|
|
14
|
+
|
|
15
|
+
## Fingerprints
|
|
16
|
+
|
|
17
|
+
Every pack has a fingerprint of what the model sees and what decides pass/fail (questions, images, system prompt, tools, expected answers, grader). Results record it, so after you edit a question the TUI marks older results `~ outdated` and reruns them instead of comparing different question sets. Editing a description, category, difficulty, reference answer, wrong answers or gate doesn't make results outdated.
|
|
18
|
+
|
|
19
|
+
## pack.toml
|
|
20
|
+
|
|
21
|
+
```toml
|
|
22
|
+
label = "Support bot" # shown in the TUI
|
|
23
|
+
group = "Customer support" # the use case it counts toward (see below)
|
|
24
|
+
description = "…"
|
|
25
|
+
grader = "reply" # default grader for its tests (see Graders)
|
|
26
|
+
system = "system.txt" # system prompt file (optional)
|
|
27
|
+
tools = "tools.yaml" # OpenAI-style tool definitions sent with every request (optional)
|
|
28
|
+
needs = ["vision"] # vision | tools | long_context | <python module>: skipped where unavailable
|
|
29
|
+
order = 30 # position in the list
|
|
30
|
+
tests = ["tests.yaml"] # optional; default: tests.* first, then every other .yaml/.csv file
|
|
31
|
+
screen = 20 # tests in a Screen run (spread across categories, one variant per group)
|
|
32
|
+
|
|
33
|
+
[certify]
|
|
34
|
+
repeat = 3 # repeats in a Certify run (default: models.toml [defaults] repeat)
|
|
35
|
+
|
|
36
|
+
[gate] # release criteria; see "Gates" below
|
|
37
|
+
min_accuracy = 0.90 # the 95% lower bound of the pass rate must reach this to PASS
|
|
38
|
+
max_critical_rate = 0.01 # needs 3/rate failure-free critical trials (rule of three): 300 here
|
|
39
|
+
max_p90_s = 30 # 90th-percentile seconds per answer, judged per machine
|
|
40
|
+
max_truncation = 0.01 # share of answers cut off by max_tokens
|
|
41
|
+
min_consistency = 0.95 # share of variant groups where every variant passes
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
`group` doubles as the **use case** in the readiness verdicts: all packs in a group must PASS for the use case to PASS. For example, a "Coding" use case could be a `coding` pack plus a `coding-sql` pack.
|
|
45
|
+
|
|
46
|
+
`needs` entries other than `vision`, `tools` and `long_context` name Python modules the pack's grading needs (for example `pandas`, when your hidden tests use it). The pack is skipped, with a note, where that module isn't installed in tuieval's Python environment.
|
|
47
|
+
|
|
48
|
+
## Tests (YAML)
|
|
49
|
+
|
|
50
|
+
```yaml
|
|
51
|
+
- id: refund-window # optional, must be unique in the pack (default: from description)
|
|
52
|
+
description: refund window # shown while running and in results
|
|
53
|
+
category: policy # optional grouping (Screen runs spread across categories)
|
|
54
|
+
difficulty: medium # easy | medium | hard: your label, shown in the TUI, never sent to the model
|
|
55
|
+
input: | # the user message
|
|
56
|
+
A customer bought shoes 20 days ago …
|
|
57
|
+
image: images/receipt.png # optional: sent as an image (needs vision)
|
|
58
|
+
grader: answer # optional: override the pack's grader
|
|
59
|
+
group: refund-window-2 # optional: variants of one case share a group (consistency gate)
|
|
60
|
+
critical: true # optional: any failure of this test is critical (disqualifying)
|
|
61
|
+
expected: 30 # grader-specific fields from here on
|
|
62
|
+
tolerance: 0
|
|
63
|
+
reference: "…\nANSWER: 30" # a correct model answer (never sent to the model)
|
|
64
|
+
wrong: ["ANSWER: 14"] # known-bad answers the grader must reject
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
**`reference` and `wrong` make a test check itself.** `tuieval selftest` grades every reference (it must pass with full score) and every wrong answer (it must fail). A test whose expected answer is wrong, or whose grader can't tell a classic mistake from the right answer, is caught before it fails a good model. Give every new test a reference, and a `wrong` entry for the mistake you most expect.
|
|
68
|
+
|
|
69
|
+
`selftest` also checks that no two tests send the same prompt, that tool tests expect tools the pack defines, that no `TODO` placeholders are left, that each request is one fresh single-turn conversation, and that the gate is reachable (see below).
|
|
70
|
+
|
|
71
|
+
## Tests (CSV)
|
|
72
|
+
|
|
73
|
+
For question-and-answer packs, a spreadsheet is often easier. Columns: `id, input, expected, tolerance, expected_text, category` (and any other test field). Empty cells are ignored.
|
|
74
|
+
|
|
75
|
+
```csv
|
|
76
|
+
id,input,expected,tolerance,expected_text,category
|
|
77
|
+
q1,"What is 15% of 80? End with ANSWER: <number>",12,0,,math
|
|
78
|
+
q2,"Capital of Peru? End with ANSWER: <city>",,,Lima,geo
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
## Graders
|
|
82
|
+
|
|
83
|
+
| grader | checks | test fields |
|
|
84
|
+
|---|---|---|
|
|
85
|
+
| `answer` | the last `ANSWER: …` line | `expected` (number or `NOT_AVAILABLE`) + `tolerance`, or `expected_text` (word, case-insensitive; a list means any of them) |
|
|
86
|
+
| `code` | last ```` ```python ```` block against hidden tests | `hidden_tests` with `# SETUP` / `# CHECK: name` blocks; score = fraction of checks passed |
|
|
87
|
+
| `rag` | answer grounded in passages, with citations | as `answer`, plus `expected_sources: [P3]`; half credit if the answer is right but the citation isn't |
|
|
88
|
+
| `tool_call` | the tool the model called | `expect_tool: {name, arguments}`, or `expect_no_tool: true` (+ `clarify: true`) |
|
|
89
|
+
| `reply` | a free-form reply | `max_words`, `no_markdown`, `must_include` (`"a\|b"` = either), `must_not_include`, `must_match`, `ends_with_question` |
|
|
90
|
+
|
|
91
|
+
⚠️ The `code` grader executes model-written code on your machine. Run code packs inside a container or VM with no credentials in the environment.
|
|
92
|
+
|
|
93
|
+
### Your own graders
|
|
94
|
+
|
|
95
|
+
A grader is a Python function registered by name. Put it in your workspace's `graders/` folder (one `.py` file each), or in a pack's own `grader.py` so the pack carries its grading with it:
|
|
96
|
+
|
|
97
|
+
```python
|
|
98
|
+
# graders/ticket_route.py
|
|
99
|
+
import json
|
|
100
|
+
|
|
101
|
+
from tuieval.graders import grader, result
|
|
102
|
+
|
|
103
|
+
|
|
104
|
+
@grader("ticket_route", template={"expected_queue": "TODO"})
|
|
105
|
+
def ticket_route(answer, test, meta):
|
|
106
|
+
"""The model replies with JSON like {"queue": "billing"}. Security-report tests are marked
|
|
107
|
+
`critical: true` in tests.yaml, so they count as critical trials."""
|
|
108
|
+
try:
|
|
109
|
+
queue = json.loads(answer).get("queue")
|
|
110
|
+
except (ValueError, AttributeError):
|
|
111
|
+
return result(False, "not a JSON object")
|
|
112
|
+
if queue != test["expected_queue"]:
|
|
113
|
+
# sending a security report anywhere but the security queue is disqualifying
|
|
114
|
+
severe = test["expected_queue"] == "security"
|
|
115
|
+
return result(False, f"routed to {queue!r}", severity="critical" if severe else None)
|
|
116
|
+
return result(True, f"routed to {queue!r}")
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
- `answer` is the model's final answer (reasoning already removed); `test` is the test's fields; `meta` has `finish` and `tool_calls`.
|
|
120
|
+
- `result(ok, reason, score=None, severity=None)`: `score` defaults to 1.0 or 0.0; `severity="critical"` marks a disqualifying failure.
|
|
121
|
+
- `critical=True` makes every test of this grader a critical trial (any of its failures *can* be critical). Otherwise only tests with `critical: true` or `expected: NOT_AVAILABLE` are.
|
|
122
|
+
- `template` is the skeleton `tuieval capture` writes for a new test of this grader.
|
|
123
|
+
|
|
124
|
+
Truncated and empty answers are failed before your grader runs. After changing a grader, `tuieval regrade` re-scores the stored answers without rerunning any model.
|
|
125
|
+
|
|
126
|
+
## Gates
|
|
127
|
+
|
|
128
|
+
A pack **PASSes** only on a full Certify run, when all of these hold:
|
|
129
|
+
|
|
130
|
+
- **zero critical failures**, over enough critical trials to show the critical rate is below `max_critical_rate` (rule of three: 300 clean trials for 1%, 60 for 5%)
|
|
131
|
+
- the **95% lower bound** of accuracy clears `min_accuracy`. Because it's the lower bound, the observed score must be higher, more so for small packs: with 30 trials, clearing an 80% bar takes about 29 correct.
|
|
132
|
+
- truncation, consistency and (per machine) p90 latency are within their budgets
|
|
133
|
+
|
|
134
|
+
`tuieval selftest` rejects a gate that a flawless certification run couldn't pass, and says what to add (tests, repeats, or a looser rate).
|
|
135
|
+
|
|
136
|
+
## What makes a failure critical
|
|
137
|
+
|
|
138
|
+
Graders mark failures that should disqualify a model whatever its accuracy:
|
|
139
|
+
|
|
140
|
+
- `answer`, `rag`: inventing an answer where the right one is `NOT_AVAILABLE`
|
|
141
|
+
- `tool_call`: an action nobody asked for (wrong tool, an unwanted or extra call)
|
|
142
|
+
- any test with `critical: true`
|
|
143
|
+
- whatever your own graders return with `severity="critical"`
|
|
144
|
+
|
|
145
|
+
## Difficulty labels
|
|
146
|
+
|
|
147
|
+
Every test has a `difficulty` of easy, medium or hard. It's for you, not the model: requests are built only from `input`, `image`, the system prompt and tools, so the label never reaches the model. The TUI shows it in the pack list (e.g. `12E 20M 8H`), next to the current question, in Recent results and Failures, and in Results → By difficulty.
|
|
148
|
+
|
|
149
|
+
A useful convention: **easy** = one rule or a direct read; **medium** = arithmetic, one trap, or two conditions; **hard** = rules in conflict, multi-step reasoning, or misleading data. Once several models have run, `tuieval items` flags labels the results contradict ("easier than labelled": a hard test every model passes; "harder than labelled": an easy test most models fail). `tuieval selftest` warns about tests without a label.
|
|
150
|
+
|
|
151
|
+
## Writing tests that separate production-ready models
|
|
152
|
+
|
|
153
|
+
- **Spec every behaviour you check.** A correct answer written only from the prompt must pass.
|
|
154
|
+
- **Test the boundaries** (exactly at a limit, one past it) and the conflicts between rules, not only typical cases.
|
|
155
|
+
- **Write each case twice** in a different surface form (layout, wording, field order) with the same `group`: a production model must not flip its answer.
|
|
156
|
+
- **Put traps where real use has them:** missing data, injected instructions, misleading visuals.
|
|
157
|
+
- **Use invented names, products and numbers**, so a model can't score from memory.
|
|
158
|
+
- **Generate volume with a script.** Critical gates need hundreds of trials; a small generator with a fixed seed that writes `generated.yaml` (references and wrong answers included) scales far better than hand-writing. Rerun it rather than editing its output.
|
|
159
|
+
- **After a few models have run, `tuieval items`** lists tests nobody fails (no signal), everybody fails (check the test), weak models beat strong ones on (inverted), or one model flips on (flaky).
|
|
160
|
+
- **Turn real failures into tests:** `tuieval capture logs/live/<file> --pack <name>` writes a skeleton to `captured.yaml` with the model's bad answer under `wrong`; fill in the TODOs.
|