tuieval 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (54) hide show
  1. tuieval-0.1.0/.gitignore +13 -0
  2. tuieval-0.1.0/CHANGELOG.md +12 -0
  3. tuieval-0.1.0/LICENSE +21 -0
  4. tuieval-0.1.0/PKG-INFO +156 -0
  5. tuieval-0.1.0/README.md +130 -0
  6. tuieval-0.1.0/docs/images/tui-setup.png +0 -0
  7. tuieval-0.1.0/docs/models.md +120 -0
  8. tuieval-0.1.0/docs/writing-packs.md +160 -0
  9. tuieval-0.1.0/packaging/homebrew/README.md +19 -0
  10. tuieval-0.1.0/packaging/homebrew/tuieval.rb +36 -0
  11. tuieval-0.1.0/pyproject.toml +47 -0
  12. tuieval-0.1.0/src/tuieval/__init__.py +2 -0
  13. tuieval-0.1.0/src/tuieval/__main__.py +5 -0
  14. tuieval-0.1.0/src/tuieval/cli.py +78 -0
  15. tuieval-0.1.0/src/tuieval/client.py +196 -0
  16. tuieval-0.1.0/src/tuieval/compare.py +511 -0
  17. tuieval-0.1.0/src/tuieval/engine.py +1827 -0
  18. tuieval-0.1.0/src/tuieval/graders/__init__.py +138 -0
  19. tuieval-0.1.0/src/tuieval/graders/answer.py +52 -0
  20. tuieval-0.1.0/src/tuieval/graders/code.py +104 -0
  21. tuieval-0.1.0/src/tuieval/graders/rag.py +33 -0
  22. tuieval-0.1.0/src/tuieval/graders/reply.py +44 -0
  23. tuieval-0.1.0/src/tuieval/graders/tool_call.py +75 -0
  24. tuieval-0.1.0/src/tuieval/machines.py +293 -0
  25. tuieval-0.1.0/src/tuieval/packs.py +239 -0
  26. tuieval-0.1.0/src/tuieval/profiles.py +98 -0
  27. tuieval-0.1.0/src/tuieval/run_evals.py +588 -0
  28. tuieval-0.1.0/src/tuieval/scaffold.py +106 -0
  29. tuieval-0.1.0/src/tuieval/selftest.py +185 -0
  30. tuieval-0.1.0/src/tuieval/templates/models.toml +137 -0
  31. tuieval-0.1.0/src/tuieval/templates/packs/answer/pack.toml +17 -0
  32. tuieval-0.1.0/src/tuieval/templates/packs/answer/system.txt +4 -0
  33. tuieval-0.1.0/src/tuieval/templates/packs/answer/tests.yaml +67 -0
  34. tuieval-0.1.0/src/tuieval/templates/packs/code/pack.toml +16 -0
  35. tuieval-0.1.0/src/tuieval/templates/packs/code/system.txt +3 -0
  36. tuieval-0.1.0/src/tuieval/templates/packs/code/tests.yaml +197 -0
  37. tuieval-0.1.0/src/tuieval/templates/packs/rag/pack.toml +16 -0
  38. tuieval-0.1.0/src/tuieval/templates/packs/rag/system.txt +9 -0
  39. tuieval-0.1.0/src/tuieval/templates/packs/rag/tests.yaml +99 -0
  40. tuieval-0.1.0/src/tuieval/templates/packs/reply/pack.toml +15 -0
  41. tuieval-0.1.0/src/tuieval/templates/packs/reply/system.txt +2 -0
  42. tuieval-0.1.0/src/tuieval/templates/packs/reply/tests.yaml +67 -0
  43. tuieval-0.1.0/src/tuieval/templates/packs/tool_call/pack.toml +18 -0
  44. tuieval-0.1.0/src/tuieval/templates/packs/tool_call/system.txt +2 -0
  45. tuieval-0.1.0/src/tuieval/templates/packs/tool_call/tests.yaml +59 -0
  46. tuieval-0.1.0/src/tuieval/templates/packs/tool_call/tools.yaml +32 -0
  47. tuieval-0.1.0/src/tuieval/tui.py +2558 -0
  48. tuieval-0.1.0/src/tuieval/tune.py +518 -0
  49. tuieval-0.1.0/src/tuieval/verdict.py +346 -0
  50. tuieval-0.1.0/src/tuieval/watch_proxy.py +260 -0
  51. tuieval-0.1.0/src/tuieval/workspace.py +27 -0
  52. tuieval-0.1.0/src/tuieval/yamlout.py +21 -0
  53. tuieval-0.1.0/tests/mock_server.py +121 -0
  54. tuieval-0.1.0/tests/test_tuieval.py +197 -0
@@ -0,0 +1,13 @@
1
+ __pycache__/
2
+ *.py[cod]
3
+ *.egg-info/
4
+ .venv/
5
+ build/
6
+ dist/
7
+ .DS_Store
8
+ # a workspace used while developing; never commit your own packs or results here
9
+ /my-evals/
10
+ /results/
11
+ /logs/
12
+ /tuning/
13
+ /packs/
@@ -0,0 +1,12 @@
1
+ # Changelog
2
+
3
+ ## 0.1.0
4
+
5
+ First public release.
6
+
7
+ - TUI and CLI for evaluating local models on your own eval packs: accuracy, speed and token use, with PASS / FAIL / INCONCLUSIVE verdicts per use case.
8
+ - Workspaces (`tuieval init`, `--workspace`, `TUIEVAL_HOME`) keep your packs, models and results apart from the tool.
9
+ - `tuieval new-pack` starter templates for the built-in graders: answer, rag, reply, tool_call, code.
10
+ - Custom graders from the workspace's `graders/` folder or a pack's own `grader.py`.
11
+ - Servers: llama.cpp (with per-machine fit check and speed tuning), any OpenAI-compatible server, OpenRouter (pinned to one provider endpoint).
12
+ - Speed tuning uses a built-in workload, so it needs no packs.
tuieval-0.1.0/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 tuieval contributors
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
tuieval-0.1.0/PKG-INFO ADDED
@@ -0,0 +1,156 @@
1
+ Metadata-Version: 2.5
2
+ Name: tuieval
3
+ Version: 0.1.0
4
+ Summary: Evaluate local LLMs on your own eval packs, in the terminal: accuracy, speed and tokens, with PASS/FAIL verdicts.
5
+ Project-URL: Homepage, https://github.com/ashe-wb/tuieval
6
+ Project-URL: Issues, https://github.com/ashe-wb/tuieval/issues
7
+ Author: tuieval contributors
8
+ License-Expression: MIT
9
+ License-File: LICENSE
10
+ Keywords: benchmark,evals,evaluation,llama.cpp,llm,local-llm,tui
11
+ Classifier: Development Status :: 4 - Beta
12
+ Classifier: Environment :: Console :: Curses
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: Intended Audience :: Science/Research
15
+ Classifier: Operating System :: MacOS
16
+ Classifier: Operating System :: POSIX :: Linux
17
+ Classifier: Programming Language :: Python :: 3
18
+ Classifier: Programming Language :: Python :: 3.11
19
+ Classifier: Programming Language :: Python :: 3.12
20
+ Classifier: Programming Language :: Python :: 3.13
21
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
22
+ Requires-Python: >=3.11
23
+ Requires-Dist: pyyaml>=6
24
+ Requires-Dist: textual<9,>=8
25
+ Description-Content-Type: text/markdown
26
+
27
+ # tuieval
28
+
29
+ **Evaluate local LLMs on your own questions, in the terminal.** Compare models (tiny, dense, MoE; llama.cpp GGUFs, LM Studio, Ollama, vLLM, OpenRouter) on what *you* use them for, measuring **accuracy, speed and token use** together, and get a **PASS / FAIL / INCONCLUSIVE** verdict per use case. Grading is automatic; there's no LLM judge.
30
+
31
+ tuieval ships with **no built-in benchmark**. Public benchmarks leak into training data and rarely match your work. You write *eval packs* (folders of your own questions with checkable answers), and tuieval runs them, grades them, and tells you which model is ready for that job on this machine.
32
+
33
+ ```bash
34
+ pipx install tuieval # or: pip install tuieval (Python 3.11+)
35
+ tuieval init my-evals && cd my-evals
36
+ tuieval new-pack my-first-pack # a pack of example questions to edit
37
+ tuieval add ~/models/Some-Model-Q4_K_M.gguf
38
+ tuieval # open the TUI
39
+ ```
40
+
41
+ Once you've added your own packs and models, the setup screen looks like this (packs on the left, models with a verdict code per use case on the right, the highlighted model's details below):
42
+
43
+ ![tuieval's setup screen with five eval packs and a dozen local models](https://raw.githubusercontent.com/ashe-wb/tuieval/main/docs/images/tui-setup.png)
44
+
45
+ ## What you get
46
+
47
+ - **A TUI** to pick packs and models, watch reasoning and answers stream live with TTFT, tokens/s, memory and a running score, and browse results.
48
+ - **Verdicts you can act on.** A pack passes only with zero critical failures over enough trials, and an accuracy whose 95% lower bound clears your bar. *INCONCLUSIVE* says what evidence is missing.
49
+ - **Two tiers:** *Screen* a sample of every pack to drop weak models fast, then *Certify* finalists on every test with repeats. Certification reuses the screening answers.
50
+ - **Honest numbers:** every request is a fresh single-turn conversation, with prompt caching off and repeat rounds in different orders. Each answer records what was sent, and the run warns about reused prompts or identical repeats.
51
+ - **Per-machine speed.** A fit check picks the largest context that fits your Mac's GPU memory, `tuieval tune` finds the fastest speed-only server flags (with a guard that rejects flags that change answers), and readiness includes a *Fast enough?* table per machine, measured or projected.
52
+ - **Tests that check themselves.** Every test carries a reference answer and known-wrong answers, and `tuieval selftest` checks the grader accepts the first and rejects the second, and that each gate is reachable at all.
53
+ - **History.** Verdict changes are appended to `results/verdicts.jsonl`, and replaced results are moved to `history/`, never overwritten.
54
+
55
+ ## Eval packs
56
+
57
+ A pack is a folder in your workspace's `packs/`:
58
+
59
+ ```
60
+ packs/support-bot/
61
+ pack.toml label, use case, grader, gate (what PASS means)
62
+ system.txt the system prompt
63
+ tests.yaml the questions, with expected answers, references and known-wrong answers
64
+ ```
65
+
66
+ ```yaml
67
+ - id: refund-window
68
+ difficulty: medium
69
+ input: A customer bought shoes 20 days ago and wants a refund. Our policy allows 30 days. Can they get one?
70
+ max_words: 40
71
+ must_include: ["yes|can"]
72
+ reference: "Yes, they're within the 30-day window, so they can get a refund."
73
+ wrong: ["No, the refund window has passed."]
74
+ ```
75
+
76
+ Built-in graders: `answer` (a number or word on an `ANSWER:` line, or a correct refusal when the data isn't there), `rag` (grounded answers with citations), `reply` (free-form replies against rules), `tool_call` (the right tool with the right arguments, or rightly none) and `code` (Python run against hidden tests). Your own grader is one Python file in the workspace's `graders/` folder.
77
+
78
+ `tuieval new-pack <name> --grader <grader>` creates a pack with working examples for any grader. **See [docs/writing-packs.md](docs/writing-packs.md)** for every field, gates, critical failures, difficulty labels, custom graders and how to write tests that separate good models from weak ones.
79
+
80
+ ## The workspace
81
+
82
+ Your packs, models and results live in a **workspace** folder, apart from the tool:
83
+
84
+ ```
85
+ my-evals/
86
+ models.toml your servers and models
87
+ packs/ your eval packs
88
+ graders/ your own graders (optional)
89
+ results/ logs/ tuning/ reports/ presets.toml written by tuieval
90
+ ```
91
+
92
+ tuieval uses the current folder, or `--workspace DIR` / `TUIEVAL_HOME`. Keep it in git or a synced folder: results from one machine count on another when the answer-changing settings match, so you can Certify on your fastest machine.
93
+
94
+ ## Models
95
+
96
+ ```bash
97
+ tuieval add ~/models/Some-Model-Q4_K_M.gguf --tags 9b,dense,q4 # llama.cpp (llama-server on PATH)
98
+ tuieval add ~/models/VL-Q4.gguf --mmproj ~/models/VL-mmproj.gguf # vision
99
+ tuieval add qwen3:8b --server local # an already-running server (set its url)
100
+ tuieval run --tier smoke --only openrouter:qwen/qwen3-32b # any OpenRouter model, no config needed
101
+ ```
102
+
103
+ See **[docs/models.md](docs/models.md)** for servers, OpenRouter endpoint pinning, machines, tuning and every option.
104
+
105
+ ## The TUI
106
+
107
+ 1. **Pick packs and models** (space ticks; type to filter models by label or tag). Each model shows a short code per use case, e.g. `C✓ S?` (✓ pass, ✗ fail, ? inconclusive; grey = from earlier results).
108
+ 2. **Pick a tier** (Smoke to check setup, Screen, Certify) and **press `s`**. The line above the buttons shows how many answers that is and roughly how long it will take. For each model, tuieval starts its server (or uses a running one), checks the right model is loaded, runs every selected pack, then stops it.
109
+ 3. **Watch the run:** progress with ETA, live score, tok/s, TTFT and memory, the current test with reasoning and answer side by side, recent results with the grader's reason. `k` skips a model, `c` cancels (finished work is kept and resumes next time). Runs started while one is going are queued.
110
+ 4. **Press `r` for results:** production readiness, verdict history, scorecard, speed & tokens (★ marks models nothing beats on both accuracy and time), per question, is the difference real?, tests that separate models, by difficulty, failures, and test quality.
111
+
112
+ Other keys: `t` tunes the ticked models' speed flags, `a` adds a model, `m` scans for GGUFs, `p` saves or loads a preset, `x` hides a model.
113
+
114
+ ## Command line
115
+
116
+ ```bash
117
+ tuieval run # Screen every model on every pack
118
+ tuieval run --tier certify --only a,b --packs support-bot,coding
119
+ tuieval run --dry-run # the plan and server commands
120
+ tuieval verdict # PASS / FAIL / INCONCLUSIVE per model and use case
121
+ tuieval report # the same with evidence, as a markdown file
122
+ tuieval history # every verdict change over time
123
+ tuieval compare --speed --pairwise --failures # scorecards
124
+ tuieval selftest # check every test's reference and wrong answers
125
+ tuieval items # tests that don't separate models or look broken
126
+ tuieval capture logs/live/<file> --pack X # turn a real failure into a new test
127
+ tuieval regrade # re-score stored answers after changing a grader
128
+ tuieval machines # this machine and others, fit and tuning per model
129
+ tuieval tune <model> # fastest speed flags for a model on this machine
130
+ tuieval watch --upstream http://localhost:8080 # show the reasoning of any app using your server
131
+ tuieval help
132
+ ```
133
+
134
+ ## Reading results
135
+
136
+ - **Start with Production readiness.** A single critical failure already means FAIL, whatever the accuracy: a model that breaks a hard rule once in 300 answers will do it in production. The Failures tab and `tuieval report` list exactly which answers failed.
137
+ - **INCONCLUSIVE is not "nearly passed".** It says what's missing (usually a Certify run, or more trials).
138
+ - **Trust "Is the difference real?" over raw percentages.** Differences of one or two tests are usually noise.
139
+ - **Read `trunc` before accuracy.** A model that runs out of tokens isn't wrong; it's thinking too long for the budget.
140
+ - **Watch TTFT for interactive use.** A model that is 5% more accurate but takes 3 s longer to start answering may be the worse choice.
141
+
142
+ ## Safety
143
+
144
+ ⚠️ The `code` grader executes model-written code on your machine. Run code packs inside a container or VM with no credentials in the environment.
145
+
146
+ ## Development
147
+
148
+ ```bash
149
+ git clone https://github.com/ashe-wb/tuieval && cd tuieval
150
+ python -m venv .venv && .venv/bin/pip install -e .
151
+ .venv/bin/python -m unittest discover -s tests # uses a mock server; no model needed
152
+ ```
153
+
154
+ ## License
155
+
156
+ MIT. See [LICENSE](LICENSE).
@@ -0,0 +1,130 @@
1
+ # tuieval
2
+
3
+ **Evaluate local LLMs on your own questions, in the terminal.** Compare models (tiny, dense, MoE; llama.cpp GGUFs, LM Studio, Ollama, vLLM, OpenRouter) on what *you* use them for, measuring **accuracy, speed and token use** together, and get a **PASS / FAIL / INCONCLUSIVE** verdict per use case. Grading is automatic; there's no LLM judge.
4
+
5
+ tuieval ships with **no built-in benchmark**. Public benchmarks leak into training data and rarely match your work. You write *eval packs* (folders of your own questions with checkable answers), and tuieval runs them, grades them, and tells you which model is ready for that job on this machine.
6
+
7
+ ```bash
8
+ pipx install tuieval # or: pip install tuieval (Python 3.11+)
9
+ tuieval init my-evals && cd my-evals
10
+ tuieval new-pack my-first-pack # a pack of example questions to edit
11
+ tuieval add ~/models/Some-Model-Q4_K_M.gguf
12
+ tuieval # open the TUI
13
+ ```
14
+
15
+ Once you've added your own packs and models, the setup screen looks like this (packs on the left, models with a verdict code per use case on the right, the highlighted model's details below):
16
+
17
+ ![tuieval's setup screen with five eval packs and a dozen local models](https://raw.githubusercontent.com/ashe-wb/tuieval/main/docs/images/tui-setup.png)
18
+
19
+ ## What you get
20
+
21
+ - **A TUI** to pick packs and models, watch reasoning and answers stream live with TTFT, tokens/s, memory and a running score, and browse results.
22
+ - **Verdicts you can act on.** A pack passes only with zero critical failures over enough trials, and an accuracy whose 95% lower bound clears your bar. *INCONCLUSIVE* says what evidence is missing.
23
+ - **Two tiers:** *Screen* a sample of every pack to drop weak models fast, then *Certify* finalists on every test with repeats. Certification reuses the screening answers.
24
+ - **Honest numbers:** every request is a fresh single-turn conversation, with prompt caching off and repeat rounds in different orders. Each answer records what was sent, and the run warns about reused prompts or identical repeats.
25
+ - **Per-machine speed.** A fit check picks the largest context that fits your Mac's GPU memory, `tuieval tune` finds the fastest speed-only server flags (with a guard that rejects flags that change answers), and readiness includes a *Fast enough?* table per machine, measured or projected.
26
+ - **Tests that check themselves.** Every test carries a reference answer and known-wrong answers, and `tuieval selftest` checks the grader accepts the first and rejects the second, and that each gate is reachable at all.
27
+ - **History.** Verdict changes are appended to `results/verdicts.jsonl`, and replaced results are moved to `history/`, never overwritten.
28
+
29
+ ## Eval packs
30
+
31
+ A pack is a folder in your workspace's `packs/`:
32
+
33
+ ```
34
+ packs/support-bot/
35
+ pack.toml label, use case, grader, gate (what PASS means)
36
+ system.txt the system prompt
37
+ tests.yaml the questions, with expected answers, references and known-wrong answers
38
+ ```
39
+
40
+ ```yaml
41
+ - id: refund-window
42
+ difficulty: medium
43
+ input: A customer bought shoes 20 days ago and wants a refund. Our policy allows 30 days. Can they get one?
44
+ max_words: 40
45
+ must_include: ["yes|can"]
46
+ reference: "Yes, they're within the 30-day window, so they can get a refund."
47
+ wrong: ["No, the refund window has passed."]
48
+ ```
49
+
50
+ Built-in graders: `answer` (a number or word on an `ANSWER:` line, or a correct refusal when the data isn't there), `rag` (grounded answers with citations), `reply` (free-form replies against rules), `tool_call` (the right tool with the right arguments, or rightly none) and `code` (Python run against hidden tests). Your own grader is one Python file in the workspace's `graders/` folder.
51
+
52
+ `tuieval new-pack <name> --grader <grader>` creates a pack with working examples for any grader. **See [docs/writing-packs.md](docs/writing-packs.md)** for every field, gates, critical failures, difficulty labels, custom graders and how to write tests that separate good models from weak ones.
53
+
54
+ ## The workspace
55
+
56
+ Your packs, models and results live in a **workspace** folder, apart from the tool:
57
+
58
+ ```
59
+ my-evals/
60
+ models.toml your servers and models
61
+ packs/ your eval packs
62
+ graders/ your own graders (optional)
63
+ results/ logs/ tuning/ reports/ presets.toml written by tuieval
64
+ ```
65
+
66
+ tuieval uses the current folder, or `--workspace DIR` / `TUIEVAL_HOME`. Keep it in git or a synced folder: results from one machine count on another when the answer-changing settings match, so you can Certify on your fastest machine.
67
+
68
+ ## Models
69
+
70
+ ```bash
71
+ tuieval add ~/models/Some-Model-Q4_K_M.gguf --tags 9b,dense,q4 # llama.cpp (llama-server on PATH)
72
+ tuieval add ~/models/VL-Q4.gguf --mmproj ~/models/VL-mmproj.gguf # vision
73
+ tuieval add qwen3:8b --server local # an already-running server (set its url)
74
+ tuieval run --tier smoke --only openrouter:qwen/qwen3-32b # any OpenRouter model, no config needed
75
+ ```
76
+
77
+ See **[docs/models.md](docs/models.md)** for servers, OpenRouter endpoint pinning, machines, tuning and every option.
78
+
79
+ ## The TUI
80
+
81
+ 1. **Pick packs and models** (space ticks; type to filter models by label or tag). Each model shows a short code per use case, e.g. `C✓ S?` (✓ pass, ✗ fail, ? inconclusive; grey = from earlier results).
82
+ 2. **Pick a tier** (Smoke to check setup, Screen, Certify) and **press `s`**. The line above the buttons shows how many answers that is and roughly how long it will take. For each model, tuieval starts its server (or uses a running one), checks the right model is loaded, runs every selected pack, then stops it.
83
+ 3. **Watch the run:** progress with ETA, live score, tok/s, TTFT and memory, the current test with reasoning and answer side by side, recent results with the grader's reason. `k` skips a model, `c` cancels (finished work is kept and resumes next time). Runs started while one is going are queued.
84
+ 4. **Press `r` for results:** production readiness, verdict history, scorecard, speed & tokens (★ marks models nothing beats on both accuracy and time), per question, is the difference real?, tests that separate models, by difficulty, failures, and test quality.
85
+
86
+ Other keys: `t` tunes the ticked models' speed flags, `a` adds a model, `m` scans for GGUFs, `p` saves or loads a preset, `x` hides a model.
87
+
88
+ ## Command line
89
+
90
+ ```bash
91
+ tuieval run # Screen every model on every pack
92
+ tuieval run --tier certify --only a,b --packs support-bot,coding
93
+ tuieval run --dry-run # the plan and server commands
94
+ tuieval verdict # PASS / FAIL / INCONCLUSIVE per model and use case
95
+ tuieval report # the same with evidence, as a markdown file
96
+ tuieval history # every verdict change over time
97
+ tuieval compare --speed --pairwise --failures # scorecards
98
+ tuieval selftest # check every test's reference and wrong answers
99
+ tuieval items # tests that don't separate models or look broken
100
+ tuieval capture logs/live/<file> --pack X # turn a real failure into a new test
101
+ tuieval regrade # re-score stored answers after changing a grader
102
+ tuieval machines # this machine and others, fit and tuning per model
103
+ tuieval tune <model> # fastest speed flags for a model on this machine
104
+ tuieval watch --upstream http://localhost:8080 # show the reasoning of any app using your server
105
+ tuieval help
106
+ ```
107
+
108
+ ## Reading results
109
+
110
+ - **Start with Production readiness.** A single critical failure already means FAIL, whatever the accuracy: a model that breaks a hard rule once in 300 answers will do it in production. The Failures tab and `tuieval report` list exactly which answers failed.
111
+ - **INCONCLUSIVE is not "nearly passed".** It says what's missing (usually a Certify run, or more trials).
112
+ - **Trust "Is the difference real?" over raw percentages.** Differences of one or two tests are usually noise.
113
+ - **Read `trunc` before accuracy.** A model that runs out of tokens isn't wrong; it's thinking too long for the budget.
114
+ - **Watch TTFT for interactive use.** A model that is 5% more accurate but takes 3 s longer to start answering may be the worse choice.
115
+
116
+ ## Safety
117
+
118
+ ⚠️ The `code` grader executes model-written code on your machine. Run code packs inside a container or VM with no credentials in the environment.
119
+
120
+ ## Development
121
+
122
+ ```bash
123
+ git clone https://github.com/ashe-wb/tuieval && cd tuieval
124
+ python -m venv .venv && .venv/bin/pip install -e .
125
+ .venv/bin/python -m unittest discover -s tests # uses a mock server; no model needed
126
+ ```
127
+
128
+ ## License
129
+
130
+ MIT. See [LICENSE](LICENSE).
Binary file
@@ -0,0 +1,120 @@
1
+ # Models, servers and machines
2
+
3
+ Everything about models lives in your workspace's `models.toml`. `tuieval init` writes one with three servers ready to use:
4
+
5
+ | server | what it is |
6
+ |---|---|
7
+ | `llama` | [llama.cpp](https://github.com/ggml-org/llama.cpp)'s `llama-server`, started by tuieval for each GGUF model, with a per-machine fit check and speed tuning |
8
+ | `local` | any OpenAI-compatible server that's already running (LM Studio, Ollama, vLLM, a remote box): just a `url` |
9
+ | `openrouter` | any model on [OpenRouter](https://openrouter.ai/models), named at run time, pinned to one provider endpoint |
10
+
11
+ ## Adding models
12
+
13
+ ```bash
14
+ tuieval add ~/models/Some-Model-Q4_K_M.gguf --tags 9b,dense,q4 # GGUF -> llama
15
+ tuieval add ~/models/Some-Model-Q4_K_M.gguf --no-think # same model, thinking off (label …-nothink)
16
+ tuieval add ~/models/VL-Q4.gguf --mmproj ~/models/VL-mmproj.gguf # vision model
17
+ tuieval add qwen3:8b --server local # a model your running server serves
18
+ tuieval scan --add # every new GGUF under model_dirs
19
+ tuieval list # models, packs, result status
20
+ ```
21
+
22
+ In the TUI, `a` adds a model and `m` scans your model folders. Each model is one `[[models]]` block:
23
+
24
+ | field | meaning |
25
+ |---|---|
26
+ | `server` | which `[servers.*]` block serves it |
27
+ | `model` | GGUF path (llama) or model id (other servers) |
28
+ | `label` | names `results/<label>/`; default: the file or repo name, lowercased. Never rename a label that has results. |
29
+ | `served_name` | the model name sent with each request (default: the label for llama, the model id elsewhere) |
30
+ | `vision`, `mmproj` | the model reads images; `mmproj` is llama's vision projector and implies `vision` |
31
+ | `thinking = false` | run with thinking off |
32
+ | `tags = [...]` | filter by them in the TUI and with `--tags` |
33
+ | `sampling = {...}` | override `[sampling]` for this model |
34
+ | `server_args = [...]` | extra server flags (treated as answer-changing) |
35
+ | `ctx` (llama) | cap the context; the fit check may lower it further per machine |
36
+ | `max_context` (other servers) | the context the server provides; packs whose prompts need more are skipped |
37
+ | `kv_type` (llama) | e.g. `"q8_0"`: a quantized KV cache, which makes it a different model (give it its own label) |
38
+ | `tools = false` | packs that need tool calling are skipped |
39
+
40
+ Results record the settings each model actually ran with, and the results screen notes any differences between models.
41
+
42
+ ## OpenRouter
43
+
44
+ OpenRouter models need no block. Name any model from openrouter.ai/models when you run, e.g. `tuieval run --tier smoke --only openrouter:qwen/qwen3-235b-a22b-2507`. In the TUI, tick **+ openrouter: any model** at the bottom of the model list and pick from the live list (filter by typing; it shows context and price), or press `a` and type `openrouter:<model id>`. A wrong id is refused with the closest matches.
45
+
46
+ - Needs `OPENROUTER_API_KEY` exported in the shell that starts tuieval. A plain `tuieval run` never includes OpenRouter models, since they cost money.
47
+ - Vision, tool calling and context size come from OpenRouter's model list; packs the model can't do are skipped.
48
+ - **One provider per model.** Left alone, OpenRouter spreads requests over providers running different quantizations. `pin_endpoint = true` picks one endpoint per model (closest to the released weights first: bf16 > fp8 > undeclared > lower; then tool and sampling support, uptime and price), sends every request only there with no fallback, keeps it while it's offered, and records it with the results. Answers from any other provider aren't counted.
49
+ - Hosted providers cache shared prompt openings and that can't be switched off. It doesn't change answers, but those answers (⟲ in the run) are left out of TTFT and prompt-speed numbers.
50
+ - To keep providers that store or train on prompts (and could learn your tests) out entirely, add `request = { provider = { data_collection = "deny" } }` under `[servers.openrouter]`.
51
+
52
+ ## Every answer is independent
53
+
54
+ Each request is a fresh single-turn conversation: the system prompt and one question, never an earlier answer. Repeats go round by round, each round asking the tests in its own fixed shuffled order, so a question never follows itself. The llama server is told not to reuse earlier prompts (`request = { cache_prompt = false }`). Each answer records the messages sent and any prompt tokens the server reports reusing, and the run warns if that's ever not 0. If a question gets the exact same long answer twice with sampling on, the run flags it as a probable cached response. `tuieval selftest` checks all of this.
55
+
56
+ ## Answer-changing vs speed-only flags
57
+
58
+ The best server flags differ per model and per machine, so tuieval splits them by what they change:
59
+
60
+ | Changes **answers**: same everywhere, part of each result's fingerprint | Changes **only speed**: tuned per model per machine |
61
+ |---|---|
62
+ | model file and quant, KV-cache type, `--jinja`, `--reasoning-format`, sampling, mmproj (`cmd` in `models.toml`) | threads, batch sizes, flash attention, cache reuse (`perf` and `[servers.llama.tune]`) |
63
+
64
+ **Quality results travel.** Tuning never invalidates results, and a result from one machine counts on another as long as the answer-changing settings match. So you can Certify on your fastest machine and only tune (and optionally Screen) on the others: sync the workspace folder between them.
65
+
66
+ ## Machines
67
+
68
+ - **Machine id** is detected automatically (`m1max-32gb`, `m4pro-24gb`, …; set `EVALS_MACHINE` to rename it). Every result records the machine, the server version and the exact speed flags used. `tuieval machines` lists this machine and every other machine that has run the evals (they record themselves in `tuning/`).
69
+ - **Fit check (Apple Silicon, llama):** before starting a GGUF, tuieval reads its header (layers, KV heads, hybrid attention layers) and picks the largest context that fits this machine's GPU memory, capped at `max_ctx`. A model that can't fit at 8k context is skipped with *doesn't fit on <machine>* instead of swapping. Packs that need more context than fits are skipped. It never quantizes the KV cache on its own, since that changes answers.
70
+ - **Per-machine settings** go under `[machines.<id>]`: `memory_headroom_gb` (GPU memory kept free for macOS, default 4) and `gpu_residency_gb` (see the stall guard).
71
+
72
+ ## Tuning
73
+
74
+ `tuieval tune <model>` (or `t` on the setup screen, for the ticked models) finds the fastest speed flags for a model on this machine:
75
+
76
+ 1. If `llama-bench` is installed, it sweeps threads, micro-batch and flash attention first (fast, no server starts).
77
+ 2. Then it starts the real server with one knob changed at a time and times a fixed **built-in** workload (three short prompts, three medium ones and one ~8k-token prompt), so tuning needs no packs and speeds are comparable between workspaces.
78
+ 3. An **output guard** rejects any option that changes greedy answers beyond noise. Candidates that make macOS swap are rejected.
79
+
80
+ Expect 8–15 server starts, about 20–30 minutes for a 27B model, once per model per machine. The result is saved in `tuning/<machine>/<model>.toml` and used by every later run there. Models without a profile run with each knob's first option and show *untuned*. A profile is marked for retuning when the model file or server version changes.
81
+
82
+ The knobs are `[servers.<name>.tune]` in `models.toml`: each knob is a list of options, each option a list of flags. Placeholders: `{p}` P-cores, `{p_minus_2}`, `{all}` all cores, `{gpu_safe_gb}` (the GPU residency limit minus 2 GB; also `_minus_1`, `_minus_2`, `_plus_1`).
83
+
84
+ ## Speed verdicts
85
+
86
+ Speed is judged separately, per machine: the same model can be production-grade in quality and still too slow on a smaller machine. Results → Production readiness has a **Fast enough?** table: p90 seconds per answer against each pack's `max_p90_s`, for every machine. It's *measured* where the model ran, and otherwise *projected* from each answer's token counts and that machine's tuned speeds. `tuieval compare --machine <id>` shows the same on the command line.
87
+
88
+ ## The stall guard (Apple Silicon)
89
+
90
+ Measured on an M1 Max 32 GB: once the system's GPU allocations pass about half of RAM, the GPU driver evicts and re-maps memory on every GPU job, the server spends 70–95% of its CPU in the kernel while the GPU idles, and servers that submit many small GPU jobs slow to a crawl. For servers with `stall_guard = true`, tuieval samples the server's own vs kernel CPU time and the system's GPU allocation every 5 s during runs and tuning. If more than 70% of its CPU goes to the kernel for a minute, the run stops with an explanation instead of crawling for hours. Finished answers are kept and resume next time. If your machine behaves differently, set `gpu_residency_gb` under `[machines.<id>]`.
91
+
92
+ ## Server options
93
+
94
+ | option | meaning |
95
+ |---|---|
96
+ | `cmd` | command that starts the server (no `cmd` = an already-running server at `url`). Placeholders: `{model}`, `{served_name}`, `{port}`, `{mmproj}`, `{ctx}`, `{kv_type}`, `{root}` (the workspace), and the tune placeholders. An argument `env:NAME=value` sets an environment variable instead. |
97
+ | `url` | an already-running server's base URL |
98
+ | `port` | the port `cmd` serves on |
99
+ | `cwd` | folder to start `cmd` in |
100
+ | `model_is_path` | `model` is a file or folder: check it exists (and, for GGUFs, that it fits) before starting |
101
+ | `vision_args` | flags appended for models with `mmproj` |
102
+ | `kv_type`, `max_ctx` | llama defaults: KV-cache type, upper bound on context |
103
+ | `ctx_flag` | the server's context flag, so long packs are skipped if a configured context is too small |
104
+ | `request` | fields added to every request (e.g. `{ cache_prompt = false }`) |
105
+ | `perf` | speed-only flags always applied |
106
+ | `tune` | speed-only knobs for `tuieval tune` |
107
+ | `tune_objective` | `total` (default: workload time) or `decode` (tokens/s after a warm-up pass) |
108
+ | `version_cmd` | prints the server version, recorded with every result |
109
+ | `health` | readiness path for servers without `/v1/models` (must return JSON with `model`) |
110
+ | `before_start` | a command run before starting the server (e.g. to free memory another process holds) |
111
+ | `env` | environment variables for the server |
112
+ | `log_facts` | `{name = "regex"}` read from the server's startup log and recorded with every run |
113
+ | `require_facts` | `{name = "regex"}` the startup log must show, or the run refuses to start (e.g. a setting that must be off for fair evals) |
114
+ | `stall_guard` | watch for GPU-driver stalls (see above) |
115
+ | `outputs_depend_on_machine` | answers depend on this machine's memory settings: results only count on the machine that produced them, and tuned flags are fingerprinted |
116
+ | `any_model`, `label_prefix` | any model the server lists can be named at run time as `<server>:<id>` (OpenRouter) |
117
+ | `pin_endpoint` | pin each model to one provider endpoint (OpenRouter) |
118
+ | `api_key_env` | environment variable holding the API key |
119
+ | `thinking_param` | `reasoning` to send `[sampling] enable_thinking` as OpenRouter's `reasoning.enabled` |
120
+ | `headers` | extra HTTP headers |
@@ -0,0 +1,160 @@
1
+ # Writing eval packs
2
+
3
+ tuieval ships with no packs: you write the questions that matter for what you use models for. A pack is a folder in your workspace's `packs/` with a `pack.toml` and one or more test files. Every folder there with a `pack.toml` shows up in the TUI. Delete or rename a folder (prefix it with `_` to hide it) and it's gone.
4
+
5
+ The quickest start is a starter template with example tests that already pass `tuieval selftest`:
6
+
7
+ ```bash
8
+ tuieval new-pack support-bot --grader reply # answer | rag | reply | tool_call | code
9
+ ```
10
+
11
+ Edit `packs/support-bot/tests.yaml`, replace the examples with your own questions, and run `tuieval selftest support-bot` after every change.
12
+
13
+ To try a different set of questions, copy a pack (`cp -r packs/support-bot packs/support-bot-v2`), edit the copy, and pick whichever you want in the TUI.
14
+
15
+ ## Fingerprints
16
+
17
+ Every pack has a fingerprint of what the model sees and what decides pass/fail (questions, images, system prompt, tools, expected answers, grader). Results record it, so after you edit a question the TUI marks older results `~ outdated` and reruns them instead of comparing different question sets. Editing a description, category, difficulty, reference answer, wrong answers or gate doesn't make results outdated.
18
+
19
+ ## pack.toml
20
+
21
+ ```toml
22
+ label = "Support bot" # shown in the TUI
23
+ group = "Customer support" # the use case it counts toward (see below)
24
+ description = "…"
25
+ grader = "reply" # default grader for its tests (see Graders)
26
+ system = "system.txt" # system prompt file (optional)
27
+ tools = "tools.yaml" # OpenAI-style tool definitions sent with every request (optional)
28
+ needs = ["vision"] # vision | tools | long_context | <python module>: skipped where unavailable
29
+ order = 30 # position in the list
30
+ tests = ["tests.yaml"] # optional; default: tests.* first, then every other .yaml/.csv file
31
+ screen = 20 # tests in a Screen run (spread across categories, one variant per group)
32
+
33
+ [certify]
34
+ repeat = 3 # repeats in a Certify run (default: models.toml [defaults] repeat)
35
+
36
+ [gate] # release criteria; see "Gates" below
37
+ min_accuracy = 0.90 # the 95% lower bound of the pass rate must reach this to PASS
38
+ max_critical_rate = 0.01 # needs 3/rate failure-free critical trials (rule of three): 300 here
39
+ max_p90_s = 30 # 90th-percentile seconds per answer, judged per machine
40
+ max_truncation = 0.01 # share of answers cut off by max_tokens
41
+ min_consistency = 0.95 # share of variant groups where every variant passes
42
+ ```
43
+
44
+ `group` doubles as the **use case** in the readiness verdicts: all packs in a group must PASS for the use case to PASS. For example, a "Coding" use case could be a `coding` pack plus a `coding-sql` pack.
45
+
46
+ `needs` entries other than `vision`, `tools` and `long_context` name Python modules the pack's grading needs (for example `pandas`, when your hidden tests use it). The pack is skipped, with a note, where that module isn't installed in tuieval's Python environment.
47
+
48
+ ## Tests (YAML)
49
+
50
+ ```yaml
51
+ - id: refund-window # optional, must be unique in the pack (default: from description)
52
+ description: refund window # shown while running and in results
53
+ category: policy # optional grouping (Screen runs spread across categories)
54
+ difficulty: medium # easy | medium | hard: your label, shown in the TUI, never sent to the model
55
+ input: | # the user message
56
+ A customer bought shoes 20 days ago …
57
+ image: images/receipt.png # optional: sent as an image (needs vision)
58
+ grader: answer # optional: override the pack's grader
59
+ group: refund-window-2 # optional: variants of one case share a group (consistency gate)
60
+ critical: true # optional: any failure of this test is critical (disqualifying)
61
+ expected: 30 # grader-specific fields from here on
62
+ tolerance: 0
63
+ reference: "…\nANSWER: 30" # a correct model answer (never sent to the model)
64
+ wrong: ["ANSWER: 14"] # known-bad answers the grader must reject
65
+ ```
66
+
67
+ **`reference` and `wrong` make a test check itself.** `tuieval selftest` grades every reference (it must pass with full score) and every wrong answer (it must fail). A test whose expected answer is wrong, or whose grader can't tell a classic mistake from the right answer, is caught before it fails a good model. Give every new test a reference, and a `wrong` entry for the mistake you most expect.
68
+
69
+ `selftest` also checks that no two tests send the same prompt, that tool tests expect tools the pack defines, that no `TODO` placeholders are left, that each request is one fresh single-turn conversation, and that the gate is reachable (see below).
70
+
71
+ ## Tests (CSV)
72
+
73
+ For question-and-answer packs, a spreadsheet is often easier. Columns: `id, input, expected, tolerance, expected_text, category` (and any other test field). Empty cells are ignored.
74
+
75
+ ```csv
76
+ id,input,expected,tolerance,expected_text,category
77
+ q1,"What is 15% of 80? End with ANSWER: <number>",12,0,,math
78
+ q2,"Capital of Peru? End with ANSWER: <city>",,,Lima,geo
79
+ ```
80
+
81
+ ## Graders
82
+
83
+ | grader | checks | test fields |
84
+ |---|---|---|
85
+ | `answer` | the last `ANSWER: …` line | `expected` (number or `NOT_AVAILABLE`) + `tolerance`, or `expected_text` (word, case-insensitive; a list means any of them) |
86
+ | `code` | last ```` ```python ```` block against hidden tests | `hidden_tests` with `# SETUP` / `# CHECK: name` blocks; score = fraction of checks passed |
87
+ | `rag` | answer grounded in passages, with citations | as `answer`, plus `expected_sources: [P3]`; half credit if the answer is right but the citation isn't |
88
+ | `tool_call` | the tool the model called | `expect_tool: {name, arguments}`, or `expect_no_tool: true` (+ `clarify: true`) |
89
+ | `reply` | a free-form reply | `max_words`, `no_markdown`, `must_include` (`"a\|b"` = either), `must_not_include`, `must_match`, `ends_with_question` |
90
+
91
+ ⚠️ The `code` grader executes model-written code on your machine. Run code packs inside a container or VM with no credentials in the environment.
92
+
93
+ ### Your own graders
94
+
95
+ A grader is a Python function registered by name. Put it in your workspace's `graders/` folder (one `.py` file each), or in a pack's own `grader.py` so the pack carries its grading with it:
96
+
97
+ ```python
98
+ # graders/ticket_route.py
99
+ import json
100
+
101
+ from tuieval.graders import grader, result
102
+
103
+
104
+ @grader("ticket_route", template={"expected_queue": "TODO"})
105
+ def ticket_route(answer, test, meta):
106
+ """The model replies with JSON like {"queue": "billing"}. Security-report tests are marked
107
+ `critical: true` in tests.yaml, so they count as critical trials."""
108
+ try:
109
+ queue = json.loads(answer).get("queue")
110
+ except (ValueError, AttributeError):
111
+ return result(False, "not a JSON object")
112
+ if queue != test["expected_queue"]:
113
+ # sending a security report anywhere but the security queue is disqualifying
114
+ severe = test["expected_queue"] == "security"
115
+ return result(False, f"routed to {queue!r}", severity="critical" if severe else None)
116
+ return result(True, f"routed to {queue!r}")
117
+ ```
118
+
119
+ - `answer` is the model's final answer (reasoning already removed); `test` is the test's fields; `meta` has `finish` and `tool_calls`.
120
+ - `result(ok, reason, score=None, severity=None)`: `score` defaults to 1.0 or 0.0; `severity="critical"` marks a disqualifying failure.
121
+ - `critical=True` makes every test of this grader a critical trial (any of its failures *can* be critical). Otherwise only tests with `critical: true` or `expected: NOT_AVAILABLE` are.
122
+ - `template` is the skeleton `tuieval capture` writes for a new test of this grader.
123
+
124
+ Truncated and empty answers are failed before your grader runs. After changing a grader, `tuieval regrade` re-scores the stored answers without rerunning any model.
125
+
126
+ ## Gates
127
+
128
+ A pack **PASSes** only on a full Certify run, when all of these hold:
129
+
130
+ - **zero critical failures**, over enough critical trials to show the critical rate is below `max_critical_rate` (rule of three: 300 clean trials for 1%, 60 for 5%)
131
+ - the **95% lower bound** of accuracy clears `min_accuracy`. Because it's the lower bound, the observed score must be higher, more so for small packs: with 30 trials, clearing an 80% bar takes about 29 correct.
132
+ - truncation, consistency and (per machine) p90 latency are within their budgets
133
+
134
+ `tuieval selftest` rejects a gate that a flawless certification run couldn't pass, and says what to add (tests, repeats, or a looser rate).
135
+
136
+ ## What makes a failure critical
137
+
138
+ Graders mark failures that should disqualify a model whatever its accuracy:
139
+
140
+ - `answer`, `rag`: inventing an answer where the right one is `NOT_AVAILABLE`
141
+ - `tool_call`: an action nobody asked for (wrong tool, an unwanted or extra call)
142
+ - any test with `critical: true`
143
+ - whatever your own graders return with `severity="critical"`
144
+
145
+ ## Difficulty labels
146
+
147
+ Every test has a `difficulty` of easy, medium or hard. It's for you, not the model: requests are built only from `input`, `image`, the system prompt and tools, so the label never reaches the model. The TUI shows it in the pack list (e.g. `12E 20M 8H`), next to the current question, in Recent results and Failures, and in Results → By difficulty.
148
+
149
+ A useful convention: **easy** = one rule or a direct read; **medium** = arithmetic, one trap, or two conditions; **hard** = rules in conflict, multi-step reasoning, or misleading data. Once several models have run, `tuieval items` flags labels the results contradict ("easier than labelled": a hard test every model passes; "harder than labelled": an easy test most models fail). `tuieval selftest` warns about tests without a label.
150
+
151
+ ## Writing tests that separate production-ready models
152
+
153
+ - **Spec every behaviour you check.** A correct answer written only from the prompt must pass.
154
+ - **Test the boundaries** (exactly at a limit, one past it) and the conflicts between rules, not only typical cases.
155
+ - **Write each case twice** in a different surface form (layout, wording, field order) with the same `group`: a production model must not flip its answer.
156
+ - **Put traps where real use has them:** missing data, injected instructions, misleading visuals.
157
+ - **Use invented names, products and numbers**, so a model can't score from memory.
158
+ - **Generate volume with a script.** Critical gates need hundreds of trials; a small generator with a fixed seed that writes `generated.yaml` (references and wrong answers included) scales far better than hand-writing. Rerun it rather than editing its output.
159
+ - **After a few models have run, `tuieval items`** lists tests nobody fails (no signal), everybody fails (check the test), weak models beat strong ones on (inverted), or one model flips on (flaky).
160
+ - **Turn real failures into tests:** `tuieval capture logs/live/<file> --pack <name>` writes a skeleton to `captured.yaml` with the model's bad answer under `wrong`; fill in the TODOs.