evalmetry 1.0.0__tar.gz → 1.1.1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {evalmetry-1.0.0 → evalmetry-1.1.1}/PKG-INFO +11 -6
- {evalmetry-1.0.0 → evalmetry-1.1.1}/README.md +10 -5
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/backend.py +30 -1
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/debug.py +1 -1
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/main.py +82 -12
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/module_stats.py +12 -1
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry.egg-info/PKG-INFO +11 -6
- {evalmetry-1.0.0 → evalmetry-1.1.1}/pyproject.toml +1 -1
- {evalmetry-1.0.0 → evalmetry-1.1.1}/LICENSE +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/__init__.py +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/adapters.py +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/benchmarks.py +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/hooks.py +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/judges.py +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/models.py +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/recorder.py +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/reducers.py +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/report.py +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry/storage.py +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry.egg-info/SOURCES.txt +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry.egg-info/dependency_links.txt +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry.egg-info/entry_points.txt +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry.egg-info/requires.txt +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/evalmetry.egg-info/top_level.txt +0 -0
- {evalmetry-1.0.0 → evalmetry-1.1.1}/setup.cfg +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: evalmetry
|
|
3
|
-
Version: 1.
|
|
3
|
+
Version: 1.1.1
|
|
4
4
|
Summary: A research toolkit for measuring and analyzing model behavior.
|
|
5
5
|
License-Expression: MIT
|
|
6
6
|
Project-URL: Homepage, https://github.com/minguinho26/evalmetry
|
|
@@ -52,9 +52,13 @@ Evalmetry evaluates Hugging Face language models with lm-eval and collects inter
|
|
|
52
52
|
|
|
53
53
|
Model evaluation supports one process on one CUDA GPU. CPU execution is for test fixtures and verification tools. Reports and saved-data readers do not require a GPU.
|
|
54
54
|
|
|
55
|
+
## Evidence and research example
|
|
56
|
+
|
|
57
|
+
Read the [1.1.0 verification report](https://github.com/minguinho26/evalmetry/blob/v1.1.1/docs/verification.md) and the [prediction-depth research showcase](https://github.com/minguinho26/evalmetry/blob/v1.1.1/docs/showcase.md). Generation parity is established at the same actual batch size. Generation is not batch-invariant.
|
|
58
|
+
|
|
55
59
|
## Install
|
|
56
60
|
|
|
57
|
-
Install from PyPI
|
|
61
|
+
Install from PyPI:
|
|
58
62
|
|
|
59
63
|
```bash
|
|
60
64
|
pip install evalmetry
|
|
@@ -62,8 +66,7 @@ pip install evalmetry
|
|
|
62
66
|
|
|
63
67
|
Evalmetry requires Python >=3.10. Use a compatible CUDA build of PyTorch. The Python import and CLI are both `evalmetry`.
|
|
64
68
|
|
|
65
|
-
To install this release explicitly, use `pip install evalmetry==1.
|
|
66
|
-
For development, clone this repository and run `python -m pip install -e .`.
|
|
69
|
+
To install this release explicitly, use `pip install evalmetry==1.1.1`. For development, clone this repository and run `python -m pip install -e .`.
|
|
67
70
|
|
|
68
71
|
## Quick start
|
|
69
72
|
|
|
@@ -81,14 +84,16 @@ Use `report` to generate a report from that run:
|
|
|
81
84
|
evalmetry report results/quickstart --output report/quickstart
|
|
82
85
|
```
|
|
83
86
|
|
|
84
|
-
The first run downloads the model and dataset if needed. Reusing a run directory resumes its recorded configuration
|
|
87
|
+
The first run downloads the model and dataset if needed. Reusing a run directory resumes its recorded configuration. A different experiment needs a different directory.
|
|
85
88
|
|
|
86
89
|
Other commands: `collect-research-data` adds optional tensors to a completed run, `debug` reads a saved module trace, and `module-stats` reads saved statistics. Traces and statistics must be enabled during collection. Use `evalmetry <command> --help` for options.
|
|
87
90
|
|
|
88
91
|
## Reading results
|
|
89
92
|
|
|
90
|
-
Evaluation results, manifests and collected signals are written under the selected output directory. Read saved signals with `evalmetry.load_signals(run_dir)
|
|
93
|
+
Evaluation results, manifests and collected signals are written under the selected output directory. Read saved signals with `evalmetry.load_signals(run_dir)`. `evalmetry.describe_schema()` describes the columns.
|
|
91
94
|
|
|
92
95
|
## License
|
|
93
96
|
|
|
94
97
|
Evalmetry is released under the MIT License.
|
|
98
|
+
|
|
99
|
+
This project is built on top of [EleutherAI's lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness).
|
|
@@ -13,9 +13,13 @@ Evalmetry evaluates Hugging Face language models with lm-eval and collects inter
|
|
|
13
13
|
|
|
14
14
|
Model evaluation supports one process on one CUDA GPU. CPU execution is for test fixtures and verification tools. Reports and saved-data readers do not require a GPU.
|
|
15
15
|
|
|
16
|
+
## Evidence and research example
|
|
17
|
+
|
|
18
|
+
Read the [1.1.0 verification report](https://github.com/minguinho26/evalmetry/blob/v1.1.1/docs/verification.md) and the [prediction-depth research showcase](https://github.com/minguinho26/evalmetry/blob/v1.1.1/docs/showcase.md). Generation parity is established at the same actual batch size. Generation is not batch-invariant.
|
|
19
|
+
|
|
16
20
|
## Install
|
|
17
21
|
|
|
18
|
-
Install from PyPI
|
|
22
|
+
Install from PyPI:
|
|
19
23
|
|
|
20
24
|
```bash
|
|
21
25
|
pip install evalmetry
|
|
@@ -23,8 +27,7 @@ pip install evalmetry
|
|
|
23
27
|
|
|
24
28
|
Evalmetry requires Python >=3.10. Use a compatible CUDA build of PyTorch. The Python import and CLI are both `evalmetry`.
|
|
25
29
|
|
|
26
|
-
To install this release explicitly, use `pip install evalmetry==1.
|
|
27
|
-
For development, clone this repository and run `python -m pip install -e .`.
|
|
30
|
+
To install this release explicitly, use `pip install evalmetry==1.1.1`. For development, clone this repository and run `python -m pip install -e .`.
|
|
28
31
|
|
|
29
32
|
## Quick start
|
|
30
33
|
|
|
@@ -42,14 +45,16 @@ Use `report` to generate a report from that run:
|
|
|
42
45
|
evalmetry report results/quickstart --output report/quickstart
|
|
43
46
|
```
|
|
44
47
|
|
|
45
|
-
The first run downloads the model and dataset if needed. Reusing a run directory resumes its recorded configuration
|
|
48
|
+
The first run downloads the model and dataset if needed. Reusing a run directory resumes its recorded configuration. A different experiment needs a different directory.
|
|
46
49
|
|
|
47
50
|
Other commands: `collect-research-data` adds optional tensors to a completed run, `debug` reads a saved module trace, and `module-stats` reads saved statistics. Traces and statistics must be enabled during collection. Use `evalmetry <command> --help` for options.
|
|
48
51
|
|
|
49
52
|
## Reading results
|
|
50
53
|
|
|
51
|
-
Evaluation results, manifests and collected signals are written under the selected output directory. Read saved signals with `evalmetry.load_signals(run_dir)
|
|
54
|
+
Evaluation results, manifests and collected signals are written under the selected output directory. Read saved signals with `evalmetry.load_signals(run_dir)`. `evalmetry.describe_schema()` describes the columns.
|
|
52
55
|
|
|
53
56
|
## License
|
|
54
57
|
|
|
55
58
|
Evalmetry is released under the MIT License.
|
|
59
|
+
|
|
60
|
+
This project is built on top of [EleutherAI's lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness).
|
|
@@ -30,6 +30,7 @@ Because that one is a copy, `check_upstream_source()` guards it: CI fails when l
|
|
|
30
30
|
from __future__ import annotations
|
|
31
31
|
|
|
32
32
|
import ast
|
|
33
|
+
import contextlib
|
|
33
34
|
import hashlib
|
|
34
35
|
import inspect
|
|
35
36
|
import os
|
|
@@ -579,7 +580,7 @@ Without this the backend behaves exactly like `HFLM`.
|
|
|
579
580
|
from tqdm import tqdm
|
|
580
581
|
|
|
581
582
|
results: list[str] = []
|
|
582
|
-
with self._recorder.session():
|
|
583
|
+
with self._recorder.session(), self._resolved_auto_batch_size():
|
|
583
584
|
for instance in tqdm(requests, disable=disable_tqdm,
|
|
584
585
|
desc="Running generate_until requests (traced)"):
|
|
585
586
|
self._recorder.expect_generation(
|
|
@@ -596,6 +597,34 @@ Without this the backend behaves exactly like `HFLM`.
|
|
|
596
597
|
results.extend(output)
|
|
597
598
|
return results
|
|
598
599
|
|
|
600
|
+
@contextlib.contextmanager
|
|
601
|
+
def _resolved_auto_batch_size(self):
|
|
602
|
+
"""Pin `batch_size="auto"` to one probed value for the duration of a generate loop.
|
|
603
|
+
|
|
604
|
+
The loop above calls `super().generate_until()` once per document, and lm-eval re-runs
|
|
605
|
+
`_detect_batch_size()` on every entry to that method while `batch_size` reads "auto".
|
|
606
|
+
The probed value is then discarded, because a one-instance call batches one instance
|
|
607
|
+
either way - so the probe is pure overhead, paid once per document. Measured on
|
|
608
|
+
SmolLM2-135M over 10 gsm8k documents at bfloat16 on an RTX A5000: the traced pass takes
|
|
609
|
+
8:07 with the probe repeated and 0:16 with it resolved once, against 0:17 for the same
|
|
610
|
+
pass at an explicit batch of 1. All three produce identical generations.
|
|
611
|
+
|
|
612
|
+
The probe is kept rather than forced to 1: it still decides what the scoring pass of a
|
|
613
|
+
mixed run may use, and skipping it would change behaviour rather than only its cost.
|
|
614
|
+
|
|
615
|
+
`HFLM.batch_size` is a read-only property over `batch_size_per_gpu`, so the pin is
|
|
616
|
+
written there and restored afterwards, leaving the model as this found it.
|
|
617
|
+
"""
|
|
618
|
+
if self.batch_size != "auto":
|
|
619
|
+
yield
|
|
620
|
+
return
|
|
621
|
+
original = self.batch_size_per_gpu
|
|
622
|
+
self.batch_size_per_gpu = self._detect_batch_size()
|
|
623
|
+
try:
|
|
624
|
+
yield
|
|
625
|
+
finally:
|
|
626
|
+
self.batch_size_per_gpu = original
|
|
627
|
+
|
|
599
628
|
def _model_generate(self, context, max_length, stop, **generation_kwargs):
|
|
600
629
|
"""Note the prompt length, then keep the token ids that came out.
|
|
601
630
|
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
This is a debugging aid, off by default, and deliberately separate from `recorder.py`.
|
|
4
4
|
The recorder exists to capture *signals* and is driven by what the reducers ask for; this exists to capture *what happened*, must keep working when the recorder is not recording (a batch-size probe, a document already on disk), and must survive the process that produced it.
|
|
5
5
|
|
|
6
|
-
CLI options are available through `evalmetry debug --help
|
|
6
|
+
CLI options are available through `evalmetry debug --help`.
|
|
7
7
|
Three things are worth repeating here, because they are what the code looks strange without.
|
|
8
8
|
|
|
9
9
|
**`always_call=True` is not used, on purpose.**
|
|
@@ -25,6 +25,7 @@ import re
|
|
|
25
25
|
import sys
|
|
26
26
|
import dataclasses
|
|
27
27
|
import functools
|
|
28
|
+
import time
|
|
28
29
|
from contextlib import contextmanager
|
|
29
30
|
from dataclasses import dataclass
|
|
30
31
|
from datetime import date, datetime, timezone
|
|
@@ -35,7 +36,7 @@ from .debug import TRACE_BUFFER_EVENTS
|
|
|
35
36
|
from .storage import FIXED_SETTINGS, SAMPLING_SEED, SCHEMA_VERSION
|
|
36
37
|
|
|
37
38
|
#: Version of this tool, recorded in every manifest.
|
|
38
|
-
TOOL_VERSION = "
|
|
39
|
+
TOOL_VERSION = "1.1.1"
|
|
39
40
|
|
|
40
41
|
#: Default number of documents collected per correctness group.
|
|
41
42
|
DEFAULT_COLLECT_LIMIT = 500
|
|
@@ -142,6 +143,7 @@ class RunConfig:
|
|
|
142
143
|
tasks: task names.
|
|
143
144
|
include_path: directories containing local benchmark YAML, JSONL and scoring code.
|
|
144
145
|
Paths inside YAML resolve relative to that YAML. See examples/custom_benchmarks.
|
|
146
|
+
num_fewshot: how many examples precede each document. It has no default because there is no value that is right for every task: lm-eval gives each task its own, and silently substituting one here would score gsm8k at 0-shot while the task itself asks for 5. A measurement has to say what it measured.
|
|
145
147
|
save_attention / save_hidden: opt-in signals, both collected in a second pass because correctness is not known during the first one.
|
|
146
148
|
debug: module tracing, off by default.
|
|
147
149
|
Absent from `identity` and `config_hash`: it changes how long a run takes, and with `sync` or `stop_on_nonfinite` whether it finishes, but not what the signals mean.
|
|
@@ -149,7 +151,7 @@ class RunConfig:
|
|
|
149
151
|
|
|
150
152
|
model_args: str
|
|
151
153
|
tasks: list[str]
|
|
152
|
-
num_fewshot: int
|
|
154
|
+
num_fewshot: int
|
|
153
155
|
limit: int | None = None
|
|
154
156
|
batch_size: int | str = 1
|
|
155
157
|
adapter: str | None = None
|
|
@@ -182,7 +184,7 @@ class RunConfig:
|
|
|
182
184
|
"""The lm-eval model argument string parsed into a dict.
|
|
183
185
|
|
|
184
186
|
Example:
|
|
185
|
-
>>> RunConfig("pretrained=Qwen/Qwen3-8B,dtype=bfloat16", []).model_kwargs()
|
|
187
|
+
>>> RunConfig("pretrained=Qwen/Qwen3-8B,dtype=bfloat16", [], num_fewshot=0).model_kwargs()
|
|
186
188
|
{'pretrained': 'Qwen/Qwen3-8B', 'dtype': 'bfloat16'}
|
|
187
189
|
"""
|
|
188
190
|
parsed: dict[str, str] = {}
|
|
@@ -261,7 +263,7 @@ class RunConfig:
|
|
|
261
263
|
Only a convenience for browsing; `report` reads the manifest and never parses this path.
|
|
262
264
|
|
|
263
265
|
Example:
|
|
264
|
-
>>> RunConfig("pretrained=Qwen/Qwen3-8B", ["xnli_ko"]).default_output("ab12cd34")
|
|
266
|
+
>>> RunConfig("pretrained=Qwen/Qwen3-8B", ["xnli_ko"], num_fewshot=0).default_output("ab12cd34")
|
|
265
267
|
'results/xnli_ko/Qwen__Qwen3-8B/2026-01-31-ab12cd34' # date varies
|
|
266
268
|
"""
|
|
267
269
|
task = "+".join(sorted(self.tasks)) or "unknown-task"
|
|
@@ -672,6 +674,60 @@ def _resolved_revision(lm: Any) -> str:
|
|
|
672
674
|
# --------------------------------------------------------------------------
|
|
673
675
|
|
|
674
676
|
|
|
677
|
+
class PhaseTimer:
|
|
678
|
+
"""How long each phase of a run took, recorded on every run rather than on request.
|
|
679
|
+
|
|
680
|
+
Not behind a flag. A run that cannot say how long its evaluation took cannot be
|
|
681
|
+
compared against another one, and reading a clock costs nothing beside the work it
|
|
682
|
+
measures - so the choice to record it is not the user's to make and not theirs to
|
|
683
|
+
forget. The numbers land in the manifest under `timings`.
|
|
684
|
+
|
|
685
|
+
CUDA queues kernels asynchronously, so a naive `perf_counter` pair reports how long
|
|
686
|
+
the work took to *launch*. Every boundary therefore synchronises first. That is done
|
|
687
|
+
at phase boundaries only, where a handful of syncs disappear into minutes of work;
|
|
688
|
+
`--debug-sync` is the per-module variant and exists to attribute time inside a trace,
|
|
689
|
+
not to measure a run.
|
|
690
|
+
|
|
691
|
+
Phases do not have to cover the whole run and may be absent: a run without a
|
|
692
|
+
collection pass simply has no `collection` key. `end_to_end` always spans the object's
|
|
693
|
+
lifetime, so it is not the sum of the others.
|
|
694
|
+
"""
|
|
695
|
+
|
|
696
|
+
def __init__(self) -> None:
|
|
697
|
+
self._phases: dict[str, float] = {}
|
|
698
|
+
self._started = time.perf_counter()
|
|
699
|
+
|
|
700
|
+
@staticmethod
|
|
701
|
+
def _synchronize() -> None:
|
|
702
|
+
"""Wait for queued CUDA work, when there is any to wait for.
|
|
703
|
+
|
|
704
|
+
Skipped unless a context already exists, so timing a CPU run does not initialise
|
|
705
|
+
CUDA just to measure nothing. Without torch at all there is no device work in
|
|
706
|
+
flight, so there is nothing to wait for either.
|
|
707
|
+
"""
|
|
708
|
+
try:
|
|
709
|
+
import torch
|
|
710
|
+
except ImportError:
|
|
711
|
+
return
|
|
712
|
+
if torch.cuda.is_available() and torch.cuda.is_initialized():
|
|
713
|
+
torch.cuda.synchronize()
|
|
714
|
+
|
|
715
|
+
@contextmanager
|
|
716
|
+
def phase(self, name: str):
|
|
717
|
+
"""Time one phase, recording it even if the phase raises."""
|
|
718
|
+
self._synchronize()
|
|
719
|
+
started = time.perf_counter()
|
|
720
|
+
try:
|
|
721
|
+
yield
|
|
722
|
+
finally:
|
|
723
|
+
self._synchronize()
|
|
724
|
+
self._phases[name] = round(time.perf_counter() - started, 3)
|
|
725
|
+
|
|
726
|
+
def finish(self) -> dict[str, float]:
|
|
727
|
+
self._synchronize()
|
|
728
|
+
return {**self._phases, "end_to_end": round(time.perf_counter() - self._started, 3)}
|
|
729
|
+
|
|
730
|
+
|
|
675
731
|
@contextmanager
|
|
676
732
|
def _tracing(config: RunConfig, model: Any, run_dir: str, name: str = "trace"):
|
|
677
733
|
"""Hook the model for the duration, or do nothing at all.
|
|
@@ -795,6 +851,7 @@ def cmd_run(config: RunConfig) -> str:
|
|
|
795
851
|
The run directory.
|
|
796
852
|
"""
|
|
797
853
|
validate_single_gpu_execution(config)
|
|
854
|
+
timer = PhaseTimer()
|
|
798
855
|
|
|
799
856
|
import lm_eval
|
|
800
857
|
|
|
@@ -811,7 +868,8 @@ def cmd_run(config: RunConfig) -> str:
|
|
|
811
868
|
unknown = set(config.signals) - {"logit_lens", "similarity"}
|
|
812
869
|
if unknown:
|
|
813
870
|
raise ValueError(f"unknown signals: {sorted(unknown)}")
|
|
814
|
-
|
|
871
|
+
with timer.phase("model_load"):
|
|
872
|
+
lm = load_model(config)
|
|
815
873
|
adapter = resolve_model_adapter(config, lm)
|
|
816
874
|
environment = check_environment(lm.model)
|
|
817
875
|
|
|
@@ -866,9 +924,10 @@ def cmd_run(config: RunConfig) -> str:
|
|
|
866
924
|
lm.attach_recorder(recorder)
|
|
867
925
|
|
|
868
926
|
# 3. 평가와 judge 채점을 완료하고 문서별 결과를 저장한다.
|
|
869
|
-
|
|
870
|
-
|
|
871
|
-
|
|
927
|
+
with timer.phase("evaluate"):
|
|
928
|
+
results, samples = _evaluate_and_save_samples(
|
|
929
|
+
config, lm, writer, run_dir, task_manager, evaluation_tasks
|
|
930
|
+
)
|
|
872
931
|
|
|
873
932
|
# Read before the collection pass, which may switch the model to eager attention: the manifest should say what the scored pass actually ran with.
|
|
874
933
|
attn_implementation = getattr(lm.model.config, "_attn_implementation", "unknown")
|
|
@@ -888,7 +947,7 @@ def cmd_run(config: RunConfig) -> str:
|
|
|
888
947
|
if config.save_attention or config.save_hidden or any(h.pass_name == "collection" for h in config.resolved_hooks):
|
|
889
948
|
# Its own trace: the collection pass forces eager attention and materialises a
|
|
890
949
|
# (heads, seq, seq) map per block, which is the heaviest thing this tool does.
|
|
891
|
-
with _tracing(config, lm.model, run_dir, name="collection"):
|
|
950
|
+
with timer.phase("collection"), _tracing(config, lm.model, run_dir, name="collection"):
|
|
892
951
|
collection_counts = _collect_research_data(
|
|
893
952
|
config, lm, adapter, writer, run_dir, hidden_layers
|
|
894
953
|
)
|
|
@@ -920,6 +979,7 @@ def cmd_run(config: RunConfig) -> str:
|
|
|
920
979
|
"evaluation_completed": True,
|
|
921
980
|
"generation_kwargs_source": "lm-eval task config",
|
|
922
981
|
"debug": config.debug.manifest_entry(),
|
|
982
|
+
"timings": timer.finish(),
|
|
923
983
|
},
|
|
924
984
|
)
|
|
925
985
|
storage.write_results(run_dir, manifest, results.get("results", {}), writer.signal_files())
|
|
@@ -1103,6 +1163,10 @@ def _collection_config_from_manifest(
|
|
|
1103
1163
|
model_config=saved_model.get("config", {}),
|
|
1104
1164
|
signals=(),
|
|
1105
1165
|
tasks=list(manifest["tasks"]),
|
|
1166
|
+
# Carried from the manifest rather than defaulted: collection replays the prompts the
|
|
1167
|
+
# evaluation stored, so this does not build them, but a config that disagreed with the
|
|
1168
|
+
# run it describes would be a trap for anyone reading it back.
|
|
1169
|
+
num_fewshot=manifest["num_fewshot"],
|
|
1106
1170
|
batch_size=1, # collection is always unbatched
|
|
1107
1171
|
adapter=saved_model.get("adapter") if saved_model else manifest.get("model_type"),
|
|
1108
1172
|
save_attention=args.save_attention,
|
|
@@ -1121,11 +1185,13 @@ def cmd_collect_research_data(args: argparse.Namespace) -> str:
|
|
|
1121
1185
|
|
|
1122
1186
|
Model and tokenizer settings come from the run's manifest, never from the command line: the two passes have to produce identical input tokens, so letting them be re-specified would be a way to get that wrong.
|
|
1123
1187
|
"""
|
|
1188
|
+
timer = PhaseTimer()
|
|
1124
1189
|
run_dir = args.run_dir
|
|
1125
1190
|
manifest = storage.read_manifest(run_dir)
|
|
1126
1191
|
config = _collection_config_from_manifest(args, manifest)
|
|
1127
1192
|
saved_model = manifest.get("custom_model", {})
|
|
1128
|
-
|
|
1193
|
+
with timer.phase("model_load"):
|
|
1194
|
+
lm = load_model(config)
|
|
1129
1195
|
if saved_model and config.model_provenance != saved_model:
|
|
1130
1196
|
raise ValueError("custom model implementation/config/checkpoint changed since evaluation")
|
|
1131
1197
|
from .models import resolve_model_adapter
|
|
@@ -1137,7 +1203,7 @@ def cmd_collect_research_data(args: argparse.Namespace) -> str:
|
|
|
1137
1203
|
else None
|
|
1138
1204
|
)
|
|
1139
1205
|
writer = storage.RunWriter(run_dir)
|
|
1140
|
-
with _tracing(config, lm.model, run_dir, name="collection"):
|
|
1206
|
+
with timer.phase("collection"), _tracing(config, lm.model, run_dir, name="collection"):
|
|
1141
1207
|
counts = _collect_research_data(
|
|
1142
1208
|
config, lm, adapter, writer, run_dir, hidden_layers,
|
|
1143
1209
|
write_steps=not manifest.get("internal_signals_available", True),
|
|
@@ -1156,6 +1222,7 @@ def cmd_collect_research_data(args: argparse.Namespace) -> str:
|
|
|
1156
1222
|
storage.write_results(run_dir, payload["manifest"], payload["results"], writer.signal_files())
|
|
1157
1223
|
collection_metadata = {
|
|
1158
1224
|
"collection_counts": counts,
|
|
1225
|
+
"collection_timings": timer.finish(),
|
|
1159
1226
|
"collection_attn_implementation": getattr(
|
|
1160
1227
|
lm.model.config, "_attn_implementation", "unknown"),
|
|
1161
1228
|
"options": options,
|
|
@@ -1328,7 +1395,10 @@ def build_parser() -> argparse.ArgumentParser:
|
|
|
1328
1395
|
run.add_argument("--tasks", required=True, help="comma separated lm-eval task names")
|
|
1329
1396
|
run.add_argument("--include-path", action="append", default=[],
|
|
1330
1397
|
help="local benchmark directory; repeat for multiple bundles")
|
|
1331
|
-
run.add_argument("--num-fewshot", type=int,
|
|
1398
|
+
run.add_argument("--num-fewshot", type=int, required=True,
|
|
1399
|
+
help="examples before each document. Required: lm-eval gives each "
|
|
1400
|
+
"task its own default and this overrides it for every task in "
|
|
1401
|
+
"the run, so the run has to state which number it used")
|
|
1332
1402
|
run.add_argument("--limit", type=int, default=None, help="cap on the number of documents")
|
|
1333
1403
|
run.add_argument("--batch-size", type=parse_batch_size, default=1,
|
|
1334
1404
|
help='positive integer, or "auto" to let lm-eval find the '
|
|
@@ -985,8 +985,19 @@ def resolve_database(path: str, pass_name: str = "trace") -> Path:
|
|
|
985
985
|
return given
|
|
986
986
|
if (given / "statistics.sqlite").is_file():
|
|
987
987
|
return given / "statistics.sqlite"
|
|
988
|
-
|
|
988
|
+
root = given / "debug" / "module_stats"
|
|
989
|
+
found = sorted((root / pass_name).glob("*/statistics.sqlite"))
|
|
989
990
|
if not found:
|
|
991
|
+
# "nothing was recorded" and "recorded, but not this pass" need different advice.
|
|
992
|
+
# A run with --module-stats but no collection pass is not misconfigured, so telling
|
|
993
|
+
# its owner to re-run with --module-stats sends them to change a flag already set.
|
|
994
|
+
recorded = sorted(directory.name for directory in root.glob("*")
|
|
995
|
+
if any(directory.glob("*/statistics.sqlite")))
|
|
996
|
+
if recorded:
|
|
997
|
+
raise FileNotFoundError(
|
|
998
|
+
f"no {pass_name} module statistics under {path}; this run recorded: "
|
|
999
|
+
f"{', '.join(recorded)}. The collection pass is written only when the run "
|
|
1000
|
+
f"also collects (--save-attention, --save-hidden or a collection hook).")
|
|
990
1001
|
raise FileNotFoundError(f"no module statistics under {path}; run with --module-stats")
|
|
991
1002
|
return found[-1]
|
|
992
1003
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: evalmetry
|
|
3
|
-
Version: 1.
|
|
3
|
+
Version: 1.1.1
|
|
4
4
|
Summary: A research toolkit for measuring and analyzing model behavior.
|
|
5
5
|
License-Expression: MIT
|
|
6
6
|
Project-URL: Homepage, https://github.com/minguinho26/evalmetry
|
|
@@ -52,9 +52,13 @@ Evalmetry evaluates Hugging Face language models with lm-eval and collects inter
|
|
|
52
52
|
|
|
53
53
|
Model evaluation supports one process on one CUDA GPU. CPU execution is for test fixtures and verification tools. Reports and saved-data readers do not require a GPU.
|
|
54
54
|
|
|
55
|
+
## Evidence and research example
|
|
56
|
+
|
|
57
|
+
Read the [1.1.0 verification report](https://github.com/minguinho26/evalmetry/blob/v1.1.1/docs/verification.md) and the [prediction-depth research showcase](https://github.com/minguinho26/evalmetry/blob/v1.1.1/docs/showcase.md). Generation parity is established at the same actual batch size. Generation is not batch-invariant.
|
|
58
|
+
|
|
55
59
|
## Install
|
|
56
60
|
|
|
57
|
-
Install from PyPI
|
|
61
|
+
Install from PyPI:
|
|
58
62
|
|
|
59
63
|
```bash
|
|
60
64
|
pip install evalmetry
|
|
@@ -62,8 +66,7 @@ pip install evalmetry
|
|
|
62
66
|
|
|
63
67
|
Evalmetry requires Python >=3.10. Use a compatible CUDA build of PyTorch. The Python import and CLI are both `evalmetry`.
|
|
64
68
|
|
|
65
|
-
To install this release explicitly, use `pip install evalmetry==1.
|
|
66
|
-
For development, clone this repository and run `python -m pip install -e .`.
|
|
69
|
+
To install this release explicitly, use `pip install evalmetry==1.1.1`. For development, clone this repository and run `python -m pip install -e .`.
|
|
67
70
|
|
|
68
71
|
## Quick start
|
|
69
72
|
|
|
@@ -81,14 +84,16 @@ Use `report` to generate a report from that run:
|
|
|
81
84
|
evalmetry report results/quickstart --output report/quickstart
|
|
82
85
|
```
|
|
83
86
|
|
|
84
|
-
The first run downloads the model and dataset if needed. Reusing a run directory resumes its recorded configuration
|
|
87
|
+
The first run downloads the model and dataset if needed. Reusing a run directory resumes its recorded configuration. A different experiment needs a different directory.
|
|
85
88
|
|
|
86
89
|
Other commands: `collect-research-data` adds optional tensors to a completed run, `debug` reads a saved module trace, and `module-stats` reads saved statistics. Traces and statistics must be enabled during collection. Use `evalmetry <command> --help` for options.
|
|
87
90
|
|
|
88
91
|
## Reading results
|
|
89
92
|
|
|
90
|
-
Evaluation results, manifests and collected signals are written under the selected output directory. Read saved signals with `evalmetry.load_signals(run_dir)
|
|
93
|
+
Evaluation results, manifests and collected signals are written under the selected output directory. Read saved signals with `evalmetry.load_signals(run_dir)`. `evalmetry.describe_schema()` describes the columns.
|
|
91
94
|
|
|
92
95
|
## License
|
|
93
96
|
|
|
94
97
|
Evalmetry is released under the MIT License.
|
|
98
|
+
|
|
99
|
+
This project is built on top of [EleutherAI's lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness).
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|