evalmetry 1.1.0__tar.gz → 1.1.1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {evalmetry-1.1.0 → evalmetry-1.1.1}/PKG-INFO +6 -7
- {evalmetry-1.1.0 → evalmetry-1.1.1}/README.md +5 -6
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/main.py +1 -1
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry.egg-info/PKG-INFO +6 -7
- {evalmetry-1.1.0 → evalmetry-1.1.1}/pyproject.toml +1 -1
- {evalmetry-1.1.0 → evalmetry-1.1.1}/LICENSE +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/__init__.py +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/adapters.py +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/backend.py +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/benchmarks.py +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/debug.py +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/hooks.py +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/judges.py +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/models.py +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/module_stats.py +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/recorder.py +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/reducers.py +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/report.py +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry/storage.py +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry.egg-info/SOURCES.txt +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry.egg-info/dependency_links.txt +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry.egg-info/entry_points.txt +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry.egg-info/requires.txt +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/evalmetry.egg-info/top_level.txt +0 -0
- {evalmetry-1.1.0 → evalmetry-1.1.1}/setup.cfg +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: evalmetry
|
|
3
|
-
Version: 1.1.
|
|
3
|
+
Version: 1.1.1
|
|
4
4
|
Summary: A research toolkit for measuring and analyzing model behavior.
|
|
5
5
|
License-Expression: MIT
|
|
6
6
|
Project-URL: Homepage, https://github.com/minguinho26/evalmetry
|
|
@@ -54,11 +54,11 @@ Model evaluation supports one process on one CUDA GPU. CPU execution is for test
|
|
|
54
54
|
|
|
55
55
|
## Evidence and research example
|
|
56
56
|
|
|
57
|
-
Read the [1.1.0 verification report](https://github.com/minguinho26/evalmetry/blob/
|
|
57
|
+
Read the [1.1.0 verification report](https://github.com/minguinho26/evalmetry/blob/v1.1.1/docs/verification.md) and the [prediction-depth research showcase](https://github.com/minguinho26/evalmetry/blob/v1.1.1/docs/showcase.md). Generation parity is established at the same actual batch size. Generation is not batch-invariant.
|
|
58
58
|
|
|
59
59
|
## Install
|
|
60
60
|
|
|
61
|
-
Install from PyPI
|
|
61
|
+
Install from PyPI:
|
|
62
62
|
|
|
63
63
|
```bash
|
|
64
64
|
pip install evalmetry
|
|
@@ -66,8 +66,7 @@ pip install evalmetry
|
|
|
66
66
|
|
|
67
67
|
Evalmetry requires Python >=3.10. Use a compatible CUDA build of PyTorch. The Python import and CLI are both `evalmetry`.
|
|
68
68
|
|
|
69
|
-
To install this release explicitly, use `pip install evalmetry==1.1.
|
|
70
|
-
For development, clone this repository and run `python -m pip install -e .`.
|
|
69
|
+
To install this release explicitly, use `pip install evalmetry==1.1.1`. For development, clone this repository and run `python -m pip install -e .`.
|
|
71
70
|
|
|
72
71
|
## Quick start
|
|
73
72
|
|
|
@@ -85,13 +84,13 @@ Use `report` to generate a report from that run:
|
|
|
85
84
|
evalmetry report results/quickstart --output report/quickstart
|
|
86
85
|
```
|
|
87
86
|
|
|
88
|
-
The first run downloads the model and dataset if needed. Reusing a run directory resumes its recorded configuration
|
|
87
|
+
The first run downloads the model and dataset if needed. Reusing a run directory resumes its recorded configuration. A different experiment needs a different directory.
|
|
89
88
|
|
|
90
89
|
Other commands: `collect-research-data` adds optional tensors to a completed run, `debug` reads a saved module trace, and `module-stats` reads saved statistics. Traces and statistics must be enabled during collection. Use `evalmetry <command> --help` for options.
|
|
91
90
|
|
|
92
91
|
## Reading results
|
|
93
92
|
|
|
94
|
-
Evaluation results, manifests and collected signals are written under the selected output directory. Read saved signals with `evalmetry.load_signals(run_dir)
|
|
93
|
+
Evaluation results, manifests and collected signals are written under the selected output directory. Read saved signals with `evalmetry.load_signals(run_dir)`. `evalmetry.describe_schema()` describes the columns.
|
|
95
94
|
|
|
96
95
|
## License
|
|
97
96
|
|
|
@@ -15,11 +15,11 @@ Model evaluation supports one process on one CUDA GPU. CPU execution is for test
|
|
|
15
15
|
|
|
16
16
|
## Evidence and research example
|
|
17
17
|
|
|
18
|
-
Read the [1.1.0 verification report](https://github.com/minguinho26/evalmetry/blob/
|
|
18
|
+
Read the [1.1.0 verification report](https://github.com/minguinho26/evalmetry/blob/v1.1.1/docs/verification.md) and the [prediction-depth research showcase](https://github.com/minguinho26/evalmetry/blob/v1.1.1/docs/showcase.md). Generation parity is established at the same actual batch size. Generation is not batch-invariant.
|
|
19
19
|
|
|
20
20
|
## Install
|
|
21
21
|
|
|
22
|
-
Install from PyPI
|
|
22
|
+
Install from PyPI:
|
|
23
23
|
|
|
24
24
|
```bash
|
|
25
25
|
pip install evalmetry
|
|
@@ -27,8 +27,7 @@ pip install evalmetry
|
|
|
27
27
|
|
|
28
28
|
Evalmetry requires Python >=3.10. Use a compatible CUDA build of PyTorch. The Python import and CLI are both `evalmetry`.
|
|
29
29
|
|
|
30
|
-
To install this release explicitly, use `pip install evalmetry==1.1.
|
|
31
|
-
For development, clone this repository and run `python -m pip install -e .`.
|
|
30
|
+
To install this release explicitly, use `pip install evalmetry==1.1.1`. For development, clone this repository and run `python -m pip install -e .`.
|
|
32
31
|
|
|
33
32
|
## Quick start
|
|
34
33
|
|
|
@@ -46,13 +45,13 @@ Use `report` to generate a report from that run:
|
|
|
46
45
|
evalmetry report results/quickstart --output report/quickstart
|
|
47
46
|
```
|
|
48
47
|
|
|
49
|
-
The first run downloads the model and dataset if needed. Reusing a run directory resumes its recorded configuration
|
|
48
|
+
The first run downloads the model and dataset if needed. Reusing a run directory resumes its recorded configuration. A different experiment needs a different directory.
|
|
50
49
|
|
|
51
50
|
Other commands: `collect-research-data` adds optional tensors to a completed run, `debug` reads a saved module trace, and `module-stats` reads saved statistics. Traces and statistics must be enabled during collection. Use `evalmetry <command> --help` for options.
|
|
52
51
|
|
|
53
52
|
## Reading results
|
|
54
53
|
|
|
55
|
-
Evaluation results, manifests and collected signals are written under the selected output directory. Read saved signals with `evalmetry.load_signals(run_dir)
|
|
54
|
+
Evaluation results, manifests and collected signals are written under the selected output directory. Read saved signals with `evalmetry.load_signals(run_dir)`. `evalmetry.describe_schema()` describes the columns.
|
|
56
55
|
|
|
57
56
|
## License
|
|
58
57
|
|
|
@@ -36,7 +36,7 @@ from .debug import TRACE_BUFFER_EVENTS
|
|
|
36
36
|
from .storage import FIXED_SETTINGS, SAMPLING_SEED, SCHEMA_VERSION
|
|
37
37
|
|
|
38
38
|
#: Version of this tool, recorded in every manifest.
|
|
39
|
-
TOOL_VERSION = "1.1.
|
|
39
|
+
TOOL_VERSION = "1.1.1"
|
|
40
40
|
|
|
41
41
|
#: Default number of documents collected per correctness group.
|
|
42
42
|
DEFAULT_COLLECT_LIMIT = 500
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: evalmetry
|
|
3
|
-
Version: 1.1.
|
|
3
|
+
Version: 1.1.1
|
|
4
4
|
Summary: A research toolkit for measuring and analyzing model behavior.
|
|
5
5
|
License-Expression: MIT
|
|
6
6
|
Project-URL: Homepage, https://github.com/minguinho26/evalmetry
|
|
@@ -54,11 +54,11 @@ Model evaluation supports one process on one CUDA GPU. CPU execution is for test
|
|
|
54
54
|
|
|
55
55
|
## Evidence and research example
|
|
56
56
|
|
|
57
|
-
Read the [1.1.0 verification report](https://github.com/minguinho26/evalmetry/blob/
|
|
57
|
+
Read the [1.1.0 verification report](https://github.com/minguinho26/evalmetry/blob/v1.1.1/docs/verification.md) and the [prediction-depth research showcase](https://github.com/minguinho26/evalmetry/blob/v1.1.1/docs/showcase.md). Generation parity is established at the same actual batch size. Generation is not batch-invariant.
|
|
58
58
|
|
|
59
59
|
## Install
|
|
60
60
|
|
|
61
|
-
Install from PyPI
|
|
61
|
+
Install from PyPI:
|
|
62
62
|
|
|
63
63
|
```bash
|
|
64
64
|
pip install evalmetry
|
|
@@ -66,8 +66,7 @@ pip install evalmetry
|
|
|
66
66
|
|
|
67
67
|
Evalmetry requires Python >=3.10. Use a compatible CUDA build of PyTorch. The Python import and CLI are both `evalmetry`.
|
|
68
68
|
|
|
69
|
-
To install this release explicitly, use `pip install evalmetry==1.1.
|
|
70
|
-
For development, clone this repository and run `python -m pip install -e .`.
|
|
69
|
+
To install this release explicitly, use `pip install evalmetry==1.1.1`. For development, clone this repository and run `python -m pip install -e .`.
|
|
71
70
|
|
|
72
71
|
## Quick start
|
|
73
72
|
|
|
@@ -85,13 +84,13 @@ Use `report` to generate a report from that run:
|
|
|
85
84
|
evalmetry report results/quickstart --output report/quickstart
|
|
86
85
|
```
|
|
87
86
|
|
|
88
|
-
The first run downloads the model and dataset if needed. Reusing a run directory resumes its recorded configuration
|
|
87
|
+
The first run downloads the model and dataset if needed. Reusing a run directory resumes its recorded configuration. A different experiment needs a different directory.
|
|
89
88
|
|
|
90
89
|
Other commands: `collect-research-data` adds optional tensors to a completed run, `debug` reads a saved module trace, and `module-stats` reads saved statistics. Traces and statistics must be enabled during collection. Use `evalmetry <command> --help` for options.
|
|
91
90
|
|
|
92
91
|
## Reading results
|
|
93
92
|
|
|
94
|
-
Evaluation results, manifests and collected signals are written under the selected output directory. Read saved signals with `evalmetry.load_signals(run_dir)
|
|
93
|
+
Evaluation results, manifests and collected signals are written under the selected output directory. Read saved signals with `evalmetry.load_signals(run_dir)`. `evalmetry.describe_schema()` describes the columns.
|
|
95
94
|
|
|
96
95
|
## License
|
|
97
96
|
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|