sharada 0.1.0.dev0__tar.gz → 0.1.0.dev1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/PKG-INFO +9 -9
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/README.md +8 -8
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/sharada/__init__.py +1 -1
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/.github/workflows/ci.yml +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/.github/workflows/release.yml +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/.gitignore +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/LICENSE +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/examples/finetune_your_own.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/examples/quickstart.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/examples/tickets.csv +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/pyproject.toml +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/sharada/calibrate.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/sharada/data.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/sharada/evaluate.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/sharada/layout.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/sharada/masking.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/sharada/model.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/sharada/policy.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/sharada/train.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/tests/test_calibration.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/tests/test_model.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/tests/test_policy.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/training/README.md +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/training/kaggle.ipynb +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/training/publish.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/training/run.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/training/sources.py +0 -0
- {sharada-0.1.0.dev0 → sharada-0.1.0.dev1}/uv.lock +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: sharada
|
|
3
|
-
Version: 0.1.0.
|
|
3
|
+
Version: 0.1.0.dev1
|
|
4
4
|
Summary: Typed decisions from text in one forward pass — options in the request, calibrated probabilities, easy to fine-tune.
|
|
5
5
|
Project-URL: Homepage, https://github.com/LenaBarretta/sharada
|
|
6
6
|
Project-URL: Article, https://lenatriestounderstand.com/notes/llm/024-rlcr/
|
|
@@ -23,7 +23,7 @@ Requires-Dist: datasets>=4; extra == 'train'
|
|
|
23
23
|
Description-Content-Type: text/markdown
|
|
24
24
|
|
|
25
25
|
<p align="center">
|
|
26
|
-
<img src="docs/logo.png" alt="Sharada" width="200">
|
|
26
|
+
<img src="https://raw.githubusercontent.com/LenaBarretta/sharada/main/docs/logo.png" alt="Sharada" width="200">
|
|
27
27
|
</p>
|
|
28
28
|
|
|
29
29
|
<h1 align="center">Sharada</h1>
|
|
@@ -100,8 +100,8 @@ Accuracy and calibration per label set are in each model card, measured on label
|
|
|
100
100
|
## Three things the architecture guarantees
|
|
101
101
|
|
|
102
102
|
The request is laid out as one sequence — text, question, then every option as a parallel branch — and a
|
|
103
|
-
mask decides who may read whom. Both are in [`layout.py`](sharada/layout.py) and
|
|
104
|
-
[`masking.py`](sharada/masking.py), and they buy three properties that hold by construction, not because
|
|
103
|
+
mask decides who may read whom. Both are in [`layout.py`](https://github.com/LenaBarretta/sharada/blob/main/sharada/layout.py) and
|
|
104
|
+
[`masking.py`](https://github.com/LenaBarretta/sharada/blob/main/sharada/masking.py), and they buy three properties that hold by construction, not because
|
|
105
105
|
training got them approximately right:
|
|
106
106
|
|
|
107
107
|
1. **The order of the options cannot matter.** Every option branch starts at the same position id, and
|
|
@@ -114,7 +114,7 @@ training got them approximately right:
|
|
|
114
114
|
3. **The text is read once.** Text tokens read only text tokens, so their states do not depend on the
|
|
115
115
|
question. Ten questions about one document are ten cheap read-outs over one encoding of it.
|
|
116
116
|
|
|
117
|
-
These are the tests in [`tests/test_model.py`](tests/test_model.py), checked on an untrained model.
|
|
117
|
+
These are the tests in [`tests/test_model.py`](https://github.com/LenaBarretta/sharada/blob/main/tests/test_model.py), checked on an untrained model.
|
|
118
118
|
|
|
119
119
|
## Fine-tune it on your own labels
|
|
120
120
|
|
|
@@ -146,7 +146,7 @@ task on the held-out part and writes a calibration passport. Useful arguments:
|
|
|
146
146
|
| `batch_size`, `lr`, `max_epochs`, `patience` | `16`, `2e-5`, `10`, `2` | |
|
|
147
147
|
|
|
148
148
|
Mixing several tasks in one `fit` is the normal case — give each one its own `task` name and each gets
|
|
149
|
-
its own temperature. See [`examples/finetune_your_own.py`](examples/finetune_your_own.py), which trains
|
|
149
|
+
its own temperature. See [`examples/finetune_your_own.py`](https://github.com/LenaBarretta/sharada/blob/main/examples/finetune_your_own.py), which trains
|
|
150
150
|
on a CSV.
|
|
151
151
|
|
|
152
152
|
## The number next to the answer is supposed to be true
|
|
@@ -190,8 +190,8 @@ policy = escalation_budget(policy, probabilities, budget=0.05)
|
|
|
190
190
|
|
|
191
191
|
## Train the base models yourself
|
|
192
192
|
|
|
193
|
-
[`training/`](training) has the whole run: a mix of public label sets in
|
|
194
|
-
[`sources.py`](training/sources.py) — intents, topics, sentiment, toxicity, spam, entailment, review
|
|
193
|
+
[`training/`](https://github.com/LenaBarretta/sharada/tree/main/training) has the whole run: a mix of public label sets in
|
|
194
|
+
[`sources.py`](https://github.com/LenaBarretta/sharada/blob/main/training/sources.py) — intents, topics, sentiment, toxicity, spam, entailment, review
|
|
195
195
|
scores — each one presented with several question phrasings, shuffled options and sampled option
|
|
196
196
|
subsets, so the model learns to read the options rather than their positions. Some label sets are held
|
|
197
197
|
out of training entirely and only measured, which is where the zero-shot numbers come from.
|
|
@@ -201,7 +201,7 @@ python training/run.py --encoder answerdotai/ModernBERT-base --out runs/base
|
|
|
201
201
|
```
|
|
202
202
|
|
|
203
203
|
It checkpoints every few hundred steps and resumes from the checkpoint if it finds one, which is what
|
|
204
|
-
makes it survive a Kaggle session; [`training/kaggle.ipynb`](training/kaggle.ipynb) is the notebook
|
|
204
|
+
makes it survive a Kaggle session; [`training/kaggle.ipynb`](https://github.com/LenaBarretta/sharada/blob/main/training/kaggle.ipynb) is the notebook
|
|
205
205
|
wrapper. One free T4: about an hour for base, about three for large.
|
|
206
206
|
|
|
207
207
|
## What it will not do
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
<p align="center">
|
|
2
|
-
<img src="docs/logo.png" alt="Sharada" width="200">
|
|
2
|
+
<img src="https://raw.githubusercontent.com/LenaBarretta/sharada/main/docs/logo.png" alt="Sharada" width="200">
|
|
3
3
|
</p>
|
|
4
4
|
|
|
5
5
|
<h1 align="center">Sharada</h1>
|
|
@@ -76,8 +76,8 @@ Accuracy and calibration per label set are in each model card, measured on label
|
|
|
76
76
|
## Three things the architecture guarantees
|
|
77
77
|
|
|
78
78
|
The request is laid out as one sequence — text, question, then every option as a parallel branch — and a
|
|
79
|
-
mask decides who may read whom. Both are in [`layout.py`](sharada/layout.py) and
|
|
80
|
-
[`masking.py`](sharada/masking.py), and they buy three properties that hold by construction, not because
|
|
79
|
+
mask decides who may read whom. Both are in [`layout.py`](https://github.com/LenaBarretta/sharada/blob/main/sharada/layout.py) and
|
|
80
|
+
[`masking.py`](https://github.com/LenaBarretta/sharada/blob/main/sharada/masking.py), and they buy three properties that hold by construction, not because
|
|
81
81
|
training got them approximately right:
|
|
82
82
|
|
|
83
83
|
1. **The order of the options cannot matter.** Every option branch starts at the same position id, and
|
|
@@ -90,7 +90,7 @@ training got them approximately right:
|
|
|
90
90
|
3. **The text is read once.** Text tokens read only text tokens, so their states do not depend on the
|
|
91
91
|
question. Ten questions about one document are ten cheap read-outs over one encoding of it.
|
|
92
92
|
|
|
93
|
-
These are the tests in [`tests/test_model.py`](tests/test_model.py), checked on an untrained model.
|
|
93
|
+
These are the tests in [`tests/test_model.py`](https://github.com/LenaBarretta/sharada/blob/main/tests/test_model.py), checked on an untrained model.
|
|
94
94
|
|
|
95
95
|
## Fine-tune it on your own labels
|
|
96
96
|
|
|
@@ -122,7 +122,7 @@ task on the held-out part and writes a calibration passport. Useful arguments:
|
|
|
122
122
|
| `batch_size`, `lr`, `max_epochs`, `patience` | `16`, `2e-5`, `10`, `2` | |
|
|
123
123
|
|
|
124
124
|
Mixing several tasks in one `fit` is the normal case — give each one its own `task` name and each gets
|
|
125
|
-
its own temperature. See [`examples/finetune_your_own.py`](examples/finetune_your_own.py), which trains
|
|
125
|
+
its own temperature. See [`examples/finetune_your_own.py`](https://github.com/LenaBarretta/sharada/blob/main/examples/finetune_your_own.py), which trains
|
|
126
126
|
on a CSV.
|
|
127
127
|
|
|
128
128
|
## The number next to the answer is supposed to be true
|
|
@@ -166,8 +166,8 @@ policy = escalation_budget(policy, probabilities, budget=0.05)
|
|
|
166
166
|
|
|
167
167
|
## Train the base models yourself
|
|
168
168
|
|
|
169
|
-
[`training/`](training) has the whole run: a mix of public label sets in
|
|
170
|
-
[`sources.py`](training/sources.py) — intents, topics, sentiment, toxicity, spam, entailment, review
|
|
169
|
+
[`training/`](https://github.com/LenaBarretta/sharada/tree/main/training) has the whole run: a mix of public label sets in
|
|
170
|
+
[`sources.py`](https://github.com/LenaBarretta/sharada/blob/main/training/sources.py) — intents, topics, sentiment, toxicity, spam, entailment, review
|
|
171
171
|
scores — each one presented with several question phrasings, shuffled options and sampled option
|
|
172
172
|
subsets, so the model learns to read the options rather than their positions. Some label sets are held
|
|
173
173
|
out of training entirely and only measured, which is where the zero-shot numbers come from.
|
|
@@ -177,7 +177,7 @@ python training/run.py --encoder answerdotai/ModernBERT-base --out runs/base
|
|
|
177
177
|
```
|
|
178
178
|
|
|
179
179
|
It checkpoints every few hundred steps and resumes from the checkpoint if it finds one, which is what
|
|
180
|
-
makes it survive a Kaggle session; [`training/kaggle.ipynb`](training/kaggle.ipynb) is the notebook
|
|
180
|
+
makes it survive a Kaggle session; [`training/kaggle.ipynb`](https://github.com/LenaBarretta/sharada/blob/main/training/kaggle.ipynb) is the notebook
|
|
181
181
|
wrapper. One free T4: about an hour for base, about three for large.
|
|
182
182
|
|
|
183
183
|
## What it will not do
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|