vegaml 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- vegaml-0.1.0/LICENSE +191 -0
- vegaml-0.1.0/PKG-INFO +251 -0
- vegaml-0.1.0/README.md +216 -0
- vegaml-0.1.0/pyproject.toml +62 -0
- vegaml-0.1.0/setup.cfg +4 -0
- vegaml-0.1.0/tests/test_api_surface.py +123 -0
- vegaml-0.1.0/tests/test_no_remote_code.py +80 -0
- vegaml-0.1.0/tests/test_packaging.py +49 -0
- vegaml-0.1.0/tests/test_ttt.py +67 -0
- vegaml-0.1.0/vegaml/__init__.py +192 -0
- vegaml-0.1.0/vegaml/_api.py +222 -0
- vegaml-0.1.0/vegaml/_common.py +1027 -0
- vegaml-0.1.0/vegaml/_vision.py +172 -0
- vegaml-0.1.0/vegaml/hub.py +33 -0
- vegaml-0.1.0/vegaml/py.typed +0 -0
- vegaml-0.1.0/vegaml/ttt.py +85 -0
- vegaml-0.1.0/vegaml.egg-info/PKG-INFO +251 -0
- vegaml-0.1.0/vegaml.egg-info/SOURCES.txt +19 -0
- vegaml-0.1.0/vegaml.egg-info/dependency_links.txt +1 -0
- vegaml-0.1.0/vegaml.egg-info/requires.txt +15 -0
- vegaml-0.1.0/vegaml.egg-info/top_level.txt +1 -0
vegaml-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,191 @@
|
|
|
1
|
+
|
|
2
|
+
Apache License
|
|
3
|
+
Version 2.0, January 2004
|
|
4
|
+
http://www.apache.org/licenses/
|
|
5
|
+
|
|
6
|
+
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
|
|
7
|
+
|
|
8
|
+
1. Definitions.
|
|
9
|
+
|
|
10
|
+
"License" shall mean the terms and conditions for use, reproduction,
|
|
11
|
+
and distribution as defined by Sections 1 through 9 of this document.
|
|
12
|
+
|
|
13
|
+
"Licensor" shall mean the copyright owner or entity authorized by
|
|
14
|
+
the copyright owner that is granting the License.
|
|
15
|
+
|
|
16
|
+
"Legal Entity" shall mean the union of the acting entity and all
|
|
17
|
+
other entities that control, are controlled by, or are under common
|
|
18
|
+
control with that entity. For the purposes of this definition,
|
|
19
|
+
"control" means (i) the power, direct or indirect, to cause the
|
|
20
|
+
direction or management of such entity, whether by contract or
|
|
21
|
+
otherwise, or (ii) ownership of fifty percent (50%) or more of the
|
|
22
|
+
outstanding shares, or (iii) beneficial ownership of such entity.
|
|
23
|
+
|
|
24
|
+
"You" (or "Your") shall mean an individual or Legal Entity
|
|
25
|
+
exercising permissions granted by this License.
|
|
26
|
+
|
|
27
|
+
"Source" form shall mean the preferred form for making modifications,
|
|
28
|
+
including but not limited to software source code, documentation
|
|
29
|
+
source, and configuration files.
|
|
30
|
+
|
|
31
|
+
"Object" form shall mean any form resulting from mechanical
|
|
32
|
+
transformation or translation of a Source form, including but
|
|
33
|
+
not limited to compiled object code, generated documentation,
|
|
34
|
+
and conversions to other media types.
|
|
35
|
+
|
|
36
|
+
"Work" shall mean the work of authorship, whether in Source or
|
|
37
|
+
Object form, made available under the License, as indicated by a
|
|
38
|
+
copyright notice that is included in or attached to the work
|
|
39
|
+
(an example is provided in the Appendix below).
|
|
40
|
+
|
|
41
|
+
"Derivative Works" shall mean any work, whether in Source or Object
|
|
42
|
+
form, that is based on (or derived from) the Work and for which the
|
|
43
|
+
editorial revisions, annotations, elaborations, or other modifications
|
|
44
|
+
represent, as a whole, an original work of authorship. For the purposes
|
|
45
|
+
of this License, Derivative Works shall not include works that remain
|
|
46
|
+
separable from, or merely link (or bind by name) to the interfaces of,
|
|
47
|
+
the Work and Derivative Works thereof.
|
|
48
|
+
|
|
49
|
+
"Contribution" shall mean any work of authorship, including
|
|
50
|
+
the original version of the Work and any modifications or additions
|
|
51
|
+
to that Work or Derivative Works thereof, that is intentionally
|
|
52
|
+
submitted to Licensor for inclusion in the Work by the copyright owner
|
|
53
|
+
or by an individual or Legal Entity authorized to submit on behalf of
|
|
54
|
+
the copyright owner. For the purposes of this definition, "submitted"
|
|
55
|
+
means any form of electronic, verbal, or written communication sent
|
|
56
|
+
to the Licensor or its representatives, including but not limited to
|
|
57
|
+
communication on electronic mailing lists, source code control systems,
|
|
58
|
+
and issue tracking systems that are managed by, or on behalf of, the
|
|
59
|
+
Licensor for the purpose of discussing and improving the Work, but
|
|
60
|
+
excluding communication that is conspicuously marked or otherwise
|
|
61
|
+
designated in writing by the copyright owner as "Not a Contribution."
|
|
62
|
+
|
|
63
|
+
"Contributor" shall mean Licensor and any individual or Legal Entity
|
|
64
|
+
on behalf of whom a Contribution has been received by Licensor and
|
|
65
|
+
subsequently incorporated within the Work.
|
|
66
|
+
|
|
67
|
+
2. Grant of Copyright License. Subject to the terms and conditions of
|
|
68
|
+
this License, each Contributor hereby grants to You a perpetual,
|
|
69
|
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
|
70
|
+
copyright license to reproduce, prepare Derivative Works of,
|
|
71
|
+
publicly display, publicly perform, sublicense, and distribute the
|
|
72
|
+
Work and such Derivative Works in Source or Object form.
|
|
73
|
+
|
|
74
|
+
3. Grant of Patent License. Subject to the terms and conditions of
|
|
75
|
+
this License, each Contributor hereby grants to You a perpetual,
|
|
76
|
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
|
77
|
+
(except as stated in this section) patent license to make, have made,
|
|
78
|
+
use, offer to sell, sell, import, and otherwise transfer the Work,
|
|
79
|
+
where such license applies only to those patent claims licensable
|
|
80
|
+
by such Contributor that are necessarily infringed by their
|
|
81
|
+
Contribution(s) alone or by combination of their Contribution(s)
|
|
82
|
+
with the Work to which such Contribution(s) was submitted. If You
|
|
83
|
+
institute patent litigation against any entity (including a
|
|
84
|
+
cross-claim or counterclaim in a lawsuit) alleging that the Work
|
|
85
|
+
or a Contribution incorporated within the Work constitutes direct
|
|
86
|
+
or contributory patent infringement, then any patent licenses
|
|
87
|
+
granted to You under this License for that Work shall terminate
|
|
88
|
+
as of the date such litigation is filed.
|
|
89
|
+
|
|
90
|
+
4. Redistribution. You may reproduce and distribute copies of the
|
|
91
|
+
Work or Derivative Works thereof in any medium, with or without
|
|
92
|
+
modifications, and in Source or Object form, provided that You
|
|
93
|
+
meet the following conditions:
|
|
94
|
+
|
|
95
|
+
(a) You must give any other recipients of the Work or
|
|
96
|
+
Derivative Works a copy of this License; and
|
|
97
|
+
|
|
98
|
+
(b) You must cause any modified files to carry prominent notices
|
|
99
|
+
stating that You changed the files; and
|
|
100
|
+
|
|
101
|
+
(c) You must retain, in the Source form of any Derivative Works
|
|
102
|
+
that You distribute, all copyright, patent, trademark, and
|
|
103
|
+
attribution notices from the Source form of the Work,
|
|
104
|
+
excluding those notices that do not pertain to any part of
|
|
105
|
+
the Derivative Works; and
|
|
106
|
+
|
|
107
|
+
(d) If the Work includes a "NOTICE" text file as part of its
|
|
108
|
+
distribution, then any Derivative Works that You distribute must
|
|
109
|
+
include a readable copy of the attribution notices contained
|
|
110
|
+
within such NOTICE file, excluding those notices that do not
|
|
111
|
+
pertain to any part of the Derivative Works, in at least one
|
|
112
|
+
of the following places: within a NOTICE text file distributed
|
|
113
|
+
as part of the Derivative Works; within the Source form or
|
|
114
|
+
documentation, if provided along with the Derivative Works; or,
|
|
115
|
+
within a display generated by the Derivative Works, if and
|
|
116
|
+
wherever such third-party notices normally appear. The contents
|
|
117
|
+
of the NOTICE file are for informational purposes only and
|
|
118
|
+
do not modify the License. You may add Your own attribution
|
|
119
|
+
notices within Derivative Works that You distribute, alongside
|
|
120
|
+
or as an addendum to the NOTICE text from the Work, provided
|
|
121
|
+
that such additional attribution notices cannot be construed
|
|
122
|
+
as modifying the License.
|
|
123
|
+
|
|
124
|
+
You may add Your own copyright statement to Your modifications and
|
|
125
|
+
may provide additional or different license terms and conditions
|
|
126
|
+
for use, reproduction, or distribution of Your modifications, or
|
|
127
|
+
for any such Derivative Works as a whole, provided Your use,
|
|
128
|
+
reproduction, and distribution of the Work otherwise complies with
|
|
129
|
+
the conditions stated in this License.
|
|
130
|
+
|
|
131
|
+
5. Submission of Contributions. Unless You explicitly state otherwise,
|
|
132
|
+
any Contribution intentionally submitted for inclusion in the Work
|
|
133
|
+
by You to the Licensor shall be under the terms and conditions of
|
|
134
|
+
this License, without any additional terms or conditions.
|
|
135
|
+
Notwithstanding the above, nothing herein shall supersede or modify
|
|
136
|
+
the terms of any separate license agreement you may have executed
|
|
137
|
+
with Licensor regarding such Contributions.
|
|
138
|
+
|
|
139
|
+
6. Trademarks. This License does not grant permission to use the trade
|
|
140
|
+
names, trademarks, service marks, or product names of the Licensor,
|
|
141
|
+
except as required for reasonable and customary use in describing the
|
|
142
|
+
origin of the Work and reproducing the content of the NOTICE file.
|
|
143
|
+
|
|
144
|
+
7. Disclaimer of Warranty. Unless required by applicable law or
|
|
145
|
+
agreed to in writing, Licensor provides the Work (and each
|
|
146
|
+
Contributor provides its Contributions) on an "AS IS" BASIS,
|
|
147
|
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
|
|
148
|
+
implied, including, without limitation, any warranties or conditions
|
|
149
|
+
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
|
|
150
|
+
PARTICULAR PURPOSE. You are solely responsible for determining the
|
|
151
|
+
appropriateness of using or redistributing the Work and assume any
|
|
152
|
+
risks associated with Your exercise of permissions under this License.
|
|
153
|
+
|
|
154
|
+
8. Limitation of Liability. In no event and under no legal theory,
|
|
155
|
+
whether in tort (including negligence), contract, or otherwise,
|
|
156
|
+
unless required by applicable law (such as deliberate and grossly
|
|
157
|
+
negligent acts) or agreed to in writing, shall any Contributor be
|
|
158
|
+
liable to You for damages, including any direct, indirect, special,
|
|
159
|
+
incidental, or consequential damages of any character arising as a
|
|
160
|
+
result of this License or out of the use or inability to use the
|
|
161
|
+
Work (including but not limited to damages for loss of goodwill,
|
|
162
|
+
work stoppage, computer failure or malfunction, or any and all
|
|
163
|
+
other commercial damages or losses), even if such Contributor
|
|
164
|
+
has been advised of the possibility of such damages.
|
|
165
|
+
|
|
166
|
+
9. Accepting Warranty or Additional Liability. While redistributing
|
|
167
|
+
the Work or Derivative Works thereof, You may choose to offer,
|
|
168
|
+
and charge a fee for, acceptance of support, warranty, indemnity,
|
|
169
|
+
or other liability obligations and/or rights consistent with this
|
|
170
|
+
License. However, in accepting such obligations, You may act only
|
|
171
|
+
on Your own behalf and on Your sole responsibility, not on behalf
|
|
172
|
+
of any other Contributor, and only if You agree to indemnify,
|
|
173
|
+
defend, and hold each Contributor harmless for any liability
|
|
174
|
+
incurred by, or claims asserted against, such Contributor by reason
|
|
175
|
+
of your accepting any such warranty or additional liability.
|
|
176
|
+
|
|
177
|
+
END OF TERMS AND CONDITIONS
|
|
178
|
+
|
|
179
|
+
Copyright 2026 Nandakishor M
|
|
180
|
+
|
|
181
|
+
Licensed under the Apache License, Version 2.0 (the "License");
|
|
182
|
+
you may not use this file except in compliance with the License.
|
|
183
|
+
You may obtain a copy of the License at
|
|
184
|
+
|
|
185
|
+
http://www.apache.org/licenses/LICENSE-2.0
|
|
186
|
+
|
|
187
|
+
Unless required by applicable law or agreed to in writing, software
|
|
188
|
+
distributed under the License is distributed on an "AS IS" BASIS,
|
|
189
|
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
190
|
+
See the License for the specific language governing permissions and
|
|
191
|
+
limitations under the License.
|
vegaml-0.1.0/PKG-INFO
ADDED
|
@@ -0,0 +1,251 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: vegaml
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Typed decisions from a frozen language model and a small physics engine: calibrated probabilities, conformal sets, 73k context, images.
|
|
5
|
+
Author: Nandakishor M
|
|
6
|
+
License-Expression: Apache-2.0
|
|
7
|
+
Project-URL: Homepage, https://github.com/NandhaKishorM/vegaml
|
|
8
|
+
Project-URL: Issues, https://github.com/NandhaKishorM/vegaml/issues
|
|
9
|
+
Keywords: typed-decisions,calibration,conformal-prediction,decision-model,long-context
|
|
10
|
+
Classifier: Development Status :: 4 - Beta
|
|
11
|
+
Classifier: Intended Audience :: Developers
|
|
12
|
+
Classifier: Intended Audience :: Science/Research
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
18
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
19
|
+
Classifier: Typing :: Typed
|
|
20
|
+
Requires-Python: >=3.10
|
|
21
|
+
Description-Content-Type: text/markdown
|
|
22
|
+
License-File: LICENSE
|
|
23
|
+
Requires-Dist: torch>=2.4
|
|
24
|
+
Requires-Dist: transformers>=5.17
|
|
25
|
+
Requires-Dist: safetensors>=0.4
|
|
26
|
+
Requires-Dist: huggingface_hub<2.0,>=0.30
|
|
27
|
+
Requires-Dist: numpy>=1.24
|
|
28
|
+
Provides-Extra: vision
|
|
29
|
+
Requires-Dist: pillow>=10.0; extra == "vision"
|
|
30
|
+
Provides-Extra: dev
|
|
31
|
+
Requires-Dist: pytest>=8.0; extra == "dev"
|
|
32
|
+
Requires-Dist: ruff>=0.6; extra == "dev"
|
|
33
|
+
Requires-Dist: tomli>=2.0; python_version < "3.11" and extra == "dev"
|
|
34
|
+
Dynamic: license-file
|
|
35
|
+
|
|
36
|
+
# vegaml
|
|
37
|
+
|
|
38
|
+
Typed decisions from a frozen language model and a small physics engine. You give it a **state** and a
|
|
39
|
+
**question whose answer type is fixed in advance**; it returns a value your code can use directly — a
|
|
40
|
+
choice, a 0–1 score, or the probability that a statement is true — each with a calibrated probability,
|
|
41
|
+
a conformal answer set and an explicit abstain flag.
|
|
42
|
+
|
|
43
|
+
**800M parameters · 73,728-token context · images · runs on your own hardware.** No text is generated
|
|
44
|
+
anywhere in the path.
|
|
45
|
+
|
|
46
|
+

|
|
47
|
+
|
|
48
|
+
```bash
|
|
49
|
+
pip install vegaml
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
<sub>Releases are cut by tagging `vX.Y.Z`: CI builds, checks the version matches the tag, and
|
|
53
|
+
publishes through PyPI Trusted Publishing — no API token exists in this repository.</sub>
|
|
54
|
+
|
|
55
|
+
## Quick start
|
|
56
|
+
|
|
57
|
+
```python
|
|
58
|
+
import vegaml
|
|
59
|
+
|
|
60
|
+
v = vegaml.load("nandakishorm/vega-08b-public-intents") # mode="engine" by default
|
|
61
|
+
|
|
62
|
+
out = v.decide(
|
|
63
|
+
{"from": "billing@acme.com", "subject": "Invoice overdue", "body": "Third notice. Pay now."},
|
|
64
|
+
{"team": {"type": "choice", "instructions": "Which team should handle this?",
|
|
65
|
+
"criteria": {"billing": "payments and invoices",
|
|
66
|
+
"technical": "product faults",
|
|
67
|
+
"sales": "new business"}},
|
|
68
|
+
"churn": {"type": "noul", "instructions": "Is this customer at risk of leaving?",
|
|
69
|
+
"criteria": {"true": "shows intent to cancel", "false": "no such signal"}}})
|
|
70
|
+
|
|
71
|
+
print(out["answers"]["team"]["choice"]) # 'billing'
|
|
72
|
+
print(out["answers"]["team"]["probs"]) # calibrated, sums to 1
|
|
73
|
+
print(out["answers"]["churn"]["abstain"]) # False -> the engine stands behind it
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
Both questions are answered from **one** read of the state.
|
|
77
|
+
|
|
78
|
+
### Two readouts, one set of features
|
|
79
|
+
|
|
80
|
+
`mode` picks how the pooled features are turned into an answer.
|
|
81
|
+
|
|
82
|
+
| mode | needs examples | carries | use it when |
|
|
83
|
+
|---|---|---|---|
|
|
84
|
+
| `"engine"` *(default)* | no, zero-shot | calibration, conformal sets, abstain | the normal case, and every published figure below |
|
|
85
|
+
| `"ttt"` | yes, 6 minimum | nothing — it answers even when it should not | you have labels for exactly this question |
|
|
86
|
+
| `"both"` | yes | both, side by side | deciding which to trust |
|
|
87
|
+
|
|
88
|
+
```python
|
|
89
|
+
examples = [({"body": "cancel my account"}, {"churn": "true"}),
|
|
90
|
+
({"body": "how do I export?"}, {"churn": "false"})] # ... 20 or so
|
|
91
|
+
|
|
92
|
+
report = v.fit(examples, questions)
|
|
93
|
+
print(report["churn"]["cv_accuracy"], report["churn"]["at_or_below_chance"])
|
|
94
|
+
|
|
95
|
+
out = v.decide(state, questions, mode="ttt")
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
`fit` returns a cross-validated accuracy per question. **Read it.** A head at or below chance has
|
|
99
|
+
learned nothing and its probabilities are noise; the engine is the honest answer there.
|
|
100
|
+
|
|
101
|
+
### Images
|
|
102
|
+
|
|
103
|
+
```python
|
|
104
|
+
from PIL import Image
|
|
105
|
+
|
|
106
|
+
v.decide_image(Image.open("invoice.png"),
|
|
107
|
+
{"kind": {"type": "choice", "instructions": "What kind of document is this?",
|
|
108
|
+
"criteria": {"invoice": "a bill", "receipt": "proof of payment",
|
|
109
|
+
"form": "a document to be filled in", "other": "none of these"}}})
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
The backbone is multimodal and its vision encoder is frozen with the rest, so a picture takes the place
|
|
113
|
+
of the state text in the same prompt and the same spans are pooled. The benchmark figures below are
|
|
114
|
+
text; the image path is functional but not covered by them.
|
|
115
|
+
|
|
116
|
+
### 73,728-token context
|
|
117
|
+
|
|
118
|
+
`max_len` is 73,728 tokens, and the architecture is built for it rather than merely permitting it:
|
|
119
|
+
|
|
120
|
+
* **Read-once prefix caching.** A state of at least 4,096 tokens is encoded **once**; every question
|
|
121
|
+
about that state continues from a copy of that prefix. Twelve questions about one 73k-token contract
|
|
122
|
+
cost roughly one read, not twelve.
|
|
123
|
+
* **A memory reader.** Long states are cut into 16 pieces, mean-pooled at the deepest feature layer and
|
|
124
|
+
fed to the engine alongside the spans, so evidence late in a long document still reaches the decision.
|
|
125
|
+
* **No silent truncation.** An input that will not fit is refused, never quietly shortened.
|
|
126
|
+
|
|
127
|
+
## Benchmarks
|
|
128
|
+
|
|
129
|
+
The figure at the top of this page is the table below. Item-paired against a live Jev 1.13.0 over
|
|
130
|
+
identical items, no tuning on the evaluation data. These are the areas where this model leads; the
|
|
131
|
+
reference leads on most others, particularly reranking and multi-step reasoning.
|
|
132
|
+
|
|
133
|
+
| task | this model | Jev 1.13.0 |
|
|
134
|
+
|---|---|---|
|
|
135
|
+
| Phishing screening, 800 emails | **75.4** acc · **252/400** caught · recall **0.63** | 61.9 · 99/400 · 0.25 |
|
|
136
|
+
| Dates and quantities (temporal_numeric) | **46.7** | 20.0 |
|
|
137
|
+
| Spam detection (enron-spam) | **1.000** | 0.920 |
|
|
138
|
+
| News topic (ag-news) | **0.955** | 0.806 |
|
|
139
|
+
| Typed scores (typed-decisions) | **0.438** | 0.395 |
|
|
140
|
+
| Calibration error, product relevance | **6.0** | 22.0 |
|
|
141
|
+
| Calibration error, overall (5,096 paired) | 9.5 | 9.3 |
|
|
142
|
+
| Median latency per decision | **267 ms** (T4 fp16) · 280 ms (Apple M, fp32) | 591 ms (hosted) |
|
|
143
|
+
|
|
144
|
+
2.5× the phishing caught at McNemar p = 3e-11, and 2.1× faster with nothing leaving the machine.
|
|
145
|
+
|
|
146
|
+
## How it works
|
|
147
|
+
|
|
148
|
+
One prompt, one forward pass, then physics.
|
|
149
|
+
|
|
150
|
+
**1 — One formatted prompt.** The state, the question and every candidate answer are written into a
|
|
151
|
+
single sequence:
|
|
152
|
+
|
|
153
|
+
```
|
|
154
|
+
Observation: {state} Measurement ({type}): {instructions} Possible outcomes: * {option 1} * {option 2} Outcome:
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
The option text is genuinely part of the model's input, not metadata kept outside it. No system prompt,
|
|
158
|
+
no chat template, no demonstrations, and the gold answer never appears — it exists only as a training
|
|
159
|
+
label.
|
|
160
|
+
|
|
161
|
+
**2 — One forward pass, five pooled vectors.** The frozen language model runs **once** over that
|
|
162
|
+
sequence. Vega then pools different token spans from the *same* hidden states: the situation span, the
|
|
163
|
+
question span, one span per answer option, and the final token. Three questions about one state mean
|
|
164
|
+
three prompts and three passes, except for the long-input case above, where the prefix is shared.
|
|
165
|
+
|
|
166
|
+
**3 — Pooled vectors become initial conditions.** The situation vector projects to a world latent and
|
|
167
|
+
gets a bounded nonlinear nudge; the question and the final token project to a probe and an impulse.
|
|
168
|
+
Together they place a particle at position `z₀` with momentum `p₀` in a 64-dimensional decision space.
|
|
169
|
+
|
|
170
|
+
**4 — Every candidate answer becomes a valley.** Each option's own features produce a Gaussian well —
|
|
171
|
+
a centre `c_k`, a depth `a_k` and a width `σ_k`:
|
|
172
|
+
|
|
173
|
+
```
|
|
174
|
+
U(z) = ½κ‖z‖² − Σ_k a_k · exp( −‖z − c_k‖² / 2σ_k² )
|
|
175
|
+
```
|
|
176
|
+
|
|
177
|
+
The quadratic term keeps the particle bounded; each well pulls it toward one answer. Score levels sit on
|
|
178
|
+
a one-dimensional rail, so ordinal neighbours are physical neighbours.
|
|
179
|
+
|
|
180
|
+
**5 — The particle rolls and settles.** Damped Hamiltonian dynamics — symplectic Euler with friction and
|
|
181
|
+
a learned state-space thermostat — for a fixed, small step budget, with early exit once it has settled.
|
|
182
|
+
A decision is a short simulation with constant cost, not a sampling loop, so it is deterministic.
|
|
183
|
+
|
|
184
|
+
**6 — Where it settles is the answer.**
|
|
185
|
+
|
|
186
|
+
```
|
|
187
|
+
E_k = ‖z_T − c_k‖² / 2σ_k² − log a_k P = softmax(−E / τ)
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
### What is new here
|
|
191
|
+
|
|
192
|
+
**The architecture.** The language model is frozen and never trained — it is perception only, read at
|
|
193
|
+
two intermediate layers and cut off above the deepest one, so a decision never pays for the layers above
|
|
194
|
+
it and the LM head never runs. Everything learned lives in a 55 MB engine whose state is a *physical*
|
|
195
|
+
one: a position and a momentum, not a logit vector.
|
|
196
|
+
|
|
197
|
+
**The training method.** The engine is trained on the settling behaviour, not on next-token likelihood:
|
|
198
|
+
a counterfactual objective pairs items that share an answer space, and a one-step world-dynamics block
|
|
199
|
+
is trained to imagine the next state from the current one, so the latent carries what happens next
|
|
200
|
+
rather than only what was said. Two low-rank adapters (rank 32) attach to seven engine projections, and
|
|
201
|
+
a per-question sigmoid gate decides **for each question independently** whether an adapter contributes.
|
|
202
|
+
The gate value and the chosen adapter are returned with every answer, so routing is auditable rather
|
|
203
|
+
than implicit.
|
|
204
|
+
|
|
205
|
+
**The calibration method.** The readout temperature is not a constant. It is predicted per decision from
|
|
206
|
+
the physical state the particle ended in:
|
|
207
|
+
|
|
208
|
+
```
|
|
209
|
+
log τ = b + w · [ log(1 + residual kinetic energy),
|
|
210
|
+
log(1 + distance to the nearest well bottom),
|
|
211
|
+
fraction of the step budget used,
|
|
212
|
+
log (number of options) ]
|
|
213
|
+
```
|
|
214
|
+
|
|
215
|
+
A particle still in motion, or stopped far from every well, is an uncertain decision and gets a hotter
|
|
216
|
+
temperature. Each adapter route carries its own calibration vector. On top of that, every answer has a
|
|
217
|
+
split-conformal set at a chosen risk level, an **abstain** flag below a fitted confidence floor, and an
|
|
218
|
+
**unbound** flag when the particle settled far from every well — the model's own way of saying the
|
|
219
|
+
question is outside what it knows.
|
|
220
|
+
|
|
221
|
+
## Security
|
|
222
|
+
|
|
223
|
+
A checkpoint is data from a third party, so the inference code lives in this package and is reviewed
|
|
224
|
+
with it. `vegaml.load()` downloads **only** `vega_config.json`, `engine.safetensors` and the adapter
|
|
225
|
+
files; nothing fetched at runtime is imported or executed, and weights load through safetensors, never
|
|
226
|
+
pickle. Tests enforce that allow-list, and CI fails on any `pickle`, `torch.load`, `eval`, `exec`,
|
|
227
|
+
`subprocess` or `sys.path` insertion reaching the shipped package.
|
|
228
|
+
|
|
229
|
+
## Repository settings worth knowing
|
|
230
|
+
|
|
231
|
+
Two things this repository cannot enforce on a private repo without a paid plan, documented here so
|
|
232
|
+
nobody mistakes a gap for a gate:
|
|
233
|
+
|
|
234
|
+
* **Branch protection.** Both classic protection and rulesets return `403 Upgrade to GitHub Pro or
|
|
235
|
+
make this repository public`, so the server enforces nothing: a force-push to `main`, a deletion,
|
|
236
|
+
or a merge over a red check are all possible. The workflows still run on every push and pull
|
|
237
|
+
request; only the enforcement is missing.
|
|
238
|
+
|
|
239
|
+
The nearest available substitute is a local hook, `.githooks/pre-push`, which refuses a force-push
|
|
240
|
+
or a deletion of `main`. Enable it in each clone with `git config core.hooksPath .githooks`. It
|
|
241
|
+
stops the accident from a configured clone and **nothing else** — not another machine, not the web
|
|
242
|
+
UI, not `--no-verify`. Treat it as a seatbelt, not a lock.
|
|
243
|
+
* **CodeQL.** Code scanning needs GitHub Advanced Security on a private repository
|
|
244
|
+
(`422 Advanced security has not been purchased`), so the CodeQL job reports why it skipped instead
|
|
245
|
+
of failing forever. Secret scanning, the dependency CVE audit and the supply-chain gate do run.
|
|
246
|
+
|
|
247
|
+
Making the repository public, or upgrading, turns both on with no change to the workflows.
|
|
248
|
+
|
|
249
|
+
## Licence
|
|
250
|
+
|
|
251
|
+
Apache 2.0.
|
vegaml-0.1.0/README.md
ADDED
|
@@ -0,0 +1,216 @@
|
|
|
1
|
+
# vegaml
|
|
2
|
+
|
|
3
|
+
Typed decisions from a frozen language model and a small physics engine. You give it a **state** and a
|
|
4
|
+
**question whose answer type is fixed in advance**; it returns a value your code can use directly — a
|
|
5
|
+
choice, a 0–1 score, or the probability that a statement is true — each with a calibrated probability,
|
|
6
|
+
a conformal answer set and an explicit abstain flag.
|
|
7
|
+
|
|
8
|
+
**800M parameters · 73,728-token context · images · runs on your own hardware.** No text is generated
|
|
9
|
+
anywhere in the path.
|
|
10
|
+
|
|
11
|
+

|
|
12
|
+
|
|
13
|
+
```bash
|
|
14
|
+
pip install vegaml
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
<sub>Releases are cut by tagging `vX.Y.Z`: CI builds, checks the version matches the tag, and
|
|
18
|
+
publishes through PyPI Trusted Publishing — no API token exists in this repository.</sub>
|
|
19
|
+
|
|
20
|
+
## Quick start
|
|
21
|
+
|
|
22
|
+
```python
|
|
23
|
+
import vegaml
|
|
24
|
+
|
|
25
|
+
v = vegaml.load("nandakishorm/vega-08b-public-intents") # mode="engine" by default
|
|
26
|
+
|
|
27
|
+
out = v.decide(
|
|
28
|
+
{"from": "billing@acme.com", "subject": "Invoice overdue", "body": "Third notice. Pay now."},
|
|
29
|
+
{"team": {"type": "choice", "instructions": "Which team should handle this?",
|
|
30
|
+
"criteria": {"billing": "payments and invoices",
|
|
31
|
+
"technical": "product faults",
|
|
32
|
+
"sales": "new business"}},
|
|
33
|
+
"churn": {"type": "noul", "instructions": "Is this customer at risk of leaving?",
|
|
34
|
+
"criteria": {"true": "shows intent to cancel", "false": "no such signal"}}})
|
|
35
|
+
|
|
36
|
+
print(out["answers"]["team"]["choice"]) # 'billing'
|
|
37
|
+
print(out["answers"]["team"]["probs"]) # calibrated, sums to 1
|
|
38
|
+
print(out["answers"]["churn"]["abstain"]) # False -> the engine stands behind it
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
Both questions are answered from **one** read of the state.
|
|
42
|
+
|
|
43
|
+
### Two readouts, one set of features
|
|
44
|
+
|
|
45
|
+
`mode` picks how the pooled features are turned into an answer.
|
|
46
|
+
|
|
47
|
+
| mode | needs examples | carries | use it when |
|
|
48
|
+
|---|---|---|---|
|
|
49
|
+
| `"engine"` *(default)* | no, zero-shot | calibration, conformal sets, abstain | the normal case, and every published figure below |
|
|
50
|
+
| `"ttt"` | yes, 6 minimum | nothing — it answers even when it should not | you have labels for exactly this question |
|
|
51
|
+
| `"both"` | yes | both, side by side | deciding which to trust |
|
|
52
|
+
|
|
53
|
+
```python
|
|
54
|
+
examples = [({"body": "cancel my account"}, {"churn": "true"}),
|
|
55
|
+
({"body": "how do I export?"}, {"churn": "false"})] # ... 20 or so
|
|
56
|
+
|
|
57
|
+
report = v.fit(examples, questions)
|
|
58
|
+
print(report["churn"]["cv_accuracy"], report["churn"]["at_or_below_chance"])
|
|
59
|
+
|
|
60
|
+
out = v.decide(state, questions, mode="ttt")
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
`fit` returns a cross-validated accuracy per question. **Read it.** A head at or below chance has
|
|
64
|
+
learned nothing and its probabilities are noise; the engine is the honest answer there.
|
|
65
|
+
|
|
66
|
+
### Images
|
|
67
|
+
|
|
68
|
+
```python
|
|
69
|
+
from PIL import Image
|
|
70
|
+
|
|
71
|
+
v.decide_image(Image.open("invoice.png"),
|
|
72
|
+
{"kind": {"type": "choice", "instructions": "What kind of document is this?",
|
|
73
|
+
"criteria": {"invoice": "a bill", "receipt": "proof of payment",
|
|
74
|
+
"form": "a document to be filled in", "other": "none of these"}}})
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
The backbone is multimodal and its vision encoder is frozen with the rest, so a picture takes the place
|
|
78
|
+
of the state text in the same prompt and the same spans are pooled. The benchmark figures below are
|
|
79
|
+
text; the image path is functional but not covered by them.
|
|
80
|
+
|
|
81
|
+
### 73,728-token context
|
|
82
|
+
|
|
83
|
+
`max_len` is 73,728 tokens, and the architecture is built for it rather than merely permitting it:
|
|
84
|
+
|
|
85
|
+
* **Read-once prefix caching.** A state of at least 4,096 tokens is encoded **once**; every question
|
|
86
|
+
about that state continues from a copy of that prefix. Twelve questions about one 73k-token contract
|
|
87
|
+
cost roughly one read, not twelve.
|
|
88
|
+
* **A memory reader.** Long states are cut into 16 pieces, mean-pooled at the deepest feature layer and
|
|
89
|
+
fed to the engine alongside the spans, so evidence late in a long document still reaches the decision.
|
|
90
|
+
* **No silent truncation.** An input that will not fit is refused, never quietly shortened.
|
|
91
|
+
|
|
92
|
+
## Benchmarks
|
|
93
|
+
|
|
94
|
+
The figure at the top of this page is the table below. Item-paired against a live Jev 1.13.0 over
|
|
95
|
+
identical items, no tuning on the evaluation data. These are the areas where this model leads; the
|
|
96
|
+
reference leads on most others, particularly reranking and multi-step reasoning.
|
|
97
|
+
|
|
98
|
+
| task | this model | Jev 1.13.0 |
|
|
99
|
+
|---|---|---|
|
|
100
|
+
| Phishing screening, 800 emails | **75.4** acc · **252/400** caught · recall **0.63** | 61.9 · 99/400 · 0.25 |
|
|
101
|
+
| Dates and quantities (temporal_numeric) | **46.7** | 20.0 |
|
|
102
|
+
| Spam detection (enron-spam) | **1.000** | 0.920 |
|
|
103
|
+
| News topic (ag-news) | **0.955** | 0.806 |
|
|
104
|
+
| Typed scores (typed-decisions) | **0.438** | 0.395 |
|
|
105
|
+
| Calibration error, product relevance | **6.0** | 22.0 |
|
|
106
|
+
| Calibration error, overall (5,096 paired) | 9.5 | 9.3 |
|
|
107
|
+
| Median latency per decision | **267 ms** (T4 fp16) · 280 ms (Apple M, fp32) | 591 ms (hosted) |
|
|
108
|
+
|
|
109
|
+
2.5× the phishing caught at McNemar p = 3e-11, and 2.1× faster with nothing leaving the machine.
|
|
110
|
+
|
|
111
|
+
## How it works
|
|
112
|
+
|
|
113
|
+
One prompt, one forward pass, then physics.
|
|
114
|
+
|
|
115
|
+
**1 — One formatted prompt.** The state, the question and every candidate answer are written into a
|
|
116
|
+
single sequence:
|
|
117
|
+
|
|
118
|
+
```
|
|
119
|
+
Observation: {state} Measurement ({type}): {instructions} Possible outcomes: * {option 1} * {option 2} Outcome:
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
The option text is genuinely part of the model's input, not metadata kept outside it. No system prompt,
|
|
123
|
+
no chat template, no demonstrations, and the gold answer never appears — it exists only as a training
|
|
124
|
+
label.
|
|
125
|
+
|
|
126
|
+
**2 — One forward pass, five pooled vectors.** The frozen language model runs **once** over that
|
|
127
|
+
sequence. Vega then pools different token spans from the *same* hidden states: the situation span, the
|
|
128
|
+
question span, one span per answer option, and the final token. Three questions about one state mean
|
|
129
|
+
three prompts and three passes, except for the long-input case above, where the prefix is shared.
|
|
130
|
+
|
|
131
|
+
**3 — Pooled vectors become initial conditions.** The situation vector projects to a world latent and
|
|
132
|
+
gets a bounded nonlinear nudge; the question and the final token project to a probe and an impulse.
|
|
133
|
+
Together they place a particle at position `z₀` with momentum `p₀` in a 64-dimensional decision space.
|
|
134
|
+
|
|
135
|
+
**4 — Every candidate answer becomes a valley.** Each option's own features produce a Gaussian well —
|
|
136
|
+
a centre `c_k`, a depth `a_k` and a width `σ_k`:
|
|
137
|
+
|
|
138
|
+
```
|
|
139
|
+
U(z) = ½κ‖z‖² − Σ_k a_k · exp( −‖z − c_k‖² / 2σ_k² )
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
The quadratic term keeps the particle bounded; each well pulls it toward one answer. Score levels sit on
|
|
143
|
+
a one-dimensional rail, so ordinal neighbours are physical neighbours.
|
|
144
|
+
|
|
145
|
+
**5 — The particle rolls and settles.** Damped Hamiltonian dynamics — symplectic Euler with friction and
|
|
146
|
+
a learned state-space thermostat — for a fixed, small step budget, with early exit once it has settled.
|
|
147
|
+
A decision is a short simulation with constant cost, not a sampling loop, so it is deterministic.
|
|
148
|
+
|
|
149
|
+
**6 — Where it settles is the answer.**
|
|
150
|
+
|
|
151
|
+
```
|
|
152
|
+
E_k = ‖z_T − c_k‖² / 2σ_k² − log a_k P = softmax(−E / τ)
|
|
153
|
+
```
|
|
154
|
+
|
|
155
|
+
### What is new here
|
|
156
|
+
|
|
157
|
+
**The architecture.** The language model is frozen and never trained — it is perception only, read at
|
|
158
|
+
two intermediate layers and cut off above the deepest one, so a decision never pays for the layers above
|
|
159
|
+
it and the LM head never runs. Everything learned lives in a 55 MB engine whose state is a *physical*
|
|
160
|
+
one: a position and a momentum, not a logit vector.
|
|
161
|
+
|
|
162
|
+
**The training method.** The engine is trained on the settling behaviour, not on next-token likelihood:
|
|
163
|
+
a counterfactual objective pairs items that share an answer space, and a one-step world-dynamics block
|
|
164
|
+
is trained to imagine the next state from the current one, so the latent carries what happens next
|
|
165
|
+
rather than only what was said. Two low-rank adapters (rank 32) attach to seven engine projections, and
|
|
166
|
+
a per-question sigmoid gate decides **for each question independently** whether an adapter contributes.
|
|
167
|
+
The gate value and the chosen adapter are returned with every answer, so routing is auditable rather
|
|
168
|
+
than implicit.
|
|
169
|
+
|
|
170
|
+
**The calibration method.** The readout temperature is not a constant. It is predicted per decision from
|
|
171
|
+
the physical state the particle ended in:
|
|
172
|
+
|
|
173
|
+
```
|
|
174
|
+
log τ = b + w · [ log(1 + residual kinetic energy),
|
|
175
|
+
log(1 + distance to the nearest well bottom),
|
|
176
|
+
fraction of the step budget used,
|
|
177
|
+
log (number of options) ]
|
|
178
|
+
```
|
|
179
|
+
|
|
180
|
+
A particle still in motion, or stopped far from every well, is an uncertain decision and gets a hotter
|
|
181
|
+
temperature. Each adapter route carries its own calibration vector. On top of that, every answer has a
|
|
182
|
+
split-conformal set at a chosen risk level, an **abstain** flag below a fitted confidence floor, and an
|
|
183
|
+
**unbound** flag when the particle settled far from every well — the model's own way of saying the
|
|
184
|
+
question is outside what it knows.
|
|
185
|
+
|
|
186
|
+
## Security
|
|
187
|
+
|
|
188
|
+
A checkpoint is data from a third party, so the inference code lives in this package and is reviewed
|
|
189
|
+
with it. `vegaml.load()` downloads **only** `vega_config.json`, `engine.safetensors` and the adapter
|
|
190
|
+
files; nothing fetched at runtime is imported or executed, and weights load through safetensors, never
|
|
191
|
+
pickle. Tests enforce that allow-list, and CI fails on any `pickle`, `torch.load`, `eval`, `exec`,
|
|
192
|
+
`subprocess` or `sys.path` insertion reaching the shipped package.
|
|
193
|
+
|
|
194
|
+
## Repository settings worth knowing
|
|
195
|
+
|
|
196
|
+
Two things this repository cannot enforce on a private repo without a paid plan, documented here so
|
|
197
|
+
nobody mistakes a gap for a gate:
|
|
198
|
+
|
|
199
|
+
* **Branch protection.** Both classic protection and rulesets return `403 Upgrade to GitHub Pro or
|
|
200
|
+
make this repository public`, so the server enforces nothing: a force-push to `main`, a deletion,
|
|
201
|
+
or a merge over a red check are all possible. The workflows still run on every push and pull
|
|
202
|
+
request; only the enforcement is missing.
|
|
203
|
+
|
|
204
|
+
The nearest available substitute is a local hook, `.githooks/pre-push`, which refuses a force-push
|
|
205
|
+
or a deletion of `main`. Enable it in each clone with `git config core.hooksPath .githooks`. It
|
|
206
|
+
stops the accident from a configured clone and **nothing else** — not another machine, not the web
|
|
207
|
+
UI, not `--no-verify`. Treat it as a seatbelt, not a lock.
|
|
208
|
+
* **CodeQL.** Code scanning needs GitHub Advanced Security on a private repository
|
|
209
|
+
(`422 Advanced security has not been purchased`), so the CodeQL job reports why it skipped instead
|
|
210
|
+
of failing forever. Secret scanning, the dependency CVE audit and the supply-chain gate do run.
|
|
211
|
+
|
|
212
|
+
Making the repository public, or upgrading, turns both on with no change to the workflows.
|
|
213
|
+
|
|
214
|
+
## Licence
|
|
215
|
+
|
|
216
|
+
Apache 2.0.
|