@vibgrate/haile 2026.916.2 → 2026.921.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +216 -0
- package/dist/arch-ui/map.js +28 -11
- package/dist/engine/haile.wasm +0 -0
- package/dist/index.d.ts +132 -1
- package/dist/index.js +11 -11
- package/package.json +17 -4
package/README.md
CHANGED
|
@@ -64,3 +64,219 @@ VG_HAILE_RELEASE=1 pnpm --filter @vibgrate/haile build # refuses without wasm
|
|
|
64
64
|
pnpm --filter @vibgrate/haile test # loader ↔ kernel contract (kernel cases need the wasm)
|
|
65
65
|
(cd crate && cargo test) # kernel unit tests, native
|
|
66
66
|
```
|
|
67
|
+
|
|
68
|
+
## Decision Engine (Stage 0 — not yet wired into any production code path)
|
|
69
|
+
|
|
70
|
+
> **Status: Stage 0.** This is package/ABI/provider scaffolding only. No
|
|
71
|
+
> production code path calls `createDecisionProvider()` yet — see
|
|
72
|
+
> [`docs/DECISION-ENGINE.md`](../../docs/DECISION-ENGINE.md) for the
|
|
73
|
+
> "as built" reference (architecture diagram, worked examples with rendered
|
|
74
|
+
> screenshots, real benchmark numbers) and
|
|
75
|
+
> [`docs/DECISION-ENGINE-WASM-SPEC.md`](../../docs/DECISION-ENGINE-WASM-SPEC.md)
|
|
76
|
+
> for the full original design (that spec originally proposed a standalone
|
|
77
|
+
> `packages/vibgrate-decision/` package; the maintainer decided afterward
|
|
78
|
+
> that this capability should instead live inside the existing
|
|
79
|
+
> `@vibgrate/haile` package, reusing its Rust crate / build / package
|
|
80
|
+
> infrastructure rather than standing up a parallel one — everything else in
|
|
81
|
+
> the spec, including the ABI shape, the six decision families, and every
|
|
82
|
+
> non-negotiable safety rule, applies unchanged). The seed-pack scoring
|
|
83
|
+
> weights in `crate/src/decision/pack_data.rs` are honest, hand-written
|
|
84
|
+
> placeholders — **not trained on a real corpus** — pending a later stacked
|
|
85
|
+
> PR's benchmark harness and a real, labelled `decision-pack.json`.
|
|
86
|
+
|
|
87
|
+
The Vibgrate Decision Engine is a **separate, additive, probabilistic**
|
|
88
|
+
capability living alongside Haile's deterministic architecture classification
|
|
89
|
+
above, in the same compiled WASM artifact. Given evidence Vibgrate already
|
|
90
|
+
knows (impact analysis, test results, verification history, Haile's own
|
|
91
|
+
abstention, …), it answers one of six small bounded questions and returns a
|
|
92
|
+
calibrated probability, never a fact. It never establishes facts, never
|
|
93
|
+
overrides deterministic evidence, and can never be mistaken for Haile's
|
|
94
|
+
`classify()` result — every `DecisionResult` carries its own
|
|
95
|
+
`"schema":"vg.decision.result.v1"` and `"source":"decision"`, fields Haile's
|
|
96
|
+
`HaileClassification` never has.
|
|
97
|
+
|
|
98
|
+
```ts
|
|
99
|
+
import { createDecisionProvider } from "@vibgrate/haile";
|
|
100
|
+
|
|
101
|
+
const decisions = createDecisionProvider();
|
|
102
|
+
// null when dist/engine/haile.wasm lacks the decision/ export family —
|
|
103
|
+
// never a stub provider. decide() itself never throws either: a
|
|
104
|
+
// missing/broken engine or a rejected request is always `null`.
|
|
105
|
+
decisions?.decide({
|
|
106
|
+
schema: "vg.decision.request.v1",
|
|
107
|
+
decision: "failure_related",
|
|
108
|
+
context: {},
|
|
109
|
+
evidence: [
|
|
110
|
+
{ id: "impact:AuthService", kind: "impact", feature: "stack_symbol_in_impact", value: true },
|
|
111
|
+
{ id: "test:AuthServiceTest", kind: "test", feature: "failing_test_covers_changed_symbol", value: true },
|
|
112
|
+
],
|
|
113
|
+
});
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
### The six decision families, one worked example each
|
|
117
|
+
|
|
118
|
+
**`task_complexity`** — recommend an initial Code Mode (advisory only; explicit `--mode`/`--model`/`--provider`, project config, and an active pin all outrank it, and `resolveMode()` still performs the hardware fit).
|
|
119
|
+
|
|
120
|
+
```json
|
|
121
|
+
// in: {"schema":"vg.decision.request.v1","decision":"task_complexity","context":{},
|
|
122
|
+
// "evidence":[{"id":"e1","kind":"change","feature":"migration_indicator","value":true},
|
|
123
|
+
// {"id":"e2","kind":"architecture","feature":"architecture_boundaries_crossed","value":true}]}
|
|
124
|
+
// out: {"schema":"vg.decision.result.v1","decision":"task_complexity","primitive":"choice",
|
|
125
|
+
// "selected":"forge","probabilities":[{"id":"spark","probability":0.0119},
|
|
126
|
+
// {"id":"flow","probability":0.1078},{"id":"forge","probability":0.8803}],
|
|
127
|
+
// "confidence":0.8803,"band":"high","abstain":false,
|
|
128
|
+
// "reasonCodes":["architecture_boundaries_crossed","migration_indicator"],
|
|
129
|
+
// "evidenceRefs":["e2","e1"],"calibrationId":"task-complexity-seed-1","source":"decision"}
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
**`failure_related`** — is a new test/compiler/runtime failure probably caused by the current edit? May route toward the bounded repair loop; may never suppress the failure.
|
|
133
|
+
|
|
134
|
+
```json
|
|
135
|
+
// in: {"schema":"vg.decision.request.v1","decision":"failure_related","context":{},
|
|
136
|
+
// "evidence":[{"id":"impact:AuthService","kind":"impact","feature":"stack_symbol_in_impact","value":true},
|
|
137
|
+
// {"id":"test:AuthServiceTest","kind":"test","feature":"failing_test_covers_changed_symbol","value":true}]}
|
|
138
|
+
// out: {"schema":"vg.decision.result.v1","decision":"failure_related","primitive":"binary",
|
|
139
|
+
// "selected":"related","probabilities":[{"id":"related","probability":0.9837},
|
|
140
|
+
// {"id":"unrelated","probability":0.0163}],"confidence":0.9837,"band":"high","abstain":false,
|
|
141
|
+
// "reasonCodes":["stack_symbol_in_impact","failing_test_covers_changed_symbol"],
|
|
142
|
+
// "evidenceRefs":["impact:AuthService","test:AuthServiceTest"],
|
|
143
|
+
// "calibrationId":"failure-related-seed-1","source":"decision"}
|
|
144
|
+
```
|
|
145
|
+
|
|
146
|
+
**`repair_action`** — after a failed verification attempt, which bounded VG Code mechanism runs next (respecting `maxRepairRounds`)? It never writes code itself.
|
|
147
|
+
|
|
148
|
+
```json
|
|
149
|
+
// in: {"schema":"vg.decision.request.v1","decision":"repair_action","context":{},
|
|
150
|
+
// "evidence":[{"id":"e1","kind":"verification","feature":"capsule_stale","value":true}]}
|
|
151
|
+
// out: {"schema":"vg.decision.result.v1","decision":"repair_action","primitive":"choice",
|
|
152
|
+
// "selected":"refresh_capsule","probabilities":[{"id":"retry_local_repair","probability":0.2123},
|
|
153
|
+
// {"id":"refresh_capsule","probability":0.577},{"id":"expand_context","probability":0.0707},
|
|
154
|
+
// {"id":"run_additional_verification","probability":0.0639},
|
|
155
|
+
// {"id":"escalate_reasoning","probability":0.0474},{"id":"ask_human","probability":0.0287}],
|
|
156
|
+
// "confidence":0.577,"band":"low","abstain":false,"reasonCodes":["capsule_stale"],
|
|
157
|
+
// "evidenceRefs":["e1"],"calibrationId":"repair-action-seed-1","source":"decision"}
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
**`completion_confidence`** — does the implementation appear to satisfy the task? Asymmetric by design: may say "not done yet"; may never bypass a failed test or a protected finding.
|
|
161
|
+
|
|
162
|
+
```json
|
|
163
|
+
// in: {"schema":"vg.decision.request.v1","decision":"completion_confidence","context":{},
|
|
164
|
+
// "evidence":[{"id":"e1","kind":"test","feature":"tests_passed_ratio","value":1},
|
|
165
|
+
// {"id":"e2","kind":"diagnostic","feature":"protected_findings_present","value":true}]}
|
|
166
|
+
// out: {"schema":"vg.decision.result.v1","decision":"completion_confidence","primitive":"binary",
|
|
167
|
+
// "selected":"incomplete","probabilities":[{"id":"complete","probability":0.0037},
|
|
168
|
+
// {"id":"incomplete","probability":0.9963}],"confidence":0.9963,"band":"high","abstain":false,
|
|
169
|
+
// "reasonCodes":["tests_passed_ratio","protected_findings_present"],
|
|
170
|
+
// "evidenceRefs":["e1","e2"],"calibrationId":"completion-confidence-seed-1","source":"decision"}
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
**`finding_attention`** — queue-prioritisation only. Never modifies severity, CVSS, EPSS, KEV status, reachability, RiskScore, DriftScore, `protected_finding`, or the review decision.
|
|
174
|
+
|
|
175
|
+
```json
|
|
176
|
+
// in: {"schema":"vg.decision.request.v1","decision":"finding_attention","context":{},
|
|
177
|
+
// "evidence":[{"id":"e1","kind":"diagnostic","feature":"severity_score","value":1},
|
|
178
|
+
// {"id":"e2","kind":"diagnostic","feature":"exploit_maturity_score","value":1}]}
|
|
179
|
+
// out: {"schema":"vg.decision.result.v1","decision":"finding_attention","primitive":"choice",
|
|
180
|
+
// "selected":"immediate","probabilities":[{"id":"immediate","probability":0.7516},
|
|
181
|
+
// {"id":"review","probability":0.1518},{"id":"normal","probability":0.092},
|
|
182
|
+
// {"id":"defer","probability":0.0046}],"confidence":0.7516,"band":"medium","abstain":false,
|
|
183
|
+
// "reasonCodes":["severity_score","exploit_maturity_score"],"evidenceRefs":["e1","e2"],
|
|
184
|
+
// "calibrationId":"finding-attention-seed-1","source":"decision"}
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
**`architecture_fallback`** — only when Haile's own classifier has explicitly abstained. Uses its own small role vocabulary (never Haile's `taxonomy::ROLES`), reads Haile's abstention only as one input feature, and — like every other decision above — always carries `"source":"decision"` so it can never be mistaken for Haile's deterministic classification.
|
|
188
|
+
|
|
189
|
+
```json
|
|
190
|
+
// in: {"schema":"vg.decision.request.v1","decision":"architecture_fallback","context":{},
|
|
191
|
+
// "evidence":[{"id":"e1","kind":"architecture","feature":"role_hint_authenticate","value":1},
|
|
192
|
+
// {"id":"e2","kind":"architecture","feature":"haile_abstained","value":true}]}
|
|
193
|
+
// out: {"schema":"vg.decision.result.v1","decision":"architecture_fallback","primitive":"choice",
|
|
194
|
+
// "selected":"authenticate","probabilities":[{"id":"authenticate","probability":0.8007},
|
|
195
|
+
// {"id":"persist","probability":0.0399},{"id":"respond","probability":0.0399},
|
|
196
|
+
// {"id":"validate","probability":0.0399},{"id":"orchestrate","probability":0.0399},
|
|
197
|
+
// {"id":"render","probability":0.0399}],"confidence":0.8007,"band":"high","abstain":false,
|
|
198
|
+
// "reasonCodes":["haile_abstained_fallback_engaged","role_hint_authenticate"],
|
|
199
|
+
// "evidenceRefs":["e2","e1"],"calibrationId":"architecture-fallback-seed-1","source":"decision"}
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
### Determinism, calibration, and abstention
|
|
203
|
+
|
|
204
|
+
Same input JSON always yields byte-identical output for a given engine/pack
|
|
205
|
+
version (no randomness, clock, or unordered iteration). Confidence bands
|
|
206
|
+
(`high`/`medium`/`low`/`abstain`) are per-decision, versioned thresholds
|
|
207
|
+
(`crate/src/decision/calibration.rs`), not one global cutoff. An honest
|
|
208
|
+
`abstain` (e.g. no evidence supplied) is a normal, fully-shaped result, not
|
|
209
|
+
an error — the host falls back to existing behaviour exactly as it does when
|
|
210
|
+
the engine itself is absent.
|
|
211
|
+
|
|
212
|
+
### Runnable examples (`examples/decision/` — one worked example each)
|
|
213
|
+
|
|
214
|
+
```bash
|
|
215
|
+
pnpm --filter @vibgrate/haile run examples:decision # all six, in sequence, real WASM output
|
|
216
|
+
pnpm --filter @vibgrate/haile exec tsx examples/decision/failure-related.example.ts # one at a time
|
|
217
|
+
```
|
|
218
|
+
|
|
219
|
+
Six small standalone scripts, one per decision family, each building a
|
|
220
|
+
realistic `DecisionRequest` with an inline comment on every evidence field,
|
|
221
|
+
calling the real `createDecisionProvider()`, and pretty-printing the
|
|
222
|
+
`DecisionResult`. `docs/generate-screenshots.mjs`
|
|
223
|
+
(`pnpm --filter @vibgrate/haile run docs:screenshots`) runs these same six
|
|
224
|
+
scripts to regenerate [`docs/DECISION-ENGINE.md`](../../docs/DECISION-ENGINE.md)'s
|
|
225
|
+
worked-example JSON and screenshots — see that doc for the full architecture
|
|
226
|
+
reference plus rendered examples.
|
|
227
|
+
|
|
228
|
+
### Tests
|
|
229
|
+
|
|
230
|
+
```bash
|
|
231
|
+
pnpm --filter @vibgrate/haile test # includes tests/decision-*.test.ts and tests/bench-harness.test.ts
|
|
232
|
+
(cd crate && cargo test) # includes crate/src/decision/**/*'s own unit tests
|
|
233
|
+
```
|
|
234
|
+
|
|
235
|
+
### Benchmarking (`bench/` — spec §17/§18/§32)
|
|
236
|
+
|
|
237
|
+
```bash
|
|
238
|
+
pnpm --filter @vibgrate/haile run bench:decision # all six decisions
|
|
239
|
+
pnpm --filter @vibgrate/haile run bench:decision:failure-related # one decision
|
|
240
|
+
pnpm --filter @vibgrate/haile run bench:decision:generate-corpus # regenerate bench/corpus/*.jsonl
|
|
241
|
+
```
|
|
242
|
+
|
|
243
|
+
`bench/evaluate.ts` loads each labelled corpus under `bench/corpus/*.jsonl`,
|
|
244
|
+
splits it by `repo_id` (spec §32 — never by random row), runs the real
|
|
245
|
+
`decision-wasm` backend (`bench/backends/decision-wasm.ts`, wrapping
|
|
246
|
+
`createDecisionProvider()` above) over the held-out test split, and prints
|
|
247
|
+
accuracy, precision, recall, F1, Brier score, negative log likelihood,
|
|
248
|
+
expected calibration error, coverage, accuracy-by-confidence-band, and
|
|
249
|
+
abstention rate — plus a naive-heuristic baseline for comparison
|
|
250
|
+
(`bench/backends/full-model.ts`; Jev is stubbed unavailable per spec §24,
|
|
251
|
+
`bench/backends/jev.ts`).
|
|
252
|
+
|
|
253
|
+
> **Every corpus row is SYNTHETIC** (`bench/corpus/README.md`) and the
|
|
254
|
+
> seed-pack weights (`crate/src/decision/pack_data.rs`) are untrained
|
|
255
|
+
> placeholders — so the numbers below are genuinely mediocre. That is the
|
|
256
|
+
> correct, honest Stage 0 result, not a bug in the harness: this corpus is
|
|
257
|
+
> for validating the benchmark machinery itself, not for judging production
|
|
258
|
+
> readiness. It is **not** the real 500–1,000-example corpus spec §31/§32
|
|
259
|
+
> requires before any production hook.
|
|
260
|
+
|
|
261
|
+
Real output from an actual run (`pnpm --filter @vibgrate/haile run bench:decision`), unedited:
|
|
262
|
+
|
|
263
|
+
```
|
|
264
|
+
=== Summary — vibgrate-decision-wasm across all six decisions ===
|
|
265
|
+
decision | n(test) | accuracy | F1(macro) | Brier | NLL | ECE | coverage | abstain%
|
|
266
|
+
task_complexity | 15 | 46.7% | 24.6% | 0.5782 | 0.8517 | 0.3216 | 73.3% | 26.7%
|
|
267
|
+
failure_related | 20 | 35.0% | 25.9% | 1.2023 | 2.6754 | 0.6201 | 100.0% | 0.0%
|
|
268
|
+
repair_action | 18 | 100.0% | 100.0% | 0.2879 | 0.6553 | 0.4757 | 66.7% | 33.3%
|
|
269
|
+
completion_confidence | 20 | 55.0% | 52.0% | 0.7553 | 1.3406 | 0.3805 | 85.0% | 15.0%
|
|
270
|
+
finding_attention | 16 | 37.5% | 34.4% | 0.5932 | 1.0679 | 0.1527 | 31.3% | 68.8%
|
|
271
|
+
architecture_fallback | 21 | 100.0% | 100.0% | 0.2486 | 0.5423 | 0.3416 | 76.2% | 23.8%
|
|
272
|
+
```
|
|
273
|
+
|
|
274
|
+
Per-decision output also prints a naive-heuristic-baseline row for
|
|
275
|
+
comparison and accuracy broken out by confidence band; run the command
|
|
276
|
+
above for the full table. `repair_action` and `architecture_fallback`
|
|
277
|
+
scoring 100% here is a property of their small synthetic test splits (18
|
|
278
|
+
and 21 rows respectively, after the repo-level split), not a general claim —
|
|
279
|
+
`task_complexity` and `failure_related`, evaluated against the same kind of
|
|
280
|
+
corpus, land far lower. Treat every number as "the harness works and these
|
|
281
|
+
are the real, unflattering seed-pack results," not as a production
|
|
282
|
+
readiness signal.
|