@vibgrate/haile 2026.916.2 → 2026.921.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -64,3 +64,219 @@ VG_HAILE_RELEASE=1 pnpm --filter @vibgrate/haile build # refuses without wasm
64
64
  pnpm --filter @vibgrate/haile test # loader ↔ kernel contract (kernel cases need the wasm)
65
65
  (cd crate && cargo test) # kernel unit tests, native
66
66
  ```
67
+
68
+ ## Decision Engine (Stage 0 — not yet wired into any production code path)
69
+
70
+ > **Status: Stage 0.** This is package/ABI/provider scaffolding only. No
71
+ > production code path calls `createDecisionProvider()` yet — see
72
+ > [`docs/DECISION-ENGINE.md`](../../docs/DECISION-ENGINE.md) for the
73
+ > "as built" reference (architecture diagram, worked examples with rendered
74
+ > screenshots, real benchmark numbers) and
75
+ > [`docs/DECISION-ENGINE-WASM-SPEC.md`](../../docs/DECISION-ENGINE-WASM-SPEC.md)
76
+ > for the full original design (that spec originally proposed a standalone
77
+ > `packages/vibgrate-decision/` package; the maintainer decided afterward
78
+ > that this capability should instead live inside the existing
79
+ > `@vibgrate/haile` package, reusing its Rust crate / build / package
80
+ > infrastructure rather than standing up a parallel one — everything else in
81
+ > the spec, including the ABI shape, the six decision families, and every
82
+ > non-negotiable safety rule, applies unchanged). The seed-pack scoring
83
+ > weights in `crate/src/decision/pack_data.rs` are honest, hand-written
84
+ > placeholders — **not trained on a real corpus** — pending a later stacked
85
+ > PR's benchmark harness and a real, labelled `decision-pack.json`.
86
+
87
+ The Vibgrate Decision Engine is a **separate, additive, probabilistic**
88
+ capability living alongside Haile's deterministic architecture classification
89
+ above, in the same compiled WASM artifact. Given evidence Vibgrate already
90
+ knows (impact analysis, test results, verification history, Haile's own
91
+ abstention, …), it answers one of six small bounded questions and returns a
92
+ calibrated probability, never a fact. It never establishes facts, never
93
+ overrides deterministic evidence, and can never be mistaken for Haile's
94
+ `classify()` result — every `DecisionResult` carries its own
95
+ `"schema":"vg.decision.result.v1"` and `"source":"decision"`, fields Haile's
96
+ `HaileClassification` never has.
97
+
98
+ ```ts
99
+ import { createDecisionProvider } from "@vibgrate/haile";
100
+
101
+ const decisions = createDecisionProvider();
102
+ // null when dist/engine/haile.wasm lacks the decision/ export family —
103
+ // never a stub provider. decide() itself never throws either: a
104
+ // missing/broken engine or a rejected request is always `null`.
105
+ decisions?.decide({
106
+ schema: "vg.decision.request.v1",
107
+ decision: "failure_related",
108
+ context: {},
109
+ evidence: [
110
+ { id: "impact:AuthService", kind: "impact", feature: "stack_symbol_in_impact", value: true },
111
+ { id: "test:AuthServiceTest", kind: "test", feature: "failing_test_covers_changed_symbol", value: true },
112
+ ],
113
+ });
114
+ ```
115
+
116
+ ### The six decision families, one worked example each
117
+
118
+ **`task_complexity`** — recommend an initial Code Mode (advisory only; explicit `--mode`/`--model`/`--provider`, project config, and an active pin all outrank it, and `resolveMode()` still performs the hardware fit).
119
+
120
+ ```json
121
+ // in: {"schema":"vg.decision.request.v1","decision":"task_complexity","context":{},
122
+ // "evidence":[{"id":"e1","kind":"change","feature":"migration_indicator","value":true},
123
+ // {"id":"e2","kind":"architecture","feature":"architecture_boundaries_crossed","value":true}]}
124
+ // out: {"schema":"vg.decision.result.v1","decision":"task_complexity","primitive":"choice",
125
+ // "selected":"forge","probabilities":[{"id":"spark","probability":0.0119},
126
+ // {"id":"flow","probability":0.1078},{"id":"forge","probability":0.8803}],
127
+ // "confidence":0.8803,"band":"high","abstain":false,
128
+ // "reasonCodes":["architecture_boundaries_crossed","migration_indicator"],
129
+ // "evidenceRefs":["e2","e1"],"calibrationId":"task-complexity-seed-1","source":"decision"}
130
+ ```
131
+
132
+ **`failure_related`** — is a new test/compiler/runtime failure probably caused by the current edit? May route toward the bounded repair loop; may never suppress the failure.
133
+
134
+ ```json
135
+ // in: {"schema":"vg.decision.request.v1","decision":"failure_related","context":{},
136
+ // "evidence":[{"id":"impact:AuthService","kind":"impact","feature":"stack_symbol_in_impact","value":true},
137
+ // {"id":"test:AuthServiceTest","kind":"test","feature":"failing_test_covers_changed_symbol","value":true}]}
138
+ // out: {"schema":"vg.decision.result.v1","decision":"failure_related","primitive":"binary",
139
+ // "selected":"related","probabilities":[{"id":"related","probability":0.9837},
140
+ // {"id":"unrelated","probability":0.0163}],"confidence":0.9837,"band":"high","abstain":false,
141
+ // "reasonCodes":["stack_symbol_in_impact","failing_test_covers_changed_symbol"],
142
+ // "evidenceRefs":["impact:AuthService","test:AuthServiceTest"],
143
+ // "calibrationId":"failure-related-seed-1","source":"decision"}
144
+ ```
145
+
146
+ **`repair_action`** — after a failed verification attempt, which bounded VG Code mechanism runs next (respecting `maxRepairRounds`)? It never writes code itself.
147
+
148
+ ```json
149
+ // in: {"schema":"vg.decision.request.v1","decision":"repair_action","context":{},
150
+ // "evidence":[{"id":"e1","kind":"verification","feature":"capsule_stale","value":true}]}
151
+ // out: {"schema":"vg.decision.result.v1","decision":"repair_action","primitive":"choice",
152
+ // "selected":"refresh_capsule","probabilities":[{"id":"retry_local_repair","probability":0.2123},
153
+ // {"id":"refresh_capsule","probability":0.577},{"id":"expand_context","probability":0.0707},
154
+ // {"id":"run_additional_verification","probability":0.0639},
155
+ // {"id":"escalate_reasoning","probability":0.0474},{"id":"ask_human","probability":0.0287}],
156
+ // "confidence":0.577,"band":"low","abstain":false,"reasonCodes":["capsule_stale"],
157
+ // "evidenceRefs":["e1"],"calibrationId":"repair-action-seed-1","source":"decision"}
158
+ ```
159
+
160
+ **`completion_confidence`** — does the implementation appear to satisfy the task? Asymmetric by design: may say "not done yet"; may never bypass a failed test or a protected finding.
161
+
162
+ ```json
163
+ // in: {"schema":"vg.decision.request.v1","decision":"completion_confidence","context":{},
164
+ // "evidence":[{"id":"e1","kind":"test","feature":"tests_passed_ratio","value":1},
165
+ // {"id":"e2","kind":"diagnostic","feature":"protected_findings_present","value":true}]}
166
+ // out: {"schema":"vg.decision.result.v1","decision":"completion_confidence","primitive":"binary",
167
+ // "selected":"incomplete","probabilities":[{"id":"complete","probability":0.0037},
168
+ // {"id":"incomplete","probability":0.9963}],"confidence":0.9963,"band":"high","abstain":false,
169
+ // "reasonCodes":["tests_passed_ratio","protected_findings_present"],
170
+ // "evidenceRefs":["e1","e2"],"calibrationId":"completion-confidence-seed-1","source":"decision"}
171
+ ```
172
+
173
+ **`finding_attention`** — queue-prioritisation only. Never modifies severity, CVSS, EPSS, KEV status, reachability, RiskScore, DriftScore, `protected_finding`, or the review decision.
174
+
175
+ ```json
176
+ // in: {"schema":"vg.decision.request.v1","decision":"finding_attention","context":{},
177
+ // "evidence":[{"id":"e1","kind":"diagnostic","feature":"severity_score","value":1},
178
+ // {"id":"e2","kind":"diagnostic","feature":"exploit_maturity_score","value":1}]}
179
+ // out: {"schema":"vg.decision.result.v1","decision":"finding_attention","primitive":"choice",
180
+ // "selected":"immediate","probabilities":[{"id":"immediate","probability":0.7516},
181
+ // {"id":"review","probability":0.1518},{"id":"normal","probability":0.092},
182
+ // {"id":"defer","probability":0.0046}],"confidence":0.7516,"band":"medium","abstain":false,
183
+ // "reasonCodes":["severity_score","exploit_maturity_score"],"evidenceRefs":["e1","e2"],
184
+ // "calibrationId":"finding-attention-seed-1","source":"decision"}
185
+ ```
186
+
187
+ **`architecture_fallback`** — only when Haile's own classifier has explicitly abstained. Uses its own small role vocabulary (never Haile's `taxonomy::ROLES`), reads Haile's abstention only as one input feature, and — like every other decision above — always carries `"source":"decision"` so it can never be mistaken for Haile's deterministic classification.
188
+
189
+ ```json
190
+ // in: {"schema":"vg.decision.request.v1","decision":"architecture_fallback","context":{},
191
+ // "evidence":[{"id":"e1","kind":"architecture","feature":"role_hint_authenticate","value":1},
192
+ // {"id":"e2","kind":"architecture","feature":"haile_abstained","value":true}]}
193
+ // out: {"schema":"vg.decision.result.v1","decision":"architecture_fallback","primitive":"choice",
194
+ // "selected":"authenticate","probabilities":[{"id":"authenticate","probability":0.8007},
195
+ // {"id":"persist","probability":0.0399},{"id":"respond","probability":0.0399},
196
+ // {"id":"validate","probability":0.0399},{"id":"orchestrate","probability":0.0399},
197
+ // {"id":"render","probability":0.0399}],"confidence":0.8007,"band":"high","abstain":false,
198
+ // "reasonCodes":["haile_abstained_fallback_engaged","role_hint_authenticate"],
199
+ // "evidenceRefs":["e2","e1"],"calibrationId":"architecture-fallback-seed-1","source":"decision"}
200
+ ```
201
+
202
+ ### Determinism, calibration, and abstention
203
+
204
+ Same input JSON always yields byte-identical output for a given engine/pack
205
+ version (no randomness, clock, or unordered iteration). Confidence bands
206
+ (`high`/`medium`/`low`/`abstain`) are per-decision, versioned thresholds
207
+ (`crate/src/decision/calibration.rs`), not one global cutoff. An honest
208
+ `abstain` (e.g. no evidence supplied) is a normal, fully-shaped result, not
209
+ an error — the host falls back to existing behaviour exactly as it does when
210
+ the engine itself is absent.
211
+
212
+ ### Runnable examples (`examples/decision/` — one worked example each)
213
+
214
+ ```bash
215
+ pnpm --filter @vibgrate/haile run examples:decision # all six, in sequence, real WASM output
216
+ pnpm --filter @vibgrate/haile exec tsx examples/decision/failure-related.example.ts # one at a time
217
+ ```
218
+
219
+ Six small standalone scripts, one per decision family, each building a
220
+ realistic `DecisionRequest` with an inline comment on every evidence field,
221
+ calling the real `createDecisionProvider()`, and pretty-printing the
222
+ `DecisionResult`. `docs/generate-screenshots.mjs`
223
+ (`pnpm --filter @vibgrate/haile run docs:screenshots`) runs these same six
224
+ scripts to regenerate [`docs/DECISION-ENGINE.md`](../../docs/DECISION-ENGINE.md)'s
225
+ worked-example JSON and screenshots — see that doc for the full architecture
226
+ reference plus rendered examples.
227
+
228
+ ### Tests
229
+
230
+ ```bash
231
+ pnpm --filter @vibgrate/haile test # includes tests/decision-*.test.ts and tests/bench-harness.test.ts
232
+ (cd crate && cargo test) # includes crate/src/decision/**/*'s own unit tests
233
+ ```
234
+
235
+ ### Benchmarking (`bench/` — spec §17/§18/§32)
236
+
237
+ ```bash
238
+ pnpm --filter @vibgrate/haile run bench:decision # all six decisions
239
+ pnpm --filter @vibgrate/haile run bench:decision:failure-related # one decision
240
+ pnpm --filter @vibgrate/haile run bench:decision:generate-corpus # regenerate bench/corpus/*.jsonl
241
+ ```
242
+
243
+ `bench/evaluate.ts` loads each labelled corpus under `bench/corpus/*.jsonl`,
244
+ splits it by `repo_id` (spec §32 — never by random row), runs the real
245
+ `decision-wasm` backend (`bench/backends/decision-wasm.ts`, wrapping
246
+ `createDecisionProvider()` above) over the held-out test split, and prints
247
+ accuracy, precision, recall, F1, Brier score, negative log likelihood,
248
+ expected calibration error, coverage, accuracy-by-confidence-band, and
249
+ abstention rate — plus a naive-heuristic baseline for comparison
250
+ (`bench/backends/full-model.ts`; Jev is stubbed unavailable per spec §24,
251
+ `bench/backends/jev.ts`).
252
+
253
+ > **Every corpus row is SYNTHETIC** (`bench/corpus/README.md`) and the
254
+ > seed-pack weights (`crate/src/decision/pack_data.rs`) are untrained
255
+ > placeholders — so the numbers below are genuinely mediocre. That is the
256
+ > correct, honest Stage 0 result, not a bug in the harness: this corpus is
257
+ > for validating the benchmark machinery itself, not for judging production
258
+ > readiness. It is **not** the real 500–1,000-example corpus spec §31/§32
259
+ > requires before any production hook.
260
+
261
+ Real output from an actual run (`pnpm --filter @vibgrate/haile run bench:decision`), unedited:
262
+
263
+ ```
264
+ === Summary — vibgrate-decision-wasm across all six decisions ===
265
+ decision | n(test) | accuracy | F1(macro) | Brier | NLL | ECE | coverage | abstain%
266
+ task_complexity | 15 | 46.7% | 24.6% | 0.5782 | 0.8517 | 0.3216 | 73.3% | 26.7%
267
+ failure_related | 20 | 35.0% | 25.9% | 1.2023 | 2.6754 | 0.6201 | 100.0% | 0.0%
268
+ repair_action | 18 | 100.0% | 100.0% | 0.2879 | 0.6553 | 0.4757 | 66.7% | 33.3%
269
+ completion_confidence | 20 | 55.0% | 52.0% | 0.7553 | 1.3406 | 0.3805 | 85.0% | 15.0%
270
+ finding_attention | 16 | 37.5% | 34.4% | 0.5932 | 1.0679 | 0.1527 | 31.3% | 68.8%
271
+ architecture_fallback | 21 | 100.0% | 100.0% | 0.2486 | 0.5423 | 0.3416 | 76.2% | 23.8%
272
+ ```
273
+
274
+ Per-decision output also prints a naive-heuristic-baseline row for
275
+ comparison and accuracy broken out by confidence band; run the command
276
+ above for the full table. `repair_action` and `architecture_fallback`
277
+ scoring 100% here is a property of their small synthetic test splits (18
278
+ and 21 rows respectively, after the repo-level split), not a general claim —
279
+ `task_complexity` and `failure_related`, evaluated against the same kind of
280
+ corpus, land far lower. Treat every number as "the harness works and these
281
+ are the real, unflattering seed-pack results," not as a production
282
+ readiness signal.