@hyperfixi/testing-framework 2.11.0 → 2.11.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@hyperfixi/testing-framework",
3
- "version": "2.11.0",
3
+ "version": "2.11.1",
4
4
  "description": "Cross-platform behavior testing suite for LokaScript applications",
5
5
  "main": "dist/index.js",
6
6
  "module": "dist/index.mjs",
@@ -57,11 +57,11 @@
57
57
  "author": "LokaScript Contributors",
58
58
  "license": "MIT",
59
59
  "dependencies": {
60
- "@hyperfixi/core": "^2.11.0",
61
- "@hyperfixi/patterns-reference": "^2.11.0",
62
- "@lokascript/compilation-service": "^2.11.0",
63
- "@lokascript/i18n": "^2.11.0",
64
- "@lokascript/semantic": "^2.11.0",
60
+ "@hyperfixi/core": "^2.11.1",
61
+ "@hyperfixi/patterns-reference": "^2.11.1",
62
+ "@lokascript/compilation-service": "^2.11.1",
63
+ "@lokascript/i18n": "^2.11.1",
64
+ "@lokascript/semantic": "^2.11.1",
65
65
  "diff": "^8.0.3",
66
66
  "esbuild": "^0.28.0",
67
67
  "happy-dom": "^20.10.6",
@@ -71,13 +71,13 @@ Run file:
71
71
  }
72
72
  ```
73
73
 
74
- **No A/B run is committed here, deliberately.** The tasks and their reference
75
- implementations were authored in the same session that would have produced the
76
- candidates, so any one-shot number from that session measures recall of
77
- just-written answers, not generation. A meaningful run needs a generator that
78
- has not seen this directory — until one exists, the honest position is a harness
79
- with no number attached, not a flattering number with a caveat. `score` is fully
80
- implemented and ready for that run.
74
+ **The committed run required an isolated generator, and waited for one.** The
75
+ tasks and their references were authored in the same session that would have
76
+ produced the candidates, so any one-shot number from that session would have
77
+ measured recall of just-written answers, not generation. The run in
78
+ [`runs/ab-2026-08-25.json`](../../runs/ab-2026-08-25.json) was generated by an
79
+ agent that had never seen this directory — see [the A/B finding](#the-ab-run-2026-08-25)
80
+ for the result and how isolation was enforced.
81
81
 
82
82
  Not a CI gate: LLM-in-the-loop is nondeterministic and this repo's gates stay
83
83
  deterministic. Half 1 **is** gated, because it has no generator in it.
@@ -160,3 +160,60 @@ reference, and the fixture markup it needs. Then run `verify-references` — a
160
160
  reference that does not parse, or produces no DOM effect, is rejected (same
161
161
  eligibility bar as R2's execution subset), because scoring against an empty
162
162
  signature would make wrong answers look right.
163
+
164
+ ### The A/B run, 2026-08-25
165
+
166
+ The first run with a generator that had never seen this directory.
167
+ [`runs/ab-2026-08-25.json`](../../runs/ab-2026-08-25.json).
168
+
169
+ | | one-shot | loop |
170
+ | ---------------- | -------- | -------- |
171
+ | parse rate | 100% | 100% |
172
+ | behavior rate | **90%** | **100%** |
173
+ | parsed-but-wrong | 0 | 0 |
174
+
175
+ **The loop closed the gap, and every point of it came through a diagnostic
176
+ Arc 3b added.** Both one-shot failures were the same shape — `set <prop> of
177
+ <target> to <value>`, which returns `ok: true` at confidence 1.0 with an
178
+ **empty handler** (`parsed: "on()"`) and produces no DOM effect at all:
179
+
180
+ ```
181
+ set @aria-expanded of #panel to 'true' → on() ⚠ UNCONSUMED_INPUT (6 tokens)
182
+ set innerHTML of #output to 'Done' → on() ⚠ UNCONSUMED_INPUT (6 tokens)
183
+ ```
184
+
185
+ Pre-3b these two rows would have been **silent** — `ok: true`, zero
186
+ diagnostics — and the loop would have had nothing to react to, so the delta
187
+ would have been 0. The `UNCONSUMED_INPUT` plumbing fix is what made them
188
+ repairable. That is the benchmark's own claim about itself, measured rather
189
+ than asserted: the loop is worth what the diagnostics are worth.
190
+
191
+ The generator rewrote both to the possessive form (`set #panel's
192
+ @aria-expanded to "true"`) after the warning, and reported a third repair the
193
+ scoreboard cannot show — it rejected `set *innerHTML of #output to "Done"`,
194
+ which parses **clean at confidence 1.0**, on IR inspection alone: `*` is the
195
+ style sigil, so the IR read `destination=#output.*innerHTML`. That is the
196
+ loop's step-2 discipline (_check the IR against your intent_) catching what no
197
+ diagnostic could.
198
+
199
+ Isolation, enforced rather than requested:
200
+
201
+ - One-shot answers were written and `chmod 444` locked **before** the first
202
+ `feedback` call, so the loop phase could not retroactively revise them.
203
+ - `feedback` is the only channel to the validator and returns only diagnostics
204
+ and the IR — never the reference, never the behavior verdict.
205
+ - Audited from the generator's transcript afterward: **zero** Read/Glob/Grep
206
+ calls, all 48 Bash calls were the two permitted wrapper scripts, zero
207
+ commands referencing the repo path.
208
+ - The generator was barred from the project docs too, not just this directory —
209
+ [AGENTS.md](../../../../AGENTS.md) quotes this benchmark's findings and names
210
+ the specific trap phrasings, so a doc-equipped generator would be reading
211
+ leaked answers.
212
+
213
+ Read the number for what it is: **n=1 generator, 20 tasks, 2 tasks moved.** It
214
+ is not a claim about models in general, and the one-shot arm is an
215
+ _undocumented_ generator — a briefed one starts higher and the delta narrows.
216
+ What is solid is the mechanism, which the probe half measures deterministically:
217
+ the failures the loop can fix are exactly the ones something makes visible.
218
+ Of the 7 answers the loop changed, 2 changed the outcome and 5 were cosmetic —
219
+ iterating on already-correct answers cost nothing and broke nothing.