@hyperfixi/testing-framework 2.11.0 → 2.11.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +6 -6
- package/src/agent-bench/README.md +64 -7
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@hyperfixi/testing-framework",
|
|
3
|
-
"version": "2.11.
|
|
3
|
+
"version": "2.11.1",
|
|
4
4
|
"description": "Cross-platform behavior testing suite for LokaScript applications",
|
|
5
5
|
"main": "dist/index.js",
|
|
6
6
|
"module": "dist/index.mjs",
|
|
@@ -57,11 +57,11 @@
|
|
|
57
57
|
"author": "LokaScript Contributors",
|
|
58
58
|
"license": "MIT",
|
|
59
59
|
"dependencies": {
|
|
60
|
-
"@hyperfixi/core": "^2.11.
|
|
61
|
-
"@hyperfixi/patterns-reference": "^2.11.
|
|
62
|
-
"@lokascript/compilation-service": "^2.11.
|
|
63
|
-
"@lokascript/i18n": "^2.11.
|
|
64
|
-
"@lokascript/semantic": "^2.11.
|
|
60
|
+
"@hyperfixi/core": "^2.11.1",
|
|
61
|
+
"@hyperfixi/patterns-reference": "^2.11.1",
|
|
62
|
+
"@lokascript/compilation-service": "^2.11.1",
|
|
63
|
+
"@lokascript/i18n": "^2.11.1",
|
|
64
|
+
"@lokascript/semantic": "^2.11.1",
|
|
65
65
|
"diff": "^8.0.3",
|
|
66
66
|
"esbuild": "^0.28.0",
|
|
67
67
|
"happy-dom": "^20.10.6",
|
|
@@ -71,13 +71,13 @@ Run file:
|
|
|
71
71
|
}
|
|
72
72
|
```
|
|
73
73
|
|
|
74
|
-
**
|
|
75
|
-
|
|
76
|
-
candidates, so any one-shot number from that session
|
|
77
|
-
just-written answers, not generation.
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
74
|
+
**The committed run required an isolated generator, and waited for one.** The
|
|
75
|
+
tasks and their references were authored in the same session that would have
|
|
76
|
+
produced the candidates, so any one-shot number from that session would have
|
|
77
|
+
measured recall of just-written answers, not generation. The run in
|
|
78
|
+
[`runs/ab-2026-08-25.json`](../../runs/ab-2026-08-25.json) was generated by an
|
|
79
|
+
agent that had never seen this directory — see [the A/B finding](#the-ab-run-2026-08-25)
|
|
80
|
+
for the result and how isolation was enforced.
|
|
81
81
|
|
|
82
82
|
Not a CI gate: LLM-in-the-loop is nondeterministic and this repo's gates stay
|
|
83
83
|
deterministic. Half 1 **is** gated, because it has no generator in it.
|
|
@@ -160,3 +160,60 @@ reference, and the fixture markup it needs. Then run `verify-references` — a
|
|
|
160
160
|
reference that does not parse, or produces no DOM effect, is rejected (same
|
|
161
161
|
eligibility bar as R2's execution subset), because scoring against an empty
|
|
162
162
|
signature would make wrong answers look right.
|
|
163
|
+
|
|
164
|
+
### The A/B run, 2026-08-25
|
|
165
|
+
|
|
166
|
+
The first run with a generator that had never seen this directory.
|
|
167
|
+
[`runs/ab-2026-08-25.json`](../../runs/ab-2026-08-25.json).
|
|
168
|
+
|
|
169
|
+
| | one-shot | loop |
|
|
170
|
+
| ---------------- | -------- | -------- |
|
|
171
|
+
| parse rate | 100% | 100% |
|
|
172
|
+
| behavior rate | **90%** | **100%** |
|
|
173
|
+
| parsed-but-wrong | 0 | 0 |
|
|
174
|
+
|
|
175
|
+
**The loop closed the gap, and every point of it came through a diagnostic
|
|
176
|
+
Arc 3b added.** Both one-shot failures were the same shape — `set <prop> of
|
|
177
|
+
<target> to <value>`, which returns `ok: true` at confidence 1.0 with an
|
|
178
|
+
**empty handler** (`parsed: "on()"`) and produces no DOM effect at all:
|
|
179
|
+
|
|
180
|
+
```
|
|
181
|
+
set @aria-expanded of #panel to 'true' → on() ⚠ UNCONSUMED_INPUT (6 tokens)
|
|
182
|
+
set innerHTML of #output to 'Done' → on() ⚠ UNCONSUMED_INPUT (6 tokens)
|
|
183
|
+
```
|
|
184
|
+
|
|
185
|
+
Pre-3b these two rows would have been **silent** — `ok: true`, zero
|
|
186
|
+
diagnostics — and the loop would have had nothing to react to, so the delta
|
|
187
|
+
would have been 0. The `UNCONSUMED_INPUT` plumbing fix is what made them
|
|
188
|
+
repairable. That is the benchmark's own claim about itself, measured rather
|
|
189
|
+
than asserted: the loop is worth what the diagnostics are worth.
|
|
190
|
+
|
|
191
|
+
The generator rewrote both to the possessive form (`set #panel's
|
|
192
|
+
@aria-expanded to "true"`) after the warning, and reported a third repair the
|
|
193
|
+
scoreboard cannot show — it rejected `set *innerHTML of #output to "Done"`,
|
|
194
|
+
which parses **clean at confidence 1.0**, on IR inspection alone: `*` is the
|
|
195
|
+
style sigil, so the IR read `destination=#output.*innerHTML`. That is the
|
|
196
|
+
loop's step-2 discipline (_check the IR against your intent_) catching what no
|
|
197
|
+
diagnostic could.
|
|
198
|
+
|
|
199
|
+
Isolation, enforced rather than requested:
|
|
200
|
+
|
|
201
|
+
- One-shot answers were written and `chmod 444` locked **before** the first
|
|
202
|
+
`feedback` call, so the loop phase could not retroactively revise them.
|
|
203
|
+
- `feedback` is the only channel to the validator and returns only diagnostics
|
|
204
|
+
and the IR — never the reference, never the behavior verdict.
|
|
205
|
+
- Audited from the generator's transcript afterward: **zero** Read/Glob/Grep
|
|
206
|
+
calls, all 48 Bash calls were the two permitted wrapper scripts, zero
|
|
207
|
+
commands referencing the repo path.
|
|
208
|
+
- The generator was barred from the project docs too, not just this directory —
|
|
209
|
+
[AGENTS.md](../../../../AGENTS.md) quotes this benchmark's findings and names
|
|
210
|
+
the specific trap phrasings, so a doc-equipped generator would be reading
|
|
211
|
+
leaked answers.
|
|
212
|
+
|
|
213
|
+
Read the number for what it is: **n=1 generator, 20 tasks, 2 tasks moved.** It
|
|
214
|
+
is not a claim about models in general, and the one-shot arm is an
|
|
215
|
+
_undocumented_ generator — a briefed one starts higher and the delta narrows.
|
|
216
|
+
What is solid is the mechanism, which the probe half measures deterministically:
|
|
217
|
+
the failures the loop can fix are exactly the ones something makes visible.
|
|
218
|
+
Of the 7 answers the loop changed, 2 changed the outcome and 5 were cosmetic —
|
|
219
|
+
iterating on already-correct answers cost nothing and broke nothing.
|