pog-mcp 0.9.2 → 0.9.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/skill/reference/measurements.md +2174 -0
|
@@ -1228,6 +1228,480 @@ N=2000 WIDE_N=1000 npx tsx … # faster; the bar is fixed, the detectable EFFE
|
|
|
1228
1228
|
|
|
1229
1229
|
---
|
|
1230
1230
|
|
|
1231
|
+
## Is a per-player scene weight a decision? (`appr-role.mts`)
|
|
1232
|
+
|
|
1233
|
+
The section above swept engine constants and found nothing that re-prices a
|
|
1234
|
+
squad choice — not because the effects were small, but structurally: every
|
|
1235
|
+
constant it can reach sits in front of BOTH `ofPoint` and `dfPoint`, so it moves
|
|
1236
|
+
the two squads together. Per-player selection weights are the one remaining
|
|
1237
|
+
candidate that creates a **difference between two squads** rather than moving
|
|
1238
|
+
both sides, and they are the axis the tactical-sheet track would ship.
|
|
1239
|
+
|
|
1240
|
+
**The lever measured here does not exist in the engine.** `*_APPR` is a table
|
|
1241
|
+
per POSITION, so two forwards in the same squad are drawn with the same weight
|
|
1242
|
+
and a manager cannot tell them apart. The probe patches `getPlayer` to multiply
|
|
1243
|
+
each candidate's weight by a per-player, per-scene number carried on the player,
|
|
1244
|
+
defaulting to 1 — nothing else — and sweeps that number. So the answer is about
|
|
1245
|
+
the axis, not about one proposed interface.
|
|
1246
|
+
|
|
1247
|
+
### The statistic, and the controls
|
|
1248
|
+
|
|
1249
|
+
Rows are in **win-equivalent share**: a win is 1, a draw is 0.5, every match
|
|
1250
|
+
counts. Deltas are **paired** — every sheet replays one seed column, so the
|
|
1251
|
+
default cell and the swept cell are matched match for match — and the seed
|
|
1252
|
+
depends on the squads and the side and never on the sheet.
|
|
1253
|
+
|
|
1254
|
+
**Bonferroni over m = 13,314 pre-registered comparisons, two-sided: |z| >= 4.62.**
|
|
1255
|
+
That m prices the SEARCHES as well as the tests. A ladder read against the
|
|
1256
|
+
FIXED shipped default costs only its rungs, because an argmax over the ladder
|
|
1257
|
+
could only ever have surfaced one of the contrasts already counted. A comparison
|
|
1258
|
+
whose BOTH sides are data-chosen — "is the optimum interior", which pits a
|
|
1259
|
+
selected rung against a selected neighbour, and "which sheet is the best
|
|
1260
|
+
response", which pits two selected sheets — is charged K(K−1)/2, the number of
|
|
1261
|
+
pairs the search could have surfaced.
|
|
1262
|
+
|
|
1263
|
+
The noise floor is the whole run's, not one section's: **the largest standard
|
|
1264
|
+
error anywhere in this run is 0.59 points of share**, which puts the
|
|
1265
|
+
significance threshold at **2.74 points** and the 80%-power effect at **3.24**.
|
|
1266
|
+
A paired standard error depends on the outcome variance and on the covariance
|
|
1267
|
+
between the two conditions, so it is not constant across sections merely because
|
|
1268
|
+
`N` is — quoting the concentration ladders' own worst cell and extending it to
|
|
1269
|
+
every later null was a review finding on this section. So every effect called
|
|
1270
|
+
large below is an order of magnitude above the run's own floor, and every "no
|
|
1271
|
+
effect" is a claim about a band 2.74 points wide.
|
|
1272
|
+
|
|
1273
|
+
| Control | Result |
|
|
1274
|
+
| --- | --- |
|
|
1275
|
+
| Mirror, both modes (one squad against itself; true z = 0) | z **+0.4** / **−0.5** |
|
|
1276
|
+
| POLICY mirror, both modes (same squad AND same sheet on both sides; true z = 0) | z **+1.0** / **+0.8** |
|
|
1277
|
+
| Patch identity — the patched engine with no sheet vs the stock engine | 1,000 matches identical, **whole `MatchResult`** |
|
|
1278
|
+
| Permutation null — the same ×32 on either of `role-433`'s two byte-identical midfielders | 1,000 matches identical, scoreline and winner |
|
|
1279
|
+
| Dead-scene null — ×1024 and ×0 on the offence table's ZP/PP1/FK1/FK2 columns | 1,000 matches identical, **whole `MatchResult`** |
|
|
1280
|
+
| Zero-weight null — ×1024 on the keeper, whose offence weight is 0 everywhere | 1,000 matches identical, **whole `MatchResult`** |
|
|
1281
|
+
| Paired alignment — the default column, replayed | 10,000 matches, largest per-match difference **0** |
|
|
1282
|
+
| Anti-focus calibration — funnelling `stars-433`'s 12-point defender must LOSE | **−11.6pp** (z −28.9) / **−14.0pp** (z −26.6) |
|
|
1283
|
+
|
|
1284
|
+
<!-- not a payoff table -->
|
|
1285
|
+
|
|
1286
|
+
**Every one of those aborts the run.** The comparison DEPTH in the right-hand
|
|
1287
|
+
column is part of the control rather than an implementation detail: where the
|
|
1288
|
+
claim is that nothing happened at all, the whole `MatchResult` is compared —
|
|
1289
|
+
every event in order, every player stat, every shootout round — because a
|
|
1290
|
+
rewrite that changed which player was drawn while landing on the same score
|
|
1291
|
+
would pass a scoreline check and leave every table below measuring the rewrite.
|
|
1292
|
+
The permutation null is the one exception and deliberately so: there the two
|
|
1293
|
+
players ARE interchanged, so the event names and per-player stats move by
|
|
1294
|
+
construction and only the outcome can be asserted.
|
|
1295
|
+
|
|
1296
|
+
Three are worth naming for what they separate. The **permutation null**
|
|
1297
|
+
separates "the weight moved the ball to a better player" from "the weight
|
|
1298
|
+
perturbed the RNG": moving the same multiplier between two players who are
|
|
1299
|
+
identical to the byte must change nothing, and it does not. The **anti-focus
|
|
1300
|
+
calibration** is the directional negative control —
|
|
1301
|
+
a harness that reported a gain from concentrating on the squad's worst player
|
|
1302
|
+
would be measuring the act of concentrating rather than who is concentrated on,
|
|
1303
|
+
and every row below would be that artefact. The **dead-scene null** doubles as a
|
|
1304
|
+
re-derivation: the probe enumerates the offensive `getPlayer` call sites FROM
|
|
1305
|
+
THE SOURCE, finds `A0 A1 A2 CA1 CA2 PP2` and nothing else, and confirms that a
|
|
1306
|
+
multiplier on the four remaining offence columns is inert — the dead table the
|
|
1307
|
+
plan describes, checked rather than quoted.
|
|
1308
|
+
|
|
1309
|
+
### The concentration curve
|
|
1310
|
+
|
|
1311
|
+
`stars-433` is the focal squad because it is the only published archetype whose
|
|
1312
|
+
players differ inside a position group — three forwards at 28-29 points and four
|
|
1313
|
+
defenders at 12. `role-433` rides along as the other half of the question: its
|
|
1314
|
+
three forwards are identical to the byte, so any movement there is cross-GROUP
|
|
1315
|
+
reallocation with the within-group term pinned at exactly zero.
|
|
1316
|
+
|
|
1317
|
+
| Focal slot (`stars-433`) | ×0 | ×1 | ×1024 | League | Knockout |
|
|
1318
|
+
| --- | --- | --- | --- | --- | --- |
|
|
1319
|
+
| slot 8 — FW, 29 pts, pass 6 shoot 9 | 33.1 | 36.7 | **68.0** | **+31.3pp** | **+35.8pp** |
|
|
1320
|
+
| slot 10 — FW, 28 pts, pass 4 shoot 10 | 33.1 | 36.7 | 64.8 | +28.1pp | +32.3pp |
|
|
1321
|
+
| slot 7 — OMF, 26 pts, alone in his group | 37.5 | 36.7 | 29.2 | −7.5pp | −8.8pp |
|
|
1322
|
+
| slot 1 — DF, 12 pts (anti-focus control) | 38.0 | 36.7 | 25.1 | −11.6pp | −14.0pp |
|
|
1323
|
+
| `role-433` slot 8 — FW, identical to its two peers | 46.9 | 48.1 | 58.9 | +10.8pp | +13.2pp |
|
|
1324
|
+
|
|
1325
|
+
> **Measured on** — `stars-433` (4-3-3, `STARS` slot weights, `ATTACKING` mix
|
|
1326
|
+
> from `squad-lib.mts`: keeper 19, four defenders at 12, midfield 17/16/26,
|
|
1327
|
+
> forwards 29/29/28) against `def-433` (4-3-3, `EVEN` weights, `DEFENSIVE` mix),
|
|
1328
|
+
> which plays the shipped weights in every cell. The two squads never change:
|
|
1329
|
+
> all 212 points, both formations, both sets of kicker flags and the whole
|
|
1330
|
+
> defensive table are byte-identical across every row. The only thing that moves
|
|
1331
|
+
> is one number in front of one player's offensive selection weight, in all six
|
|
1332
|
+
> reachable scenes at once. The last row swaps the focal squad for `role-433`
|
|
1333
|
+
> (`EVEN` weights, `ATTACKING` mix), whose forwards are identical to each other.
|
|
1334
|
+
> **Sample** — 10,000 matches per cell, home and away, SHA-256 seeds; 20,668,000
|
|
1335
|
+
> matches in the run, about four minutes.
|
|
1336
|
+
> **Mode** — both, one column each. League is draws-allowed (`gameFlg 0`),
|
|
1337
|
+
> Knockout is winner-guaranteed. The League and Knockout columns are the paired
|
|
1338
|
+
> delta from the ×1 cell; the ×0/×1/×1024 columns are League shares.
|
|
1339
|
+
> **Effect** — 31.3 percentage points of win-equivalent share in League and 35.8
|
|
1340
|
+
> in Knockout, between two squads that are the same eleven players. That is
|
|
1341
|
+
> nearly three times the largest thing the engine-constant sweep could move
|
|
1342
|
+
> (11.8pp) and larger than any squad-construction axis measured anywhere in this
|
|
1343
|
+
> file. Concentrating on the WRONG player is worth −11.6pp, so the axis is
|
|
1344
|
+
> double-edged rather than free. Even with the within-group term removed by
|
|
1345
|
+
> construction — `role-433`'s identical forwards — pure cross-group
|
|
1346
|
+
> reallocation is still worth 10.8pp.
|
|
1347
|
+
> **Source** — `appr-role.mts`, section (1).
|
|
1348
|
+
|
|
1349
|
+
**The best focal player is not the best finisher.** Slot 8 beats slot 10 by
|
|
1350
|
+
3.2pp in League and 3.5pp in Knockout, and slot 10 is the one with `shoot` 10.
|
|
1351
|
+
That much is measured. **Why is not** — and it cannot be, on a fixed points
|
|
1352
|
+
budget. The two forwards differ in five things at once: `pass` 6 against 4,
|
|
1353
|
+
`dribble` 9 against 10, `shoot` 9 against 10, `defense` 5 against 4, and one
|
|
1354
|
+
point of total; slot 10 also carries the penalty-kicker flag. Moving a single
|
|
1355
|
+
attribute in isolation is impossible when every point spent has to come from
|
|
1356
|
+
another attribute of the same player, so no cell in this probe separates them.
|
|
1357
|
+
|
|
1358
|
+
The candidate mechanism, offered as a hypothesis rather than a result, is the
|
|
1359
|
+
`style = attrs.pass` branch: `playArea2` and `playArea1` both decide
|
|
1360
|
+
pass-versus-dribble-versus-long-shot from `pass` alone, so a low-pass carrier
|
|
1361
|
+
ends possessions with long shots instead of advancing them. What supports it
|
|
1362
|
+
without a cross-player comparison is the scene table below — the same player's
|
|
1363
|
+
multiplier reverses sign between the midfield scene and the box scenes — and
|
|
1364
|
+
what would settle it is an engine change, not another squad.
|
|
1365
|
+
|
|
1366
|
+
### Does the anti-concentration term bite?
|
|
1367
|
+
|
|
1368
|
+
`atkAttrSum` is the selected player's POSITION GROUP total minus his own, so the
|
|
1369
|
+
better the player selected, the less team-aggregate support stands behind him —
|
|
1370
|
+
and `atkRate` varies by scene, so the size of that penalty does too. The
|
|
1371
|
+
question the plan asks is whether that produces a non-flat optimum.
|
|
1372
|
+
|
|
1373
|
+
It is real, and it is not enough. One point of a selected player's total is
|
|
1374
|
+
worth **+1/3** on `ofPoint` directly and costs at most **rate × 0.1 ×
|
|
1375
|
+
teamPow/100** of group support. The largest `O_*_RATE` anywhere in the shipped
|
|
1376
|
+
tables is 1.00, so the loss coefficient never exceeds 0.100 against a gain
|
|
1377
|
+
coefficient of 0.333. The probe prints that arithmetic from the source tables
|
|
1378
|
+
and then measures it against a COUNTERFACTUAL engine with the subtraction
|
|
1379
|
+
deleted:
|
|
1380
|
+
|
|
1381
|
+
| `stars-433` slot 8, all scenes | ×1 | ×8 | ×1024 |
|
|
1382
|
+
| --- | --- | --- | --- |
|
|
1383
|
+
| shipped, League | 36.7 | 49.2 | 68.0 |
|
|
1384
|
+
| subtraction deleted, League | 36.9 | 52.3 | 73.4 |
|
|
1385
|
+
| shipped, Knockout | 34.6 | 49.6 | 70.4 |
|
|
1386
|
+
| subtraction deleted, Knockout | 35.2 | 52.9 | 75.7 |
|
|
1387
|
+
|
|
1388
|
+
> **Measured on** — the `stars-433` vs `def-433` fixture above, played twice:
|
|
1389
|
+
> once on the shipped engine and once on an engine identical except that all
|
|
1390
|
+
> five `atkAttrSum = <group> - <player>.total` sites — including the home-team
|
|
1391
|
+
> DMF path that reads the ball carrier rather than the selected player — are
|
|
1392
|
+
> replaced by `atkAttrSum = <group>`. Nothing else differs; the ×1 row is the
|
|
1393
|
+
> same squads with no sheet at all, so the gap between the two ×1 cells is what
|
|
1394
|
+
> the term is worth to a game nobody has concentrated.
|
|
1395
|
+
> **Sample** — 10,000 matches per cell, home and away, SHA-256 seeds.
|
|
1396
|
+
> **Mode** — both, two rows each.
|
|
1397
|
+
> **Effect** — 5.4 points of share in League (z 13.4) and 5.3 in Knockout
|
|
1398
|
+
> (z 11.2) is what the term is worth AT MAXIMUM CONCENTRATION; at the shipped ×1
|
|
1399
|
+
> it is worth 0.2 (z 0.5) and 0.7 (z 1.4), both inside the sample's own error.
|
|
1400
|
+
> So it takes 5.2 points off a concentration payoff that would otherwise reach
|
|
1401
|
+
> 36.5 — about a seventh — and changes nothing at all about who you should
|
|
1402
|
+
> concentrate on. **0 of the 50 ladders in the one-opponent sections have a
|
|
1403
|
+
> significant interior optimum: 48 put their argmax at a ladder END and the
|
|
1404
|
+
> other 2 peak on an interior rung the sample cannot separate from its
|
|
1405
|
+
> neighbours,** which is undecided rather than monotone and is counted
|
|
1406
|
+
> separately for that reason. Monotonicity is a different question from where
|
|
1407
|
+
> the argmax is — a ladder can dip significantly in the middle and still peak at
|
|
1408
|
+
> the top — so every adjacent rung is tested too: **134 of the 134 ladders swept
|
|
1409
|
+
> anywhere in this run have no significant step against their own direction.**
|
|
1410
|
+
> The curve is steeper without the term and the same shape.
|
|
1411
|
+
> **Source** — `appr-role.mts`, section (2).
|
|
1412
|
+
|
|
1413
|
+
The place where the term is not small is the **lone player in his group**. A
|
|
1414
|
+
4-3-3's single OMF gets `omfAttr − his own total` = exactly zero support, while
|
|
1415
|
+
a forward in a three-man group carries 4.56 points of it into every A1
|
|
1416
|
+
selection. That is most of why slot 7 loses in the table above despite holding
|
|
1417
|
+
26 points: funnelling the attack through him costs the whole group term. It is
|
|
1418
|
+
also a fact about FORMATION rather than about roles — the same player in a 4-4-2
|
|
1419
|
+
would have a partner to be supported by.
|
|
1420
|
+
|
|
1421
|
+
### Scene-dependence: the same player, opposite instructions
|
|
1422
|
+
|
|
1423
|
+
| `stars-433` sheet on slot 7 (the 26-point OMF) | League | Knockout |
|
|
1424
|
+
| --- | --- | --- |
|
|
1425
|
+
| default — no sheet | 36.7 | 34.6 |
|
|
1426
|
+
| ×32 in A2 only | 39.0 (+2.3pp) | 37.5 (+2.9pp) |
|
|
1427
|
+
| ×0 in A0 and A1 only | 38.4 (+1.7pp) | 36.5 (+2.0pp) |
|
|
1428
|
+
| **both — ×32 in A2, ×0 in A0/A1** | **41.4 (+4.7pp)** | **40.1 (+5.5pp)** |
|
|
1429
|
+
| the same ×32 in EVERY scene | 30.7 (−6.0pp) | 27.6 (−7.0pp) |
|
|
1430
|
+
|
|
1431
|
+
> **Measured on** — the `stars-433` vs `def-433` fixture above, with every sheet
|
|
1432
|
+
> applied to ONE slot: the 26-point attacking midfielder, pass 8 dribble 8 shoot
|
|
1433
|
+
> 7, the only OMF in the formation. Every other slot on both squads is
|
|
1434
|
+
> untouched, and the opponent plays the shipped weights throughout.
|
|
1435
|
+
> **Sample** — 10,000 matches per cell, home and away, SHA-256 seeds.
|
|
1436
|
+
> **Mode** — both, one column each; every row clears the corrected bar (|z| 5.8
|
|
1437
|
+
> to 14.4).
|
|
1438
|
+
> **Effect** — 10.7 percentage points of share in League and 12.5 in Knockout
|
|
1439
|
+
> separate the best and worst way of spending the SAME multiplier on the SAME
|
|
1440
|
+
> player: +4.7pp when its sign follows the scene and −6.0pp when it does not.
|
|
1441
|
+
> A per-player slider would have produced the losing row. The scene signs are
|
|
1442
|
+
> stable across modes: A2 positive, A0/A1/CA1/PP2 negative, CA2 inside the noise.
|
|
1443
|
+
> **Source** — `appr-role.mts`, section (3).
|
|
1444
|
+
|
|
1445
|
+
This is the strongest evidence for the `style = attrs.pass` hypothesis above,
|
|
1446
|
+
and it is stronger for being WITHIN one player: the OMF is the squad's best
|
|
1447
|
+
passer, and his multiplier pays in midfield (A2) and costs in the box (A0/A1),
|
|
1448
|
+
where the forwards' `shoot` and their group support both beat him. No squad
|
|
1449
|
+
attribute changes between those rows — only which scene the weight applies to —
|
|
1450
|
+
so the confound that blocks the cross-player comparison is absent here. It is
|
|
1451
|
+
still not proof: A0/A1 differ from A2 in what they reward as well as in who
|
|
1452
|
+
carries, and separating those needs an engine change. **No scene of the
|
|
1453
|
+
forward's own ladder points the other way**, so
|
|
1454
|
+
scene-dependence is a property of the player rather than of the axis: slot 8
|
|
1455
|
+
gives one distinct argmax in both modes, six positive signs in League and five
|
|
1456
|
+
plus one below the bar in Knockout, and not a single negative one.
|
|
1457
|
+
|
|
1458
|
+
### Opponent-dependence, and the round robin that decides it
|
|
1459
|
+
|
|
1460
|
+
| Sheet, `stars-433` | vs `role` | vs `def-433` | vs `def-532` | vs `shootonly` | vs `passonly` | vs `flat` |
|
|
1461
|
+
| --- | --- | --- | --- | --- | --- | --- |
|
|
1462
|
+
| default | 37.2 | 36.7 | 34.5 | 42.5 | 36.1 | 47.4 |
|
|
1463
|
+
| **carrier** — slot 8 ×32, all scenes | **60.7** | **60.6** | **52.7** | **71.4** | **62.6** | **75.6** |
|
|
1464
|
+
| finisher — slot 10 ×32, all scenes | 58.6 | 58.4 | 52.3 | 68.1 | 59.1 | 72.3 |
|
|
1465
|
+
| playmaker — slot 7 ×32 A2, ×0 A0/A1 | 42.8 | 41.4 | 37.1 | 50.9 | 44.1 | 55.6 |
|
|
1466
|
+
| sheet — playmaker plus slot 10 ×32 in A0/A1 | 54.6 | 51.3 | 44.0 | 67.6 | 57.7 | 71.8 |
|
|
1467
|
+
| tier — slots 7-10 ×8, all scenes | 50.7 | 49.3 | 43.8 | 62.2 | 52.7 | 68.2 |
|
|
1468
|
+
| spread — the four 12-point defenders ×4 | 28.6 | 30.3 | 30.3 | 30.4 | 26.4 | 33.9 |
|
|
1469
|
+
|
|
1470
|
+
> **Measured on** — `stars-433` against six published archetypes built by
|
|
1471
|
+
> `squad-lib.mts` — `role-433`, `def-433`, `def-532`, `shootonly-433`,
|
|
1472
|
+
> `passonly-433`, `flat-433` — every one of them playing the shipped weights.
|
|
1473
|
+
> The focal squad's 212 points never change across the whole grid; only its
|
|
1474
|
+
> sheet does. `tier` is the coordinated policy: four players lifted together at
|
|
1475
|
+
> a smaller multiplier rather than one player at a large one.
|
|
1476
|
+
> **Sample** — 10,000 matches per cell, home and away, SHA-256 seeds; 84 cells
|
|
1477
|
+
> across the two modes.
|
|
1478
|
+
> **Mode** — League shown; the Knockout grid is in the probe output, where
|
|
1479
|
+
> `carrier` runs 54.0 to 78.6 and every margin over `default` is wider. Its
|
|
1480
|
+
> ranking matches this one in five of its six rows — the sixth is the `def-532`
|
|
1481
|
+
> swap in Effect below.
|
|
1482
|
+
> **Effect** — 41.7 percentage points of share separate the best sheet from the
|
|
1483
|
+
> worst against the same opponent (`flat-433`: 75.6 vs 33.9), and the ordering
|
|
1484
|
+
> of the seven sheets is identical in eleven of the twelve opponent-by-mode
|
|
1485
|
+
> cells. The twelfth is `def-532` in Knockout, where `carrier` and `finisher`
|
|
1486
|
+
> swap by 0.1 points of share — inside the sample's own error, which is why the
|
|
1487
|
+
> band there names both. `carrier` sits in every cell's leader band; `def-532`,
|
|
1488
|
+
> the deepest defensive shape in the grid, is the only opponent that pulls
|
|
1489
|
+
> anything level with it. The distinct raw argmax count is **2**, and one of the
|
|
1490
|
+
> two is never beaten anywhere.
|
|
1491
|
+
> **Source** — `appr-role.mts`, section (4).
|
|
1492
|
+
|
|
1493
|
+
That grid holds the opponent's sheet at the default. The next one lets the
|
|
1494
|
+
opponent choose too, on the SAME squad, so the sheet is the only difference:
|
|
1495
|
+
|
|
1496
|
+
| Round robin, `stars-433` both sides | beats significantly (League / KO) | in the unbeaten band |
|
|
1497
|
+
| --- | --- | --- |
|
|
1498
|
+
| carrier | 5 of 6 / 5 of 6 | yes |
|
|
1499
|
+
| sheet | 4 of 6 / 3 of 6 | yes |
|
|
1500
|
+
| finisher | 3 of 6 / 3 of 6 | no |
|
|
1501
|
+
| tier | 3 of 6 / 3 of 6 | no |
|
|
1502
|
+
| playmaker | 2 of 6 / 2 of 6 | no |
|
|
1503
|
+
| default | 1 of 6 / 1 of 6 | no |
|
|
1504
|
+
| spread | 0 of 6 / 0 of 6 | no |
|
|
1505
|
+
|
|
1506
|
+
> **Measured on** — `stars-433` against `stars-433`: the same eleven players,
|
|
1507
|
+
> the same 212 points and the same kicker flags on both sides of every fixture,
|
|
1508
|
+
> with the sheet as the only difference. Twenty-one fixtures per mode.
|
|
1509
|
+
> **Sample** — 10,000 matches per fixture, home and away, SHA-256 seeds.
|
|
1510
|
+
> **Mode** — both; the two modes agree row for row.
|
|
1511
|
+
> **Effect** — 76.0% of the share for `carrier` against `default` in League and
|
|
1512
|
+
> 78.3% in Knockout — a 52-to-57 point win-rate advantage from the weights
|
|
1513
|
+
> alone. **Significant 3-cycles: 0 in both modes.** The unbeaten set — the
|
|
1514
|
+
> sheets nothing significantly beats — holds `carrier` and `sheet`, 1.4 points
|
|
1515
|
+
> of share apart in League and 1.5 in Knockout, under the corrected bar. Read
|
|
1516
|
+
> that as a PARTIAL order and not a total one: 3 of the 21 pairs in League and 4
|
|
1517
|
+
> in Knockout are inside the bar, so "no cycles and a two-member unbeaten set"
|
|
1518
|
+
> leaves some pairs unresolved rather than ranking all seven. What it does
|
|
1519
|
+
> settle is the direction: nothing here beats `carrier`, and no sheet outside
|
|
1520
|
+
> the unbeaten pair is unbeaten.
|
|
1521
|
+
> **Source** — `appr-role.mts`, section (5).
|
|
1522
|
+
|
|
1523
|
+
### Crossing the two — the grid the verdict actually rests on
|
|
1524
|
+
|
|
1525
|
+
The two grids above are one-dimensional slices: one varies the opponent's SQUAD
|
|
1526
|
+
with its sheet pinned to the default, the other varies its SHEET with its squad
|
|
1527
|
+
pinned. A best response can be conditional on the product while both slices
|
|
1528
|
+
share a leader, so neither settles the roadmap question on its own. Multiplying
|
|
1529
|
+
them does:
|
|
1530
|
+
|
|
1531
|
+
Each cell is my best-response BAND against that opponent squad playing that
|
|
1532
|
+
sheet, and — after the bar — the argmax rung of the full concentration ladder
|
|
1533
|
+
swept inside the same cell, so "more is better up to the corner" is checked
|
|
1534
|
+
where the best response is rather than only against one fixed opponent.
|
|
1535
|
+
|
|
1536
|
+
| My best response \| ladder argmax | vs `default` | vs `carrier` | vs `finisher` | vs `playmaker` | vs `sheet` | vs `tier` | vs `spread` |
|
|
1537
|
+
| --- | --- | --- | --- | --- | --- | --- | --- |
|
|
1538
|
+
| `role-433` | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 |
|
|
1539
|
+
| `def-433` | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 | C/F \| ×1024 | C \| ×1024 | C \| ×1024 |
|
|
1540
|
+
| **`def-532`** | **C/F** \| ×1024 | **C/F** \| ×1024 | **C/F** \| ×1024 | **C/F** \| ×1024 | **C/F** \| ×1024 | **C/F** \| ×1024 | **C/F** \| ×1024 |
|
|
1541
|
+
| `shootonly-433` | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 |
|
|
1542
|
+
| `passonly-433` | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 |
|
|
1543
|
+
| `flat-433` | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 | C \| ×1024 |
|
|
1544
|
+
|
|
1545
|
+
> **Measured on** — `stars-433` against all six opponent squads, each of them
|
|
1546
|
+
> playing each of the seven sheets, with all seven of my own sheets AND the full
|
|
1547
|
+
> nine-rung concentration ladder tried against every combination. C is
|
|
1548
|
+
> `carrier`, F is `finisher`; a cell names every sheet the sample cannot
|
|
1549
|
+
> separate from that cell's leader. A second focal squad runs the same grid:
|
|
1550
|
+
> `role-433`, the hard case, since its three forwards are identical to the byte
|
|
1551
|
+
> and only cross-group reallocation is left — against the other five squads over
|
|
1552
|
+
> a reduced sheet set (`default`, `carrier`, `spread`), 30 further cells.
|
|
1553
|
+
> **Sample** — 10,000 matches per cell, home and away, SHA-256 seeds: 114
|
|
1554
|
+
> squad-by-sheet-by-mode cells, 14,700,000 matches in this section alone.
|
|
1555
|
+
> **Mode** — League shown. The Knockout grid differs in three `def-433` cells
|
|
1556
|
+
> and nowhere else: it ties `carrier` with `finisher` against `carrier`,
|
|
1557
|
+
> `finisher` and `tier`, where League resolves, and resolves against `sheet`,
|
|
1558
|
+
> where League ties. `def-532` is C/F in all fourteen of its cells in both modes.
|
|
1559
|
+
> **Effect** — 0 cells out of 114 exclude `carrier` from the best-response band,
|
|
1560
|
+
> and the ladder's argmax is the TOP rung in all 84 cells where it was swept: 84
|
|
1561
|
+
> ends, 0 interior optima, 0 unresolved, and 84 with no significant step
|
|
1562
|
+
> downward at any rung along the way. The distinct raw best-response count is
|
|
1563
|
+
> 2, and the three cells where `finisher` edges ahead on the raw number are all
|
|
1564
|
+
> `def-532` in Knockout, by 0.1 to 0.2 points of share against a 2.74-point
|
|
1565
|
+
> threshold. Every cell of the second focal squad's grid reads C/F, which is the
|
|
1566
|
+
> permutation control showing through rather than a conditional answer:
|
|
1567
|
+
> `role-433`'s slot 8 and slot 10 are the same player twice, differing only in
|
|
1568
|
+
> the penalty-kicker flag. So one sheet is a best response against every
|
|
1569
|
+
> squad-and-sheet combination measured, and concentrating harder is better in
|
|
1570
|
+
> every one of them.
|
|
1571
|
+
> **Source** — `appr-role.mts`, section (6).
|
|
1572
|
+
|
|
1573
|
+
### The other half of the table — defensive selection weights
|
|
1574
|
+
|
|
1575
|
+
`D_*_APPR` picks who contests each scene, and the same multiplier applies there.
|
|
1576
|
+
It is a comparably large lever and it points somewhere surprising.
|
|
1577
|
+
|
|
1578
|
+
| Defensive focal slot | ×1 | ×1024 | League | Knockout |
|
|
1579
|
+
| --- | --- | --- | --- | --- |
|
|
1580
|
+
| `stars-433` slot 8 — FW, defense 5, **29 pts** | 37.2 | **52.0** | **+14.8pp** | **+18.7pp** |
|
|
1581
|
+
| `stars-433` slot 1 — DF, defense 5, 12 pts | 37.2 | 37.8 | +0.6pp | +0.9pp |
|
|
1582
|
+
| `def-433` slot 2 — DF, **defense 9**, 20 pts | 52.4 | 57.1 | +4.7pp | +5.3pp |
|
|
1583
|
+
| `def-433` slot 8 — FW, defense 2, 19 pts | 52.4 | 17.7 | −34.7pp | −35.5pp |
|
|
1584
|
+
|
|
1585
|
+
> **Measured on** — `stars-433` and `def-433`, each against `role-433` playing
|
|
1586
|
+
> the shipped weights. The multiplier is applied to the DEFENSIVE table in every
|
|
1587
|
+
> scene; the keeper's A0 weight is the one entry left alone, and the offensive
|
|
1588
|
+
> table is untouched in all four rows. Both focal squads keep their 212 points
|
|
1589
|
+
> and their formations across every cell.
|
|
1590
|
+
> **Sample** — 10,000 matches per cell, home and away, SHA-256 seeds.
|
|
1591
|
+
> **Mode** — both, one column each. As in the concentration table, the ×1 and
|
|
1592
|
+
> ×1024 columns are League shares and the League/Knockout columns are the paired
|
|
1593
|
+
> delta from that mode's own ×1 cell.
|
|
1594
|
+
> **Effect** — 14.8 percentage points of share in League and 18.7 in Knockout
|
|
1595
|
+
> for funnelling defensive duty onto `stars-433`'s biggest FORWARD, against
|
|
1596
|
+
> +0.6pp for its nominal defender. Every defensive point is
|
|
1597
|
+
> `defense + cond/2 + total/3`, so a 29-point forward with defense 5 outscores a
|
|
1598
|
+
> 12-point defender with defense 5 by more than five points before any group
|
|
1599
|
+
> term. On `def-433`, whose totals are flat, the answer reverts to the defender
|
|
1600
|
+
> with the highest `defense` (+4.7pp) and the forward becomes catastrophic
|
|
1601
|
+
> (−34.7pp). **0 of the 8 defensive ladders have a significant interior
|
|
1602
|
+
> optimum; 7 peak at an end and 1 peaks inside without separating from a
|
|
1603
|
+
> neighbour.**
|
|
1604
|
+
> **Source** — `appr-role.mts`, section (7).
|
|
1605
|
+
|
|
1606
|
+
### Degrees of freedom this probe LEFT free
|
|
1607
|
+
|
|
1608
|
+
The keeper section below was wrong three times in one day and the cause was a
|
|
1609
|
+
different uncontrolled degree of freedom each time. So, explicitly, and knowing
|
|
1610
|
+
that the verdict under this list is a NEGATIVE one — the things left free are
|
|
1611
|
+
the ways it could be wrong:
|
|
1612
|
+
|
|
1613
|
+
- **`cond`.** Fixed at 5 for all 22 players in every cell, as production does
|
|
1614
|
+
today. **This is the load-bearing one.** Concentration here has no cost
|
|
1615
|
+
because involvement has no cost; an involvement-driven form system is exactly
|
|
1616
|
+
the change that would put a price on the ×1024 column, and it would invalidate
|
|
1617
|
+
every row above rather than shift it.
|
|
1618
|
+
- **Growth buffs.** Every squad is tenure 0, so `teamPow` is exactly 100.
|
|
1619
|
+
Production raises effective `total`, which scales BOTH the individual `total/3`
|
|
1620
|
+
and the group term the anti-concentration section measures — so the 5.4pp that
|
|
1621
|
+
section reports is a tenure-0 figure.
|
|
1622
|
+
- **Manager response.** The opponent grid is fixed and nobody re-solves a squad
|
|
1623
|
+
against a sheet. A squad built KNOWING it will funnel one player is a
|
|
1624
|
+
different 212-point problem, and it is not measured here: the co-optimisation
|
|
1625
|
+
of points and weights would only make the effect larger, which is not the
|
|
1626
|
+
direction that would rescue the axis.
|
|
1627
|
+
- **The sheet space.** Seven hand-built sheets and one nine-rung ladder per
|
|
1628
|
+
slot. That is a sample of a space with eleven slots × six scenes of free
|
|
1629
|
+
parameters, not an argmax over it. "No cycle among these seven" and "one best
|
|
1630
|
+
response over the crossed grid" are both claims about these seven — a
|
|
1631
|
+
conditional best response living outside them would read as an answer here.
|
|
1632
|
+
- **The focal squads.** Two, and only one of them (`stars-433`) has players who
|
|
1633
|
+
differ inside a position group at all, so the within-group half of this axis
|
|
1634
|
+
is measured on one roster. The crossed grid runs `role-433` as a second focal
|
|
1635
|
+
squad, but over a reduced opponent-sheet set and with a formation and a set of
|
|
1636
|
+
kicker flags it shares with the first. **Which** player to funnel is already
|
|
1637
|
+
known to be squad-conditional — the defensive table above reverses its answer
|
|
1638
|
+
between `stars-433` and `def-433` — and nothing here says the best focal SLOT
|
|
1639
|
+
is the same on a roster neither squad resembles. What the grid does support is
|
|
1640
|
+
narrower and is what the verdict uses: whether to concentrate at all does not
|
|
1641
|
+
depend on the opponent.
|
|
1642
|
+
- **Set-piece slots.** `squad-lib` pins the FK taker to slot 9 and the penalty
|
|
1643
|
+
taker to slot 10. Slot 10 appears above as a focal player, so its rows carry
|
|
1644
|
+
kicker duty as well as its weight — slot 8, the headline row, carries neither.
|
|
1645
|
+
- **Formation.** Every focal squad is a 4-3-3. The lone-OMF result is a fact
|
|
1646
|
+
about a group of one, and a 4-4-2 would not reproduce it.
|
|
1647
|
+
- **Interactions.** One sheet at a time against the default. Nothing here says
|
|
1648
|
+
what a role sheet does while an engine constant is also moved.
|
|
1649
|
+
- **The defensive keeper slot.** Section (7) leaves the keeper's A0 weight at
|
|
1650
|
+
100 in every cell. Whether a manager should be able to keep his keeper OUT of
|
|
1651
|
+
the 1v1 is a question this probe does not ask.
|
|
1652
|
+
|
|
1653
|
+
### The verdict, and what it means for tactical sheets
|
|
1654
|
+
|
|
1655
|
+
**Detectable: overwhelmingly.** Up to 35.8 points of win-equivalent share on one
|
|
1656
|
+
squad whose 212 points never change, and 76.0% head to head against the same
|
|
1657
|
+
eleven players on the default sheet. Every other section of this file measures a
|
|
1658
|
+
lever by changing the players; this one changes nothing except which of them the
|
|
1659
|
+
ball reaches, and moves more than any of them. The engine-constant sweep — the
|
|
1660
|
+
only other section that holds a squad fixed — moved 11.8pp at its widest and
|
|
1661
|
+
re-priced nothing.
|
|
1662
|
+
|
|
1663
|
+
**A decision: no.** Every diagnostic points the same way. 0 of 50 ladders in the
|
|
1664
|
+
one-opponent sections have a significant interior optimum — 48 peak at an end,
|
|
1665
|
+
the remaining 2 peak inside without separating from a neighbour — 84 of 84
|
|
1666
|
+
ladders swept INSIDE a crossed cell peak at the top rung with nothing
|
|
1667
|
+
unresolved, and all 134 ladders in the run are free of any significant step
|
|
1668
|
+
against their own direction. The anti-concentration term is real but worth about a seventh of the
|
|
1669
|
+
payoff, and it changes the slope without changing the answer. The round robin
|
|
1670
|
+
finds zero 3-cycles in either mode and an unbeaten set of two sheets 1.4 points
|
|
1671
|
+
of share apart. And the crossed grid — every opponent squad against every
|
|
1672
|
+
opponent sheet, on two focal squads, 114 cells — never once excludes `carrier`
|
|
1673
|
+
from the best-response band. **Concentration is monotone: more is better, up to
|
|
1674
|
+
the corner, against every squad and every sheet measured.**
|
|
1675
|
+
|
|
1676
|
+
So the axis is not a decision *yet*, and the word doing the work is `cond`. The
|
|
1677
|
+
reason concentration is free is that involvement costs nothing — the eleventh
|
|
1678
|
+
selection of a player is exactly as good as his first, because `cond` is a
|
|
1679
|
+
constant. That is what a form system changes, and it is the one change that
|
|
1680
|
+
could put an interior optimum on these ladders. **Roles before form would ship
|
|
1681
|
+
the largest monotone answer in the game.**
|
|
1682
|
+
|
|
1683
|
+
Two things in the tables above survive that verdict and are worth keeping:
|
|
1684
|
+
|
|
1685
|
+
- **Scene shape is where the interesting half lives.** The same multiplier on
|
|
1686
|
+
the same player is worth +4.7pp or −6.0pp depending only on which scenes it
|
|
1687
|
+
applies to, and the best focal player is the best PASSER rather than the best
|
|
1688
|
+
finisher. Whatever a sheet eventually exposes, a single "involvement" slider
|
|
1689
|
+
would throw both away.
|
|
1690
|
+
- **The defensive table is a second lever of the same size and nobody has looked
|
|
1691
|
+
at it.** On a squad with uneven totals the best defender to funnel duty to is
|
|
1692
|
+
the biggest FORWARD, because `total/3` is in every defensive point. That is
|
|
1693
|
+
either a feature to expose or a formula to reconsider, and it is decided by
|
|
1694
|
+
the same track.
|
|
1695
|
+
|
|
1696
|
+
Re-run after any engine change:
|
|
1697
|
+
|
|
1698
|
+
```bash
|
|
1699
|
+
npx tsx packages/mcp/skill/reference/probes/appr-role.mts # ~21M matches, ~4 min
|
|
1700
|
+
N=2000 npx tsx … # ~50s; the bar is fixed, the detectable EFFECT moves
|
|
1701
|
+
```
|
|
1702
|
+
|
|
1703
|
+
---
|
|
1704
|
+
|
|
1231
1705
|
## Note for the maintainers
|
|
1232
1706
|
|
|
1233
1707
|
> **2026-08-19: this note was rewritten twice in one day.** First it said the
|
|
@@ -1317,3 +1791,1703 @@ If a trigger fires, re-run the probes above BEFORE touching anything, then
|
|
|
1317
1791
|
decide between the options in #604. Whatever is chosen, the engine change and
|
|
1318
1792
|
both documents move in the SAME change — `SKILL.md` is what agents believe, and
|
|
1319
1793
|
a stale sentence there makes them exclude squads that are perfectly legal.
|
|
1794
|
+
|
|
1795
|
+
|
|
1796
|
+
|
|
1797
|
+
|
|
1798
|
+
## How much depth does rotation need, and what does depth buy? (`depth-cap.mts`)
|
|
1799
|
+
|
|
1800
|
+
Squad order is a lever nobody priced. That is the headline; the depth numbers are below it.
|
|
1801
|
+
|
|
1802
|
+
Reordering the SAME eleven — no reserves, no purchase — is worth **+5.7pp** on a squad
|
|
1803
|
+
whose outfield players differ, and nothing (+0.1…+0.9pp) on three squads whose outfield
|
|
1804
|
+
weights are near-uniform. The array is sorted by `slotIndex`, a required field on the team
|
|
1805
|
+
payload, so the manager sets it. It is **not** an API-only lever: dropping one card onto
|
|
1806
|
+
another is classified as a `swap` in `FormationPitch`, which calls `handleDrop` →
|
|
1807
|
+
`swapSlots(prev, df, i)`, and submission serializes that array straight back as
|
|
1808
|
+
`slotIndex: i`. **Create-team only, though**: the owner edit view renders the pitch
|
|
1809
|
+
`readOnly` (`interactive = !readOnly`) and its own file header says it has no save button,
|
|
1810
|
+
so the drop handler wired there never fires. One drag reaches this at team creation; an
|
|
1811
|
+
existing team cannot be reordered that way.
|
|
1812
|
+
|
|
1813
|
+
**The mechanism is unidentified, and both candidates are back open.**
|
|
1814
|
+
|
|
1815
|
+
| candidate | test | result |
|
|
1816
|
+
| --- | --- | --- |
|
|
1817
|
+
| home-DMF `this.kiten` bug | split ONE reordering by side | both sides large — **not rejected** |
|
|
1818
|
+
| zone-press fallback | clone ablation | tight interval, but at fixed POSITION — **withdrawn** |
|
|
1819
|
+
|
|
1820
|
+
Seeing both sides move does not clear the home-DMF path. It shows the effect is not
|
|
1821
|
+
home-EXCLUSIVE, which is compatible with the defect contributing to the home number while
|
|
1822
|
+
something else produces the away one — and permuting changes the whole slot order, leaving
|
|
1823
|
+
every index-sensitive path live. Rejecting it needs the defect disabled and the same
|
|
1824
|
+
permutations rerun, which is an engine change and out of scope here.
|
|
1825
|
+
|
|
1826
|
+
The zone-press story looked settled and was not. Every defensive appearance weight at
|
|
1827
|
+
`SCENE.ZP` is `0.0`, so `weightedPick`'s empty-index fallback fires on every zone press and
|
|
1828
|
+
hands the job to `players[first non-GK]` — all true, and none of it is the effect. The
|
|
1829
|
+
ablation: make the four DFs **three identical clones plus one marked player** of the same
|
|
1830
|
+
total but a different split. Position and weight sequences are then identical whatever the
|
|
1831
|
+
arrangement, so every non-ZP draw selects the same way, while ZP deterministically takes
|
|
1832
|
+
the first slot. Moving the marked player in and out of that slot is worth −0.1…+0.3pp, and the Bonferroni-corrected simultaneous interval across
|
|
1833
|
+
the three placements is **[−1.8, +2.0]pp** — inside ±5pp on both sides. A near-zero point
|
|
1834
|
+
estimate is not a rejection on its own, and neither is an upper bound alone: three
|
|
1835
|
+
placements all at −10pp would still have an upper bound under 5pp. Equivalence needs the
|
|
1836
|
+
whole interval.
|
|
1837
|
+
|
|
1838
|
+
The first version of that ablation built the probe squad uniform at 19 apiece — the same
|
|
1839
|
+
trap this note accuses §0.3 of, reproduced in the experimental design. A uniform squad is
|
|
1840
|
+
where ordering effects vanish by construction, so it could not have distinguished anything.
|
|
1841
|
+
Rebuilt heterogeneous (totals 16–26), with clones only inside the DF group.
|
|
1842
|
+
|
|
1843
|
+
What IS established is narrower: the effect appears only when the **position sequence**
|
|
1844
|
+
changes. Swapping two same-position players moves nothing (−0.2…+0.2pp); swapping across
|
|
1845
|
+
positions produces the whole effect (−1.5…**+5.6pp**). So what matters is not which player
|
|
1846
|
+
sits at an index but which position does — which points at the arrangement of unequal
|
|
1847
|
+
weights in `weightedPick`'s cumulative scan, and no further.
|
|
1848
|
+
|
|
1849
|
+
Two lessons are worth carrying out of this. **An index-dependent path is invisible in a
|
|
1850
|
+
homogeneous squad** — which is why `01-engine-expansion.md` §0.3 demoted this lever as a
|
|
1851
|
+
"weak-seed-era ghost" after 30,000 matches across ONE squad family. And **a clean interval
|
|
1852
|
+
proves nothing if the dimension you manipulated is not the one that moves**: the clone
|
|
1853
|
+
ablation held position fixed while the real permutation changes it, so its tight interval
|
|
1854
|
+
rejected zone-press-via-attributes and nothing else. That withdrawal is recorded below.
|
|
1855
|
+
|
|
1856
|
+
The search is 120 random reorderings with no dedup, not an exhaustive pass over 11!, so
|
|
1857
|
+
+5.7pp is what the search **found** — a floor, not a ceiling.
|
|
1858
|
+
|
|
1859
|
+
Every number in this section is measured against the same league basket. The first version
|
|
1860
|
+
ranked each ordering against the ORIGINAL ordering — the same objective error already fixed
|
|
1861
|
+
on the purchase axis, still living in this function. Fixing one and not the other happened
|
|
1862
|
+
twice in this probe, with common random numbers as well as with the objective.
|
|
1863
|
+
|
|
1864
|
+
### Depth — how many reserves, and what they buy
|
|
1865
|
+
|
|
1866
|
+
Pools of core-11 plus k generated reserves; every candidate eleven is CONSTRUCTED (I-a's
|
|
1867
|
+
assignment rules) and run through `validateTeam`, because sums alone do not decide
|
|
1868
|
+
legality. Reserves come from the production generator (`progressive-free-agent.ts`, sums
|
|
1869
|
+
uniform on [16,28]) and are **unminted** — §3-13 makes mutability a function of
|
|
1870
|
+
`innateFixed`, and freezing name/position rejects combinations that are legal.
|
|
1871
|
+
|
|
1872
|
+
Four cores, because one is not enough to see the shape. The first version of this
|
|
1873
|
+
measurement used `EVEN/FLAT` alone and concluded "one sum-matched reserve is enough for
|
|
1874
|
+
everybody". That squad is the WORST template by `band-worstcase`, and it sits with both
|
|
1875
|
+
scarcity caps empty (0 tens and 0 eights against a 3/5 budget), so every reserve fits.
|
|
1876
|
+
Cores that sit AT the caps behave nothing like it.
|
|
1877
|
+
|
|
1878
|
+
Two walls decide whether a reserve can ever enter:
|
|
1879
|
+
|
|
1880
|
+
| core | totals | outside FA's [16,28] | tens | eights |
|
|
1881
|
+
| --- | --- | --- | --- | --- |
|
|
1882
|
+
| `flat` | 19–20 | **0** | 0 | 0 |
|
|
1883
|
+
| `shaped` | 12–29 | **6 of 11** | 3 | 5 |
|
|
1884
|
+
| `gkheavy` | 18–29 | 1 | 3 | 5 |
|
|
1885
|
+
| `gkmin` | 10–21 | 1 | 2 | 5 |
|
|
1886
|
+
|
|
1887
|
+
Scarcity headroom is the first wall — a core already holding 3 tens and 5 eights blocks
|
|
1888
|
+
any reserve carrying one, whatever its sum. The reachable-sum range is the second: the
|
|
1889
|
+
generator cannot produce a sum outside [16,28], so six of `shaped`'s eleven players have
|
|
1890
|
+
no matchable reserve that could exist.
|
|
1891
|
+
|
|
1892
|
+
Reserves needed for a legal alternative eleven in 90% of pools (Wilson 95% lower bound,
|
|
1893
|
+
300 pools per N — a point estimate moves the threshold between runs):
|
|
1894
|
+
|
|
1895
|
+
| core | random FA sums | sum-matched |
|
|
1896
|
+
| --- | --- | --- |
|
|
1897
|
+
| `flat` | 7 | **1** |
|
|
1898
|
+
| `shaped` | 6 | **5** |
|
|
1899
|
+
| `gkheavy` | >7 | **2** |
|
|
1900
|
+
| `gkmin` | 6 | **1** |
|
|
1901
|
+
|
|
1902
|
+
Sum-matching still beats random sums everywhere. What changed is that the count it buys
|
|
1903
|
+
runs from 1 to 5 depending on the core.
|
|
1904
|
+
|
|
1905
|
+
### What this axis does NOT measure
|
|
1906
|
+
|
|
1907
|
+
Every match here runs at `cond: 5` with no workload carried between fixtures. So this
|
|
1908
|
+
measures **how much a new attribute vector improves one fresh XI** — market quality — and
|
|
1909
|
+
not **what depth buys by resting tired players**, which is the mechanism H exists to cap.
|
|
1910
|
+
Simulating that needs a recovery rate and a neutral point, and those are C-2's to decide.
|
|
1911
|
+
**H's maximum stays open.** What follows is a market-quality result, and it is worth having
|
|
1912
|
+
on its own.
|
|
1913
|
+
|
|
1914
|
+
**Two errors point opposite ways, so this is not an upper bound.** The buyer is fully
|
|
1915
|
+
informed — it sees all 30 listings' vectors, while production hides attributes and even
|
|
1916
|
+
`total` until discovery stage 4 (`player-visibility-policy.ts`) and charges globally visible
|
|
1917
|
+
trials to reveal one, with no purchase reservation. That pushes the estimate up. But the
|
|
1918
|
+
buyer also does not search the attainable optimum: enumeration stops at `ENUM_CAP`, at most
|
|
1919
|
+
`ALT_CAP` subsets and `CAND_BUDGET` orderings survive, and the published run hit the cap 423
|
|
1920
|
+
times. That pushes it down. The number is conditional on both, not a bound in either
|
|
1921
|
+
direction [Codex P2, 2026-08-31].
|
|
1922
|
+
|
|
1923
|
+
### The grant bounds the market — rotation quality does not follow it
|
|
1924
|
+
|
|
1925
|
+
Two corrections turned this section around. The grant sizes were hardcoded
|
|
1926
|
+
(`shaped: 5, gkheavy: 2`) while the proposed rule says `headroom < 4 → 6` — the document
|
|
1927
|
+
was recommending one policy and measuring another. And the opponent basket held four bare,
|
|
1928
|
+
default-order cores, which are far weaker than the post-grant baselines they were meant to
|
|
1929
|
+
represent. Grants now come from the classifier, and every basket opponent gets its own
|
|
1930
|
+
grant and optimizes its ordering.
|
|
1931
|
+
|
|
1932
|
+
30 pools, selection on 30 seeds, validation 1500, reporting 1500 (untouched). `buy 0` does
|
|
1933
|
+
not search, so it is structurally 0.0.
|
|
1934
|
+
|
|
1935
|
+
| core | headroom | free | baseline vs bare core | rotation quality (within 5pp, lower) | buy 0 | 1 | 2 | 3 |
|
|
1936
|
+
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
|
|
1937
|
+
| `flat` | 8 | 1 | +3.5 | 87% (70%) | 0.0 | **+17.5** | **+23.0** | **+23.9** |
|
|
1938
|
+
| `shaped` | 0 | 6 | +27.5 | 77% (59%) | 0.0 | **+6.5** | **+11.1** | **+12.8** |
|
|
1939
|
+
| `gkheavy` | 0 | 6 | +7.7 | 57% (39%) | 0.0 | +0.8 | +2.2 | +3.6 |
|
|
1940
|
+
| `gkmin` | 1 | 6 | +10.8 | 80% (63%) | 0.0 | 0.0 | +3.1 | +3.8 |
|
|
1941
|
+
|
|
1942
|
+
### Vary the grant, hold everything else — the controlled version
|
|
1943
|
+
|
|
1944
|
+
An earlier version read the policy table sideways and claimed two things: that the grant
|
|
1945
|
+
size bounds the purchase advantage (`gkmin` on six buys nothing, `flat` on one pays
|
|
1946
|
+
+14.6pp), and that a bigger grant costs rotation quality (`gkheavy` 87% → 57%). **In both
|
|
1947
|
+
comparisons something else moved with the grant** — the core in the first, the whole
|
|
1948
|
+
experimental setup in the second, since 87% came from an earlier run with a different
|
|
1949
|
+
basket and procedure. So the grant was swept with the core, basket, pools and seeds held
|
|
1950
|
+
fixed.
|
|
1951
|
+
|
|
1952
|
+
Grants are drawn **nested** — one max-length list per pool seed, sliced — so a smaller grant
|
|
1953
|
+
is a SUBSET of a larger one, and the evaluation column is shared across grant sizes. Drawing
|
|
1954
|
+
fresh reserves per grant would move the reserve identity along with the grant, which is the
|
|
1955
|
+
same control failure the sweep exists to fix.
|
|
1956
|
+
|
|
1957
|
+
| core | buy-1 @1 | @2 | @4 | @6 | quality @1 | @2 | @4 | @6 |
|
|
1958
|
+
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
|
|
1959
|
+
| `flat` | **+15.4** | +5.8 | +6.0 | **+5.8** | 90% | 70% | 67% | 67% |
|
|
1960
|
+
| `shaped` | **+17.0** | **+19.1** | **+15.6** | +5.8 | 17% | 40% | 63% | **73%** |
|
|
1961
|
+
| `gkheavy` | +4.6 | +5.2 | +0.0 | +0.0 | 63% | 83% | 80% | 57% |
|
|
1962
|
+
| `gkmin` | **+9.4** | +2.5 | +0.0 | +0.0 | 73% | 90% | 77% | 83% |
|
|
1963
|
+
|
|
1964
|
+
**The first claim is established on two cores — and not in the direction the point
|
|
1965
|
+
estimates suggested.** The table above is point estimates, and reading it across is how the
|
|
1966
|
+
earlier version reached "it falls on all four, past the gate on three". That was said without
|
|
1967
|
+
a test: the second claim had a paired one attached and the first did not, though the nested
|
|
1968
|
+
pools make the tool obvious — subtract the two grants' purchase advantage within each pool.
|
|
1969
|
+
|
|
1970
|
+
| core | paired mean Δ (6 − 1) | simultaneous interval | verdict |
|
|
1971
|
+
| --- | --- | --- | --- |
|
|
1972
|
+
| `flat` | **−8.7pp** | [−14.1, −3.3] | **reduces it** |
|
|
1973
|
+
| `gkmin` | **−6.4pp** | [−10.9, −1.9] | **reduces it** |
|
|
1974
|
+
| `shaped` | −7.2pp | [−14.9, 0.5] | undecided |
|
|
1975
|
+
| `gkheavy` | −2.5pp | [−6.5, 1.5] | undecided |
|
|
1976
|
+
|
|
1977
|
+
**The critical value is t.** The standard deviation is ESTIMATED from those same 30 paired
|
|
1978
|
+
deltas, so a normal critical value is anti-conservative at that sample size — and `shaped`'s
|
|
1979
|
+
upper bound sat at a rounded −0.0, meaning the verdict turned on the choice. The
|
|
1980
|
+
Bonferroni-corrected t with 29 degrees of freedom (2.663 against the normal 2.498) carries
|
|
1981
|
+
that core back across zero: **two cores established, not three**. The quantile comes from a
|
|
1982
|
+
Cornish-Fisher expansion, and the probe checks it against three published t-table entries at
|
|
1983
|
+
startup and exits on a mismatch — a wrong approximation would give every interval below a
|
|
1984
|
+
quietly wrong width.
|
|
1985
|
+
|
|
1986
|
+
`flat` is in that list — the core the earlier version singled out as the one that "responds
|
|
1987
|
+
least" because of its headroom of 8. Paired, its reduction is −8.7pp, the LARGEST of the
|
|
1988
|
+
four. **The endpoints hide it entirely**: `flat` and `shaped` both land on +5.8, which says
|
|
1989
|
+
nothing, because they start from 15.4 and 17.0.
|
|
1990
|
+
|
|
1991
|
+
**The two results do not stand on the same cores.** `shaped` is where the rotation-quality
|
|
1992
|
+
claim is established (p<0.001 below) and where the purchase-advantage reduction is not;
|
|
1993
|
+
`flat` and `gkmin` are the reverse. That split is why the two claims have to be carried
|
|
1994
|
+
separately rather than as one "grants fix the market" sentence. What bounds the market is not "how much you may buy" but how much you are
|
|
1995
|
+
given.
|
|
1996
|
+
|
|
1997
|
+
**But "it falls below the gate" is withdrawn.** A median under 5 is not a gate verdict —
|
|
1998
|
+
the verdict this probe uses everywhere else is the fraction of pools above 5pp with a Wilson
|
|
1999
|
+
upper bound, and applied here almost every bound covers 50%: `flat` 93/81/76/74%, `shaped`
|
|
2000
|
+
88/88/92/76%, `gkheavy` 74/74/**32**/54%, `gkmin` 83/66/54/**43**% across grants
|
|
2001
|
+
1/2/4/6 (family = grants × cores = 16). At six reserves the only core established below the
|
|
2002
|
+
gate is `gkmin`. Reducing the advantage and pushing it under the gate are two claims, and
|
|
2003
|
+
the earlier version put them in one sentence.
|
|
2004
|
+
|
|
2005
|
+
Which cores respond changed when the sampling keys were fixed — hashing the ENUMERATION
|
|
2006
|
+
INDEX meant each grant size searched a different slice, blurring exactly this.
|
|
2007
|
+
|
|
2008
|
+
**The second holds on one core only, and every other direction has now evaporated.** The point estimates
|
|
2009
|
+
look erratic, and two versions of this document read them opposite ways without a test. The
|
|
2010
|
+
pools are nested, so the same pools' success indicators pair across grant sizes (`b` =
|
|
2011
|
+
succeeds only at the small grant, `c` = only at the large one), tested exactly:
|
|
2012
|
+
|
|
2013
|
+
| core | b | c | Δ (6 − 1) | corrected p | verdict |
|
|
2014
|
+
| --- | --- | --- | --- | --- | --- |
|
|
2015
|
+
| `shaped` | 1 | 18 | **+57%p** | **<0.001** | **bigger grant is better** |
|
|
2016
|
+
| `flat` | 10 | 3 | −23%p | 0.37 | undecided |
|
|
2017
|
+
| `gkheavy` | 8 | 6 | −7%p | 1.00 | undecided |
|
|
2018
|
+
| `gkmin` | 3 | 6 | +10%p | 1.00 | undecided |
|
|
2019
|
+
|
|
2020
|
+
The test is exact McNemar, Bonferroni-corrected across cores. `flat` looked significant in
|
|
2021
|
+
the opposite direction until two things were fixed: a normal approximation excluded zero by
|
|
2022
|
+
a hair on ten discordant pairs, and the sampling keys moved with the grant. Exact test plus
|
|
2023
|
+
stable keys, and it is gone (b=10, c=3, p=0.37). **Adopting a test is not the same as
|
|
2024
|
+
checking it fits the data** — the lesson that point estimates cannot carry a direction had
|
|
2025
|
+
just been learned, and the instrument brought in to fix it went unexamined.
|
|
2026
|
+
|
|
2027
|
+
`shaped`'s result matches the mechanism cleanly. With zero headroom, a small grant leaves
|
|
2028
|
+
its alternatives threadbare (quality 17%), so more grant is what makes rotation usable — 18
|
|
2029
|
+
of 30 pools fail at one reserve and succeed at six, and one goes the other way.
|
|
2030
|
+
|
|
2031
|
+
So: **for cores without headroom, a larger grant genuinely makes rotation usable, and
|
|
2032
|
+
nothing in this data shows it costing anything anywhere.** The "fairness versus experience"
|
|
2033
|
+
trade-off two earlier versions of this note described, in opposite directions, is not
|
|
2034
|
+
there.
|
|
2035
|
+
|
|
2036
|
+
Sampling limit, stated: the enumeration cap (4000) was hit **423 times**, so those calls
|
|
2037
|
+
sampled a shuffled DFS prefix rather than the full legal set. **That is the aggregate for
|
|
2038
|
+
the whole strength invocation** and cannot be attributed to any one grant size: a single
|
|
2039
|
+
counter collects basket construction, every grant's baseline and buy-1 search, and the later
|
|
2040
|
+
K sweep. Writing "with six-reserve pools" charged the whole bucket to one experiment
|
|
2041
|
+
[Codex P2, 2026-08-31].
|
|
2042
|
+
|
|
2043
|
+
### The ordering lever against policy-realistic opponents
|
|
2044
|
+
|
|
2045
|
+
| core | best of 120 reorderings | paired upper bound | vs the gate | ZP: best → first slot restored |
|
|
2046
|
+
| --- | --- | --- | --- | --- |
|
|
2047
|
+
| `flat` | −0.7 | 0.9 | below | lower bound −2.4 — undecidable |
|
|
2048
|
+
| `shaped` | **+6.5** | **8.7** | **undecided** | **6.5 → 0.9** (the gain lives in that slot) |
|
|
2049
|
+
| `gkheavy` | −0.4 | 1.6 | below | lower bound −2.3 — undecidable |
|
|
2050
|
+
| `gkmin` | +1.8 | 3.9 | below | lower bound −0.3 — undecidable |
|
|
2051
|
+
|
|
2052
|
+
**That upper bound is now measured; for one version it was assumed.** The standard error was
|
|
2053
|
+
hard-coded as `sqrt(2*0.25/n)`, and three cores were classified below the gate on that
|
|
2054
|
+
assumption. It is wrong twice over: candidate and baseline share each seed, so the
|
|
2055
|
+
observations are PAIRS rather than two independent samples, and `n` counted seeds × 2 for
|
|
2056
|
+
home/away when those two share a seed as well — that correlation is exactly why a squad
|
|
2057
|
+
against itself scores 0.00. Computing the interval from per-seed paired differences instead
|
|
2058
|
+
(weighting ladder and knockout within a seed, since they share it too) narrowed the bounds
|
|
2059
|
+
against the assumed version of the same run. The variance pairing removes exceeded the
|
|
2060
|
+
inflation from double-counting sides, so the old assumption erred toward permissive. The
|
|
2061
|
+
four verdicts are unchanged, but observation carries them now, and the point estimates do
|
|
2062
|
+
not move at all — same seeds, same matches.
|
|
2063
|
+
|
|
2064
|
+
**"It shrinks against policy opponents" is withdrawn.** That read set an EARLIER run's
|
|
2065
|
+
bare-opponent +6.4pp beside this run's policy figure, and the earlier run selected its basket
|
|
2066
|
+
on a ladder-only objective — the two are not comparable. Scoring the same permutation against
|
|
2067
|
+
the four bare cores inside THIS run gives `flat` −2.5, `shaped` +4.9, `gkheavy` −0.1, `gkmin`
|
|
2068
|
+
+1.2, against policy figures of −0.7 / +6.5 / −0.4 / +1.8. The direction is the opposite of
|
|
2069
|
+
what was written.
|
|
2070
|
+
|
|
2071
|
+
It does not license the reverse claim either: the permutation was CHOSEN against the policy
|
|
2072
|
+
basket, so doing better there is expected, and the bare column is an out-of-objective
|
|
2073
|
+
evaluation. What stands is that moving to policy opponents did not shrink the lever.
|
|
2074
|
+
|
|
2075
|
+
Against opponents who also hold their
|
|
2076
|
+
grant and optimize their order, `shaped` reorders for **+6.5pp** — and a point estimate near
|
|
2077
|
+
the gate settles nothing in either direction, so with a paired upper bound of 8.7 that core
|
|
2078
|
+
is **undecided**, not below. The other three are below it (upper bounds 0.9–3.9), and two of
|
|
2079
|
+
them have negative point estimates, so reordering is not a gain there at all.
|
|
2080
|
+
|
|
2081
|
+
What survives is sharper, and now it is controlled. `shaped`'s gain is **localized to the
|
|
2082
|
+
first non-GK slot**:
|
|
2083
|
+
|
|
2084
|
+
| manipulation | `shaped` |
|
|
2085
|
+
| --- | --- |
|
|
2086
|
+
| best reordering, untouched | **+6.5pp** |
|
|
2087
|
+
| restore the **first non-GK slot**'s occupant to the baseline's | **+0.9pp** (nine-tenths gone) |
|
|
2088
|
+
| leave the first slot alone, swap the other slot with a SAME-POSITION one | **+7.5pp** (intact) |
|
|
2089
|
+
|
|
2090
|
+
All three come from **one run on one seed column** — mixing in a value from a small side
|
|
2091
|
+
run is how the same quantity ends up with two numbers.
|
|
2092
|
+
|
|
2093
|
+
Restoring moves two slots, so on its own it cannot say which one mattered — the second
|
|
2094
|
+
slot or the interaction would do just as well. The paired control settles it: leave the
|
|
2095
|
+
first slot alone and the gain survives whole.
|
|
2096
|
+
|
|
2097
|
+
**The control has to be position-matched.** A first version swapped in an arbitrary non-GK
|
|
2098
|
+
slot, which lets the second slot receive a player of a different position than the
|
|
2099
|
+
restoration puts there — two different lineups, so a surviving gain could come from that
|
|
2100
|
+
choice rather than from the first slot. `getPoint` branches on position, so position is the
|
|
2101
|
+
dimension to match; with no same-position partner the localization is left open. In
|
|
2102
|
+
`shaped` the restoration puts a FORWARD in that slot, so the control uses a forward slot,
|
|
2103
|
+
and the gain holds at +7.5pp. **The localization is decided on one core only**: on the other
|
|
2104
|
+
three the ordering gain's own paired lower bound is below zero (−2.4, −2.3, −0.3), so asking
|
|
2105
|
+
where it lives is not a question. That door was held by a magic number (`|gain| < 1.5`) even
|
|
2106
|
+
after the measured interval existed, and it let `gkmin` through to print "ZP is likely" on a
|
|
2107
|
+
1.8pp gain with a 3.9 upper bound. The zone-press fallback selects exactly that
|
|
2108
|
+
slot, deterministically. Confirming it still needs the path disabled in the engine, but its
|
|
2109
|
+
standing as the leading candidate is now earned rather than asserted.
|
|
2110
|
+
|
|
2111
|
+
**Both earlier rejections are withdrawn.** The home-DMF split showed only that the effect is
|
|
2112
|
+
not home-EXCLUSIVE. And the clone ablation held POSITION fixed while varying attributes —
|
|
2113
|
+
but the real permutation changes the position of that slot, and `getPoint` branches on
|
|
2114
|
+
position into different rate tables. **A clean interval proves nothing if the dimension you
|
|
2115
|
+
manipulated is not the dimension that moves.** Settling it needs the path disabled in the
|
|
2116
|
+
engine and the same permutations replayed.
|
|
2117
|
+
|
|
2118
|
+
### Which grant does an arbitrary core need? Headroom, and reachable sums
|
|
2119
|
+
|
|
2120
|
+
Four synthetic cores cannot support "the core's state sets the grant" — nothing maps an
|
|
2121
|
+
arbitrary legal core to a number. 30 random legal cores (4 formations × 6 mixes × random
|
|
2122
|
+
weights, keeping only those whose eleven is the sole legal XI), 600 pools per N, threshold
|
|
2123
|
+
Bonferroni-corrected across `cores × N` = 210:
|
|
2124
|
+
|
|
2125
|
+
| scarcity headroom = (3 − tens) + (5 − eights) | cores | grant needed (min · median · max) |
|
|
2126
|
+
| --- | --- | --- |
|
|
2127
|
+
| **0** (caps saturated) | 15 / 30 | 3 · 4 · **6** |
|
|
2128
|
+
| 1 | 3 / 30 | 3 · 4 · 5 |
|
|
2129
|
+
| 2 | 3 / 30 | 2 · 2 · 3 |
|
|
2130
|
+
| 3 | 1 / 30 | 4 · 4 · 4 |
|
|
2131
|
+
| 4 · 5 | 3 / 30 | 1 · 1 · 1 |
|
|
2132
|
+
| 7 · 8 | 5 / 30 | 1 · 1 · 1 |
|
|
2133
|
+
|
|
2134
|
+
All 30 reach the 90% threshold by N ≤ 18, and the need falls as headroom rises — but the
|
|
2135
|
+
rule is the CONJUNCTION below, not headroom alone: quoting headroom by itself lets back in
|
|
2136
|
+
the very no-rotation case the reachability check exists to prevent.
|
|
2137
|
+
|
|
2138
|
+
The rule is `headroom ≥ 4 AND at least two totals inside [16,28] → 1, else 6`. Headroom
|
|
2139
|
+
alone is not enough, because it ignores the second wall this same measurement found: a legal
|
|
2140
|
+
core of four 29s, five 14s and two 13s has headroom 4 but **no** total the FA generator can
|
|
2141
|
+
reach, so no matched reserve exists and one player cannot preserve 212. Writing a finding
|
|
2142
|
+
down is not the same as putting it in the rule. On this sample (reachable counts 2–11) the
|
|
2143
|
+
rule under-grants **0 of 30** cores; reachable counts of 0–1 are unmeasured and fall to the
|
|
2144
|
+
larger grant.
|
|
2145
|
+
|
|
2146
|
+
**The boundary is not pinned**: headroom 1, 2 and 3 have 3, 3 and 1 cores, and headroom 3
|
|
2147
|
+
(n=1) comes out above headroom 2, so it is not monotone there.
|
|
2148
|
+
|
|
2149
|
+
**Half of the sampled cores sit at the caps — inside this sampler.** 15 of the 30 have zero
|
|
2150
|
+
headroom, and that frequency must not be read as the league's. The sampler is 4 formations
|
|
2151
|
+
crossed with SIX HAND-AUTHORED mix templates and random weights, then deterministically
|
|
2152
|
+
repaired until the scarcity caps hold: it excludes most otherwise-legal attribute layouts and
|
|
2153
|
+
is structurally prone to manufacturing saturated squads. Establishing the real frequency
|
|
2154
|
+
needs production team data or a generator argued to be representative, and neither is here.
|
|
2155
|
+
|
|
2156
|
+
What stands is conditional: within this sampler half the cores have zero headroom, and
|
|
2157
|
+
`flat` (headroom 8) is the uncommon shape among them. A grant policy must not be justified
|
|
2158
|
+
with "half the league looks like this" — but "cores with zero headroom exist and one reserve
|
|
2159
|
+
is not enough for them" is fully supported, and that is what the classifier rests on.
|
|
2160
|
+
|
|
2161
|
+
### Rotation quality — the other half of the minimum
|
|
2162
|
+
|
|
2163
|
+
Existence does not settle whether rotation is usable. The best alternative whose **member
|
|
2164
|
+
set actually differs** — same eleven in another order is not rotation, and membership is
|
|
2165
|
+
compared by asset id because two generated reserves can share a name — plays the champion.
|
|
2166
|
+
**The numbers live in the policy table above**, where grants come from the classifier; an
|
|
2167
|
+
earlier version of this section carried a second table built on the pre-classifier grants
|
|
2168
|
+
(5/2/1) and reported different results for the same cores.
|
|
2169
|
+
|
|
2170
|
+
The verdict is the tail, not the median. An earlier version read the medians (−0.3 to
|
|
2171
|
+
−1.7pp) and concluded that resting a player costs a point or two — while the same rows
|
|
2172
|
+
carried −19pp at the bottom, and pools where no alternative was found had been dropped from
|
|
2173
|
+
the summary entirely.
|
|
2174
|
+
|
|
2175
|
+
Dropping the failures to summarize the successes is the third instance of one error shape
|
|
2176
|
+
in this probe: rejection-sampling shortfalls, then pools with no alternative, then this.
|
|
2177
|
+
|
|
2178
|
+
**The conclusion moved five times, and each time one arm had been modelled weaker than it
|
|
2179
|
+
is** — the baseline four times (worst-template core; only alternatives allowed to explore
|
|
2180
|
+
ordering, plus dropped no-alternative pools; no free grant; single grant draw with the
|
|
2181
|
+
baseline missing from the candidate set), the treatment once (random grants standing in for
|
|
2182
|
+
a buyer who chooses). The sixth round found something worse than a weak arm: **the axis was
|
|
2183
|
+
measuring a different quantity than the one the unit asked for.** Both failures have the
|
|
2184
|
+
same check — ask what each side was not allowed to do, and ask whether the thing you varied
|
|
2185
|
+
is the thing the question is about.
|
|
2186
|
+
|
|
2187
|
+
Controls, both of which must hold before any row above is read: the core eleven alone
|
|
2188
|
+
yields exactly one legal XI (all four cores), and a squad against itself scores exactly
|
|
2189
|
+
0.00pp by home/away symmetry (all four cores).
|
|
2190
|
+
|
|
2191
|
+
```
|
|
2192
|
+
ONLY=existence POOLS=300 N_MAX=18 npx tsx packages/mcp/skill/reference/probes/depth-cap.mts
|
|
2193
|
+
ONLY=cores CORE_SAMPLES=30 POOLS=600 N_MAX=18 \
|
|
2194
|
+
npx tsx packages/mcp/skill/reference/probes/depth-cap.mts
|
|
2195
|
+
ONLY=strength STRENGTH_POOLS=30 SELECT_SEEDS=30 STRENGTH_SEEDS=1500 ALT_CAP=24 \
|
|
2196
|
+
CAND_BUDGET=48 STRENGTH_NS=0,1,2,3 PERM_TRIES=120 MARKET_SIZE=30 ENUM_CAP=4000 \
|
|
2197
|
+
ZP_SEEDS=20000 VALIDATE_SEEDS=1500 GRANT_SWEEP=1,2,4,6 \
|
|
2198
|
+
npx tsx packages/mcp/skill/reference/probes/depth-cap.mts
|
|
2199
|
+
```
|
|
2200
|
+
|
|
2201
|
+
The two axes cost two orders of magnitude apart, which is why `ONLY` exists.
|
|
2202
|
+
|
|
2203
|
+
## How big can the condition band be? (`band-worstcase.mts`)
|
|
2204
|
+
|
|
2205
|
+
§3-2 recommended "start at 0.5, never exceed 1.0" and flagged that those numbers were
|
|
2206
|
+
themselves unvalidated: they came from one squad template playing itself, with growth and
|
|
2207
|
+
role concentration excluded. This sweep runs the missing comparison, and the
|
|
2208
|
+
recommendation does not survive it.
|
|
2209
|
+
|
|
2210
|
+
### Method
|
|
2211
|
+
|
|
2212
|
+
Six legal squads, all 36 ordered matchups (mirrors included), four growth combinations
|
|
2213
|
+
(0/0, max/max, 0/max, max/0), a Δ grid, and both modes. The band is a TEAM quantity, so
|
|
2214
|
+
Δ is applied to every player in the treated XI, matching how §2 and `cond-effect.mts`
|
|
2215
|
+
measure it.
|
|
2216
|
+
|
|
2217
|
+
30,000 seeds, and **each seed's home and away runs are averaged into one observation**
|
|
2218
|
+
before any variance is computed — they share an RNG seed, so counting them separately
|
|
2219
|
+
understates nothing but invents variance where a control has none. Bonferroni family 1440.
|
|
2220
|
+
|
|
2221
|
+
Two controls, kept separate on purpose:
|
|
2222
|
+
|
|
2223
|
+
- **Plumbing check** (Δ=0, 288 cells): the treated team is the baseline team, so this is
|
|
2224
|
+
structurally zero. It verifies `prep` and seed alignment, and is NOT a statistical
|
|
2225
|
+
control — calling it one is how a degenerate cell gets mistaken for evidence.
|
|
2226
|
+
- **Side-swap accounting** (24 cells): a squad against itself at equal growth must score
|
|
2227
|
+
exactly 0.5 win-equivalent. It does.
|
|
2228
|
+
|
|
2229
|
+
**The maximum of many noisy estimates is biased upward**, so the worst cell is read as an
|
|
2230
|
+
interval, and a safety claim ("this band stays under the bar") uses the UPPER bound.
|
|
2231
|
+
|
|
2232
|
+
```
|
|
2233
|
+
SEEDS=30000 npx tsx packages/mcp/skill/reference/probes/band-worstcase.mts # ~20 min
|
|
2234
|
+
```
|
|
2235
|
+
|
|
2236
|
+
### Result — in WIN-RATE points, on the exposure-weighted mix
|
|
2237
|
+
|
|
2238
|
+
Two conversions have to happen before any number meets the 5pp bar, and getting either
|
|
2239
|
+
wrong moves the answer by more than the measurement does.
|
|
2240
|
+
|
|
2241
|
+
**Unit.** `outcome` scores a draw as 0.5, so the probe produces win-equivalent SHARE, while
|
|
2242
|
+
the gate is written in win-rate advantage points. The canonical conversion lives in
|
|
2243
|
+
`constant-sweep.mts` as `WIN_RATE_PER_SHARE` and is exactly 2, always:
|
|
2244
|
+
`share − 0.5 = (w − l) / 2n`, so `(w − l) / n` is twice it regardless of how the draws
|
|
2245
|
+
fall. Comparing share against a 5pp bar silently applies a 10pp threshold.
|
|
2246
|
+
|
|
2247
|
+
**Weighting.** The gate is defined over an exposure-weighted mode mix, not per mode, and
|
|
2248
|
+
the mix is the repository's own: `constant-sweep.mts`'s `KO_EXPOSURE` =
|
|
2249
|
+
`1.33 / (1.33 + 12 + 3)` = **8.1%** — 1.33 knockout matches against **12 ladder matches
|
|
2250
|
+
and the cup's own three group-stage matches**, which are played in regulation mode and so
|
|
2251
|
+
belong in the denominator. An earlier version of this entry used 24 ladder matches and
|
|
2252
|
+
omitted the group stage, giving 5.2%; the source file's comment warns about exactly that,
|
|
2253
|
+
noting it overweights the knockout evidence by about two points. Weighting is applied PER
|
|
2254
|
+
CELL, since the worst regulation cell and the worst knockout cell need not be the same
|
|
2255
|
+
matchup.
|
|
2256
|
+
|
|
2257
|
+
| Δ (band) | regulation worst | knockout worst | **weighted upper** | cells over bar |
|
|
2258
|
+
| --- | --- | --- | --- | --- |
|
|
2259
|
+
| 0.2 **←ceiling** | +3.61 | +5.48 | **4.18** | 0/144 |
|
|
2260
|
+
| 0.3 | +5.36 | +8.06 | **6.07** | 24/144 |
|
|
2261
|
+
| 0.4 | +7.19 | +10.81 | **8.01** | 83/144 |
|
|
2262
|
+
| 0.5 | +8.88 | +13.47 | **9.81** | 113/144 |
|
|
2263
|
+
| 1.0 | +17.50 | +26.55 | **18.84** | 144/144 |
|
|
2264
|
+
|
|
2265
|
+
**Only 0.2 clears the gate** (weighted upper 4.18, 0 of 144). 0.3 is already over at 6.07,
|
|
2266
|
+
and 1.0 puts all 144 cells over. The weighted slope is ≈17.7 pp per cond point, so the bar
|
|
2267
|
+
is crossed near **0.28**; 0.2 is the largest 0.1 grid point beneath it.
|
|
2268
|
+
|
|
2269
|
+
**And the verdict barely depends on that weight.** Following what `constant-sweep` does,
|
|
2270
|
+
the gate is solved across a range rather than at one point (values are weighted upper
|
|
2271
|
+
bounds):
|
|
2272
|
+
|
|
2273
|
+
| KO weight | Δ0.2 | Δ0.3 | Δ0.4 | Δ0.5 | Δ1.0 |
|
|
2274
|
+
| --- | --- | --- | --- | --- | --- |
|
|
2275
|
+
| 5.0% | ✅ 4.18 | ❌ 6.06 | ❌ 8.00 | ❌ 9.78 | ❌ 18.77 |
|
|
2276
|
+
| 8.1% **(canonical)** | ✅ 4.18 | ❌ 6.07 | ❌ 8.01 | ❌ 9.81 | ❌ 18.84 |
|
|
2277
|
+
| 10.0% | ✅ 4.18 | ❌ 6.08 | ❌ 8.02 | ❌ 9.83 | ❌ 18.89 |
|
|
2278
|
+
| 15.0% | ✅ 4.21 | ❌ 6.11 | ❌ 8.10 | ❌ 9.80 | ❌ 19.01 |
|
|
2279
|
+
|
|
2280
|
+
0.2 passes and 0.3 fails at every weight from 5% to 15% — so even the wrong exposure figure
|
|
2281
|
+
gives the same answer, which is a fact only the sweep can establish.
|
|
2282
|
+
|
|
2283
|
+
The plan's "start at 0.5, never exceed 1.0" is therefore about **2.5× the safe value**.
|
|
2284
|
+
|
|
2285
|
+
**0.2 is provisional.** The role×condition interaction is unmeasured and can move it either
|
|
2286
|
+
way — see below — so do not treat it as a fixed value.
|
|
2287
|
+
|
|
2288
|
+
### The document already had this in its own table
|
|
2289
|
+
|
|
2290
|
+
§2 measured cond per point by template: `role-433` +5.0/+9.2, `bal-442` +4.4/+9.8,
|
|
2291
|
+
**`flat-433` +8.2/+13.5**, `def-532` +1.9/+10.5. §3-2 sized the band from `role-433` —
|
|
2292
|
+
0.5 × 9.2 = 4.6 pp, apparently under the bar — while the same table's `flat-433` gives
|
|
2293
|
+
0.5 × 13.5 = 6.75 pp.
|
|
2294
|
+
|
|
2295
|
+
This sweep reproduces that row independently: 8.18 mirror / 8.75 mixed in regulation,
|
|
2296
|
+
13.27 in knockout. A flat squad is worst because its judgments cluster near the engine's
|
|
2297
|
+
thresholds, so a small condition difference flips the most outcomes.
|
|
2298
|
+
|
|
2299
|
+
Two things are new. Mixed matchups and growth do NOT meaningfully raise the worst case —
|
|
2300
|
+
almost every worst cell is the flat mirror at growth 0/0, with one exception where a mixed
|
|
2301
|
+
pair beat the mirror by 7%. And the response is close to linear, in WIN-RATE
|
|
2302
|
+
points: **≈17.8 per cond point in regulation, ≈27.0 in knockout, a ratio of ≈1.5** — the
|
|
2303
|
+
five grid points give 17.5–18.0 and 26.6–27.4 respectively. (An earlier revision quoted
|
|
2304
|
+
8.5 and 13.4; those were share values left behind when this entry switched units.)
|
|
2305
|
+
|
|
2306
|
+
Growth turned out not to be an amplifier at all. That was one of the three axes §3-2 asked
|
|
2307
|
+
for, and the answer is negative.
|
|
2308
|
+
|
|
2309
|
+
### What this does not cover
|
|
2310
|
+
|
|
2311
|
+
Role concentration — the third axis §3-2 named — is not measured here. It requires patching
|
|
2312
|
+
the engine to carry per-player, per-scene weight sheets (`appr-role.mts` does this), which
|
|
2313
|
+
is expensive, and rejecting 0.3 and above does not need it.
|
|
2314
|
+
|
|
2315
|
+
**Do not read 0.3 as an upper estimate.** #728's +31.3 pp is the STANDALONE effect of role
|
|
2316
|
+
concentration, not the role×condition interaction, and a main effect does not fix the sign
|
|
2317
|
+
of an interaction. The explanation above cuts the other way if anything: if condition
|
|
2318
|
+
matters most where judgments cluster near thresholds, concentrating roles could move those
|
|
2319
|
+
judgments AWAY from thresholds and reduce sensitivity. The axis can move 0.3 in either
|
|
2320
|
+
direction, and it has to be measured before the value is adopted — 0.3 is what two axes
|
|
2321
|
+
give, not the answer for three.
|
|
2322
|
+
|
|
2323
|
+
## What a same-sum vector transfer is worth (`attr-vs-cond.mts`)
|
|
2324
|
+
|
|
2325
|
+
F-6 lets a registration submission redistribute each player's four attributes **within
|
|
2326
|
+
that player's own sum**. Two core players with equal sums can therefore submit each
|
|
2327
|
+
other's vectors: every check passes, and the strong build lands on whichever ID the
|
|
2328
|
+
manager picks. §3-12 answers with a per-registration cap on how far a vector may
|
|
2329
|
+
move from its own previous one, which slows a transfer to `ceil(distance / CAP)`
|
|
2330
|
+
registrations rather than blocking it; load is never moved. This probe measures what the transfer would have been worth.
|
|
2331
|
+
|
|
2332
|
+
### Method
|
|
2333
|
+
|
|
2334
|
+
Mirror match — both sides the same squad, the treatment on one side only, so the
|
|
2335
|
+
baseline is exactly 50% and the untreated cell IS the true-zero control. Seeds depend on
|
|
2336
|
+
(scenario, target, index) only, never on the arm (CRN). Win-equivalent share. **600,000
|
|
2337
|
+
seeds** — that is 600,000 independent clusters, 1,200,000 simulations, since each seed
|
|
2338
|
+
is reused for its two mirrored sides before they are averaged. Bonferroni family 72,
|
|
2339
|
+
equivalence margin ±0.25 pp.
|
|
2340
|
+
|
|
2341
|
+
**The two sides of a seed are averaged into one observation before any variance is
|
|
2342
|
+
computed** — they share an RNG seed. Counting them separately is wrong, and the control
|
|
2343
|
+
proves it: with identical teams the two scores sum to exactly 1 every seed, so
|
|
2344
|
+
seed-level variance is zero, yet that estimator reported ±0.17. After the fix the
|
|
2345
|
+
control reads ±0.00.
|
|
2346
|
+
|
|
2347
|
+
```
|
|
2348
|
+
SEEDS=600000 npx tsx packages/mcp/skill/reference/probes/attr-vs-cond.mts # the tables below
|
|
2349
|
+
npx tsx packages/mcp/skill/reference/probes/attr-vs-cond.mts # default 3000 — smoke only
|
|
2350
|
+
```
|
|
2351
|
+
|
|
2352
|
+
### Within a position, a swap moves almost nothing
|
|
2353
|
+
|
|
2354
|
+
| same-position swap | Δ(pp) | equivalence CI (±0.25) | equivalent? |
|
|
2355
|
+
| --- | --- | --- | --- |
|
|
2356
|
+
| DF regulation | -0.01 ±0.11 | [-0.12, +0.10] | yes |
|
|
2357
|
+
| DF knockout | +0.02 ±0.15 | [-0.12, +0.17] | yes |
|
|
2358
|
+
| DMF regulation | -0.00 ±0.11 | [-0.12, +0.11] | yes |
|
|
2359
|
+
| DMF knockout | +0.02 ±0.15 | [-0.13, +0.17] | yes |
|
|
2360
|
+
|
|
2361
|
+
A swap leaves the team's attribute multiset unchanged; only the labelling moves.
|
|
2362
|
+
|
|
2363
|
+
### Across positions it does — and that is a legal edit
|
|
2364
|
+
|
|
2365
|
+
F-6 freezes the four attributes but NOT a core player's position, so equal-sum players
|
|
2366
|
+
in different positions may exchange vectors and be deployed in the exchanged roles.
|
|
2367
|
+
|
|
2368
|
+
| cross-position swap | Δ(pp) | equivalence CI (±0.25) | equivalent? |
|
|
2369
|
+
| --- | --- | --- | --- |
|
|
2370
|
+
| DMF knockout | -0.57 ±0.12 | [-0.69, -0.45] | **NO** |
|
|
2371
|
+
| DMF regulation | -0.42 ±0.09 | [-0.50, -0.33] | **NO** |
|
|
2372
|
+
| FW knockout | -0.03 ±0.12 | [-0.15, +0.08] | yes |
|
|
2373
|
+
| FW regulation | +0.00 ±0.09 | [-0.08, +0.09] | yes |
|
|
2374
|
+
|
|
2375
|
+
The DMF cells (the probe pairs DMF↔FW here; the FW target pairs FW↔OMF) sit well
|
|
2376
|
+
outside the margin, at the same magnitude as one `cond+1`
|
|
2377
|
+
in the same cell. "The engine is nearly indifferent to the assignment" holds only
|
|
2378
|
+
*within* a position. The sign here is negative because the probe swaps in one
|
|
2379
|
+
direction; a launderer picks the favourable one. The magnitude is the point.
|
|
2380
|
+
|
|
2381
|
+
This is why §3-12 rejects transfers rather than pricing them: once the swap itself
|
|
2382
|
+
moves value, no load-based charge can cancel it — you would have to price the build
|
|
2383
|
+
movement too, and that varies by pair and by position combination.
|
|
2384
|
+
|
|
2385
|
+
### What laundering buys where pricing WAS plausible
|
|
2386
|
+
|
|
2387
|
+
Measured against **the compliant manager's best option** — fielding the strong build
|
|
2388
|
+
while tired, not resting it. Getting that baseline wrong makes any remedy look
|
|
2389
|
+
ineffective.
|
|
2390
|
+
|
|
2391
|
+
| threat/433 | laundering gain | cond+1 | compliant best |
|
|
2392
|
+
| --- | --- | --- | --- |
|
|
2393
|
+
| DF regulation | +0.80 ±0.13 | +0.93 | +5.26 |
|
|
2394
|
+
| DF knockout | +1.07 ±0.18 | +1.32 | +6.23 |
|
|
2395
|
+
| DMF regulation | +0.62 ±0.14 | +0.58 | +3.70 |
|
|
2396
|
+
| DMF knockout | +0.80 ±0.19 | +0.83 | +4.18 |
|
|
2397
|
+
|
|
2398
|
+
The gain (+0.62 to +1.07 pp) tracks the same-cell `cond+1` (+0.58 to +1.32 pp): within a
|
|
2399
|
+
position, laundering buys exactly one band of freshness.
|
|
2400
|
+
|
|
2401
|
+
### A trap worth keeping
|
|
2402
|
+
|
|
2403
|
+
An early **20-seed smoke run** returned **exactly 0.00** for the swap in every cell, with
|
|
2404
|
+
intervals of ±18 to ±27 pp. Read as a true zero it ends the investigation. It was an
|
|
2405
|
+
inert operation: the squad builders give equal sums equal vectors, so the swap produced
|
|
2406
|
+
an identical team. Eight cells reading 0.00 to the decimal while the interval is ±18 is a
|
|
2407
|
+
symptom, not a result.
|
|
2408
|
+
|
|
2409
|
+
(The ±18 belongs to that 20-seed smoke run and is not in tension with the ±0.17 above,
|
|
2410
|
+
which is the 300,000-seed run under the uncorrected estimator. Both figures are real;
|
|
2411
|
+
only the run differs.) The probe now skips
|
|
2412
|
+
vector-identical pairs and prints what it skipped (35 cells).
|
|
2413
|
+
|
|
2414
|
+
The exit code covers the control only. A cell failing equivalence is a **result** here,
|
|
2415
|
+
not a harness fault — §3-12's rule applies whether or not the swap moves value.
|
|
2416
|
+
|
|
2417
|
+
### What this cannot decide
|
|
2418
|
+
|
|
2419
|
+
One squad family, one tired/fresh gap (4.5 vs 5.5), and one cross-position pair. It does
|
|
2420
|
+
not measure how often the rejection blocks an edit a manager legitimately wants; §3-12
|
|
2421
|
+
bounds the geometry (four registrations to reach L1 distance 2, never closer), but the
|
|
2422
|
+
distribution of wanted edits is unmeasured. C-2's stress cases own that.
|
|
2423
|
+
|
|
2424
|
+
## What is one point of condition worth? (`cond-effect.mts`)
|
|
2425
|
+
|
|
2426
|
+
`cond` is pinned to 5 in every real match today, so this measures a lever that is
|
|
2427
|
+
not switched on yet. Task E of #732 needs it: the form-system plan proposes a
|
|
2428
|
+
ladder band of 0.5 rising to a 1.0 ceiling, calibrated under mirror matchups with
|
|
2429
|
+
growth and roles ignored.
|
|
2430
|
+
|
|
2431
|
+
**This section reports the grid and the tests. It deliberately does not convert
|
|
2432
|
+
them into calibration guidance** — the confounds below are large enough that any
|
|
2433
|
+
such reading would outrun the evidence, and six review rounds of this probe were
|
|
2434
|
+
spent withdrawing exactly those readings.
|
|
2435
|
+
|
|
2436
|
+
### Method
|
|
2437
|
+
|
|
2438
|
+
Rows are **win-equivalent share** for the condition-deficient side (win 1, draw
|
|
2439
|
+
0.5, loss 0). `Δcond` is the deficit: the short side plays at `5 − Δ`.
|
|
2440
|
+
|
|
2441
|
+
The seed depends on (scenario, direction, index) and NOT on mode, growth or Δ, so
|
|
2442
|
+
every cell within a scenario replays the identical fixture and any two form a
|
|
2443
|
+
paired contrast. Within-cell figures are per-seed paired differences against that
|
|
2444
|
+
cell's own Δ=0, with 95% CIs.
|
|
2445
|
+
|
|
2446
|
+
Cross-cell contrasts — growth (each configuration against `0/0`) and mode
|
|
2447
|
+
(knockout against league) — are per-seed **difference-in-differences** carrying
|
|
2448
|
+
**Bonferroni-adjusted** intervals over the WHOLE reported family (m=90 — 54
|
|
2449
|
+
growth plus 36 mode contrasts across all three scenarios — α=0.05, z=3.4524),
|
|
2450
|
+
because the interesting cells are selected after inspecting all of them. Counting
|
|
2451
|
+
the family per-scenario would leave the published family-wide error near 0.09. Growth additionally
|
|
2452
|
+
gets an explicit **±0.5pp equivalence margin**; mode is tested only for direction,
|
|
2453
|
+
since the useful question differs.
|
|
2454
|
+
|
|
2455
|
+
Controls: a true zero requires byte-identical teams (mirror, symmetric growth,
|
|
2456
|
+
Δ=0), and its z uses DECIDED outcomes only — a draw sits on the null mean and adds
|
|
2457
|
+
no variance. A failed control aborts the run nonzero rather than printing a
|
|
2458
|
+
warning beside normal-looking tables.
|
|
2459
|
+
|
|
2460
|
+
Growth uses the imported production `computeBuff`. `MAX_BUFF` is unreachable:
|
|
2461
|
+
core players are created at `DEFAULT_POTENTIAL_BP = 5000` (P=0.50) and generated
|
|
2462
|
+
free agents draw from [3000, 8500), so for these core-shaped squads the feasible
|
|
2463
|
+
ceiling is **+2 total per player**, not the +5 the constant names.
|
|
2464
|
+
|
|
2465
|
+
8,000 seeds × 2 directions (16,000 fixtures per cell).
|
|
2466
|
+
|
|
2467
|
+
|
|
2468
|
+
### League (gameFlg 0)
|
|
2469
|
+
|
|
2470
|
+
| scenario | growth | Δ=0 share | Δ=0.5 | Δ=1.0 | Δ=2.0 |
|
|
2471
|
+
| --- | --- | --- | --- | --- | --- |
|
|
2472
|
+
| mirror | 0/0 | 50.00 | −2.50 ±0.28 | −5.18 ±0.39 | −10.10 ±0.50 |
|
|
2473
|
+
| mirror | max/max | 49.89 | −2.40 ±0.27 | −4.78 ±0.37 | −9.76 ±0.48 |
|
|
2474
|
+
| mirror | 0/max | 37.86 | −2.53 ±0.27 | −4.81 ±0.37 | −8.95 ±0.46 |
|
|
2475
|
+
| mirror | max/0 | 62.31 | −2.34 ±0.27 | −4.92 ±0.37 | −10.00 ±0.49 |
|
|
2476
|
+
| gap:base-short | 0/0 | 70.88 | −3.52 ±0.34 | −7.16 ±0.46 | −14.68 ±0.62 |
|
|
2477
|
+
| gap:base-short | max/max | 70.97 | −3.42 ±0.34 | −6.81 ±0.46 | −13.88 ±0.61 |
|
|
2478
|
+
| gap:base-short | 0/max | 54.66 | −3.73 ±0.37 | −7.47 ±0.50 | −13.99 ±0.63 |
|
|
2479
|
+
| gap:base-short | max/0 | 84.33 | −2.56 ±0.28 | −5.41 ±0.39 | −11.60 ±0.53 |
|
|
2480
|
+
| gap:other-short | 0/0 | 29.06 | −3.41 ±0.34 | −6.53 ±0.44 | −11.79 ±0.54 |
|
|
2481
|
+
| gap:other-short | max/max | 28.70 | −3.18 ±0.32 | −6.18 ±0.43 | −11.28 ±0.53 |
|
|
2482
|
+
| gap:other-short | 0/max | 16.14 | −2.14 ±0.26 | −4.16 ±0.34 | −7.13 ±0.42 |
|
|
2483
|
+
| gap:other-short | max/0 | 45.04 | −3.90 ±0.36 | −7.84 ±0.49 | −14.43 ±0.61 |
|
|
2484
|
+
|
|
2485
|
+
### Knockout (gameFlg 1)
|
|
2486
|
+
|
|
2487
|
+
| scenario | growth | Δ=0 share | Δ=0.5 | Δ=1.0 | Δ=2.0 |
|
|
2488
|
+
| --- | --- | --- | --- | --- | --- |
|
|
2489
|
+
| mirror | 0/0 | 50.41 | −4.83 ±0.51 | −9.53 ±0.67 | −18.29 ±0.82 |
|
|
2490
|
+
| mirror | max/max | 50.39 | −4.77 ±0.52 | −9.56 ±0.67 | −18.89 ±0.81 |
|
|
2491
|
+
| mirror | 0/max | 30.34 | −4.32 ±0.47 | −8.36 ±0.61 | −14.98 ±0.71 |
|
|
2492
|
+
| mirror | max/0 | 69.75 | −3.96 ±0.48 | −8.32 ±0.64 | −18.06 ±0.80 |
|
|
2493
|
+
| gap:base-short | 0/0 | 74.63 | −4.21 ±0.44 | −8.66 ±0.60 | −18.09 ±0.77 |
|
|
2494
|
+
| gap:base-short | max/max | 74.67 | −3.72 ±0.44 | −7.96 ±0.60 | −17.29 ±0.77 |
|
|
2495
|
+
| gap:base-short | 0/max | 54.77 | −4.60 ±0.49 | −9.34 ±0.66 | −17.69 ±0.80 |
|
|
2496
|
+
| gap:base-short | max/0 | 89.04 | −2.45 ±0.34 | −5.59 ±0.47 | −12.52 ±0.63 |
|
|
2497
|
+
| gap:other-short | 0/0 | 25.38 | −3.84 ±0.43 | −7.29 ±0.55 | −12.91 ±0.65 |
|
|
2498
|
+
| gap:other-short | max/max | 24.61 | −3.59 ±0.42 | −6.94 ±0.54 | −12.33 ±0.65 |
|
|
2499
|
+
| gap:other-short | 0/max | 11.14 | −2.17 ±0.31 | −3.82 ±0.39 | −6.33 ±0.46 |
|
|
2500
|
+
| gap:other-short | max/0 | 45.27 | −5.30 ±0.49 | −10.42 ±0.65 | −18.50 ±0.79 |
|
|
2501
|
+
|
|
2502
|
+
### What the tests support
|
|
2503
|
+
|
|
2504
|
+
**Condition matters, at every Δ, in every cell.** All 72 within-cell
|
|
2505
|
+
effects are negative, and 72 of 72 have 95% intervals
|
|
2506
|
+
that exclude zero. At Δ=1 in the league the
|
|
2507
|
+
balanced mirror costs 5.18pp; §3-10's published
|
|
2508
|
+
"cond 1점 = +5.0pp" is consistent with that cell on this unit.
|
|
2509
|
+
|
|
2510
|
+
**Knockout differs from league.** Of 36 adjusted cross-mode contrasts,
|
|
2511
|
+
29 exclude zero and 7 do not; 1 of those 29 points the
|
|
2512
|
+
other way (knockout cheaper). The direction is therefore not uniform, and no
|
|
2513
|
+
single multiplier describes it.
|
|
2514
|
+
|
|
2515
|
+
**Growth: nothing is established either way.** Of 54 adjusted contrasts,
|
|
2516
|
+
**0** meet the ±0.5pp equivalence margin and **18** exclude zero. So this
|
|
2517
|
+
grid establishes NEITHER direction: growth is detectable in a third of the
|
|
2518
|
+
contrasts, yet no contrast is tight enough to call it negligible at a margin that
|
|
2519
|
+
could move a 0.5–1.0 band. Both "growth can be ignored" and "growth is bounded by
|
|
2520
|
+
X" are unsupported here. An earlier revision reported a large asymmetric effect;
|
|
2521
|
+
that was an artefact of an unattainable +4 buff and is withdrawn.
|
|
2522
|
+
|
|
2523
|
+
### What this grid CANNOT decide — read before using any of it for task E
|
|
2524
|
+
|
|
2525
|
+
- **Role concentration is absent.** §3-10's +31.3pp came from patching the
|
|
2526
|
+
engine's `*_APPR` tables (`appr-role.mts`), a probe-level mechanism not
|
|
2527
|
+
replicated here. Task E is not closed until that axis is run.
|
|
2528
|
+
- **Growth is a uniform scalar, not a per-player vector.** Production growth
|
|
2529
|
+
differs per player, so two squads can be equally grown in aggregate while
|
|
2530
|
+
concentrating buffs on different positions. Same question as role
|
|
2531
|
+
concentration, asked about a different quantity.
|
|
2532
|
+
- **Strength gaps come from squad construction, not from a varied axis.** The
|
|
2533
|
+
two `gap` scenarios ARE lopsided under symmetric growth (~71/29 in league,
|
|
2534
|
+
~75/25 in knockout), so strength-gap evidence here does not depend on growth
|
|
2535
|
+
asymmetry — but those baselines are produced by swapping in a different BUILD,
|
|
2536
|
+
not by moving a strength dial while holding everything else fixed. Any pattern
|
|
2537
|
+
across baselines is therefore an association among four hand-picked
|
|
2538
|
+
configurations.
|
|
2539
|
+
- **The most extreme baselines are only reachable via asymmetric growth.** In
|
|
2540
|
+
this grid nothing sits far outside ~71/29 without uneven growth, so at those
|
|
2541
|
+
ends the two causes cannot be told apart. Reaching an extreme baseline with
|
|
2542
|
+
symmetric growth needs a build pair with a larger intrinsic gap.
|
|
2543
|
+
- **One build pair, one formation pair, and a uniform deficit** — every player on
|
|
2544
|
+
the short side drops by the same Δ. A real fatigued squad is uneven.
|
|
2545
|
+
|
|
2546
|
+
|
|
2547
|
+
## What load does a player actually accrue? (`engagement-load.mts` + `c2-team-load.ts`)
|
|
2548
|
+
|
|
2549
|
+
**This section does NOT close C-2.** An earlier version said it did; that was wrong. `τ` was
|
|
2550
|
+
never measured — the probe's 24h default was simply used — and the neutral point was left
|
|
2551
|
+
unchosen between its two branches. An F-1 implementer following that decision would have made
|
|
2552
|
+
a probe default the production recovery rate. And `τ` cannot be measured from this data: F-1
|
|
2553
|
+
does not exist, so production has no fatigue and there is no observed recovery curve to fit.
|
|
2554
|
+
What measurement CAN do is price the choice, which is what the τ sweep does.
|
|
2555
|
+
|
|
2556
|
+
**The gate is not a per-mode ceiling — it is defined on the exposure-weighted mix.**
|
|
2557
|
+
`band-worstcase.mts` holds the canonical statement (`01-engine-expansion.md:319-329` puts
|
|
2558
|
+
~5pp on that mix; `measurements.md` pins "not per mode"), and **E's 0.2 came from that gate**.
|
|
2559
|
+
This section had been holding ladder and knockout to 5pp SEPARATELY — calibrating against the
|
|
2560
|
+
weighted gate while testing against a stricter per-mode one. With knockout at **8.1%** of
|
|
2561
|
+
exposure, that difference reverses the conclusion.
|
|
2562
|
+
|
|
2563
|
+
Three things are established:
|
|
2564
|
+
|
|
2565
|
+
- **Carrying E's ceiling across exceeds NOTHING at τ=24h.** Each team state is turned into a
|
|
2566
|
+
fresh and a tired variant and both are played against the SAME REAL OPPONENTS on the same
|
|
2567
|
+
seed, with the modes mixed by exposure: at τ=24h, 66 cells give **0 over, 3 established
|
|
2568
|
+
below, 63 undecided**. The short τ exceed badly — **32 of 52** at 6h, **12 of 63** at 12h.
|
|
2569
|
+
Only three established below is the price of the guarantee, not a smaller effect: the cell is
|
|
2570
|
+
now (team × curve × STATE RULE) with six opponents inside it, so the family is **1,836**. **"none over" and "all established below" are different claims**: the first says
|
|
2571
|
+
exceeding cannot be shown, the second says clearing can. τ=24h is the first; getting the
|
|
2572
|
+
second means lowering the amplitude. The earlier round played the tired team only against its
|
|
2573
|
+
own fresh CLONE — a mirror matchup, not what the gate is defined on. **The modes still
|
|
2574
|
+
diverge** — knockout point estimates run about **1.5×**
|
|
2575
|
+
ladder's, which is why holding each mode to 5pp separately showed 38 of 78 over. That is
|
|
2576
|
+
what the earlier "almost every team, in knockout" conclusion was: **the wrong bar**.
|
|
2577
|
+
- **So what this section delivers is a safe amplitude**: re-evaluating ALL 66 cells at each
|
|
2578
|
+
candidate under a simultaneous correction and then validating the choice on INDEPENDENT
|
|
2579
|
+
seeds gives **`b` ∈ 0.1988–0.2014** at τ=24h, worst cell 3.86pp. **All four τ values have
|
|
2580
|
+
such an amplitude** (0.16–0.21), so fairness closes on the AMPLITUDE, not on the choice of τ.
|
|
2581
|
+
That is **0.07×** the value derived by putting E's 0.2 on the median player. Testing one cell with a pointwise
|
|
2582
|
+
interval leaves winner's-curse bias: on a small sample a search that passed at 4.11pp
|
|
2583
|
+
validated at 10.48pp.
|
|
2584
|
+
- **Whether rotation pays — not at the amplitude that would ship.** At `g=0.4` the prize is
|
|
2585
|
+
0.59pp [0.25, 0.93] and the paired difference is −0.43pp [−0.76, −0.10]: the prize does exceed
|
|
2586
|
+
the cost there. But **that is not where the decision lives.** A rested player can only recover
|
|
2587
|
+
the deficit the shipped amplitude creates, and the **safe amplitude produces a deficit of
|
|
2588
|
+
0.135** (τ=24h, the largest across all validated cells — taking it from the cell with the
|
|
2589
|
+
largest BOUND instead understates what can be recovered). At that deficit the prize is
|
|
2590
|
+
**0.16pp [−0.08, 0.41]** against a cheapest cost of **0.16pp [−0.36, 0.67]**, paired difference
|
|
2591
|
+
**−0.00pp [−0.47, 0.46] — undecided**, and the same at all four τ (the eight repricing calls
|
|
2592
|
+
publish a prize interval and a paired-difference interval apiece, so all **16 intervals** split
|
|
2593
|
+
one 5% budget). The cost of a one-point-weaker slot runs 0.16 to
|
|
2594
|
+
2.21pp depending on WHICH point moves, with selection and publication on SEPARATE seed columns
|
|
2595
|
+
(the selection column's own spread across the 16 is 2.57pp).
|
|
2596
|
+
**Which one is cheapest is not established**: measuring all 16 on the publication column too,
|
|
2597
|
+
its minimum is `dribble→shoot` (−0.01pp), not the selection's `dribble→pass` (0.16pp). The
|
|
2598
|
+
comparison keeps the selection's pick — taking the publication column's minimum would bias that
|
|
2599
|
+
column — so the cheapest substitute costs somewhere in **0 to 0.2pp** and which one it is stays
|
|
2600
|
+
open. Every interval here is distribution-free (empirical Bernstein), which is why this section
|
|
2601
|
+
runs at 60,000 seeds.
|
|
2602
|
+
|
|
2603
|
+
**The measurement was rebuilt five times and the conclusion moved each time.** A synthetic
|
|
2604
|
+
squad grid gave "this axis is too small" (prize 0.46pp). Replaying production matchups gave
|
|
2605
|
+
a longer tail and a prize the same order as one ability point. Integrating the production
|
|
2606
|
+
TIMELINE gives the numbers above. The first two are withdrawn.
|
|
2607
|
+
|
|
2608
|
+
### λ_team × multiplier cannot produce L
|
|
2609
|
+
|
|
2610
|
+
That product multiplies two upper quantiles, so it puts `L_top` at 63 — while integrating
|
|
2611
|
+
the real timeline puts the peak `L` at **p95 21.3** (observed max 50.5). The two quantiles do
|
|
2612
|
+
not arrive on the same player-day. And bucketing by day discards the SPACING, which is
|
|
2613
|
+
everything to an EWMA: cup fixtures cluster into a few hours after opening, so eight closely
|
|
2614
|
+
spaced calls do not equilibrate like eight spread over a day.
|
|
2615
|
+
|
|
2616
|
+
So the layer is gone. `matches.input_snapshot` (#694) freezes the exact two teams handed to
|
|
2617
|
+
the engine, so the last 30 days' **1,201 real kickoff matchups** are replayed in time order
|
|
2618
|
+
and the decay is integrated over them:
|
|
2619
|
+
|
|
2620
|
+
L ← L·exp(−Δt/τ) + (that match's engagements)
|
|
2621
|
+
|
|
2622
|
+
Neither `λ` nor a multiplier appears. Each snapshot carries its own `gameFlg`, so the mode
|
|
2623
|
+
mix is right by construction — removing an estimate beats getting it right.
|
|
2624
|
+
|
|
2625
|
+
Three further corrections: only **credited** appearances accumulate (a human ladder away side
|
|
2626
|
+
and an AI cup entrant carry zero F-1 load, and an earlier version pooled them into the human
|
|
2627
|
+
multiplier); each matchup is replayed **five times and averaged**, because a single replay's
|
|
2628
|
+
integer counts read within-match noise as between-player heterogeneity; and the synthetic
|
|
2629
|
+
grid is kept as a control only.
|
|
2630
|
+
|
|
2631
|
+
**τ is therefore not a free parameter.** An earlier version used the equilibrium form and
|
|
2632
|
+
concluded τ cancels. On the timeline τ sets the load itself: every table below is τ=24h and
|
|
2633
|
+
another value needs another integration.
|
|
2634
|
+
|
|
2635
|
+
### Stage 1 — team call frequency, 30 completed UTC days
|
|
2636
|
+
|
|
2637
|
+
| humans, 32 teams | median | p90 | p99 | max |
|
|
2638
|
+
| --- | --- | --- | --- | --- |
|
|
2639
|
+
| **calendar days** (zeros included, 865 team-days) | **0** | 8 | 10 | 21 |
|
|
2640
|
+
| active days (zeros excluded, 270 team-days) | 6 | 8 | 19 | 21 |
|
|
2641
|
+
| **per-team window total** (32 teams, median 30 observed days) | **0** | **183** | — | **199** |
|
|
2642
|
+
|
|
2643
|
+
**A team-day median cannot say how many teams sat out.** An earlier version read the
|
|
2644
|
+
calendar median of 0 as "half the human teams played nothing"; a median of 0 over team-days
|
|
2645
|
+
only says half the team-DAYS were empty — every team could have played on a minority of its
|
|
2646
|
+
days, and late-created teams contribute fewer days besides. Counting each team's total over
|
|
2647
|
+
its own observed window instead: **18 of 32 human teams (56.3%) played nothing** in the
|
|
2648
|
+
window (bots: 919/1,536 = 59.8%).
|
|
2649
|
+
|
|
2650
|
+
**The grid includes each team's creation day — that was removed and then put back.** Removing
|
|
2651
|
+
it avoids counting a partial day as a full team-day of zero, but it also **deletes every daily
|
|
2652
|
+
fill-bot match**: fill-bots are created the morning their cup opens, play only that day
|
|
2653
|
+
(AGENTS.md L712-717 — a deployment with zero humans still runs a 48-bot cup daily), and
|
|
2654
|
+
contribute nothing but zeros on the window's other days. Losing the numerator outright is far
|
|
2655
|
+
worse than one partial day in the denominator. Putting it back moves the bot cohort a long
|
|
2656
|
+
way — active team-days 664 → **1,224**, "never played in the window" 96.3% → **59.8%** — and
|
|
2657
|
+
leaves the human quantiles alone, since only 8 human team-days (42 calls) are creation days.
|
|
2658
|
+
|
|
2659
|
+
**That count also exposes what the median hid — the distribution is bimodal.** Median 0, p90
|
|
2660
|
+
**183**, max 199. Human teams either barely play or play throughout; **there is no typical
|
|
2661
|
+
team in between.** So §3-1's neutral point at "today's typical load" has no typical to
|
|
2662
|
+
anchor to: wherever it goes, one end sits in the wrong place. On the calendar reading it is
|
|
2663
|
+
0, which sets `a` = 0 and puts everyone who plays into penalty — the mirror of the all-bonus
|
|
2664
|
+
failure §3-1 warns about; on the active-day reading (6) the 56% who sit out all land on the
|
|
2665
|
+
bonus side. Not chosen here.
|
|
2666
|
+
|
|
2667
|
+
Buckets are anchored to UTC explicitly: `date_trunc('day', NOW())` follows the session
|
|
2668
|
+
timezone, so a non-UTC database silently measures partial days under a "completed UTC" label.
|
|
2669
|
+
|
|
2670
|
+
**"Has a snapshot" counts only what the repository would accept.** An earlier version checked
|
|
2671
|
+
that `home.team` and `away.team` were JSON objects, which lets `{"team": {}}` or a squad
|
|
2672
|
+
without eleven players count toward 100% coverage — while `parseMatchInputSnapshot` throws on
|
|
2673
|
+
them (fail-closed, #694), so the exporter either drops the row silently or ships it and
|
|
2674
|
+
`playOne` dies. The coverage query now applies the same predicate as `isSideSnapshot`, plus
|
|
2675
|
+
the one condition only this consumer needs (exactly eleven players, since the probe indexes
|
|
2676
|
+
slots 0–10), and the coverage and export queries share the same SQL fragment — if they
|
|
2677
|
+
diverge, coverage says 100% while the export loses rows. A row that passes SQL but fails the
|
|
2678
|
+
JS parse is a contradiction, so the script exits rather than dropping it. Tightening changed
|
|
2679
|
+
nothing here: the same 1,201 matches, byte-identical export. "Fully covered" is now verified
|
|
2680
|
+
rather than assumed.
|
|
2681
|
+
|
|
2682
|
+
### F-1's predicate: counting actors only drops half, and concentrates the rest
|
|
2683
|
+
|
|
2684
|
+
Production replay, 1,201 matchups over 5 replays, 8,283 credited human player-matches:
|
|
2685
|
+
|
|
2686
|
+
| predicate | team total | p95 | median | p95/median | zero-load share |
|
|
2687
|
+
| --- | --- | --- | --- | --- | --- |
|
|
2688
|
+
| **A.** hyoka credit sites, unsigned | 6.60 | 1.60 | 0.40 | **4.00** | 13% |
|
|
2689
|
+
| **B.** actor only | 8.97 | 2.00 | 0.80 | **2.50** | 14% |
|
|
2690
|
+
| **C.** actor + responder | **17.88** | 3.20 | 1.60 | **2.00** | **4%** |
|
|
2691
|
+
| **D.** C without the ZP fallback | 16.61 | 2.80 | 1.40 | 2.00 | 4% |
|
|
2692
|
+
|
|
2693
|
+
The pool is credited HUMAN team-sides only — λ is a human-cohort figure, so the pool has to
|
|
2694
|
+
be as well. **This table alone uses expected RATES** (the 5-replay mean per player-match).
|
|
2695
|
+
Counting realized integers instead would put a player whose rate is 0.4 into "no load" on
|
|
2696
|
+
every draw that came up 0, and that player does accrue load across matches; the question here
|
|
2697
|
+
is whether a player is STRUCTURALLY load-free, not how often they were picked in one draw.
|
|
2698
|
+
The timeline integration asks the opposite question and uses integers.
|
|
2699
|
+
Counting actors only drops 49.8% of the engagement — (17.88−8.97)/17.88 — and what
|
|
2700
|
+
remains is MORE concentrated: p95/median 2.50 against 2.00 (4.00 at the hyoka sites), with a
|
|
2701
|
+
zero-load share three and a half times higher. Establishing §3-10's premise across the whole eleven needs responders counted.
|
|
2702
|
+
|
|
2703
|
+
**The earlier "65% accrue nothing" is withdrawn.** That came from single-replay integer
|
|
2704
|
+
counts; on expected rates it is 14%. Within-match noise had been read as between-player
|
|
2705
|
+
structure. The argument survives in a different form.
|
|
2706
|
+
|
|
2707
|
+
The keeper still accrues almost nothing — it enters a contest only at the A0 one-on-one and
|
|
2708
|
+
at penalties — so no rule here makes one tired, and F-1 has to fill that blank deliberately.
|
|
2709
|
+
|
|
2710
|
+
Predicate D now subtracts fallbacks on BOTH sides of the contest: zone press picks its ACTOR
|
|
2711
|
+
through the fallback too, and an earlier version recorded only the responder's, so "C without
|
|
2712
|
+
the ZP fallback" was removing half of what it claimed.
|
|
2713
|
+
|
|
2714
|
+
**A and C rose again in this pass** — A from 5.72 to **6.60**, C from 16.54 to **17.88** —
|
|
2715
|
+
because two sites were not being counted at all:
|
|
2716
|
+
|
|
2717
|
+
- **The goal-scorer's credit does not go through `applyHyoka`.** `scoreGoal`
|
|
2718
|
+
(`engine.ts:1581`) writes `atk.hyoka[i] += 25 + rand(50)` DIRECTLY. Defining A as "the
|
|
2719
|
+
current credit sites" and then counting only `applyHyoka` drops every scorer.
|
|
2720
|
+
- **The defender picked after a failed dribble was uncounted.** `engine.ts:653` selects it
|
|
2721
|
+
with `getPlayer(def, scene, false)`, which never passes through `getPoint`, so it sat
|
|
2722
|
+
outside the responder tally — yet that player commits or avoids the foul and can be booked.
|
|
2723
|
+
In the production distribution: **2,894 selections, 13.2% of all defensive selections,
|
|
2724
|
+
7.1% of C**. If C means "defensive selections", this is one.
|
|
2725
|
+
|
|
2726
|
+
**B did not move by a single digit** (8.97 → 8.97), which is the control that says both fixes
|
|
2727
|
+
touched only what they were meant to.
|
|
2728
|
+
|
|
2729
|
+
### The parameters
|
|
2730
|
+
|
|
2731
|
+
E capped the ATTAINABLE spread at 0.2, so the derivation fixes that and solves for `b`.
|
|
2732
|
+
`L` is measured, not estimated — timeline integration at τ=24h over the 143 credited human players:
|
|
2733
|
+
|
|
2734
|
+
143 human players over an **11-day window with 100% snapshot coverage** (2026-08-21 onward,
|
|
2735
|
+
1,201 matches), τ=24h.
|
|
2736
|
+
|
|
2737
|
+
**"the match that happened" is verified per match, not inferred from a date.** Calling a
|
|
2738
|
+
stored-seed replay that only holds if the engine replaying it is the engine that ran it, and
|
|
2739
|
+
when that breaks fixture consistency breaks with it — a different knockout winner leaves a
|
|
2740
|
+
team that cannot exist in that timeline still accruing load, exactly why the derived-seed
|
|
2741
|
+
timelines were dropped.
|
|
2742
|
+
|
|
2743
|
+
That boundary was once taken from `git log`'s COMMIT date. A commit date is not a deployment:
|
|
2744
|
+
it can ship after the following midnight, be rolled back, or roll out gradually. So the
|
|
2745
|
+
inference is gone and the replay is compared against what production STORED instead — reproduce
|
|
2746
|
+
it and the match happens the same way under today's engine, otherwise drop it. The result is
|
|
2747
|
+
**1,201 of 1,201 (100%)**: the two `engine.ts` changes inside the window (`74af7010`
|
|
2748
|
+
golden-goal control flow, `b587f5d4` scene tags) did not change these matches' outcomes. The
|
|
2749
|
+
209 matches cut by date were fine.
|
|
2750
|
+
|
|
2751
|
+
**The comparison runs on three axes, each closer than the last to what this probe counts:**
|
|
2752
|
+
|
|
2753
|
+
| axis | what it covers | this window |
|
|
2754
|
+
| --- | --- | --- |
|
|
2755
|
+
| final score | coarsest — "same score, different path" passes | 1,201/1,201 |
|
|
2756
|
+
| event-sequence digest (`seq|minute|type|player|player2`) | the REPORTED path | 1,201/1,201 |
|
|
2757
|
+
| stored **per-player rating** | a projection of the hyoka accumulator | 2,402/2,402 sides · 26,422 player-cells · 0 mismatches |
|
|
2758
|
+
|
|
2759
|
+
The second axis alone is not enough: selections that go through `getPoint` and emit NO event
|
|
2760
|
+
at all (a completed pass, a successful dribble) still land in the `off`/`def` arrays being
|
|
2761
|
+
integrated. The third covers that hole — `player_match_stats.rating` is
|
|
2762
|
+
`clamp(round(hyoka/100), 1, 10)`, so it is a projection of the hyoka accumulator, and hyoka is
|
|
2763
|
+
moved by exactly those selection sites (a successful dribble calls `applyHyoka(atk, kiten, 1)`).
|
|
2764
|
+
|
|
2765
|
+
**That axis carries little information, and that is measured too.** 97.4% of the 26,422 cells
|
|
2766
|
+
are `4`, only three values occur, entropy is **0.176 bits per cell**, and the 2,402 eleven-vectors
|
|
2767
|
+
take only **57 distinct values**. So "26,422 cells agree" must not be read as 26,422 independent
|
|
2768
|
+
tests. What it actually constrains is whether hyoka moved enough to push some player across a
|
|
2769
|
+
rounding boundary — a weak constraint, but one independent of the event axis.
|
|
2770
|
+
|
|
2771
|
+
**`scene` and `detail` are deliberately excluded from the digest.** Both are presentation
|
|
2772
|
+
metadata and both changed inside the window (#725 added scene tags on 08-22); including them
|
|
2773
|
+
would reject the 08-21–08-22 matches for a reason unrelated to the estimand. `player` already
|
|
2774
|
+
carries the side.
|
|
2775
|
+
|
|
2776
|
+
**And the functions that produce the selection path did not change in the window — checkable at
|
|
2777
|
+
the source level.** Six commits touched `packages/engine/src` inside it, but only two touched
|
|
2778
|
+
simulation (the rest are test files and `validator.ts`): `74af7010` (golden-goal termination)
|
|
2779
|
+
and `b587f5d4` (scene tags on events). Neither touches `getPlayer`, `getDefPlayer`,
|
|
2780
|
+
`getPoint`, or `weightedPick`:
|
|
2781
|
+
|
|
2782
|
+
```
|
|
2783
|
+
git log --since=2026-08-14 --oneline origin/main -- packages/engine/src/
|
|
2784
|
+
git show b587f5d4 -- packages/engine/src/engine.ts # all pushEventAt / scoreGoal arguments
|
|
2785
|
+
```
|
|
2786
|
+
|
|
2787
|
+
**What remains unverified.** Production has NEVER persisted the `off`/`def` tally, so for a
|
|
2788
|
+
window that has already passed, "persist the tally and compare it" is not available at all.
|
|
2789
|
+
What survives all three axes is a revision that leaves the events AND the ratings identical
|
|
2790
|
+
while selecting different players — which requires hyoka credits to cancel exactly. Cutting the
|
|
2791
|
+
window by deployment date is not the alternative: a commit date is not a deployment, and that
|
|
2792
|
+
inference is precisely what this verification replaced.
|
|
2793
|
+
|
|
2794
|
+
**Engagements are integrated as integers.** An earlier version replayed each matchup 5 times,
|
|
2795
|
+
AVERAGED the counts, and fed those fractional expectations into the EWMA. That is wrong:
|
|
2796
|
+
every published quantity — per-player peak, p95, the gate deficits — is a NONLINEAR
|
|
2797
|
+
functional, so `E[f(X)] ≠ f(E[X])`, and averaging erases within-match selection streaks and
|
|
2798
|
+
narrows the load vector. A narrower vector passes the gate more easily, so the error runs in
|
|
2799
|
+
exactly the direction that says "safe". What runs now is **one realized timeline**, replaying
|
|
2800
|
+
each match's stored `matches.rng_seed` (**1,201 of 1,201**) — that pass is not a different draw
|
|
2801
|
+
of the same matchup but the match that happened.
|
|
2802
|
+
|
|
2803
|
+
A middle version ran five timelines, four of them on derived seeds, and that was wrong too:
|
|
2804
|
+
the fixture list comes from production and is fixed, so a derived seed that changes a knockout
|
|
2805
|
+
winner leaves a team that **could not exist** in that timeline still accruing load in the next
|
|
2806
|
+
round. Those four were inconsistent histories, and they were feeding the published medians and
|
|
2807
|
+
the worst-state search. The cost of dropping them is that the across-draw spread is no longer
|
|
2808
|
+
measurable here — it needs bracket progression regenerated, which this probe does not do.
|
|
2809
|
+
|
|
2810
|
+
Fixing the averaging moved the vectors up: max deficits rose 5–14% per cell and the per-match
|
|
2811
|
+
increment rose **31–33%**. The safe amplitude moved about 2% (measured then against the
|
|
2812
|
+
per-mode gate, so the absolute values differ from today's) — the conclusion was insensitive to
|
|
2813
|
+
the error, but that is only knowable after re-measuring.
|
|
2814
|
+
|
|
2815
|
+
**Which instant you read `L` at decides the policy.** Three values differ:
|
|
2816
|
+
|
|
2817
|
+
| instant | median | p95 | observed max |
|
|
2818
|
+
| --- | --- | --- | --- |
|
|
2819
|
+
| **at appearance** (what cond reads, 7,139 reads) | **8.46** | **26.21** | 55.27 |
|
|
2820
|
+
| at the window end (an arbitrary clock time) | 7.03 | 16.24 | — |
|
|
2821
|
+
| path maximum (after its own match's engagements) | 19.35 | 38.17 | — |
|
|
2822
|
+
|
|
2823
|
+
**The load event is not placed at kickoff.** F-1 pins the decay reference to the time the
|
|
2824
|
+
match's condition is FIRST SNAPSHOTTED, which is when the scheduler picks the match up —
|
|
2825
|
+
closer to `completed_at`, since simulation takes milliseconds. An earlier version substituted
|
|
2826
|
+
`kickoff_at` and justified it with "the lag inside the window is under a minute". **That was
|
|
2827
|
+
false**: the window's last day has a median lag of **428 minutes** and the overall p99 is
|
|
2828
|
+
**508** (35% of τ=24h) — some cups are processed well after their scheduled time.
|
|
2829
|
+
|
|
2830
|
+
So it is measured rather than argued. The published `claim` axis uses **1,144** cup
|
|
2831
|
+
`completed_at` values plus **57** playoff `finalized_at` values: playoff writes the result as
|
|
2832
|
+
`live` first, then that latter column records the completion transition. A per-row
|
|
2833
|
+
`claimSource` proves which hand supplied it; if either the completion time or its source is
|
|
2834
|
+
missing, the probe refuses the whole window instead of falling back to kickoff [Codex P2,
|
|
2835
|
+
2026-09-04]. Integrating on the explicit `kickoff` alternative instead gives median 8.66 and
|
|
2836
|
+
p95 26.35 — **2.3%** and **0.5%** away.
|
|
2837
|
+
|
|
2838
|
+
**That comparison holds the warm-up boundary fixed.** Left to itself each variant recomputes
|
|
2839
|
+
`collectFrom` from its own first event — and switching the time base, or dropping the
|
|
2840
|
+
long-lag matches, is exactly what moves that first event. The printed difference would then be
|
|
2841
|
+
a clock effect PLUS a changed sample rather than the substitution alone. The published run's
|
|
2842
|
+
boundary is passed in so every variant collects from the same absolute instant. Here the
|
|
2843
|
+
correction moved the median by 0.01 (8.65 → 8.66): small, but measured rather than assumed.
|
|
2844
|
+
|
|
2845
|
+
**Neither completion timestamp is the first claim.** The daemon PRESERVES a failed attempt's
|
|
2846
|
+
snapshot, resets the match to `scheduled`, and picks it up later, while cup `completed_at`
|
|
2847
|
+
records the attempt that finally succeeded. Playoff `finalized_at` is coarser still: it closes
|
|
2848
|
+
the live window roughly two to three minutes after the result was computed. The store has no
|
|
2849
|
+
first-snapshot column, so it cannot be recovered; excluding the **104** matches with more than
|
|
2850
|
+
60 minutes of lag gives median 8.66 and p95 27.26, **+2.4%** and **+4.0%**. The uncertainty on
|
|
2851
|
+
this axis is 2–4%, and F-1 persisting the first snapshot time
|
|
2852
|
+
removes it.
|
|
2853
|
+
|
|
2854
|
+
**That this axis is a real claim clock is now proved by the ledger, not by lag** [Codex P2,
|
|
2855
|
+
2026-09-02]. Two hands write `completed_at`: a live completion stamps the moment it finished,
|
|
2856
|
+
and `repairMissingCompletionPositions` stamps the moment it was DISCOVERED. Nothing on the row
|
|
2857
|
+
says which. An earlier version separated them by size — "more than 24 hours of lag means
|
|
2858
|
+
backfill" — and that is the wrong axis: the repair also runs during rolling deployments (its
|
|
2859
|
+
own comment says so), so an old replica's just-finished match discovered minutes later carries
|
|
2860
|
+
a small lag and passes that test. Magnitude cannot tell the two hands apart.
|
|
2861
|
+
|
|
2862
|
+
For `completed_at`, the `data_backfills` ledger can. `runOnce` returns at the `<name>:closed` marker BEFORE doing
|
|
2863
|
+
any work, so once that marker is written the repair never runs again — which makes "closed at
|
|
2864
|
+
T" a proof that every `completed_at` at or after T was written by a live completion. In this
|
|
2865
|
+
store the marker was written at **2026-08-17T09:56:13Z**. The repair itself ran at 06:37Z that
|
|
2866
|
+
day: of the 1,502 completed competitive matches kicking off between 08-03 and 08-16, every one
|
|
2867
|
+
that carries a `completed_at` carries it **inside 06:37:21.9–23.5Z**, a 1.6-second window —
|
|
2868
|
+
which is the fingerprint the repair's own `base + rn milliseconds` produces. The window opens on
|
|
2869
|
+
08-21 and its earliest claim is 08-21T04:00:08Z. The 57 playoff rows use a separate durable
|
|
2870
|
+
proof: `finalizePlayoffMatch` writes `status='complete'` and `finalized_at` in the SAME update,
|
|
2871
|
+
and the completion backfill never edits that column. The export carries both timestamp and
|
|
2872
|
+
source. The probe checks that all 1,144 `completed_at` values postdate the ledger close, that
|
|
2873
|
+
`finalized_at` was not assigned to a cup, and that every row has both fields. With no proof it
|
|
2874
|
+
does not drop rows, it **refuses to compute**: dropping them would punch a hole in the window
|
|
2875
|
+
and bring the same bias back wearing a different face.
|
|
2876
|
+
|
|
2877
|
+
**The EWMA is warmed before observations are collected.** Truncating the window (coverage,
|
|
2878
|
+
engine boundary) leaves players carrying load they arrived with, while the integration starts
|
|
2879
|
+
them at `L=0` — treating the boundary as a league-wide fatigue reset. An earlier version merely
|
|
2880
|
+
listed this as unmeasured; measured, it is real: discarding the first **2τ** as warm-up raised
|
|
2881
|
+
the at-appearance median from 7.32 to 7.76 (+6%) — **that run's numbers**; the current run
|
|
2882
|
+
prints only the warmed value, and with the predicate corrected it is **8.46**. The cold start was biasing load DOWN,
|
|
2883
|
+
which makes the gate easier. The collection window is 8.8 days and the residual bias is
|
|
2884
|
+
`exp(−2) ≈ 0.135` — recorded, not claimed gone.
|
|
2885
|
+
|
|
2886
|
+
Calibration uses **at appearance**. Condition is frozen at claim time (§7), so what a player
|
|
2887
|
+
actually carries is the `L` read as they enter a match. The window end is a state at an
|
|
2888
|
+
instant, not a state on entry — and **nobody ever experiences the path maximum**, because it
|
|
2889
|
+
includes the match's own engagements and no one reads condition after that. Earlier versions
|
|
2890
|
+
used each player's own last kickoff (which leaves a player who stopped un-cooled) and then the
|
|
2891
|
+
window end (right for "now", wrong for "on entry").
|
|
2892
|
+
|
|
2893
|
+
| `L₀/L_top` | `b` | `a` (neutral) | team spread | one-player deficit | prize* | x per match | in pp |
|
|
2894
|
+
| --- | --- | --- | --- | --- | --- | --- | --- |
|
|
2895
|
+
| 1× | 0.820 | 0.200 | 0.200 | 0.410 | 0.602pp | 0.029 | 0.043pp |
|
|
2896
|
+
| 2× | 1.439 | 0.200 | 0.200 | 0.480 | 0.705pp | 0.046 | 0.068pp |
|
|
2897
|
+
| 4× | 2.679 | 0.200 | 0.200 | 0.536 | 0.788pp | 0.063 | 0.093pp |
|
|
2898
|
+
|
|
2899
|
+
\* deficit × slope (1.47 pp/cond); those deficits sit above the largest directly measured one
|
|
2900
|
+
(0.40) and the response is concave, so they are biased high. The measured figure is 0.59pp,
|
|
2901
|
+
and at the safe amplitude the whole-vector team-level figure is 3.86pp.
|
|
2902
|
+
|
|
2903
|
+
Repricing at each safe amplitude uses the largest deficit across all independently validated
|
|
2904
|
+
cells, not the deficit of the cell with the largest confidence bound:
|
|
2905
|
+
|
|
2906
|
+
| τ | safe `b` | deficit | prize there | cheapest cost | paired cost − prize | verdict |
|
|
2907
|
+
| --- | --- | --- | --- | --- | --- | --- |
|
|
2908
|
+
| 6h | 0.1633 | 0.112 | 0.13pp [−0.10, 0.37] | 0.16pp | 0.02pp [−0.44, 0.49] | undecided |
|
|
2909
|
+
| 12h | 0.1639 | 0.110 | 0.13pp [−0.10, 0.37] | 0.16pp | 0.03pp [−0.44, 0.50] | undecided |
|
|
2910
|
+
| **24h** | **0.1988** | **0.135** | **0.16pp [−0.08, 0.41]** | **0.16pp** | **−0.00pp [−0.47, 0.46]** | **undecided** |
|
|
2911
|
+
| 48h | 0.2061 | 0.137 | 0.16pp [−0.08, 0.41] | 0.16pp | −0.00pp [−0.47, 0.46] | undecided |
|
|
2912
|
+
|
|
2913
|
+
The eight repricing calls publish a prize interval and a paired-difference interval apiece.
|
|
2914
|
+
All 16 intervals split one 5% family budget; each two-sided interval splits its share again
|
|
2915
|
+
between its two tails [Codex P2, 2026-09-02; 2026-09-04].
|
|
2916
|
+
|
|
2917
|
+
`x`, the number §3-1 asked to be stated: one more match costs a top-load player **0.029–0.063
|
|
2918
|
+
condition points (0.043–0.093pp)**. No "N matches per day" qualifier is needed — τ and the
|
|
2919
|
+
real spacing are already inside `L`. The increment is the p95 of the **realized integer**
|
|
2920
|
+
distribution (4 engagements; median 1) — one match is one draw, so the p95 of the expected
|
|
2921
|
+
RATE ("a busy player's average match") answers a different question. That distinction alone
|
|
2922
|
+
raised these two columns 31–33%.
|
|
2923
|
+
|
|
2924
|
+
**The pp column is a CONVERSION, not a measurement.** The 1.47 slope is a SECANT taken at
|
|
2925
|
+
`g = 0.4` (that one slot's loss divided by 0.4), and win probability is nonlinear in condition
|
|
2926
|
+
and depends on role, attributes and opponent. So the same deficits were measured directly, on
|
|
2927
|
+
the same slot and the same published seed column:
|
|
2928
|
+
|
|
2929
|
+
| τ | per-match increment (cond) | **measured directly** | secant conversion | ratio |
|
|
2930
|
+
| --- | --- | --- | --- | --- |
|
|
2931
|
+
| 6h | 0.189 | **0.21pp** [−0.05, 0.47] | 0.278pp | 1.3× |
|
|
2932
|
+
| 12h | 0.124 | **0.14pp** [−0.10, 0.38] | 0.182pp | 1.3× |
|
|
2933
|
+
| **24h** | **0.063** | **0.07pp** [−0.15, 0.28] | **0.093pp** | **1.3×** |
|
|
2934
|
+
| 48h | 0.031 | **0.06pp** [−0.14, 0.26] | 0.046pp | 0.8× |
|
|
2935
|
+
|
|
2936
|
+
**None of the four resolves at this budget** — every interval covers zero once the alpha is
|
|
2937
|
+
split across the 16 intervals from eight repricing calls — so the pp column has to be read as a conversion
|
|
2938
|
+
throughout. Comparing point estimates, the secant runs about 1.3× the direct measurement, but
|
|
2939
|
+
that comparison is itself not established. The sweep table's per-match column is the same
|
|
2940
|
+
conversion.
|
|
2941
|
+
|
|
2942
|
+
### τ cannot be measured — only priced
|
|
2943
|
+
|
|
2944
|
+
`τ` is the recovery rate §6 C-2 asks for. F-1 does not exist, so production has no fatigue
|
|
2945
|
+
and there is no observed curve to fit: `τ` is a DESIGN choice, not a measurement. What
|
|
2946
|
+
measurement can do is re-run everything at each candidate and report what follows. Each row
|
|
2947
|
+
below re-integrates the timeline, re-runs that τ's full cell set, and re-searches for the safe
|
|
2948
|
+
amplitude (one run, one snapshot).
|
|
2949
|
+
|
|
2950
|
+
| τ(h) | `L` med | `L` p95 | published `b` (4×) | per match (pp) | over gate | **safe `b`** | worst at its low end | `b` proven over |
|
|
2951
|
+
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
|
|
2952
|
+
| 6 | 2.68 | 11.80 | 3.716 | 0.278 | **32/52** | 0.1633–0.1669 | 3.64pp | 1.8581 |
|
|
2953
|
+
| 12 | 4.19 | 16.18 | 3.291 | 0.182 | **12/63** | 0.1639–0.1671 | 3.42pp | 1.6455 |
|
|
2954
|
+
| **24** | **8.46** | **26.21** | **2.679** | **0.093** | **0/66** | **0.1988–0.2014** | **3.86pp** | — |
|
|
2955
|
+
| 48 | 17.55 | 45.94 | 2.294 | 0.046 | **0/61** | 0.2061–0.2084 | 3.56pp | — |
|
|
2956
|
+
|
|
2957
|
+
Three readings. **`L` is near-linear in τ** (2.68 → 17.55), so §3-6's `L* = λτ` holds on the
|
|
2958
|
+
timeline too and changing τ means rebuilding the whole table. **Longer τ makes a match cheaper
|
|
2959
|
+
and the field fairer** — marginal cost 0.278 → 0.046pp, cells over **32 → 0** — because a
|
|
2960
|
+
longer memory dilutes both the single match and the heterogeneity between players; a short τ
|
|
2961
|
+
buys "one match hurts" at the price of unfairness, and here that price is steep. **All four
|
|
2962
|
+
have a safe amplitude**, so fairness is closed by the AMPLITUDE, not by τ; τ is the "how much
|
|
2963
|
+
does one match hurt" dial.
|
|
2964
|
+
|
|
2965
|
+
**And the bisection errs in one direction under noise.** One unlucky evaluation above 5 near
|
|
2966
|
+
the threshold means everything above it is never revisited, while an unlucky PASS is caught by
|
|
2967
|
+
the independent-seed validation — the asymmetry is structural. This run's τ=24h trace shows the
|
|
2968
|
+
scale (4.97pp at b=0.1988, 5.14pp at b=0.2014, 1.3% higher). The published amplitude is
|
|
2969
|
+
therefore a LOWER bound on the threshold, and the error runs toward the safe side. And the
|
|
2970
|
+
right-hand endpoint is NOT a confidence limit: "not established below" is not "over" — that
|
|
2971
|
+
would need a LOWER bound above 5. Across this sweep two candidates cleared that bar and survived an independent-seed confirmation
|
|
2972
|
+
(τ=6h at b=1.8581 and τ=12h at b=1.6455), and only those thresholds acquire an upper bound.
|
|
2973
|
+
|
|
2974
|
+
**Whether the weight decides the verdict is measured too.** Re-counting the same cells at
|
|
2975
|
+
knockout weights of 5%, 8.1% (canonical), 10% and 15%: at τ=24h cells over run **0·0·0·0** —
|
|
2976
|
+
the verdict does not hang on that number at all. The canonical weight is used; where the
|
|
2977
|
+
dependence grows is recorded rather than assumed away.
|
|
2978
|
+
|
|
2979
|
+
The safe `b` is a **bracket**, not a point: the largest candidate that passed and the smallest
|
|
2980
|
+
that failed. And that bracket is SEARCH GRANULARITY, not a confidence interval — each
|
|
2981
|
+
candidate's bound is itself estimated from 6,000 seeds PER OPPONENT. This run's lower endpoints
|
|
2982
|
+
increase (0.163 · 0.164 · 0.199 · 0.206), but the procedure does not test ordering across τ, so
|
|
2983
|
+
that shape must not be read as structural. The
|
|
2984
|
+
bisection errs in ONE DIRECTION under noise: a single unlucky evaluation above 5 near the
|
|
2985
|
+
threshold means everything above it is never revisited, while an unlucky pass is caught by the
|
|
2986
|
+
independent-seed validation. τ=24h's trace shows the scale — 4.97pp at b=0.1988 against 5.14pp
|
|
2987
|
+
at b=0.2014, 1.3% higher. So the published value is a LOWER BOUND on the threshold, and the
|
|
2988
|
+
direction it errs in is the safe one. **The right endpoint is not an upper bound either** —
|
|
2989
|
+
"not established below" is not "over". What is established is that every τ in this range has an
|
|
2990
|
+
amplitude in the reported safe range that puts all cells under the gate, and that cells over fall 32 → 0
|
|
2991
|
+
with τ.
|
|
2992
|
+
|
|
2993
|
+
**The sweep caught two defects in the gate** that a single τ could never have shown. First,
|
|
2994
|
+
capping a degenerate sample's bound at infinity made cells where the amplitude is small enough
|
|
2995
|
+
that the two variants coincide fail the search outright — τ=6, 12 and 48 all reported "no safe
|
|
2996
|
+
amplitude" for that reason, and τ=24 only appeared to succeed because it happened to land
|
|
2997
|
+
ABOVE that point. The earlier τ=24 result was luck, not structure. A rule-of-three bound
|
|
2998
|
+
(`c + (−ln α / n)·(max − c)`) fixes it and all four then resolve. Second, five bisection steps
|
|
2999
|
+
were too coarse: τ=6 stalled at 5.35pp on b=0.122 and that was printed as "none exists", which
|
|
3000
|
+
is search depth, not a finding. Ten steps now, with an early exit on the first failing cell to
|
|
3001
|
+
pay for them.
|
|
3002
|
+
|
|
3003
|
+
### Testing the vector against the gate instead of borrowing a ceiling
|
|
3004
|
+
|
|
3005
|
+
E applies the same Δ to all eleven; fatigue makes a vector. An earlier version averaged
|
|
3006
|
+
engagement by ARRAY INDEX across every formation and applied that to a synthetic 4-3-3 — but
|
|
3007
|
+
index is slot order, so the average at slot s mixes unrelated roles and is nobody's load
|
|
3008
|
+
vector. Worse, it described the result as "played against its fresh copy" while the code put
|
|
3009
|
+
each variant against a rotating synthetic field and subtracted field-relative win rates.
|
|
3010
|
+
Match outcomes are nonlinear and opponent-dependent; that difference is not a head-to-head.
|
|
3011
|
+
|
|
3012
|
+
Corrected: each team state is turned into a fresh and a tired variant, and BOTH are played
|
|
3013
|
+
against EACH OF THE SIX FIXED REAL-OPPONENT BASKET MEMBERS as its own cell — never averaged or
|
|
3014
|
+
screened away — on the same seed,
|
|
3015
|
+
then differenced. **Each seed runs BOTH home and away; those two
|
|
3016
|
+
differences are averaged before they become one observation** [Codex P2, 2026-09-04]. Alternating
|
|
3017
|
+
orientation by seed index would require home and away to be identically distributed, which the
|
|
3018
|
+
engine's home-only DMF branch disproves. Playing the tired side against its own fresh clone is a mirror matchup, and the
|
|
3019
|
+
gate is explicitly defined by VARYING the opponent, because the engine's thresholds make the
|
|
3020
|
+
condition response matchup-dependent. The opponent family is every distinct XI actually fielded
|
|
3021
|
+
in the window, folded on a hash of the fielded eleven rather than on `teamId` — folding on
|
|
3022
|
+
`teamId` keeps only each team's first snapshot and discards every later XI. The model's
|
|
3023
|
+
neutral offset `a` is applied to both sides, since the equation is `5 + a − bL/(L+L₀)` and
|
|
3024
|
+
omitting `a` overstates the gap by exactly that.
|
|
3025
|
+
|
|
3026
|
+
**The opponent is a CELL AXIS — neither averaged nor screened.** The canonical gate
|
|
3027
|
+
(`band-worstcase.mts:161`) loops `for (A) for (B) for (growth) for (mode)`, so both the count of
|
|
3028
|
+
cells over and the choice of worst live on that grid. This got it wrong twice:
|
|
3029
|
+
|
|
3030
|
+
1. **Averaged.** Drawing one pool member per seed and averaging them into a single cell dilutes
|
|
3031
|
+
an effect that appears only against a particular opponent, and a mixture mean is never above
|
|
3032
|
+
the worst, so the error ran toward "safe".
|
|
3033
|
+
2. **Then screened.** Testing all 62 and confirming only the winner leaves the discarded 61 with
|
|
3034
|
+
NO guarantee and a family that never counts the opponent axis. Worse, that screen was NOISE.
|
|
3035
|
+
This run measures one term's SD at a median of **24.0pp** (a win/draw/loss difference), so at
|
|
3036
|
+
120 screen seeds an opponent mean has SE 2.19pp and the max−min of 62 draws is **≈9.4pp from
|
|
3037
|
+
noise alone**. The 10.85pp range the previous round read as "the opponent axis matters" was
|
|
3038
|
+
SMALLER than what noise produces.
|
|
3039
|
+
|
|
3040
|
+
So screening is gone and the basket is fixed: the 62 fielded XIs sorted by team strength, six
|
|
3041
|
+
taken at even quantiles. Deterministic, never touching an outcome, and the family counts the
|
|
3042
|
+
whole axis (cell = state × curve, with six opponents inside; family = (78+78+78+72) × 6 =
|
|
3043
|
+
**1,836**). Publication and out-of-basket exploration do not divide seeds across the basket — each opponent gets the full 12,000 — because
|
|
3044
|
+
a degenerate cell's bound floors at `(−ln(α/2)/n)·SER_MAX`, so dividing `n` by the basket size
|
|
3045
|
+
widens that floor by the same factor; dividing it produced "no safe amplitude" at all four τ,
|
|
3046
|
+
which is arithmetic rather than a finding.
|
|
3047
|
+
|
|
3048
|
+
**Bounding all 62 under 5pp is not affordable** — at that SD a 1.5pp half-width per opponent
|
|
3049
|
+
needs ≈4,092 seeds each, tens of times the cell budget across 62. So the basket is a SAMPLE of
|
|
3050
|
+
the 62, the same kind of limit the canonical gate accepts by using a designed grid. The question
|
|
3051
|
+
it leaves open — does an opponent outside the basket exceed? — is asked separately below.
|
|
3052
|
+
|
|
3053
|
+
**What that fold actually changed was measured.** Across 2,402 sides there are 95 `teamId`s but
|
|
3054
|
+
only **62 distinct XIs**. The old pool held 95 objects, of which only **47** were distinct
|
|
3055
|
+
lineups — 48 were duplicates. And **35 of the 95 teams fielded more than one XI** inside the
|
|
3056
|
+
window, every later one discarded. The new pool CONTAINS the old one and adds 15 lineups.
|
|
3057
|
+
(Snapshot `cond` is 5 for everyone in this window — F-1 does not exist yet — so "fielded XI"
|
|
3058
|
+
is squad identity here, not condition.)
|
|
3059
|
+
|
|
3060
|
+
**Taking only the top three was arbitrary — all 13 teams that played are measured.** Two
|
|
3061
|
+
things were wrong at once. Teams were ranked by *peak single load* while states were chosen by
|
|
3062
|
+
*summed deficit* — two currencies, and putting them on one **changed two of the top three**.
|
|
3063
|
+
With the currency fixed, the cut turned out not to exist:
|
|
3064
|
+
|
|
3065
|
+
```
|
|
3066
|
+
70363723 5.93 · 30178fc1 4.62 · 6097e735 4.45 · 704b8e45 4.34 · 6aec8bd8 4.31
|
|
3067
|
+
118f1f13 4.20 · 6a37ed0a 4.20 · 2aa911db 4.12 · 2b865779 4.01 · 23baeba2 3.74
|
|
3068
|
+
930b4313 3.58 · 796efbfe 3.34 · e8dc63e1 0.15 (summed deficit, max over curves)
|
|
3069
|
+
```
|
|
3070
|
+
|
|
3071
|
+
Second (4.62) through twelfth (3.34) sit close together, so "the worst team" would
|
|
3072
|
+
have been decided by which three I picked. Nothing is cut now, and the simultaneous correction
|
|
3073
|
+
widened accordingly. (τ=24h below; other τ in the sweep above.)
|
|
3074
|
+
|
|
3075
|
+
| team | `L₀` | state (UTC) | max deficit | ladder | knockout | **weighted** | lo | hi | 5pp |
|
|
3076
|
+
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
|
|
3077
|
+
| `70363723` | 1× | 2026-08-25 12:58 | 0.524 | 5.00 | 7.86 | **5.24** | 2.59 | 7.88 | undecided |
|
|
3078
|
+
| `70363723` | 1× | 2026-08-27 09:57 | 0.556 | 5.10 | 7.33 | **5.28** | 2.60 | 7.95 | undecided |
|
|
3079
|
+
| `70363723` | 2× | 2026-08-25 11:55 | 0.684 | 6.30 | 9.09 | **6.52** | 3.72 | 9.33 | undecided |
|
|
3080
|
+
| `70363723` | 2× | 2026-08-27 09:57 | 0.739 | 6.32 | 8.86 | **6.53** | 3.76 | 9.29 | undecided |
|
|
3081
|
+
| `70363723` | 4× | 2026-08-25 11:55 | 0.834 | 6.78 | 10.16 | **7.05** | 4.16 | 9.94 | undecided |
|
|
3082
|
+
| `70363723` | 4× | 2026-08-27 09:57 | 0.925 | 6.71 | 10.45 | **7.02** | 4.14 | 9.89 | undecided |
|
|
3083
|
+
| `30178fc1` | 1× | 2026-08-31 12:56 | 0.404 | 4.55 | 7.14 | **4.76** | 2.16 | 7.35 | undecided |
|
|
3084
|
+
| `30178fc1` | 1× | 2026-08-27 07:26 | 0.434 | 5.47 | 7.74 | **5.65** | 3.00 | 8.31 | undecided |
|
|
3085
|
+
| `30178fc1` | 2× | 2026-08-31 12:56 | 0.471 | 5.24 | 8.66 | **5.52** | 2.84 | 8.19 | undecided |
|
|
3086
|
+
| `30178fc1` | 2× | 2026-08-27 07:26 | 0.518 | 5.34 | 7.91 | **5.55** | 2.84 | 8.26 | undecided |
|
|
3087
|
+
| `30178fc1` | 4× | 2026-08-31 12:56 | 0.524 | 6.25 | 8.27 | **6.41** | 3.67 | 9.16 | undecided |
|
|
3088
|
+
| `30178fc1` | 4× | 2026-08-27 07:26 | 0.588 | 5.93 | 8.35 | **6.13** | 3.36 | 8.89 | undecided |
|
|
3089
|
+
| `6097e735` | 1× | 2026-08-28 07:24 | 0.415 | 5.31 | 6.83 | **5.44** | 2.75 | 8.12 | undecided |
|
|
3090
|
+
| `6097e735` | 1× | 2026-08-31 12:56 | 0.441 | 4.22 | 5.88 | **4.36** | 1.75 | 6.96 | undecided |
|
|
3091
|
+
| `6097e735` | 2× | 2026-08-28 07:24 | 0.487 | 5.42 | 7.34 | **5.57** | 2.83 | 8.31 | undecided |
|
|
3092
|
+
| `6097e735` | 2× | 2026-08-31 12:56 | 0.530 | 4.61 | 6.84 | **4.79** | 2.11 | 7.47 | undecided |
|
|
3093
|
+
| `6097e735` | 4× | 2026-08-28 07:24 | 0.546 | 5.46 | 7.89 | **5.66** | 2.84 | 8.47 | undecided |
|
|
3094
|
+
| `6097e735` | 4× | 2026-08-31 12:56 | 0.605 | 5.04 | 7.67 | **5.26** | 2.54 | 7.97 | undecided |
|
|
3095
|
+
| `704b8e45` | 1× | 2026-08-28 07:26 | 0.470 | 4.47 | 6.35 | **4.62** | 2.05 | 7.19 | undecided |
|
|
3096
|
+
| `704b8e45` | 2× | 2026-08-28 07:26 | 0.578 | 5.14 | 6.71 | **5.27** | 2.61 | 7.93 | undecided |
|
|
3097
|
+
| `704b8e45` | 4× | 2026-08-28 07:26 | 0.673 | 5.19 | 7.49 | **5.38** | 2.65 | 8.10 | undecided |
|
|
3098
|
+
| `6aec8bd8` | 1× | 2026-08-23 07:26 | 0.438 | 5.52 | 6.53 | **5.60** | 2.90 | 8.29 | undecided |
|
|
3099
|
+
| `6aec8bd8` | 1× | 2026-08-26 07:26 | 0.466 | 5.02 | 6.25 | **5.12** | 2.44 | 7.80 | undecided |
|
|
3100
|
+
| `6aec8bd8` | 2× | 2026-08-23 07:26 | 0.524 | 5.95 | 7.61 | **6.08** | 3.30 | 8.87 | undecided |
|
|
3101
|
+
| `6aec8bd8` | 2× | 2026-08-26 07:26 | 0.571 | 5.70 | 7.40 | **5.84** | 3.10 | 8.58 | undecided |
|
|
3102
|
+
| `6aec8bd8` | 4× | 2026-08-23 07:26 | 0.596 | 6.19 | 7.85 | **6.32** | 3.50 | 9.14 | undecided |
|
|
3103
|
+
| `6aec8bd8` | 4× | 2026-08-26 07:26 | 0.663 | 5.94 | 7.38 | **6.06** | 3.28 | 8.83 | undecided |
|
|
3104
|
+
| `118f1f13` | 1× | 2026-08-30 07:24 | 0.468 | 3.78 | 6.70 | **4.02** | 1.54 | 6.50 | undecided |
|
|
3105
|
+
| `118f1f13` | 2× | 2026-08-30 07:24 | 0.576 | 4.42 | 7.42 | **4.66** | 2.09 | 7.24 | undecided |
|
|
3106
|
+
| `118f1f13` | 4× | 2026-08-30 07:24 | 0.669 | 5.18 | 7.98 | **5.41** | 2.78 | 8.04 | undecided |
|
|
3107
|
+
| `6a37ed0a` | 1× | 2026-08-25 07:26 | 0.441 | 4.17 | 6.13 | **4.33** | 1.76 | 6.89 | undecided |
|
|
3108
|
+
| `6a37ed0a` | 1× | 2026-08-28 07:26 | 0.459 | 4.04 | 5.30 | **4.14** | 1.59 | 6.70 | undecided |
|
|
3109
|
+
| `6a37ed0a` | 2× | 2026-08-25 07:26 | 0.530 | 4.94 | 7.10 | **5.12** | 2.46 | 7.77 | undecided |
|
|
3110
|
+
| `6a37ed0a` | 2× | 2026-08-28 07:26 | 0.560 | 4.42 | 6.39 | **4.58** | 1.94 | 7.22 | undecided |
|
|
3111
|
+
| `6a37ed0a` | 4× | 2026-08-25 07:26 | 0.605 | 5.50 | 7.58 | **5.67** | 2.99 | 8.36 | undecided |
|
|
3112
|
+
| `6a37ed0a` | 4× | 2026-08-28 07:26 | 0.647 | 4.75 | 7.67 | **4.99** | 2.35 | 7.62 | undecided |
|
|
3113
|
+
| `2aa911db` | 1× | 2026-08-29 07:26 | 0.404 | 4.04 | 5.57 | **4.17** | 1.61 | 6.72 | undecided |
|
|
3114
|
+
| `2aa911db` | 1× | 2026-08-30 07:26 | 0.452 | 3.92 | 5.66 | **4.07** | 1.52 | 6.61 | undecided |
|
|
3115
|
+
| `2aa911db` | 2× | 2026-08-29 07:26 | 0.471 | 4.24 | 6.51 | **4.43** | 1.86 | 6.99 | undecided |
|
|
3116
|
+
| `2aa911db` | 2× | 2026-08-30 07:26 | 0.548 | 4.63 | 6.02 | **4.74** | 2.15 | 7.32 | undecided |
|
|
3117
|
+
| `2aa911db` | 4× | 2026-08-29 07:26 | 0.524 | 4.93 | 6.97 | **5.10** | 2.45 | 7.74 | undecided |
|
|
3118
|
+
| `2aa911db` | 4× | 2026-08-30 07:26 | 0.630 | 5.06 | 6.86 | **5.20** | 2.55 | 7.86 | undecided |
|
|
3119
|
+
| `2b865779` | 1× | 2026-08-26 07:26 | 0.453 | 4.33 | 6.60 | **4.51** | 1.92 | 7.11 | undecided |
|
|
3120
|
+
| `2b865779` | 2× | 2026-08-26 07:26 | 0.549 | 5.02 | 7.92 | **5.26** | 2.58 | 7.94 | undecided |
|
|
3121
|
+
| `2b865779` | 4× | 2026-08-26 07:26 | 0.632 | 4.90 | 8.32 | **5.18** | 2.49 | 7.88 | undecided |
|
|
3122
|
+
| `23baeba2` | 1× | 2026-08-31 12:56 | 0.405 | 4.06 | 5.37 | **4.17** | 1.16 | 7.18 | undecided |
|
|
3123
|
+
| `23baeba2` | 1× | 2026-08-29 07:06 | 0.409 | 3.61 | 5.93 | **3.80** | 1.25 | 6.34 | undecided |
|
|
3124
|
+
| `23baeba2` | 2× | 2026-08-31 12:56 | 0.472 | 4.85 | 6.32 | **4.97** | 1.85 | 8.08 | undecided |
|
|
3125
|
+
| `23baeba2` | 2× | 2026-08-29 07:06 | 0.478 | 4.30 | 5.85 | **4.43** | 1.81 | 7.05 | undecided |
|
|
3126
|
+
| `23baeba2` | 4× | 2026-08-31 12:56 | 0.525 | 4.75 | 6.79 | **4.91** | 1.73 | 8.10 | undecided |
|
|
3127
|
+
| `23baeba2` | 4× | 2026-08-29 07:06 | 0.534 | 5.04 | 5.46 | **5.08** | 2.42 | 7.73 | undecided |
|
|
3128
|
+
| `930b4313` | 1× | 2026-08-23 07:01 | 0.383 | 3.77 | 5.64 | **3.92** | 1.38 | 6.46 | undecided |
|
|
3129
|
+
| `930b4313` | 1× | 2026-08-24 07:14 | 0.429 | 3.62 | 6.07 | **3.82** | 1.29 | 6.35 | undecided |
|
|
3130
|
+
| `930b4313` | 2× | 2026-08-23 07:01 | 0.439 | 4.14 | 6.92 | **4.37** | 1.78 | 6.95 | undecided |
|
|
3131
|
+
| `930b4313` | 2× | 2026-08-24 07:14 | 0.510 | 4.19 | 5.71 | **4.32** | 1.76 | 6.87 | undecided |
|
|
3132
|
+
| `930b4313` | 4× | 2026-08-23 07:01 | 0.482 | 4.85 | 6.92 | **5.02** | 2.41 | 7.62 | undecided |
|
|
3133
|
+
| `930b4313` | 4× | 2026-08-24 07:14 | 0.577 | 4.35 | 6.38 | **4.51** | 1.94 | 7.09 | undecided |
|
|
3134
|
+
| `796efbfe` | 1× | 2026-08-25 07:06 | 0.351 | 3.59 | 6.89 | **3.86** | 1.44 | 6.28 | undecided |
|
|
3135
|
+
| `796efbfe` | 1× | 2026-08-28 06:58 | 0.367 | 2.98 | 4.89 | **3.13** | 0.80 | 5.46 | undecided |
|
|
3136
|
+
| `796efbfe` | 2× | 2026-08-25 07:06 | 0.392 | 3.65 | 8.03 | **4.01** | 1.53 | 6.48 | undecided |
|
|
3137
|
+
| `796efbfe` | 2× | 2026-08-28 06:58 | 0.415 | 3.06 | 4.94 | **3.21** | 0.83 | 5.60 | undecided |
|
|
3138
|
+
| `796efbfe` | 4× | 2026-08-25 07:06 | 0.422 | 4.03 | 7.65 | **4.32** | 1.80 | 6.84 | undecided |
|
|
3139
|
+
| `796efbfe` | 4× | 2026-08-28 06:58 | 0.451 | 3.21 | 5.15 | **3.37** | 0.99 | 5.75 | undecided |
|
|
3140
|
+
| `e8dc63e1` | 1× | 2026-08-24 02:05 | 0.017 | 0.20 | 0.25 | **0.20** | -1.39 | 1.79 | **below** |
|
|
3141
|
+
| `e8dc63e1` | 2× | 2026-08-24 02:05 | 0.015 | 0.15 | 0.25 | **0.17** | -1.35 | 1.72 | **below** |
|
|
3142
|
+
| `e8dc63e1` | 4× | 2026-08-24 02:05 | 0.015 | 0.15 | 0.03 | **0.17** | -1.34 | 1.69 | **below** |
|
|
3143
|
+
|
|
3144
|
+
**None of the 66 cells is over**; 3 are established below and 63 undecided. So many undecided is
|
|
3145
|
+
what the guarantee cost — the family is **1,836** = (78+78+78+72) × 6 opponents.
|
|
3146
|
+
|
|
3147
|
+
**The state axis had the same defect.** The previous round chose each team's worst state by a
|
|
3148
|
+
low-seed screen — exactly what had just been thrown out on the opponent axis: discarded states
|
|
3149
|
+
carry no guarantee, the family never counts that axis, and the screen itself is noise. The state
|
|
3150
|
+
is now chosen by two DETERMINISTIC rules — largest summed deficit (how tired the team is overall)
|
|
3151
|
+
and largest single-slot deficit (a vector concentrated on one place) — and BOTH are measured. At
|
|
3152
|
+
τ=24h the two rules pointed at different states in **27 of 39** (team, L₀) pairs.
|
|
3153
|
+
|
|
3154
|
+
**The previous round's "1 over" is withdrawn.** It came from screening 62 opponents, then
|
|
3155
|
+
correcting the winner against a family of 153 — the search over 62 was never counted. Corrected
|
|
3156
|
+
for it (α = 0.05/18972), nothing at τ=24h is established over. **"none over" and "all established
|
|
3157
|
+
below" remain different claims** — the first says exceedance cannot be established, the second
|
|
3158
|
+
says non-exceedance can — and the 63 undecided are neither.
|
|
3159
|
+
|
|
3160
|
+
**At the short τ, however, an exceeding matchup does exist.** Since the basket is a sample of the
|
|
3161
|
+
62, "does an opponent outside it exceed?" is a separate question, asked separately: screen all 62
|
|
3162
|
+
per cell under common random numbers (120 seeds), measure the winner on DIFFERENT seeds at full
|
|
3163
|
+
count (12,000), and correct for having searched 62.
|
|
3164
|
+
|
|
3165
|
+
| τ | (cell, opponent) established over | worst |
|
|
3166
|
+
| --- | --- | --- |
|
|
3167
|
+
| 6h | **23** | `30178fc1` 4× — weighted 13.32pp, lower bound 9.59pp |
|
|
3168
|
+
| 12h | **5** | `6aec8bd8` 4× — weighted 9.81pp, lower bound 6.36pp |
|
|
3169
|
+
| 24h · 48h | 0 | — |
|
|
3170
|
+
|
|
3171
|
+
Inside the basket, the point estimates for one state spread by **1.98pp** median (4.14pp max) —
|
|
3172
|
+
the real measure of how much the opponent moves the verdict, and next to the ≈9.4pp noise range
|
|
3173
|
+
above it shows why a screen cannot rank them.
|
|
3174
|
+
|
|
3175
|
+
This does NOT replace the verdict above: screening is drawn to noise, so it does not establish
|
|
3176
|
+
"we found the worst". What it establishes is an EXISTENCE claim — at τ=6h and 12h the published
|
|
3177
|
+
amplitude fails against a real matchup from this window. At τ=24h and 48h not even that stands. A cell is (team 13 × curve 3), and the simultaneous family
|
|
3178
|
+
is the SUM of each τ's cell bounds times the opponent basket, **(78+78+78+72) × 6 = 1,836**
|
|
3179
|
+
(τ=48h loses one team to the longer warm-up) — the document's claim spans every τ, and the counts differ per τ. A
|
|
3180
|
+
cell is (team × curve) and the two modes are mixed **inside** it by exposure, because that is
|
|
3181
|
+
how the gate is defined. The per-mode figures sit beside it because they show WHERE the effect
|
|
3182
|
+
lives: knockout point estimates run about 1.5× ladder's. Holding each mode to 5pp separately
|
|
3183
|
+
made 38 of 78 look over, and that was the earlier "almost every team, in knockout" conclusion.
|
|
3184
|
+
**It is withdrawn — it was measured against the wrong bar.**
|
|
3185
|
+
|
|
3186
|
+
**Choosing the state by summed deficit was also wrong.** The sum ignores WHICH slot is tired,
|
|
3187
|
+
and influence differs by slot (§3-16's ZP fallback slot is the example) and by mode. Every one
|
|
3188
|
+
of a team's states is now measured on 200 seeds with COMMON RANDOM NUMBERS across states, and
|
|
3189
|
+
the worst weighted edge wins. That choice differed from the summed-deficit pick in **29 of 39**
|
|
3190
|
+
cells.
|
|
3191
|
+
|
|
3192
|
+
**But "we found the worst" is not claimable — that was measured too.** The gap between first
|
|
3193
|
+
and second place is **0.16pp** median (min 0.00), while the gap between first and last is
|
|
3194
|
+
**4.41pp** median (min 0.00). The range is some 27 times the margin, so the screen carries
|
|
3195
|
+
real signal and **only the top is tied** — the difference between what was chosen and the true
|
|
3196
|
+
worst is about the margin, ≈0.2pp, which cannot move a 5pp gate. It also reframes the
|
|
3197
|
+
disagreement with the summed-deficit pick: not "the proxy was badly wrong" but "the top states
|
|
3198
|
+
are tied, so any rule picks one of several equals".
|
|
3199
|
+
|
|
3200
|
+
The selection comes from an outcome, so winner's-curse bias remains. That is why the states are
|
|
3201
|
+
**re-screened at the chosen amplitude** before validation, then measured on independent seeds:
|
|
3202
|
+
freezing the selection at the published amplitude would leave validation unable to see a state
|
|
3203
|
+
that is worse at a much smaller `b`, because the response is nonlinear. The bisection itself
|
|
3204
|
+
runs on the frozen selection as a SEARCH HEURISTIC; only the published verdict is exact at its
|
|
3205
|
+
own amplitude, and if validation fails the amplitude halves and the whole step repeats.
|
|
3206
|
+
|
|
3207
|
+
**Those repeats spend a divided alpha budget.** Up to five independently seeded confirmations
|
|
3208
|
+
that accept the FIRST to pass, each spending the full `0.05/1836`, would inflate the chance of
|
|
3209
|
+
falsely declaring all cells below the gate by as much as five times. Attempt `i` gets
|
|
3210
|
+
`1/2^(i+1)` of the budget, which sums to under one, so the union bound preserves the family-wise
|
|
3211
|
+
rate while the first attempt — usually the only one — pays a factor of two rather than five.
|
|
3212
|
+
The correction moved only the confirmation bound; the safe amplitude itself is unchanged,
|
|
3213
|
+
since the bisection spends the search budget.
|
|
3214
|
+
|
|
3215
|
+
**And the family is not this τ's cells alone.** The document claims a safe amplitude exists for
|
|
3216
|
+
EVERY τ, so spending a separate 5% per τ would let the chance that at least one bound misses
|
|
3217
|
+
run well above 5%. The simultaneous correction is over the SUM of each τ's cell bounds times the basket (**1,836**), counted rather than
|
|
3218
|
+
multiplied — the burn-in is proportional to τ, so a team can have post-warmup states at a short
|
|
3219
|
+
τ and none at a long one, and here τ=48h does lose one. The cell count printed beside the tally
|
|
3220
|
+
stays this τ's, because labelling the family while the tally sums to something else is the next
|
|
3221
|
+
accident.
|
|
3222
|
+
|
|
3223
|
+
**The vector has to be a SYNCHRONOUS state.** An earlier version took each player's maximum
|
|
3224
|
+
over all appearances independently and assembled the eleven — which puts one player's day-3
|
|
3225
|
+
peak beside another's day-9 peak in an XI that **never existed**. That is the same mistake as
|
|
3226
|
+
averaging across timelines, transposed onto the time axis, and it runs the other way: it
|
|
3227
|
+
inflates the vector and overstates the exceedance. In the three-team version, fixing it moved
|
|
3228
|
+
13 → **12** cells over and the worst cell down with it (the safe amplitude `b` was
|
|
3229
|
+
unchanged).
|
|
3230
|
+
|
|
3231
|
+
**The "last lineup" machinery disappeared as a side effect.** The lineup now comes from the
|
|
3232
|
+
same state as the loads, so loads can never be applied to a player who has since left — the
|
|
3233
|
+
mismatch that machinery existed to prevent is now unrepresentable. Each state is the worst
|
|
3234
|
+
OBSERVED one per team and `L₀` (**649** post-warm-up candidates = that team's appearances),
|
|
3235
|
+
chosen on summed deficit, before a single match is played. `b` is a common factor of the
|
|
3236
|
+
deficits so it cannot reorder them — the bisection tests the same state throughout. `L₀` can
|
|
3237
|
+
and does reorder them: `70363723`'s worst kickoff differs at 1×, 2× and 4×.
|
|
3238
|
+
|
|
3239
|
+
Two things had been hiding that. The gate rebuilt the
|
|
3240
|
+
curve per team, which pins each team's MEDIAN deficit to 0.20 — the median column read
|
|
3241
|
+
exactly 0.200 everywhere and I took that for the design working, when it meant no published
|
|
3242
|
+
parameter set was being tested and the most-loaded team's deficit was suppressed. And it ran
|
|
3243
|
+
ladder only; knockout bites more than twice as hard.
|
|
3244
|
+
|
|
3245
|
+
### So what amplitude is safe?
|
|
3246
|
+
|
|
3247
|
+
Exceeding the gate does not mean the mechanism fails; it means the amplitude is too large.
|
|
3248
|
+
Bisect for the largest `b` whose worst cell stays under, measuring EVERY cell at each step
|
|
3249
|
+
rather than assuming linearity — a single proportional guess landed at 5.15, still grazing it.
|
|
3250
|
+
|
|
3251
|
+
| `b` | cell hit / worst | upper | lower | max deficit | verdict |
|
|
3252
|
+
| --- | --- | --- | --- | --- | --- |
|
|
3253
|
+
| 1.3394 | `70363723` 1× | 13.38 | 3.51 | 0.856 | not established |
|
|
3254
|
+
| 0.6697 | `70363723` 1× | 8.66 | -0.12 | 0.428 | not established |
|
|
3255
|
+
| 0.3349 | `70363723` 1× | 5.70 | -1.67 | 0.214 | not established |
|
|
3256
|
+
| 0.1674 | `23baeba2` 1× | 4.84 | -2.28 | 0.083 | **below** |
|
|
3257
|
+
| 0.2511 | `70363723` 1× | 5.28 | -2.09 | 0.160 | not established |
|
|
3258
|
+
| 0.2093 | `30178fc1` 1× | 5.03 | -2.12 | 0.111 | not established |
|
|
3259
|
+
| 0.1884 | `6aec8bd8` 1× | 4.91 | -2.19 | 0.101 | **below** |
|
|
3260
|
+
| 0.1988 | `30178fc1` 1× | 4.97 | -2.12 | 0.098 | **below** |
|
|
3261
|
+
| 0.2041 | `30178fc1` 1× | 5.06 | -2.06 | 0.101 | not established |
|
|
3262
|
+
| 0.2014 | `30178fc1` 1× | 5.14 | -2.02 | 0.107 | not established |
|
|
3263
|
+
| **0.1988** (validated on independent seeds) | `6aec8bd8` 1× | **3.86** | — | 0.106 | **all cells below** |
|
|
3264
|
+
|
|
3265
|
+
**`b` ∈ 0.1988–0.2014**, with all 66 cells established below the gate on independent seeds at
|
|
3266
|
+
the lower end — after re-screening the states at that amplitude — and the upper end being the
|
|
3267
|
+
smallest candidate that failed.
|
|
3268
|
+
|
|
3269
|
+
**Do not read that bracket as precision.** The fail rows are not monotone even though the
|
|
3270
|
+
deficit falls with `b`, and the cell that trips changes between them: each bound is estimated
|
|
3271
|
+
from 6,000 seeds and near 5 the estimation error exceeds the spacing between candidates, so
|
|
3272
|
+
which candidate looks worse is decided by the seeds. The bracket is search granularity, not a
|
|
3273
|
+
confidence interval — which is why validation re-measures at twice the sample (12,000 per opponent) on a
|
|
3274
|
+
separate seed column. Note
|
|
3275
|
+
that the worst cell moves: at the published amplitude it is `L₀`=4×, but at low `b` it is
|
|
3276
|
+
**1×**, because a smaller `L₀` saturates sooner and so produces a larger deficit for the same
|
|
3277
|
+
`b`. Testing only the cell that was worst at the starting amplitude would have missed that. E's 0.2 was measured with a uniform Δ across all eleven, and
|
|
3278
|
+
putting it on the MEDIAN player is what let the top of a real vector (max deficit 0.02–0.88)
|
|
3279
|
+
run past the gate. That value is the worst over all 13 teams that played, but NOT over every appearance: the
|
|
3280
|
+
state comes from two deterministic rules (largest summed deficit, largest single-slot deficit)
|
|
3281
|
+
applied to the 649 post-warm-up candidates, and only those one or two states per (team, curve)
|
|
3282
|
+
carry a bound. The remaining appearances are unmeasured — the same kind of limit the opponent
|
|
3283
|
+
basket carries. What is gone is selection on an OUTCOME; what is not gone is that this is a
|
|
3284
|
+
sample of the observed states, and everything outside this window is unseen entirely.
|
|
3285
|
+
|
|
3286
|
+
**The bound's support constant was wrong, too.** One `ser` term averages the home and away
|
|
3287
|
+
versions of `(win(fresh) − win(tired))`, then multiplies by `100 · 2`. Each difference lies in
|
|
3288
|
+
`[−1, 1]`, so their average still has support `[−1, 1]` and the term's support is **±200**: both
|
|
3289
|
+
orientations can attain the same extreme. The constant once said 100. Halving the support credits
|
|
3290
|
+
an unseen outcome with only half its possible effect, which can print an UNSAFE amplitude as below
|
|
3291
|
+
5pp. With it corrected to 200, a degenerate cell's
|
|
3292
|
+
bound has a floor independent of `b`: at SD = 0 the surviving empirical-Bernstein term is
|
|
3293
|
+
`3·R·ln(3/δ)/n` with `R = 2·SER_MAX = 400` and `δ = α/4` (α split once across the two modes and
|
|
3294
|
+
again across the two sides), `α = 0.05/family`. If that floor exceeds 5 the bisection cannot
|
|
3295
|
+
pass at ANY amplitude — not a finding, a seed shortage — so the probe PRINTS it (2.60pp at
|
|
3296
|
+
6,000 seeds per opponent and family 1,836 in this run). Cut the seeds and it climbs past 5 and
|
|
3297
|
+
the search reports "none found"; without that line one would read it as a result.
|
|
3298
|
+
|
|
3299
|
+
**Student-t was the wrong interval here.** Each seed observation averages the two orientations
|
|
3300
|
+
and therefore takes one of nine values, from `−200` through `200` in 50pp steps. It is not a
|
|
3301
|
+
normal sample, and t coverage is exact only under normality. The tail this section uses is
|
|
3302
|
+
`α = 0.05/1836 ≈ 2.7e-5`; ordinary large-sample
|
|
3303
|
+
accuracy says nothing at that extreme, and an anti-conservative cell interval lets an unsafe
|
|
3304
|
+
amplitude pass. The gate now uses an **empirical Bernstein** bound (Maurer–Pontil), which
|
|
3305
|
+
assumes only boundedness and adapts to the variance — far tighter than Hoeffding here, where
|
|
3306
|
+
almost every observation is 0:
|
|
3307
|
+
|
|
3308
|
+
μ ≤ x̄ + SD·√(2 ln(3/δ)/n) + 3R ln(3/δ)/n, R = support width = 400
|
|
3309
|
+
|
|
3310
|
+
The price is width. At this family, a half-width of 2.9pp needs **12,000 seeds per opponent**;
|
|
3311
|
+
at 6,000 it is 4.9pp, which would leave the 5pp gate no room at all. So the seed budget per
|
|
3312
|
+
opponent went up 4×, the safe amplitude fell from 0.15 to 0.13, and only 3 of 66 cells are
|
|
3313
|
+
established below at the published amplitude. Every one of those is the cost of not borrowing
|
|
3314
|
+
normality.
|
|
3315
|
+
|
|
3316
|
+
It also subsumes the degenerate case: with SD = 0 the second term survives, so the interval
|
|
3317
|
+
never collapses to zero width and the separate rule-of-three guard is gone — that guard had
|
|
3318
|
+
been a second argument resting on its own iid assumption.
|
|
3319
|
+
|
|
3320
|
+
**Degenerate samples no longer need special handling.** Under the old t interval, a cell where
|
|
3321
|
+
every seed returned the same value had SD 0, so the interval collapsed to zero width and any
|
|
3322
|
+
mean at all printed as "established below the gate" — no variance means unmeasurable, not
|
|
3323
|
+
certain, and a small-sample smoke run produced exactly that false clear. A separate
|
|
3324
|
+
rule-of-three bound was bolted on for it. Under empirical Bernstein the range term
|
|
3325
|
+
`3R ln(3/δ)/n` survives at SD = 0, so the interval never collapses in the first place: no
|
|
3326
|
+
guard, no special case, and one fewer argument resting on its own iid assumption.
|
|
3327
|
+
|
|
3328
|
+
### Stage 3 — a multi-wallet feeder is exempt, not merely advantaged
|
|
3329
|
+
|
|
3330
|
+
F-1's side rule writes nothing for a ladder away side that is human; `playoff-finalize.ts`
|
|
3331
|
+
calls `applyResult` for both; `selectOpponentByStrength` draws from the five strength-nearest
|
|
3332
|
+
in-division; the cooldown is initiator-only. Observed at-appearance `L` reaches 55 for an
|
|
3333
|
+
honest player while a team taking the same match count as the away side accumulates zero. That
|
|
3334
|
+
is exemption, not a ratio — bounded only by what bounds the feature, since the prize is 0.59pp at g=0.4 (0.16pp at the τ=24h safe amplitude)
|
|
3335
|
+
and the team-level gap is 3.86pp at the safe amplitude, under the gate.
|
|
3336
|
+
|
|
3337
|
+
### What this does not settle
|
|
3338
|
+
|
|
3339
|
+
Fallback provenance is now bound to (team, index) and cleared the moment it is consumed.
|
|
3340
|
+
Predicate D ("C without the ZP fallback") depends on whether `weightedPick` took the fallback
|
|
3341
|
+
branch; holding that as a bare index and never clearing it let a FIXED actor — an FK kicker,
|
|
3342
|
+
PP1's forward, a penalty taker, none of which go through `weightedPick` — inherit the previous
|
|
3343
|
+
selection's mark and lose a legitimate involvement. There is a real producer: `engine.ts:653`
|
|
3344
|
+
picks a defender on a failed dribble and no `getPoint` ever consumes it. D moved 15.25 →
|
|
3345
|
+
15.27 after the fix; A, B and C were unchanged.
|
|
3346
|
+
|
|
3347
|
+
**The same site then produced something larger.** The mark was left dangling only because that
|
|
3348
|
+
selection was not being COUNTED at all; it is now (see the predicate table). Counting it also
|
|
3349
|
+
consumes the fallback mark at that point, so nothing can be left dangling — the cause is gone
|
|
3350
|
+
rather than a second guard added on top of it.
|
|
3351
|
+
|
|
3352
|
+
That fix created an adjacent defect, and the document is what caught it. Counting actor-side
|
|
3353
|
+
fallbacks gave one counter TWO consumers: predicate D subtracts from `off + def` and needs
|
|
3354
|
+
both, while the ZP-fallback table subtracts from `def` alone and needs only the responder's.
|
|
3355
|
+
Sharing one field made the table subtract actor fallbacks too, putting "slot 1 minus fallback"
|
|
3356
|
+
BELOW the other defenders (0.33 against 0.91) — which does not merely change a number, it
|
|
3357
|
+
breaks that table's conclusion. Splitting into `fbOff`/`fbDef` restored 0.90 against 0.91.
|
|
3358
|
+
How it surfaced matters: the probe disagreed with the document, and the document was right.
|
|
3359
|
+
The reflex was to update the "stale" document to match the new run.
|
|
3360
|
+
|
|
3361
|
+
A team with zero load is not a cell. A team whose only appearance in the window is its first
|
|
3362
|
+
has eleven zero pre-match loads, so the tired side IS the fresh side, every seed returns the
|
|
3363
|
+
same value, and SD is 0 — the degenerate guard then sets that cell's bound to infinity and
|
|
3364
|
+
**one such team would block the amplitude search entirely**. The gate table skipped those
|
|
3365
|
+
states while `allCells` did not. A deficit of exactly zero is not evidence of clearing the
|
|
3366
|
+
gate either, so such teams are excluded outright and the output says how many. None occur in
|
|
3367
|
+
this window.
|
|
3368
|
+
|
|
3369
|
+
The prize-versus-cost ordering AT THE AMPLITUDE THAT WOULD SHIP — it splits at `g=0.4`
|
|
3370
|
+
(−0.43pp [−0.76, −0.10]) but not at the safe amplitude's own deficit (prize 0.16pp, cheapest
|
|
3371
|
+
cost 0.16pp, paired difference −0.00pp [−0.47, 0.46]), and that second one is the decision. The neutral-point fork. Maximum LEGAL role concentration (the
|
|
3372
|
+
replay is what happened, not what the rules permit). Player identity is the snapshot's `actorIds` playerId, keyed GLOBALLY rather than per team —
|
|
3373
|
+
a minted player sold mid-window would otherwise get two chains starting from zero, understating
|
|
3374
|
+
the buyer's vector. A missing actor map or even one player that fails to join is rejected by the
|
|
3375
|
+
replay-completeness gate; there is no name-key fallback. **Two code paths went
|
|
3376
|
+
untested by this data**: no playerId appears under two teams in the window (0 of 143), and
|
|
3377
|
+
only 1 of 754 human team-sides was uncredited — credit now gates the increment alone, since
|
|
3378
|
+
condition is read on entry whether or not load is charged, and ownership moves on every human
|
|
3379
|
+
appearance. Both are right by construction, which is not the same as having been seen to
|
|
3380
|
+
run. All four τ values `{6, 12, 24, 48}h` are integrated separately, with the gate and safe-amplitude
|
|
3381
|
+
search rerun on that τ's own cells rather than converted from 24h. At τ=24h the published run has
|
|
3382
|
+
66 exposure-mixed cells: 0 established above the gate, 3 established below it, and 63 undecided.
|
|
3383
|
+
Snapshot coverage turned out to be a step function — 0% before 2026-08-21 and 100% after — so
|
|
3384
|
+
the window is 11 fully covered days rather than a 37% sample, and the script finds that step
|
|
3385
|
+
itself. The cost is that nothing with a period longer than 11 days is visible.
|
|
3386
|
+
|
|
3387
|
+
### Reproduction
|
|
3388
|
+
|
|
3389
|
+
```
|
|
3390
|
+
# ENGINE_SINCE is OPTIONAL — the probe verifies each match by replaying it against
|
|
3391
|
+
# the stored score, event digest and per-player ratings, so there is no date to infer.
|
|
3392
|
+
C2_ENV_FILE=<main checkout>/.env.deploy DAYS=30 WINDOW_END=2026-09-01 \
|
|
3393
|
+
EXPORT_SNAPSHOTS=/tmp/prod-tl.json \
|
|
3394
|
+
pnpm --filter @sws26/api exec tsx scripts/c2-team-load.ts
|
|
3395
|
+
MATCHES=3000 SNAPSHOTS=/tmp/prod-tl.json REPLAYS=5 TAU_SWEEP=6,12,24,48 \
|
|
3396
|
+
LAMBDA_MED=6 LAMBDA_MAX=21 \
|
|
3397
|
+
npx tsx packages/mcp/skill/reference/probes/engagement-load.mts
|
|
3398
|
+
```
|
|
3399
|
+
|
|
3400
|
+
**The window's end is now stated explicitly, so these figures can be rebuilt.** The default is
|
|
3401
|
+
still "the last completed UTC day", so an unqualified re-export extends it by a day — but
|
|
3402
|
+
`WINDOW_END` makes that date the window's (exclusive) end. The numbers in this section come
|
|
3403
|
+
from **08-21 through 08-31, 1,201 matches**, and re-exporting with `WINDOW_END=2026-09-01`
|
|
3404
|
+
reproduces that `matches` array **byte for byte** (verified) — which came for free from
|
|
3405
|
+
collapsing a window boundary that had been spelled out in eighteen separate places into one
|
|
3406
|
+
definition. The window's length remains a limitation.
|
|
3407
|
+
|
|
3408
|
+
**"Fully covered" cannot be counted from completed rows alone.** A day with rows still
|
|
3409
|
+
unfinished, or whose cup bracket has not completed, is not fully covered — knockout rows are
|
|
3410
|
+
inserted only after the preceding round finishes, so **rounds not yet inserted never appear in
|
|
3411
|
+
the denominator** and counting completed rows makes such a day look like 100%. Admitting it
|
|
3412
|
+
drops the later engagements and biases `L` and the safe amplitude DOWNWARD, the unsafe
|
|
3413
|
+
direction. The unfinished-row count and the open-cup count are now printed per day, and a day
|
|
3414
|
+
counts as fully covered only when both are zero. No day in this window is affected.
|
|
3415
|
+
|
|
3416
|
+
**And that day's own cup has to exist — any completed cup will not do** [Codex P2, 2026-09-02
|
|
3417
|
+
×2; 2026-09-04]. A day whose cup was never created still has ladder and playoff rows, so it
|
|
3418
|
+
passes all four checks above. Requiring one completed cup is still not enough. The first repair
|
|
3419
|
+
still BUILT each group from match kickoff day and merely required the cup-ID date to equal that
|
|
3420
|
+
key. If `SEASON_OPEN_AT` pushes the entire cup into the next UTC day, every `wc-<yesterday>` row
|
|
3421
|
+
sits in today's group and even yesterday's completed cup contributes nothing to yesterday's
|
|
3422
|
+
`done_cups`.
|
|
3423
|
+
|
|
3424
|
+
The cup calendar is now built INDEPENDENTLY from the bracket ID's `wc-YYYY-MM-DD` suffix, and
|
|
3425
|
+
competitive rows are attached to that cup's own date. Thus, even if every fixture kicks off the
|
|
3426
|
+
next day, the `wc-<yesterday>` rows do not invent activity today: they fill yesterday's
|
|
3427
|
+
denominator and snapshot count beside yesterday's completed bracket. The two calendars are
|
|
3428
|
+
`FULL JOIN`ed, so a bracket-only or match-only hole stays visible. A completed bracket with empty
|
|
3429
|
+
`match_ids`, or a day with zero match rows, is not evidence and cannot pass. Only legacy or
|
|
3430
|
+
hand-made IDs that carry no date, plus standalone playoffs, fall back to fixture UTC date like
|
|
3431
|
+
`cupDateOfFixture`. **The export start, claim check, and postcondition share that same competition-
|
|
3432
|
+
day expression**: fixing coverage alone would let the kickoff filter pull a previous cup's
|
|
3433
|
+
after-midnight rows back into the new fully covered window. Neither the verdict nor one byte of
|
|
3434
|
+
the 1,201-match production export changed; only the delayed-schedule boundary did.
|
|
3435
|
+
|
|
3436
|
+
**The window has to close on both clocks** [Codex P2, 2026-09-02; 2026-09-04]. Load events are placed at the
|
|
3437
|
+
claim time while the window was cut on kickoff — so a match kicking off just before midnight and
|
|
3438
|
+
processed just after is dropped from the export while the coverage table still counts that day as
|
|
3439
|
+
100%. That team's engagement vanishes silently and `L` drops, the unsafe direction again. Such a
|
|
3440
|
+
match is no longer counted and reported: the window's end is pulled back a day at a time until
|
|
3441
|
+
both clocks close. Crucially, the trim now groups by COMPETITION DAY and checks that day's maximum
|
|
3442
|
+
kickoff and maximum claim together. Otherwise, pulling the end back to today's midnight for a
|
|
3443
|
+
claim can leave a `wc-<yesterday>` row retained by competition date but silently removed by its
|
|
3444
|
+
today kickoff. If either clock of a retained competition day crosses, the end retreats another
|
|
3445
|
+
day, so a cup stays whole or leaves whole. The postcondition does not pre-filter on the effective
|
|
3446
|
+
kickoff end: it counts `kickoff >= end OR claim >= end`, which makes that omission visible. Neither
|
|
3447
|
+
clock crosses in this production window, so the trim is zero days.
|
|
3448
|
+
|
|
3449
|
+
**The export file is private too** [Codex P2, 2026-09-04]. An engine team inside the snapshot
|
|
3450
|
+
contains the effective `total` from which paid potential can be inferred. Writing
|
|
3451
|
+
`/tmp/prod-tl.json` with Node defaults commonly creates 0644, while merely passing
|
|
3452
|
+
`{ mode: 0o600 }` does not repair an EXISTING 0644 destination. The exporter now writes an
|
|
3453
|
+
exclusive 0600 temporary file in the same directory, applies `chmod(0600)`, and atomically renames
|
|
3454
|
+
it over the destination. Both a new file and an existing-file replacement are owner-only; a failed
|
|
3455
|
+
temporary is removed.
|
|
3456
|
+
|
|
3457
|
+
**Replay verification replaces the engine boundary.** The probe replays every match from its
|
|
3458
|
+
stored seed and compares three things — the **stored score**, the **event-sequence digest**, and
|
|
3459
|
+
the **stored per-player ratings** — then drops the ones that do not reproduce and prints the
|
|
3460
|
+
rate. A rate below 100% means the window crossed an engine deployment — no date to infer, and no
|
|
3461
|
+
`ENGINE_SINCE` needed (given, it is only a pre-filter).
|
|
3462
|
+
A missing event digest, missing ratings, a missing `actorIds` map, or even one of the eleven player
|
|
3463
|
+
cells failing to join makes that evidence axis **unverifiable**, not equal; the other axes cannot admit that row.
|
|
3464
|
+
Rejecting any such row then trips the whole-window completeness guard below, so the probe neither
|
|
3465
|
+
counts absent evidence as agreement nor integrates a window with a missing increment.
|
|
3466
|
+
|
|
3467
|
+
**And a window with holes is not integrated — the probe exits.** If some matches were rejected,
|
|
3468
|
+
their increments would vanish while every later appearance stayed, so those players would start
|
|
3469
|
+
from an artificially low `L` and `Lmed`, the gate vectors and the safe amplitude would all be
|
|
3470
|
+
understated. Engine changes are input-dependent, so interspersed failures are a normal
|
|
3471
|
+
possibility rather than a clean release boundary.
|
|
3472
|
+
|
|
3473
|
+
`LOAD_PRED` selects the integration predicate (default `C`). C is published, and the probe also
|
|
3474
|
+
integrates at D at every τ and prints the difference — F-1 has not chosen, so the calibration
|
|
3475
|
+
has to carry which predicate it is conditional on.
|
|
3476
|
+
|
|
3477
|
+
Without `TAU_SWEEP` only `TAU_H` runs. **Calling a single-τ run a "decision" makes a probe
|
|
3478
|
+
default the production recovery rate** — that is why the sweep exists.
|
|
3479
|
+
|
|
3480
|
+
`REPLAYS` now applies only to the predicate table, where each fixture is replayed
|
|
3481
|
+
independently. The timeline is integrated once, from each match's stored `rng_seed`.
|
|
3482
|
+
|
|
3483
|
+
**"`tsc` does not see these probes" is wrong, and it was checked.** #719 added
|
|
3484
|
+
`skill/reference/probes/**/*.mts` to `packages/mcp/tsconfig.typecheck.json`, and
|
|
3485
|
+
`packages/api/tsconfig.typecheck.json` already includes `scripts`. The local gate covers both
|
|
3486
|
+
files. The opposite claim sat here because a pre-#719 fact was copied forward unverified —
|
|
3487
|
+
not harmless, since believing there is no static gate changes how one writes.
|
|
3488
|
+
|
|
3489
|
+
What is genuinely uncovered is the throwaway EDIT scripts. A backtick terminating a template
|
|
3490
|
+
literal early broke one seven times; the last was a markdown table's `` `70363723` `` breaking
|
|
3491
|
+
the edit script itself. No tsconfig includes those. The remedy is asserting the match count on
|
|
3492
|
+
every replacement, so a broken script dies before it touches the file — which is exactly what
|
|
3493
|
+
happened. A smoke run is a behavioural check, not the parse gate.
|