pog-mcp 0.9.15 → 0.9.17
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/skill/reference/measurements.md +618 -121
package/package.json
CHANGED
|
@@ -664,18 +664,18 @@ sets it — there is no `fair:` outside the engine's own tests, and every engine
|
|
|
664
664
|
input is built from an explicit field list — so it is a global constant in
|
|
665
665
|
practice, and setting it on both squads IS the rule change.
|
|
666
666
|
|
|
667
|
-
**
|
|
668
|
-
|
|
669
|
-
slot"
|
|
670
|
-
|
|
671
|
-
|
|
667
|
+
**The zone-press coefficient waited for a selector.** While every candidate
|
|
668
|
+
weight at that scene was `0.0`, `weightedPick` fell back to "the first non-GK
|
|
669
|
+
slot" and a coefficient in front of an undefined selection measured nothing.
|
|
670
|
+
#730 fixed the selector (`zp-selector.mts` below), so the axis became sweepable
|
|
671
|
+
and it is swept here, over 1.5..4.5.
|
|
672
672
|
|
|
673
673
|
### The statistic, and why not the obvious one
|
|
674
674
|
|
|
675
675
|
Rows below are in **win-equivalent share**: a win is 1, a draw is 0.5, a loss is
|
|
676
676
|
0, and every match counts. The obvious alternative — win rate among DECIDED
|
|
677
677
|
matches — conditions on something the swept constants themselves move. Free-kick
|
|
678
|
-
occurrence takes the regulation draw rate from 57.
|
|
678
|
+
occurrence takes the regulation draw rate from 57.7% to 38.0% across its range,
|
|
679
679
|
so a decided-only percentage mixes "who wins more" with "which matches got
|
|
680
680
|
decided at all". It also inflates the number the activation gate is read
|
|
681
681
|
against: an 11/80/9 record reads 55% decided, clearing a +5pp gate, while the
|
|
@@ -722,7 +722,7 @@ to count rather than a failure:
|
|
|
722
722
|
|
|
723
723
|
| Measured every run | Result |
|
|
724
724
|
| --- | --- |
|
|
725
|
-
| Knockout draws, across both baskets | **0 in
|
|
725
|
+
| Knockout draws, across both baskets | **0 in 10,880,000** knockout matches |
|
|
726
726
|
|
|
727
727
|
That draw count is a claim this section originally got wrong in the other
|
|
728
728
|
direction. It said knockout mode "has no draws by construction". It nearly does:
|
|
@@ -741,9 +741,10 @@ without — a paired test cannot detect misaligned columns from its own output,
|
|
|
741
741
|
a freshly patched shipped-value engine is required to reproduce the default
|
|
742
742
|
column match for match, in order.
|
|
743
743
|
|
|
744
|
-
|
|
745
|
-
matches each,
|
|
746
|
-
|
|
744
|
+
6,800,000 matches in the narrow basket (5 archetypes, 10 fixtures, 10,000
|
|
745
|
+
matches each, 34 conditions, both modes) plus 14,960,000 in the wide one
|
|
746
|
+
(11 archetypes, 55 fixtures, 4,000 matches each, the same 34 conditions).
|
|
747
|
+
**Bonferroni over m = 17,160 pre-registered comparisons, two-sided: |z| >= 4.68.**
|
|
747
748
|
That m prices the SEARCHES, not just the tests. A gate cell does not spend one
|
|
748
749
|
hypothesis, and BOTH of its sides are data-chosen. The baseline side picks a
|
|
749
750
|
leader out of K and builds a band around it — noisy standings could have named
|
|
@@ -770,10 +771,10 @@ power, not a minimum detectable effect. The 80%-power figure adds z(0.80):
|
|
|
770
771
|
|
|
771
772
|
| median / worst | significance threshold | 80%-power effect |
|
|
772
773
|
| --- | --- | --- |
|
|
773
|
-
| narrow, regulation |
|
|
774
|
-
| narrow, knockouts | 1.
|
|
775
|
-
| wide, regulation | 1.
|
|
776
|
-
| wide, knockouts | 2.
|
|
774
|
+
| narrow, regulation | 0.8 / 2.2pp | 0.9 / 2.6pp |
|
|
775
|
+
| narrow, knockouts | 1.5 / 3.1pp | 1.8 / 3.6pp |
|
|
776
|
+
| wide, regulation | 1.5 / 4.0pp | 1.7 / 4.7pp |
|
|
777
|
+
| wide, knockouts | 2.5 / 4.9pp | 3.0 / **5.7pp** |
|
|
777
778
|
|
|
778
779
|
Read the worst column, not the median, since "every test is powered enough" is a
|
|
779
780
|
claim about the tail — and read it honestly: **the worst wide-knockout fixture
|
|
@@ -796,10 +797,12 @@ gate, and says so.
|
|
|
796
797
|
That only settles anything if the band is SHARP — noisy standings widen it, and
|
|
797
798
|
a wide enough band absorbs every new leader, which would be a power limitation
|
|
798
799
|
wearing the costume of a result. So the probe reports the band's own
|
|
799
|
-
discrimination
|
|
800
|
-
ones, with its edge **1.
|
|
801
|
-
|
|
802
|
-
|
|
800
|
+
discrimination at its WORST over the searched weights: it holds **2 of the 5**
|
|
801
|
+
narrow squads and **3 of the 11** wide ones, with its edge **1.51pp** and
|
|
802
|
+
**2.29pp** from the leader. Two of five is not a band that has resolved the
|
|
803
|
+
basket — it is why this section's verdict is UNRESOLVED against a wide band
|
|
804
|
+
rather than "no contest", and it is the number a rerun with a further-apart
|
|
805
|
+
basket would move.
|
|
803
806
|
|
|
804
807
|
### The table
|
|
805
808
|
|
|
@@ -807,8 +810,11 @@ Basket: `role-433`, `def-433`, `def-532`, `shootonly-433`, `flat-433` — five
|
|
|
807
810
|
outfield templates, the shared `squad-lib.mts` builders, so a row here and a row
|
|
808
811
|
in the round-robin above are the same squads.
|
|
809
812
|
|
|
810
|
-
- **`sig`** — fixture deltas clearing the corrected bar, out of
|
|
811
|
-
|
|
813
|
+
- **`sig`** — fixture deltas clearing the corrected bar, out of **10 fixtures x
|
|
814
|
+
that axis's swept values** in the narrow basket and **55 x the same** in the
|
|
815
|
+
wide one. Most axes sweep four, so most rows read out of 40 and 220; set-piece
|
|
816
|
+
`dBonus` sweeps five to cover 3..8 end to end, so its rows read out of 50 and
|
|
817
|
+
275.
|
|
812
818
|
Both are printed, because an axis can be classified on the wide basket's
|
|
813
819
|
evidence alone and a table that hid it would publish a verdict whose decisive
|
|
814
820
|
measurement is invisible.
|
|
@@ -821,28 +827,93 @@ in the round-robin above are the same squads.
|
|
|
821
827
|
|
|
822
828
|
| Axis (shipped) | Mode | sig (narrow) | sig (wide) | rev | Largest narrow delta | gate | What that row leaves FREE |
|
|
823
829
|
| --- | --- | --- | --- | --- | --- | --- | --- |
|
|
824
|
-
| **FK occurrence** `rand(10)<1` | reg |
|
|
825
|
-
| | ko |
|
|
826
|
-
| **Post play** `rand(10)<1.5` | reg |
|
|
827
|
-
| | ko |
|
|
828
|
-
| **Counter** `THRESHOLD=5` | reg |
|
|
829
|
-
| | ko | **0/40** |
|
|
830
|
-
| **`fair`** default 5 | reg | 13/40 |
|
|
831
|
-
| | ko | 7/40 |
|
|
832
|
-
|
|
833
|
-
|
|
834
|
-
|
|
835
|
-
|
|
836
|
-
|
|
837
|
-
|
|
838
|
-
|
|
839
|
-
**
|
|
840
|
-
|
|
841
|
-
|
|
842
|
-
|
|
843
|
-
|
|
844
|
-
|
|
845
|
-
|
|
830
|
+
| **FK occurrence** `rand(10)<1` | reg | 23/40 | 98/220 | 0 | **+11.8pp** `def-532` vs `flat-433` @x4 | none | the FK taker: every basket squad kicks from slot 9, so this is "more free kicks", never "more free kicks AND a better taker" |
|
|
831
|
+
| | ko | 19/40 | 92/220 | 0 | +8.5pp `def-532` vs `flat-433` @x4 | none | shootout exposure — it is a property of the PAIRING, and the basket fixes both halves |
|
|
832
|
+
| **Post play** `rand(10)<1.5` | reg | 16/40 | 52/220 | 0 | −5.0pp `role-433` vs `flat-433` @x4 | none | the PP2 shooter is drawn by `getPlayer`, so the row confounds "more post plays" with "who the draw lands on" |
|
|
833
|
+
| | ko | 7/40 | 33/220 | 0 | −4.1pp `def-532` vs `shootonly-433` @x4 | none | same, plus FW count: 433 and 532 field three and two forwards, and `fwIdx.length` gates the branch |
|
|
834
|
+
| **Counter** `THRESHOLD=5` | reg | 2/40 | 5/220 | 0 | −1.8pp `shootonly-433` vs `flat-433` @0% | none | the four sampled thresholds span 0% to 100%, but changing this constant reroutes RNG consumption, so outcomes need not interpolate between them |
|
|
835
|
+
| | ko | **0/40** | 3/220 | 0 | +1.5pp (below bar) | none | as above. This is the control: it never produced a gate in either basket, which is the only behaviour that matters here |
|
|
836
|
+
| **`fair`** default 5 | reg | 13/40 | 33/220 | **1** | **+11.4pp** `shootonly-433` vs `flat-433` @22 | none | `fair` is uniform across all 22 players here; per-player or per-position `fair` is a different (and unmeasured) axis |
|
|
837
|
+
| | ko | 7/40 | 28/220 | 0 | +10.7pp `shootonly-433` vs `flat-433` @22 | none | keeper `defense`, which every extra penalty routes through — held fixed inside each template |
|
|
838
|
+
| **Set-piece `dBonus`** 5 | reg | 15/50 | 56/275 | 0 | −1.7pp `def-532` vs `flat-433` @8 | none | it widens the DEFENDER's roll at FK1/FK2 only, and the taker is slot 9 in every basket squad. Five swept values, not four: 3..8 end to end with the default column carrying 5 |
|
|
839
|
+
| | ko | 10/50 | 74/275 | 0 | +1.4pp `role-433` vs `shootonly-433` @8 | none | same, plus the shootout, which this axis does not touch (`takePkShot` has its own width) |
|
|
840
|
+
| **Base `dBonus`** 8 | reg | 25/40 | 43/220 | 0 | +2.9pp `def-532` vs `flat-433` @6 | none | the widest-reaching of the seven — it is the defender roll for every non-set-piece contest — and still nothing re-prices |
|
|
841
|
+
| | ko | 14/40 | 41/220 | 0 | +3.7pp `def-532` vs `shootonly-433` @6 | none | as above |
|
|
842
|
+
| **PK attack noise** `rand(16)` | reg | 9/40 | 25/220 | 0 | +0.9pp `role-433` vs `def-532` @8 | none | in regulation it reaches only awarded penalties, which is why the row is nearly flat |
|
|
843
|
+
| | ko | 25/40 | 102/220 | 0 | **+9.7pp** `role-433` vs `def-532` @8 | none | the shootout routes every kick through it, so knockouts feel it and regulation barely does — the largest mode split in the table, and still no gate |
|
|
844
|
+
| **Zone press** `ratio * 3` | reg | **0/40** | **0/220** | 0 | −0.8pp `def-433` vs `flat-433` @4.5 | none | the axis #730 unblocked, and the only one in the file that moves NOTHING: zero significant deltas in either basket, in either mode |
|
|
845
|
+
| | ko | **0/40** | **0/220** | 0 | +1.0pp `def-433` vs `def-532` @1.5 | none | as above. INERT rather than MARGINS — even the counter control clears the bar somewhere |
|
|
846
|
+
|
|
847
|
+
**Nothing here becomes a candidate.** Fourteen of the sixteen rows move margins
|
|
848
|
+
somewhere — the control included, once the wide basket is counted — and the one
|
|
849
|
+
pairing that re-prices anything does so against a squad that was already tied for
|
|
850
|
+
best. The two that move nothing at all are zone press, in both modes. The pass
|
|
851
|
+
set at the plan's own activation gate is empty over both baskets.
|
|
852
|
+
|
|
853
|
+
**The pool is now measured end to end, and nothing in it clears the gate.** An
|
|
854
|
+
earlier run covered only four
|
|
855
|
+
axes and said so: schedule none of those four, and go measure the other three
|
|
856
|
+
before concluding anything about the pool. Those three — set-piece `dBonus`,
|
|
857
|
+
base `dBonus` and penalty attack noise — are in the table above and land where
|
|
858
|
+
the first four did.
|
|
859
|
+
|
|
860
|
+
And the pool is now COMPLETE: the zone-press coefficient, the one parameter the
|
|
861
|
+
plan named but forbade "until #730 decides the selector", is swept here too — the
|
|
862
|
+
selector was decided, so the prohibition lapsed. It comes back INERT, the only
|
|
863
|
+
axis in the file that moves no fixture in either basket in either mode.
|
|
864
|
+
|
|
865
|
+
Two axes are covered END TO END, because the plan gives them ranges and the sweep
|
|
866
|
+
spans them: set-piece `dBonus` over 3..8 (five swept values plus the default 5)
|
|
867
|
+
and base `dBonus` over 6..10 (four plus the default 8). **Penalty attack noise
|
|
868
|
+
and zone press are not**: §8 3.1 names their constants — `rand(16)` and
|
|
869
|
+
`ratio * 3` — and no range, so this run picked ±25% and ±50% around each and the
|
|
870
|
+
conclusions are bounded to 8..24 and 1.5..4.5. A rule proposing a value outside
|
|
871
|
+
those spans is unmeasured, and these do not interpolate — changing them reroutes
|
|
872
|
+
the RNG stream. **No axis produced a candidate that clears the gate: the largest
|
|
873
|
+
gain reachable at ANY exposure weight, across all eight axes, both baskets and
|
|
874
|
+
20,146 solved interior weights, is 0.00pp against a 5pp threshold.**
|
|
875
|
+
|
|
876
|
+
**And that is a failure to distinguish, not a demonstration that nothing is
|
|
877
|
+
there.** The distinction is the whole reading of this section, so it is worth
|
|
878
|
+
being exact about. A gate cell is scored by pairing the swept optimum against the
|
|
879
|
+
stale one; every one of the 99 cells per basket formed NO PAIR, because the
|
|
880
|
+
swept optimum was already inside the default band. A cell that forms no pair
|
|
881
|
+
contributes a zero gain, and 99 zeros is what produces the 0.00pp above — so the
|
|
882
|
+
headline number is the ABSENCE of a resolvable contest, not the presence of a
|
|
883
|
+
measured null. Whether an axis re-prices the decision is therefore **UNRESOLVED**
|
|
884
|
+
on this evidence.
|
|
885
|
+
|
|
886
|
+
What would resolve it is a sharper band. At its worst the default band holds 2 of
|
|
887
|
+
the 5 narrow squads, with the edge 1.51pp from the leader, and 3 of 11 in the wide
|
|
888
|
+
basket at 2.29pp: a band that wide absorbs any new leader before the gate can
|
|
889
|
+
weigh it. A basket whose members sit further apart would shrink it. A lower bar
|
|
890
|
+
would not — that would only relabel the same inability.
|
|
891
|
+
|
|
892
|
+
One thing this is definitely NOT: "these constants do nothing". PK attack noise
|
|
893
|
+
moves a knockout fixture by 9.7pp and base `dBonus` moves a regulation one by
|
|
894
|
+
2.9pp; margins move everywhere, and the wide basket below changes the top build
|
|
895
|
+
in 12 of 33 knockout cells. What none of that reaches is the gate.
|
|
896
|
+
|
|
897
|
+
**Scope, stated once and plainly.** Every row above is measured at `cond` 5 on
|
|
898
|
+
both sides, and production has not been that since #794 turned the form system
|
|
899
|
+
on. `getPoint` reads condition in both its individual and its team term, so a
|
|
900
|
+
swept threshold can interact with it; this run cannot see that interaction. So
|
|
901
|
+
the claim this section supports is **no pool axis clears the gate under neutral
|
|
902
|
+
conditions**, not "under production conditions" — closing that gap means
|
|
903
|
+
re-running the pool over representative frozen condition states, which is a
|
|
904
|
+
separate measurement and is not done here.
|
|
905
|
+
|
|
906
|
+
**What that means for 3.1.** The pool produced no schedulable candidate, and
|
|
907
|
+
this run cannot say whether that is because none exists or because its basket
|
|
908
|
+
cannot resolve one. So 3.1's premise — that a weekly rule rotates which build is
|
|
909
|
+
worth fielding — is **not supported by anything measured here, and not refuted
|
|
910
|
+
either.** Two measurements would move it: a basket whose members sit further
|
|
911
|
+
apart, which is what the unformed gates are asking for, and a rerun over
|
|
912
|
+
production's condition states. Building the feature on this evidence would be
|
|
913
|
+
building on an unresolved gate; shelving it permanently on this evidence would
|
|
914
|
+
be over-reading the same gate in the other direction. That is a product call
|
|
915
|
+
rather than a
|
|
916
|
+
measurement one, and it is recorded on #733.
|
|
846
917
|
|
|
847
918
|
The verdict is deliberately absent from the rows above: it belongs to the AXIS,
|
|
848
919
|
not to a mode. The next section is where it is issued.
|
|
@@ -899,7 +970,7 @@ nothing, and no amount of simulation shrinks that. Nor can the difference be
|
|
|
899
970
|
paired away: the leader and its weakest opponent are argmax and argmin picks
|
|
900
971
|
that can change IDENTITY between the two conditions, so there is no fixed pair
|
|
901
972
|
to difference. What this probe does instead is a 2-sigma DIRECTION screen at
|
|
902
|
-
+3.0. It refuses far more than a corrected test would — that would sit near 6.
|
|
973
|
+
+3.0. It refuses far more than a corrected test would — that would sit near 6.6
|
|
903
974
|
— which is the safe direction for a gate whose job is to refuse. It is a
|
|
904
975
|
heuristic, it says so in the probe's own legend, and it is not the plan's
|
|
905
976
|
sentence. The plan does not put numbers on the feel bands, so the ones
|
|
@@ -917,10 +988,14 @@ new leader drawn FROM the band scores zero by construction.
|
|
|
917
988
|
|
|
918
989
|
| Axis | Best value | Gate (ko weight 0.081) | dom | domAbs | draw | goals | Blocked by |
|
|
919
990
|
| --- | --- | --- | --- | --- | --- | --- | --- |
|
|
920
|
-
| FK occurrence | x0.25 | **0.00pp** | +0.
|
|
921
|
-
| Post play | x0.25 | **0.00pp** |
|
|
922
|
-
| Counter (control) | 0% | **0.00pp** | +0.
|
|
923
|
-
| `fair` | 1 | **0.00pp** |
|
|
991
|
+
| FK occurrence | x0.25 | **0.00pp** | +0.1 | 0.3 | **+5.0pp** | −16% | gain — and it would also fail the draw band |
|
|
992
|
+
| Post play | x0.25 | **0.00pp** | +0.7 | 1.1 | −1.1pp | +3% | gain |
|
|
993
|
+
| Counter (control) | 0% | **0.00pp** | +0.4 | 0.8 | +1.4pp | −6% | gain |
|
|
994
|
+
| `fair` | 1 | **0.00pp** | −0.2 | 0.2 | +3.7pp | −12% | gain |
|
|
995
|
+
| Set-piece `dBonus` | 3 | **0.00pp** | +0.6 | 0.2 | −2.4pp | +8% | gain |
|
|
996
|
+
| Base `dBonus` | 6 | **0.00pp** | −0.2 | 0.2 | **−5.1pp** | +22% | gain |
|
|
997
|
+
| PK attack noise | 8 | **0.00pp** | +1.1 | 0.1 | +1.1pp | −4% | gain |
|
|
998
|
+
| Zone press | x0.5 | **0.00pp** | +1.1 | 0.2 | −1.3pp | +5% | gain — and it is INERT, not MARGINS |
|
|
924
999
|
|
|
925
1000
|
`dom` is the CHANGE in the leader's weakest margin; `domAbs` is that margin
|
|
926
1001
|
itself, in the eleven-archetype basket. Both block, and the two are not
|
|
@@ -928,7 +1003,7 @@ interchangeable: a rule that leaves an already-dominant squad exactly where it
|
|
|
928
1003
|
stood scores a `dom` near zero while failing the requirement outright, which is
|
|
929
1004
|
why the absolute column exists at all. `domAbs` is the one carrying the plan's
|
|
930
1005
|
sentence — at or past the corrected bar, the same |z| the rest of this section
|
|
931
|
-
uses, 4.
|
|
1006
|
+
uses, 4.68 at these sample sizes, one squad significantly beats every other,
|
|
932
1007
|
which is what a dominant build is. `dom` is the direction screen described
|
|
933
1008
|
above, at a threshold this probe declares rather than quotes.
|
|
934
1009
|
|
|
@@ -957,6 +1032,10 @@ table is not four verdicts.)
|
|
|
957
1032
|
| Post play | 0.00pp | 0.00pp | 0.00pp |
|
|
958
1033
|
| Counter (control) | 0.00pp | 0.00pp | 0.00pp |
|
|
959
1034
|
| `fair` | 0.00pp | 0.00pp | 0.00pp |
|
|
1035
|
+
| Set-piece `dBonus` | 0.00pp | 0.00pp | 0.00pp |
|
|
1036
|
+
| Base `dBonus` | 0.00pp | 0.00pp | 0.00pp |
|
|
1037
|
+
| PK attack noise | 0.00pp | 0.00pp | 0.00pp |
|
|
1038
|
+
| Zone press | 0.00pp | 0.00pp | 0.00pp |
|
|
960
1039
|
|
|
961
1040
|
Those three points do not settle the interior on their own: each squad's
|
|
962
1041
|
weighted score is linear in the weight, the leader is their upper envelope, and
|
|
@@ -979,11 +1058,11 @@ on its own grid answers a different question. Evaluating every boundary, and
|
|
|
979
1058
|
inside every resulting interval its midpoint plus both ONE-SIDED limits (the
|
|
980
1059
|
gain jumps where the leader changes, so a shared endpoint reports the wrong
|
|
981
1060
|
squad), plus the weights where the two baskets' gain lines cross — that is where
|
|
982
|
-
`min(narrow, wide)` peaks when their slopes oppose — comes to
|
|
1061
|
+
`min(narrow, wide)` peaks when their slopes oppose — comes to 20,146 probe
|
|
983
1062
|
weights across every value:
|
|
984
1063
|
|
|
985
1064
|
> **The largest gain reachable at ANY exposure weight is 0.00pp**, against a 5pp
|
|
986
|
-
> threshold. At all
|
|
1065
|
+
> threshold. At all 20,146 probe weights the swept optimum was already inside the
|
|
987
1066
|
> default band in at least one basket.
|
|
988
1067
|
|
|
989
1068
|
That statement deliberately carries no significance claim, and does not need
|
|
@@ -1017,8 +1096,8 @@ on every run instead of implying it did.
|
|
|
1017
1096
|
|
|
1018
1097
|
**And the near miss is not a near miss.** FK x0.25 fails the gain gate outright —
|
|
1019
1098
|
0.00pp, because the build it promotes was already inside the default band — and
|
|
1020
|
-
it would fail the draw band too, at +5.
|
|
1021
|
-
is fine on the weighted season (+0.
|
|
1099
|
+
it would fail the draw band too, at +5.0pp against a +/-5pp limit. Its dominance
|
|
1100
|
+
is fine on the weighted season (+0.1 change, 0.3 absolute), which is worth
|
|
1022
1101
|
stating precisely: an earlier version of this section read dominance off the two
|
|
1023
1102
|
pure modes and reported +4.4, and that was an artefact of not weighting. The
|
|
1024
1103
|
candidate that looked closest to shippable still fails, but on the gain and the
|
|
@@ -1027,8 +1106,9 @@ feel band, not on concentration.
|
|
|
1027
1106
|
### What the control calibrates
|
|
1028
1107
|
|
|
1029
1108
|
The counter trigger is in this sweep to check that the bar rejects things, and
|
|
1030
|
-
under a paired test it is **not** inert:
|
|
1031
|
-
bar
|
|
1109
|
+
under a paired test it is **not** inert: 2 of its 80 narrow fixture deltas clear
|
|
1110
|
+
the bar (both in regulation; knockouts give 0 of 40), and 8 of 440 across the
|
|
1111
|
+
wide basket's two modes. They are small — and no
|
|
1032
1112
|
value of it ever produced a gate, in either basket, at any exposure weight. That
|
|
1033
1113
|
is the useful result:
|
|
1034
1114
|
|
|
@@ -1038,22 +1118,34 @@ is the useful result:
|
|
|
1038
1118
|
|
|
1039
1119
|
It also supplies a floor. In the NARROW basket — same N, same ten fixtures on
|
|
1040
1120
|
both sides, the only comparison here that is apples to apples — the largest
|
|
1041
|
-
movement this known non-lever produces is **
|
|
1121
|
+
movement this known non-lever produces is **1.8pp**. Anything a candidate does
|
|
1042
1122
|
that is not comfortably above that is not distinguishable from what a knob
|
|
1043
1123
|
nobody would ship already does:
|
|
1044
1124
|
|
|
1045
|
-
| Axis | Mode | Largest fixture move | vs the control's
|
|
1125
|
+
| Axis | Mode | Largest fixture move | vs the control's 1.8pp |
|
|
1046
1126
|
| --- | --- | --- | --- |
|
|
1047
|
-
| FK occurrence | reg | 11.8pp | **6.
|
|
1048
|
-
| FK occurrence | ko | 8.
|
|
1049
|
-
| `fair` | reg |
|
|
1050
|
-
| `fair` | ko | 10.
|
|
1051
|
-
|
|
|
1052
|
-
| Post play |
|
|
1053
|
-
|
|
1054
|
-
|
|
1055
|
-
|
|
1056
|
-
|
|
1127
|
+
| FK occurrence | reg | 11.8pp | **6.6x** |
|
|
1128
|
+
| FK occurrence | ko | 8.5pp | **4.8x** |
|
|
1129
|
+
| `fair` | reg | 11.4pp | **6.4x** |
|
|
1130
|
+
| `fair` | ko | 10.7pp | **6.0x** |
|
|
1131
|
+
| PK attack noise | ko | 9.7pp | **5.5x** |
|
|
1132
|
+
| Post play | reg | 5.0pp | 2.8x |
|
|
1133
|
+
| Post play | ko | 4.1pp | 2.3x |
|
|
1134
|
+
| Base `dBonus` | ko | 3.7pp | 2.1x |
|
|
1135
|
+
| Base `dBonus` | reg | 2.9pp | 1.6x |
|
|
1136
|
+
| Set-piece `dBonus` | reg | 1.7pp | **0.9x — inside the control's own noise** |
|
|
1137
|
+
| Set-piece `dBonus` | ko | 1.4pp | **0.8x — inside the control's own noise** |
|
|
1138
|
+
| PK attack noise | reg | 0.9pp | **0.5x — inside the control's own noise** |
|
|
1139
|
+
| Zone press | ko | 1.0pp | **0.6x — inside the control's own noise** |
|
|
1140
|
+
| Zone press | reg | 0.8pp | **0.5x — inside the control's own noise** |
|
|
1141
|
+
|
|
1142
|
+
The probe flags a row as indistinguishable from the control below **1.5x**.
|
|
1143
|
+
Five of the fourteen candidate rows are below it — the control is the floor, so
|
|
1144
|
+
it is not one of them — and all five belong to the axes this run added: set-piece
|
|
1145
|
+
`dBonus` in both modes, PK noise in regulation, and zone press in both modes move
|
|
1146
|
+
a fixture LESS than a knob nobody would ship. PK noise in knockouts is the
|
|
1147
|
+
opposite case — 5.5x, fourth largest behind FK regulation at 6.6x and `fair` at
|
|
1148
|
+
6.4x and 6.0x — which is the shootout, and still no gate.
|
|
1057
1149
|
|
|
1058
1150
|
The control is swept in BOTH baskets — that is what the gate needs — but this
|
|
1059
1151
|
ratio is deliberately narrow-only on both sides. It compares a MAXIMUM over a
|
|
@@ -1065,43 +1157,53 @@ these ratios by a factor of two between runs that changed no measurement.
|
|
|
1065
1157
|
### Per axis, what actually moved
|
|
1066
1158
|
|
|
1067
1159
|
**FK occurrence is the strongest axis, and what it does is CONCENTRATE.** From
|
|
1068
|
-
2.5% to 40% the regulation draw rate runs 57.
|
|
1069
|
-
goals per match 0.
|
|
1160
|
+
2.5% to 40% the regulation draw rate runs 57.7% -> 52.6% (shipped) -> 38.0% and
|
|
1161
|
+
goals per match 0.658 -> 0.787 -> 1.290. Every large delta at x4 has the same
|
|
1070
1162
|
shape: the differentiated squads pull away from `flat-433` (`role-433` vs
|
|
1071
|
-
`flat-433` 68.
|
|
1072
|
-
|
|
1073
|
-
|
|
1163
|
+
`flat-433` 68.1% of the win-equivalent share at the shipped rate, +11.4pp of it
|
|
1164
|
+
at x4). Turning it DOWN
|
|
1165
|
+
compresses the field instead — `flat-433`'s own share rises 0.335 -> 0.368 while
|
|
1166
|
+
every other squad's falls. So the axis is a skill-expression dial:
|
|
1074
1167
|
more free kicks means more of the match decided by who invested, less means more
|
|
1075
1168
|
of it decided by nothing. Turning it down therefore costs on both feel gates at
|
|
1076
1169
|
once — more draws AND fewer goals — and in knockouts alone it also raises the
|
|
1077
|
-
leader's weakest margin from z
|
|
1170
|
+
leader's weakest margin from z 5.2 to 6.6, which points at concentration rather
|
|
1078
1171
|
than rotation. Read that as the knockout-only reading it is: on the
|
|
1079
1172
|
exposure-weighted season the same change is +0.4, well inside the direction
|
|
1080
1173
|
screen, and the activation table above reports the weighted number.
|
|
1081
1174
|
|
|
1082
1175
|
**`fair` is the only axis that ROTATES demand rather than amplifying it.**
|
|
1083
|
-
Raising it from
|
|
1084
|
-
|
|
1085
|
-
`def-433` stand still — more fouls means
|
|
1086
|
-
those are converted by `shoot`. It carries the
|
|
1087
|
-
vs `def-433` goes from **47.
|
|
1088
|
-
`fair 22`. Read that precisely — the DELTA is significant, the new level is a
|
|
1176
|
+
Raising it from the shipped 22.7% to 100% moves `shootonly-433`'s
|
|
1177
|
+
win-equivalent share from 0.451 to 0.498 in regulation and 0.456 to 0.505 in
|
|
1178
|
+
knockouts, while `role-433` and `def-433` stand still or fall — more fouls means
|
|
1179
|
+
more free kicks and penalties, and those are converted by `shoot`. It carries the
|
|
1180
|
+
run's single reversal: `role-433` vs `def-433` goes from **47.2%** at the shipped
|
|
1181
|
+
default to **50.3%** at `fair 22`. Read that precisely — the DELTA is significant, the new level is a
|
|
1089
1182
|
coin flip. The rule turned "the defensive lean is better" into "there is no
|
|
1090
1183
|
difference", which is a re-pricing; it did not turn it into "the attacking lean
|
|
1091
1184
|
is better".
|
|
1092
1185
|
|
|
1093
1186
|
`fair` is also the only axis whose regulation dominance column does not move
|
|
1094
1187
|
(+0.0). At the midpoint (`fair 11`, 50%) it still moves four fixtures
|
|
1095
|
-
significantly while taking the draw rate DOWN, 52.
|
|
1096
|
-
0.
|
|
1188
|
+
significantly while taking the draw rate DOWN, 52.6% -> 48.4%, and goals up,
|
|
1189
|
+
0.787 -> 0.903 — the one cell in the whole sweep whose feel side-effects point
|
|
1097
1190
|
in a direction anyone would ask for.
|
|
1098
1191
|
|
|
1099
|
-
**
|
|
1100
|
-
|
|
1101
|
-
|
|
1102
|
-
|
|
1103
|
-
|
|
1104
|
-
|
|
1192
|
+
**Set-piece `dBonus` is the weakest axis that moves anything at all.** Zone press
|
|
1193
|
+
is weaker and moves nothing: it is the only INERT row here, 0 significant deltas
|
|
1194
|
+
in either basket. Set-piece `dBonus` does clear the bar in 15 narrow and 56 wide
|
|
1195
|
+
regulation cells, but its largest move is 0.9x the control's in regulation and
|
|
1196
|
+
0.8x in knockouts — a knob that gates 3.9% of goals moves a fixture MORE than it
|
|
1197
|
+
does. The feel barely registers either —
|
|
1198
|
+
its regulation draw rate spans 50.2% to 55.9% across the whole 3..8 range — yet
|
|
1199
|
+
it still changes the top of the
|
|
1200
|
+
eleven-archetype knockout table in two of its ten wide-basket cells, which is
|
|
1201
|
+
the whole lesson of the wide basket restated: moving a table's top is cheap,
|
|
1202
|
+
and moving the decision is what the gate asks about.
|
|
1203
|
+
|
|
1204
|
+
**Post play needs its whole range to reach 2.8x.** It gets there only at x4
|
|
1205
|
+
(60% of box entries), and the feel barely registers on the way (draw rate
|
|
1206
|
+
+4.4pp, goals −0.096).
|
|
1105
1207
|
|
|
1106
1208
|
**The counter trigger never gates.** Swept from "counters never happen" to
|
|
1107
1209
|
"every eligible turnover becomes one", goals per match moved 0.036, the basket's
|
|
@@ -1117,8 +1219,7 @@ The narrow basket's biggest weakness is that five archetypes may simply not
|
|
|
1117
1219
|
contain the alternative a rule rewards. So EVERY swept value of every axis — the
|
|
1118
1220
|
control included, so it is checked on the same surface as the candidates — is
|
|
1119
1221
|
re-asked over the ELEVEN archetypes the round-robin publishes, 55 fixtures,
|
|
1120
|
-
4,000 matches each. That is 34
|
|
1121
|
-
is simple enough to state:
|
|
1222
|
+
4,000 matches each. That is 68 cells, 34 per mode; the probe prints them all.
|
|
1122
1223
|
|
|
1123
1224
|
Standings here are the **win-equivalent share** — the same 1 / 0.5 / 0 statistic
|
|
1124
1225
|
the gate measures, so the leader they name is the squad the gate is then
|
|
@@ -1126,26 +1227,47 @@ evaluated on. Ranking by league points instead is not a monotonic
|
|
|
1126
1227
|
transformation of it when draw rates differ, and these constants move draw rates
|
|
1127
1228
|
by fifteen points.
|
|
1128
1229
|
|
|
1129
|
-
| Mode | Shipped rule's top four | Last |
|
|
1230
|
+
| Mode | Shipped rule's top four | Last | Swept cells whose top changed |
|
|
1130
1231
|
| --- | --- | --- | --- |
|
|
1131
|
-
| reg | `gkmin-433` .
|
|
1132
|
-
| ko | `role-433` .
|
|
1133
|
-
|
|
1134
|
-
|
|
1135
|
-
|
|
1136
|
-
|
|
1137
|
-
|
|
1138
|
-
|
|
1139
|
-
|
|
1140
|
-
|
|
1141
|
-
|
|
1142
|
-
|
|
1143
|
-
|
|
1144
|
-
|
|
1145
|
-
|
|
1146
|
-
|
|
1147
|
-
|
|
1148
|
-
|
|
1232
|
+
| reg | `gkmin-433` .601, `def-532` .586, `def-433` .586, `role-433` .557 | `flat-433` .365 | **1 of 33** (`fair 22`, and not significantly) |
|
|
1233
|
+
| ko | `role-433` .603, `gkheavy-433` .597, `def-433` .595, `def-532` .572 | `flat-433` .258 | **12 of 33**, two of them significant |
|
|
1234
|
+
|
|
1235
|
+
Regulation is almost immovable: `gkmin-433` leads 32 of the 33 swept cells, and
|
|
1236
|
+
the one exception (`fair 22`, where `def-532` reads .592 against `gkmin-433`'s
|
|
1237
|
+
.590) is not
|
|
1238
|
+
significant. `flat-433` is last in all 68 cells.
|
|
1239
|
+
|
|
1240
|
+
**Knockouts are where the newly swept axes show up.** Twelve cells change the
|
|
1241
|
+
top build, and the two that do so significantly are `fk x0.25` (+9.8pp, z 6.2)
|
|
1242
|
+
and `pknoise 8` (+11.4pp, z 7.2), both promoting `gkheavy-433` over `role-433`;
|
|
1243
|
+
`pknoise 12` (+7.3pp, z 4.7) misses the corrected bar. Base `dBonus` and set-piece `dBonus` change the
|
|
1244
|
+
top in six more cells, none significantly.
|
|
1245
|
+
|
|
1246
|
+
**Read the width axes as mean AND variance, not variance alone.** `randFloat(w)`
|
|
1247
|
+
is uniform on [0, w), so its mean is w/2: taking base `dBonus` from 8 to 10 does
|
|
1248
|
+
not just widen the defender's roll, it raises that defender's expected score from
|
|
1249
|
+
4 to 5, and narrowing it to 6 lowers the mean to 3. The kicker's PK roll moves the
|
|
1250
|
+
same way in the other direction. So the pattern — `fk 0.25`/`0.5`, `pknoise 8`/`12`
|
|
1251
|
+
and `dbonus 9`/`10` promoting the 29-point keeper, while `dbonus 6`/`7` and
|
|
1252
|
+
`pknoise 24` hand `def-433` the lead — is consistent with "whose contest is made
|
|
1253
|
+
decisive" AND with a straight shift in who is favoured on average, and this run
|
|
1254
|
+
cannot separate the two. A rule designer reading a mechanism off these rows would
|
|
1255
|
+
be guessing; separating them needs rolls centred on their shipped mean, which is
|
|
1256
|
+
a different experiment from the one the plan's parameter names.
|
|
1257
|
+
|
|
1258
|
+
**None of it produces a gate.** The promoted build was already inside the
|
|
1259
|
+
default band every time — that is why the gate column reads 0.00pp for all seven
|
|
1260
|
+
axes while this table moves as much as it does. Changing which build is at the
|
|
1261
|
+
TOP of a table is not the same as re-pricing the DECISION, and this section is
|
|
1262
|
+
the clearest demonstration of the difference in the file.
|
|
1263
|
+
|
|
1264
|
+
**The FK direction is the pair worth reading in full**, because the same
|
|
1265
|
+
promotion arrives three separate ways and this is the one with a shipped-rule
|
|
1266
|
+
baseline to compare against:
|
|
1267
|
+
|
|
1268
|
+
> `gkheavy-433` vs `role-433`, knockout: **50.9% (z +1.1) at the shipped FK rate
|
|
1269
|
+
> -> 54.9% (z +6.2) at FK x0.25.** A coin flip becomes a significant win. Those
|
|
1270
|
+
> are shares; a win-rate advantage is exactly twice the share advantage over 50%.
|
|
1149
1271
|
|
|
1150
1272
|
Mechanically that is coherent: fewer free kicks means fewer set-piece goals,
|
|
1151
1273
|
more matches level at full time, more shootouts, and the 29-point keeper the
|
|
@@ -1169,31 +1291,32 @@ since none of these constants is scheduled.
|
|
|
1169
1291
|
|
|
1170
1292
|
And it is worth **0.0pp at the activation gate** — in the mode with the LEAST
|
|
1171
1293
|
exposure, since knockouts are cup-only and the cup runs once a day for the
|
|
1172
|
-
ladder's top 48. Two things reduce it from +
|
|
1294
|
+
ladder's top 48. Two things reduce it from +9.8pp to nothing, and both are in
|
|
1173
1295
|
the activation section above. `gkheavy-433` is inside the shipped rule's
|
|
1174
1296
|
unresolved top band, so a manager could already have been on it for free. And
|
|
1175
|
-
the same cell moves the draw rate +5.
|
|
1176
|
-
is allowed to move it, so even a real gain there would have been
|
|
1297
|
+
the same cell moves the draw rate +5.0pp, at the edge of the +/-5pp band a rule
|
|
1298
|
+
change is allowed to move it, so even a real gain there would have been
|
|
1299
|
+
argued about before it shipped.
|
|
1177
1300
|
|
|
1178
1301
|
Incidentally, the shipped rows above are an 11-archetype round-robin at 4,000
|
|
1179
1302
|
matches per fixture — roughly seven times the 300 seeds behind the published
|
|
1180
1303
|
table — and they agree with it where that table says it is decidable:
|
|
1181
1304
|
`gkmin-433` clear at the top of regulation, `role-433` at the top of knockouts,
|
|
1182
|
-
`flat-433` at the bottom of
|
|
1305
|
+
`flat-433` at the bottom of every one of the 68 cells. The middle band still reorders between
|
|
1183
1306
|
the two runs, as that section says it must — and note that the bottom is where
|
|
1184
1307
|
the metric matters: on a points table `stars-433` is above `flat-433`, and on
|
|
1185
1308
|
the win-equivalent share the two swap in a couple of cells. The shipped `role-433` vs `def-433`
|
|
1186
1309
|
fixture also independently reproduces the Skill's "solidity beats aggression"
|
|
1187
|
-
pair at 10,000 matches: the attacking lean takes 47.
|
|
1188
|
-
share in regulation and
|
|
1310
|
+
pair at 10,000 matches: the attacking lean takes 47.2% of the win-equivalent
|
|
1311
|
+
share in regulation and 52.6% in knockouts.
|
|
1189
1312
|
|
|
1190
|
-
**Do not read that 47.
|
|
1313
|
+
**Do not read that 47.2% against the Skill's 56.5%** — they are different
|
|
1191
1314
|
statistics on the same fact. The Skill quotes DECIDED matches, this section
|
|
1192
1315
|
quotes the win-equivalent share, and half of all regulation matches are draws,
|
|
1193
1316
|
so the share is compressed toward 50 exactly as the round-robin correction above
|
|
1194
|
-
warns. The knockout figures (
|
|
1195
|
-
because knockout draws are vanishingly rare — zero in the
|
|
1196
|
-
this run played, though not structurally impossible.
|
|
1317
|
+
warns. The knockout figures (52.6% here, 53.6% there) are directly comparable,
|
|
1318
|
+
because knockout draws are vanishingly rare — zero in the 10,880,000 knockout
|
|
1319
|
+
matches this run played, though not structurally impossible.
|
|
1197
1320
|
|
|
1198
1321
|
### Degrees of freedom this sweep LEFT free
|
|
1199
1322
|
|
|
@@ -1211,12 +1334,18 @@ uncontrolled degree of freedom each time. So, explicitly:
|
|
|
1211
1334
|
and the zone press, the team-contribution terms and possession all read it.
|
|
1212
1335
|
- **Set-piece slots.** `squad-lib` pins the FK taker to slot 9 and the penalty
|
|
1213
1336
|
taker to slot 10 for every candidate. The FK rows are the RATE axis only.
|
|
1214
|
-
- **Slot order.** Identical across the basket
|
|
1215
|
-
fallback
|
|
1337
|
+
- **Slot order.** Identical across the basket. It used to carry a confound —
|
|
1338
|
+
the zone-press fallback made the first non-GK slot a hidden lever — and #730
|
|
1339
|
+
removed it, so the zone-press rows here are not contaminated by it. Held
|
|
1340
|
+
constant either way, and not measured.
|
|
1216
1341
|
- **Interactions.** One axis moves at a time. Nothing here says what FK x4 does
|
|
1217
1342
|
while `fair` is 22, and a pool built from two axes at once is untested.
|
|
1218
|
-
- **`cond`.** Fixed at 5 everywhere,
|
|
1219
|
-
|
|
1343
|
+
- **`cond`.** Fixed at 5 everywhere, and production no longer is: #794 turned
|
|
1344
|
+
the form system on, so competitive matches now arrive with per-player
|
|
1345
|
+
conditions derived and frozen at claim time. `getPoint` reads `cond` in both
|
|
1346
|
+
the individual and the team term, so a swept threshold can interact with it,
|
|
1347
|
+
and **every row here is a neutral-condition row.** That is the single largest
|
|
1348
|
+
scope limit on this section's conclusion — see the verdict, which states it.
|
|
1220
1349
|
- **Opponent population.** Archetypes, not the live ladder — and production bots
|
|
1221
1350
|
are a single build, which is neither.
|
|
1222
1351
|
|
|
@@ -1703,6 +1832,374 @@ N=2000 npx tsx … # ~50s; the bar is fixed, the detectable EFFECT moves
|
|
|
1703
1832
|
|
|
1704
1833
|
---
|
|
1705
1834
|
|
|
1835
|
+
## Does the best BUILD depend on who you play? (`opp-conditional.mts`)
|
|
1836
|
+
|
|
1837
|
+
Recorded 2026-09-09 — docs/tactics/01 track 0.2 ②, #729. Track 2.4 shipped a
|
|
1838
|
+
kickoff lock and a scouting surface that shows the opponent's last fielded
|
|
1839
|
+
eleven (#726). The plan's own condition for that surface having a DECISION
|
|
1840
|
+
value rather than only an information value is written in 01 §5: across a grid
|
|
1841
|
+
of opponent builds, take my best response to each — *if the distinct argmax is
|
|
1842
|
+
1, scouting's decision value is 0 and 2.4's reason to exist is gone.* Until this
|
|
1843
|
+
run that measurement had not been made; every reversal in this file is
|
|
1844
|
+
conditional on the mode or on my own attack strength, not on the opponent.
|
|
1845
|
+
|
|
1846
|
+
The section above asked the same question of a different axis — which SHEET is
|
|
1847
|
+
best against each opponent, with my 212 points fixed — and found one sheet in
|
|
1848
|
+
every cell's band. This one asks it of the axis a manager actually controls
|
|
1849
|
+
today, the 212-point build.
|
|
1850
|
+
|
|
1851
|
+
### The statistic, the three tests, and the controls
|
|
1852
|
+
|
|
1853
|
+
The eleven published archetypes from `strategy-probe.mts`, built by the shared
|
|
1854
|
+
`squad-lib.mts` so they are the round-robin's squads by construction, play each
|
|
1855
|
+
other in both modes at 10,000 matches per fixture (5,000 seeds, home and away,
|
|
1856
|
+
SHA-256 keyed on both labels and the side). Each unordered pair is simulated
|
|
1857
|
+
once and read from both sides, so the matrix is antisymmetric by construction.
|
|
1858
|
+
The diagonal is played too: a manager who knows the opponent's archetype can
|
|
1859
|
+
field the same one, so that fixture is a real response with a real answer (50
|
|
1860
|
+
by symmetry). Its 22 cells have a known truth, but they are REPORTED data — the
|
|
1861
|
+
same-build response in their own column, feeding the tests — so nothing aborts on
|
|
1862
|
+
them. The controls are separate fixtures on their own seed keys (below).
|
|
1863
|
+
|
|
1864
|
+
For every opponent and mode, my candidates are all eleven builds. The LEADER is
|
|
1865
|
+
the candidate with the highest win-equivalent share (draws half); the BAND is
|
|
1866
|
+
every candidate the sample cannot separate from the leader under the corrected
|
|
1867
|
+
bar; a cell is RESOLVED when its band has one member. Two candidates against the
|
|
1868
|
+
same opponent play different seed columns, so every comparison is UNPAIRED.
|
|
1869
|
+
|
|
1870
|
+
Bonferroni is priced before the run over the search a band performs: one
|
|
1871
|
+
comparison per unordered pair per (opponent, mode) cell, 11·10/2 = 55 pairs over
|
|
1872
|
+
22 cells — **m = 1,210, two-sided, |z| ≥ 4.10**. Enumerating every pair already
|
|
1873
|
+
covers whichever pair the data selects as leader and runner-up, so that selection
|
|
1874
|
+
costs nothing extra, and the two-sided tail already covers both directions; an
|
|
1875
|
+
earlier version charged 2 × 55 AND halved α, pricing direction twice for a bar of
|
|
1876
|
+
4.26. A too-wide bar is not the safe side here — it widens every band and every
|
|
1877
|
+
equivalence bound, and this grid has boundary-sensitive readings. Worst standard
|
|
1878
|
+
error anywhere in the run 0.71 points of share, so the significance threshold is
|
|
1879
|
+
**2.90 points of share (5.8pp of win rate)** and the 80%-power effect
|
|
1880
|
+
**3.49 (7.0pp)**.
|
|
1881
|
+
|
|
1882
|
+
Three tests are reported, and their provenance is not the same. **T1 was
|
|
1883
|
+
declared before the run.** **T2 and T3 were not**: T2 was added after the first
|
|
1884
|
+
run had been read, because "in every band" is a non-rejection and a
|
|
1885
|
+
non-rejection is not an equivalence; T3 was added after T2's result had been
|
|
1886
|
+
read, because "T1 fires and T2 fails" is not a magnitude claim either — T2
|
|
1887
|
+
failing only says no build was SHOWN within the gate everywhere, and every true
|
|
1888
|
+
effect could still be under it. On the run the tables below come from, T2 and
|
|
1889
|
+
T3 are therefore **post-hoc readings of intervals that were already priced**
|
|
1890
|
+
(the Bonferroni family covers every pairwise contrast they use), and the honest
|
|
1891
|
+
name for that is exploratory. So all three were then re-evaluated on a
|
|
1892
|
+
**confirmatory run** — the same fixtures under a fresh seed namespace
|
|
1893
|
+
(`SEED_NAMESPACE=confirm-1`, disjoint seeds, same N), with T2 and T3 fixed in
|
|
1894
|
+
the code before it was started. Its verdicts are reported in their own table
|
|
1895
|
+
below; the numbers in the sections in between are the first run's. The tests
|
|
1896
|
+
are reported **independently**; a difference can be statistically resolved and
|
|
1897
|
+
still smaller than the practical margin, so T1 and T2 can hold at once:
|
|
1898
|
+
|
|
1899
|
+
- **T1 — conditionality DETECTED**: no build is in every opponent's band; each
|
|
1900
|
+
is significantly beaten as a best response by some build against some
|
|
1901
|
+
opponent. A rejection-based positive claim.
|
|
1902
|
+
- **T2 — equivalence ESTABLISHED**: some build is non-inferior to the TRUE best
|
|
1903
|
+
response against every opponent — for that build, the corrected upper bound
|
|
1904
|
+
of (candidate − it) over *every other candidate*, not only the sample leader
|
|
1905
|
+
(the two differ: a bound against the leader alone is trivially satisfied by
|
|
1906
|
+
the leader itself, while one noisy non-leader can deny the simultaneous
|
|
1907
|
+
bound — Knockout's `role-433` and `def-433` columns qualify nobody for that
|
|
1908
|
+
reason),
|
|
1909
|
+
is under the margin in all eleven columns. A positive equivalence claim.
|
|
1910
|
+
- **T3 — conditionality BEYOND THE MARGIN**: every build is beaten, against some
|
|
1911
|
+
opponent, by more than the margin with the corrected LOWER bound of (leader −
|
|
1912
|
+
it). The positive magnitude claim, read from the same priced contrasts at the
|
|
1913
|
+
other end.
|
|
1914
|
+
- The readings: T1 with T3 → **conditional beyond the margin**; T1 with T2 →
|
|
1915
|
+
conditional but within it; T1 alone → conditional, magnitude
|
|
1916
|
+
unresolved; T2 without T1 → **equivalent**; neither T1 nor T2 →
|
|
1917
|
+
**UNDECIDED** at this sample.
|
|
1918
|
+
|
|
1919
|
+
**The margin is 5pp of win rate WITHIN a mode, and that is NOT the repository's
|
|
1920
|
+
gate.** It is deliberately the same number — below it, re-solving stops being
|
|
1921
|
+
worth an agent's trouble — but the gate proper (`01-engine-expansion.md` §5,
|
|
1922
|
+
implemented by `constant-sweep.mts`) is specified on the **exposure-weighted
|
|
1923
|
+
mode mix**, and the modes are nowhere near equally exposed. A qualifier's day is
|
|
1924
|
+
about 1.33 knockout matches against 12 ladder plus the cup's own 3 group matches,
|
|
1925
|
+
which are played in regulation: **8.1% knockout**, and 0% outside the season's
|
|
1926
|
+
top 48. `constant-sweep.mts` says in as many words that issuing a verdict per
|
|
1927
|
+
mode would call cleared what the mix has not.
|
|
1928
|
+
|
|
1929
|
+
**Which way that cuts is not decided here, and the arithmetic must not pretend
|
|
1930
|
+
otherwise.** T3 is EXISTENTIAL: each build meets SOME opponent that beats it by
|
|
1931
|
+
more than the margin. Multiplying that by the 8.1% mode exposure would assume the
|
|
1932
|
+
adverse opponent is met in every knockout fixture — a bracket may hold it rarely
|
|
1933
|
+
or not at all — so the product is not a lower bound on the mix, and the weighted
|
|
1934
|
+
regret can be anywhere from zero upward. This grid also never plays one locked
|
|
1935
|
+
build across the weighted opponent-and-mode mix, which is what the gate actually
|
|
1936
|
+
asks. So clearing the repository gate is **not established**
|
|
1937
|
+
here, and **not ruled out** either. What is established is that the best response
|
|
1938
|
+
is opponent-dependent *within a mode* by more than a within-mode 5pp — the
|
|
1939
|
+
prerequisite for scouting to be worth anything.
|
|
1940
|
+
|
|
1941
|
+
Controls, every one of which aborts the run: each of the eleven builds must
|
|
1942
|
+
reproduce a pinned signature over every field the engine reads from a player —
|
|
1943
|
+
`name`, position, the four attributes, `total`, `cond`, `fair` and both kicker
|
|
1944
|
+
flags — for all eleven players (slot totals alone would let ratios or a scarcity repair
|
|
1945
|
+
drift under the same labels); no two builds byte-identical; **132 true-zero cells**
|
|
1946
|
+
within a family-adjusted |z| < 3.55 of 50 (worst 2.3, `stars-433` on the
|
|
1947
|
+
`stars-433|bal-442` League control key) —
|
|
1948
|
+
correcting matters here: at an uncorrected two-sided 3σ each true null exceeds
|
|
1949
|
+
with probability ≈0.27%, so across 132 streams a valid rerun would fail about
|
|
1950
|
+
three times in ten (1 − 0.9973^132 ≈ 30%). All 132 are CONTROL FIXTURES of their own,
|
|
1951
|
+
never cells read back out of the grid: 22 on a same-label key `a|a` and 110 on a
|
|
1952
|
+
two-distinct-label key, the shape every published cell uses. Both halves matter.
|
|
1953
|
+
A seeding defect keyed on the pair of labels — which the retired FNV one was —
|
|
1954
|
+
passes a same-label-only control while biasing every cell that control exists to
|
|
1955
|
+
validate. And every one carries a `zero-ctl:` prefix, so no control shares a
|
|
1956
|
+
stream with a reported cell: the controls abort at their α, and a rerun under a
|
|
1957
|
+
fresh namespace selected on a control that shared those streams would be picked
|
|
1958
|
+
partly on the estimates themselves. And a
|
|
1959
|
+
published fact held — `cycle-probe.mts` reports `role-433` beating each of
|
|
1960
|
+
its seven other candidates in Knockout, and here the same eight squads say so
|
|
1961
|
+
under the same statistic, smallest decided z 3.8 — played on a `drift-ctl:` seed
|
|
1962
|
+
family disjoint from every reported cell, because a guard that aborts on the
|
|
1963
|
+
reported cells' own estimates would accept only runs reproducing them. That last control is a
|
|
1964
|
+
**compatibility** test, not a demand for renewed significance: at
|
|
1965
|
+
`cycle-probe`'s N=4000 the weakest of the seven pairs would replicate |z| ≥ 3.32
|
|
1966
|
+
with roughly even odds even if nothing had changed, so a rerun that required it
|
|
1967
|
+
would be aborting on its own sampling noise. Instead `role-433`'s decided win
|
|
1968
|
+
rate against each of the seven from the default-namespace run is pinned in the
|
|
1969
|
+
probe, and a run fails only when its own rate is incompatible with the pin at
|
|
1970
|
+
the Bonferroni-7 bound |z| ≤ 2.69 — what "the squads changed or the engine
|
|
1971
|
+
moved" looks like at any N. Worst |z| against the pins: 0.0 in the default
|
|
1972
|
+
namespace (the pins are that family's own default values) and 0.8 in
|
|
1973
|
+
`confirm-1`.
|
|
1974
|
+
The Knockout column also reproduces the cheap keeper's published collapse
|
|
1975
|
+
(`gkmin-433` at 27.8% against `def-532`, inside the 27–42% this file already
|
|
1976
|
+
states), and the League column reproduces the two pairings `cycle-probe.mts`
|
|
1977
|
+
could not separate (`gkmin-433` 50.6% vs `def-433` and 49.2% vs `def-532`,
|
|
1978
|
+
neither significant here either).
|
|
1979
|
+
|
|
1980
|
+
### League — UNDECIDED: T1 does not fire, T2 does not pass
|
|
1981
|
+
|
|
1982
|
+
| Opponent | Leader | Share | Runner-up | Gap (z) | Band | Within 5pp of the true best |
|
|
1983
|
+
| --- | --- | --- | --- | --- | --- | --- |
|
|
1984
|
+
| `flat-433` | **`gkmin-433`** | 72.0 | `def-433` | +2.5 (5.7) | `gkmin-433` | `gkmin-433` |
|
|
1985
|
+
| `role-433` | `gkmin-433` | 54.0 | `def-532` | +1.0 (2.2) | `def-433` / `gkmin-433` / `def-532` | `gkmin-433` |
|
|
1986
|
+
| `def-433` | `gkmin-433` | 50.6 | `def-433` | +0.2 (0.5) | `def-433` / `gkmin-433` / `def-532` | `def-433` / `gkmin-433` / `def-532` |
|
|
1987
|
+
| `stars-433` | **`gkmin-433`** | 68.1 | `def-532` | +4.0 (7.3) | `gkmin-433` | `gkmin-433` |
|
|
1988
|
+
| `gkheavy-433` | `gkmin-433` | 60.8 | `def-433` | +1.9 (4.1) | `def-433` / `gkmin-433` | `gkmin-433` |
|
|
1989
|
+
| `gkmin-433` | `def-532` | 50.8 | `gkmin-433` | +0.8 (1.8) | `def-433` / `gkmin-433` / `def-532` | `def-532` |
|
|
1990
|
+
| `shootonly-433` | **`gkmin-433`** | 64.3 | `def-433` | +3.7 (7.2) | `gkmin-433` | `gkmin-433` |
|
|
1991
|
+
| `passonly-433` | `def-433` | 64.7 | `gkmin-433` | +1.8 (3.9) | `def-433` / `gkmin-433` | `def-433` |
|
|
1992
|
+
| `def-532` | `def-532` | 50.1 | `def-433` | +0.1 (0.3) | `def-433` / `gkmin-433` / `def-532` | `def-433` / `gkmin-433` / `def-532` |
|
|
1993
|
+
| `bal-442` | `gkmin-433` | 58.8 | `def-433` | +1.2 (2.6) | `def-433` / `gkmin-433` | `gkmin-433` |
|
|
1994
|
+
| `atk-352` | `gkmin-433` | 63.0 | `def-433` | +1.4 (3.0) | `def-433` / `gkmin-433` | `gkmin-433` |
|
|
1995
|
+
|
|
1996
|
+
> **Measured on** — every one of the eleven published archetypes as the
|
|
1997
|
+
> opponent, with all eleven as my candidates (the opponent's own build among
|
|
1998
|
+
> them: its EXPECTED share is 50 by symmetry, and the sampled cell — 50.1 for
|
|
1999
|
+
> `def-532` here — enters the tests like any other, a known truth on reported
|
|
2000
|
+
> data rather than a control). Bold
|
|
2001
|
+
> leaders are RESOLVED cells: the band holds one build. The last column is T2's
|
|
2002
|
+
> per-cell reading — which builds the sample shows within 2.5 points of share of
|
|
2003
|
+
> the TRUE best response: the corrected upper bound of (candidate − it) is under
|
|
2004
|
+
> the margin for EVERY other candidate in the column, not just for the sample
|
|
2005
|
+
> leader. That is why a column can list nobody even though its leader is
|
|
2006
|
+
> trivially zero behind itself — one noisy candidate denies the simultaneous
|
|
2007
|
+
> bound.
|
|
2008
|
+
> **Sample** — 10,000 matches per fixture, home and away, SHA-256 seeds; 66
|
|
2009
|
+
> fixtures in this mode.
|
|
2010
|
+
> **Mode** — League (`gameFlg 0`, draws count half).
|
|
2011
|
+
> **Effect** — the raw leader is `gkmin-433` in 8 of 11 columns and 3 cells
|
|
2012
|
+
> resolve, all three to `gkmin-433`. **T1 does not fire**: `gkmin-433` is in the
|
|
2013
|
+
> band of all eleven opponents, so conditionality is NOT detected — but see the
|
|
2014
|
+
> confirmation below, where a disjoint sample of the same fixtures puts it
|
|
2015
|
+
> outside one band and T1 does fire. That instability is itself the finding
|
|
2016
|
+
> here. **T2 does not pass**: no build — `gkmin-433` included — is shown within
|
|
2017
|
+
> 5pp of every other candidate in every column; against its own archetype the
|
|
2018
|
+
> best sample response is `def-532` and the corrected upper bound of that
|
|
2019
|
+
> 0.8-point gap exceeds the margin, and against `passonly-433` the leader is
|
|
2020
|
+
> `def-433`. So
|
|
2021
|
+
> the League reading is **UNDECIDED**: this sample did not detect an opponent
|
|
2022
|
+
> that re-prices the build, and it cannot rule out re-pricing of up to its own
|
|
2023
|
+
> power — 7.0pp of win rate on the worst cell — either. One build,
|
|
2024
|
+
> `gkmin-433`, is never beaten by MORE than the gate in any column (T3's
|
|
2025
|
+
> per-build reading) — which is consistent with equivalence and is not it. "Field `gkmin-433` regardless" is consistent with the grid; the grid
|
|
2026
|
+
> does not establish it.
|
|
2027
|
+
> **Source** — `opp-conditional.mts`, sections (2) and (3), `REG`.
|
|
2028
|
+
|
|
2029
|
+
### Knockout — CONDITIONAL BEYOND THE WITHIN-MODE MARGIN: T1 fires, T3 holds
|
|
2030
|
+
|
|
2031
|
+
| Opponent | Leader | Share | Runner-up | Gap (z) | Shootout reach | Band | Within 5pp of the true best |
|
|
2032
|
+
| --- | --- | --- | --- | --- | --- | --- | --- |
|
|
2033
|
+
| `flat-433` | `def-433` | 82.8 | `def-532` | +0.3 (0.6) | 21% | `role-433` / `def-433` / `def-532` | `def-433` |
|
|
2034
|
+
| `role-433` | `gkheavy-433` | 50.6 | `role-433` | +0.3 (0.5) | 26% | `role-433` / `gkheavy-433` | — |
|
|
2035
|
+
| `def-433` | `role-433` | 52.4 | `gkheavy-433` | +0.5 (0.7) | 32% | `role-433` / `def-433` / `gkheavy-433` | `role-433` |
|
|
2036
|
+
| `stars-433` | **`gkmin-433`** | 69.9 | `def-532` | +4.8 (7.3) | 5% | `gkmin-433` | `gkmin-433` |
|
|
2037
|
+
| `gkheavy-433` | `gkmin-433` | 52.0 | `gkheavy-433` | +1.6 (2.2) | 28% | `role-433` / `gkheavy-433` / `gkmin-433` | `gkmin-433` |
|
|
2038
|
+
| `gkmin-433` | **`def-532`** | 72.2 | `def-433` | +8.7 (13.3) | 24% | `def-532` | `def-532` |
|
|
2039
|
+
| `shootonly-433` | `role-433` | 63.9 | `def-433` | +2.3 (3.3) | 18% | `role-433` / `def-433` | `role-433` |
|
|
2040
|
+
| `passonly-433` | `def-433` | 67.8 | `gkheavy-433` | +0.5 (0.7) | 25% | `role-433` / `def-433` / `gkheavy-433` | `def-433` |
|
|
2041
|
+
| `def-532` | **`gkheavy-433`** | 59.9 | `role-433` | +3.0 (4.3) | 47% | `gkheavy-433` | `gkheavy-433` |
|
|
2042
|
+
| `bal-442` | **`gkheavy-433`** | 58.9 | `role-433` | +4.0 (5.7) | 31% | `gkheavy-433` | `gkheavy-433` |
|
|
2043
|
+
| `atk-352` | `gkheavy-433` | 62.3 | `role-433` | +1.7 (2.5) | 26% | `role-433` / `gkheavy-433` | `gkheavy-433` |
|
|
2044
|
+
|
|
2045
|
+
> **Measured on** — the same eleven archetypes, the same construction. Shootout
|
|
2046
|
+
> reach is the share of that opponent's matches, averaged over all eleven
|
|
2047
|
+
> candidates, that went to penalties — DESCRIPTIVE, see below.
|
|
2048
|
+
> **Sample** — 10,000 matches per fixture, home and away, SHA-256 seeds; 66
|
|
2049
|
+
> fixtures in this mode.
|
|
2050
|
+
> **Mode** — Knockout (`gameFlg 1`: extra time and shootout, every match
|
|
2051
|
+
> decided).
|
|
2052
|
+
> **Effect** — five distinct raw leaders and **4 of 11 cells resolve, to three
|
|
2053
|
+
> different builds**: `gkmin-433` against `stars-433` (z 7.3 over the
|
|
2054
|
+
> runner-up), `def-532` against `gkmin-433` (z 13.3 — the cheap keeper's
|
|
2055
|
+
> published collapse, seen from the other side), and `gkheavy-433` against both
|
|
2056
|
+
> `def-532` (z 4.3) and `bal-442` (z 5.7). **T1 fires**: no build is in every
|
|
2057
|
+
> band. `role-433` and `gkheavy-433` share the most bands at seven each, and
|
|
2058
|
+
> neither survives: `role-433` is significantly excluded from all four resolved
|
|
2059
|
+
> cells, `gkheavy-433` from two.
|
|
2060
|
+
> The exclusions are carried by resolved cells at z 4.3 to 13.3, not by noise.
|
|
2061
|
+
> **T2 does not pass**: no build is within 5pp of every other candidate in every
|
|
2062
|
+
> column — and the `role-433` column names nobody at all once the bound is taken
|
|
2063
|
+
> over every candidate rather than the sample leader.
|
|
2064
|
+
> **T3 holds**: every one of the eleven builds is beaten by MORE than 5pp of win
|
|
2065
|
+
> rate, with the corrected lower bound, against at least one opponent. The
|
|
2066
|
+
> `gkmin-433` column alone does it for the other ten — `def-532`'s 72.2 there
|
|
2067
|
+
> clears the margin against every one of them, the closest being `def-433` at
|
|
2068
|
+
> 63.5 — and `def-532` itself is beaten beyond the margin elsewhere, most
|
|
2069
|
+
> clearly by `gkheavy-433` in its own column (59.9 against `def-532`'s 49.5).
|
|
2070
|
+
> So the Knockout
|
|
2071
|
+
> reading is **CONDITIONAL BEYOND THE WITHIN-MODE MARGIN**: the best build is
|
|
2072
|
+
> opponent-conditional, and by more than 5pp of win rate *inside Knockout* — a
|
|
2073
|
+
> claim T3 makes, not one inferred from T2 failing. Whether that clears the
|
|
2074
|
+
> repository's exposure-weighted gate is neither established nor ruled out here:
|
|
2075
|
+
> 5pp is a lower bound, and this grid never plays a locked build across the mix.
|
|
2076
|
+
> **Source** — `opp-conditional.mts`, sections (2) and (3), `KO`.
|
|
2077
|
+
|
|
2078
|
+
### Confirmation on a fresh seed namespace
|
|
2079
|
+
|
|
2080
|
+
Same eleven squads, same N, seeds prefixed `confirm-1|` so no fixture shares a
|
|
2081
|
+
match with the run above; T2 and T3 were in the code before it started. Every
|
|
2082
|
+
control holds (worst true-zero |z| 2.4, `def-433` on the
|
|
2083
|
+
`def-433|gkheavy-433` League control key; pinned-rate guard worst 1.2), and
|
|
2084
|
+
Knockout reproduces exactly — **League does not**:
|
|
2085
|
+
|
|
2086
|
+
| run | mode | T1 | T2 | T3 | reading | resolved leaders | in every band | never beaten beyond the margin |
|
|
2087
|
+
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
|
|
2088
|
+
| tables above | League | no | no | no | UNDECIDED | `gkmin-433` | `gkmin-433` | `gkmin-433` |
|
|
2089
|
+
| `confirm-1` | League | **fires** | no | no | CONDITIONAL, MAGNITUDE UNRESOLVED | `gkmin-433`, `def-433` | none | `gkmin-433`, `def-532` |
|
|
2090
|
+
| tables above | Knockout | **fires** | no | **holds** | CONDITIONAL BEYOND THE MARGIN (within mode) | `gkmin-433`, `def-532`, `gkheavy-433` | none | none |
|
|
2091
|
+
| `confirm-1` | Knockout | **fires** | no | **holds** | CONDITIONAL BEYOND THE MARGIN (within mode) | `gkmin-433`, `def-532`, `gkheavy-433` | none | none |
|
|
2092
|
+
|
|
2093
|
+
**League's T1 is not reproducible at this sample size, and that is the League
|
|
2094
|
+
result.** T1 asks whether any build sits in every opponent's band; `gkmin-433`
|
|
2095
|
+
does in the reported run and in `confirm-2`, and does not in `confirm-1` — three
|
|
2096
|
+
disjoint samples of the same eleven fixtures, two landing one way and one the
|
|
2097
|
+
other. The `passonly-433` column is where it turns: `gkmin-433` is inside that
|
|
2098
|
+
band at z 3.9 in the reported run and in `confirm-2`, and outside it at z 4.3 in
|
|
2099
|
+
`confirm-1`. Which side of T1 a sample lands on is decided by noise at N=10,000,
|
|
2100
|
+
so nothing about League should be read as "conditionality was detected" or "was
|
|
2101
|
+
not"; the honest statement is UNDECIDED, and a larger N is what would settle it.
|
|
2102
|
+
The Knockout verdict carries none of this: T1 fires and T3 holds in all three,
|
|
2103
|
+
with the same three resolved leaders.
|
|
2104
|
+
|
|
2105
|
+
The confirmatory runs are what the T2/T3 verdicts rest on; the first run's
|
|
2106
|
+
tables remain the published numbers because they are the run every figure in
|
|
2107
|
+
this section was read from, and mixing them would leave no run owning them.
|
|
2108
|
+
|
|
2109
|
+
### What the grid does and does not say about 2.4
|
|
2110
|
+
|
|
2111
|
+
**What it says.** In Knockout the answer to "which build should I field" changes
|
|
2112
|
+
with the opponent, by amounts the sample resolves (3.0 to 8.7 points of share
|
|
2113
|
+
between the leader and the runner-up in the four resolved cells: 4.8, 8.7, 3.0
|
|
2114
|
+
and 4.0). In League
|
|
2115
|
+
the same grid neither finds such an opponent nor rules one out. The plan's
|
|
2116
|
+
sentence in 01 §5 was written mode-blind; measured, the axis is a decision in
|
|
2117
|
+
one mode and undecided in the other.
|
|
2118
|
+
|
|
2119
|
+
**What it does not say — and an earlier version of this section said it.** It
|
|
2120
|
+
does not say that the cup lock is where that decision lives. The lock (#726)
|
|
2121
|
+
freezes one eleven when the cup opens and reuses it for every fixture of the
|
|
2122
|
+
day, group stage (`gameFlg 0`) and knockout rounds (`gameFlg 1`) alike. A
|
|
2123
|
+
manager cannot learn a knockout opponent and then switch to the build this grid
|
|
2124
|
+
says beats it; the decision the lock actually forces is *one build against a
|
|
2125
|
+
bracket of opponents in sequence*, and a per-opponent best response is not that
|
|
2126
|
+
object. What this grid establishes is the prerequisite: that in the mode the
|
|
2127
|
+
knockout rounds are played in, a per-opponent best response EXISTS to be
|
|
2128
|
+
informed by. Whether a single locked build should be chosen differently given
|
|
2129
|
+
the bracket — and whether the scouting surface changes that choice — needs a
|
|
2130
|
+
bracket-level experiment, and this file does not contain one. As #729 itself
|
|
2131
|
+
notes, a distinct-argmax count of 1 would not have meant "remove 2.4" either:
|
|
2132
|
+
the lock is a commitment device on its own.
|
|
2133
|
+
|
|
2134
|
+
**What produced the conditional responses is not identified.** The three builds
|
|
2135
|
+
that are unique best responses somewhere — `gkmin-433`, `gkheavy-433`,
|
|
2136
|
+
`def-532` — differ from the rest of the grid in keeper budget, but not only in
|
|
2137
|
+
that: `def-532` carries an ordinary keeper and changes formation and back-line
|
|
2138
|
+
allocation, and `GK_MIN`/`GK_HEAVY` redistribute the keeper's points across the
|
|
2139
|
+
whole outfield. The keeper section of this file supplies a mechanism that is
|
|
2140
|
+
consistent with the pattern — in Knockout the keeper's total is
|
|
2141
|
+
outfield-conditional through shootout exposure, and the table's shootout reach
|
|
2142
|
+
runs from 5% (`stars-433`) to 47% (`def-532`) across opponents — but consistency
|
|
2143
|
+
is not attribution. Isolating the keeper would take a keeper-only,
|
|
2144
|
+
fixed-outfield contrast, which this grid does not run.
|
|
2145
|
+
|
|
2146
|
+
Two secondary readings survive at the descriptive level. In League the
|
|
2147
|
+
cheap-keeper build is the raw leader in 8 of 11 columns, which is this file's
|
|
2148
|
+
"smaller keeper total is better in matches that allow draws" seen across the
|
|
2149
|
+
whole grid at once. And `role-433` — the Knockout dominant in the 8-candidate
|
|
2150
|
+
round robin above — is the raw leader in only 2 of 11 Knockout columns once the
|
|
2151
|
+
keeper-budget archetypes are in the grid, and is the unique best response to
|
|
2152
|
+
nobody. Dominance in a round robin and being the best response to each opponent
|
|
2153
|
+
are different questions, and the second is the one scouting asks.
|
|
2154
|
+
|
|
2155
|
+
### What this grid CANNOT decide
|
|
2156
|
+
|
|
2157
|
+
- **The build space.** Eleven hand-built archetypes, the same set on both axes.
|
|
2158
|
+
That is a sample, not an argmax over 212-point squads. A conditional best
|
|
2159
|
+
response can only appear among builds that exist in the grid, and a build
|
|
2160
|
+
that beats every leader here could exist outside it. Every claim above is a
|
|
2161
|
+
claim about these eleven.
|
|
2162
|
+
- **Nobody re-optimises against a KNOWN opponent.** Every candidate is a
|
|
2163
|
+
pre-built archetype. A squad rebuilt with the opponent's eleven in hand is a
|
|
2164
|
+
different 212-point problem — and it is the search the scouting surface would
|
|
2165
|
+
actually enable. This grid measures whether the choice AMONG fixed builds is
|
|
2166
|
+
conditional, which is the weaker, prerequisite question.
|
|
2167
|
+
- **The bracket.** The lock forces one build against a sequence of opponents;
|
|
2168
|
+
the grid asks about one opponent at a time. The object the lock decides is
|
|
2169
|
+
not measured here.
|
|
2170
|
+
- **Mechanism.** Formation, keeper budget and outfield allocation move together
|
|
2171
|
+
across these archetypes. Which of them re-prices the build is not separated.
|
|
2172
|
+
- **`cond` fixed at 5, tenure 0.** As everywhere in this file. Production
|
|
2173
|
+
applies growth before the engine and, since #794, an involvement-driven
|
|
2174
|
+
condition; neither is in these squads.
|
|
2175
|
+
- **Sheets at the default.** The section above showed the sheet axis moves more
|
|
2176
|
+
than any build change does; this grid holds it at the shipped weights on both
|
|
2177
|
+
sides.
|
|
2178
|
+
- **Power.** 8 of 11 League cells and 7 of 11 Knockout cells are unresolved,
|
|
2179
|
+
and League's T1 flips between two disjoint samples at this N.
|
|
2180
|
+
T2's margin is 5pp within the mode; the worst cell's 80%-power effect is 7.0pp
|
|
2181
|
+
of win rate, so this sample could not have established equivalence at that margin
|
|
2182
|
+
even where it holds — and T2's bound is taken over every candidate, so a
|
|
2183
|
+
weak candidate with a large standard error can deny it on its own. A larger
|
|
2184
|
+
N is the way to move League out of UNDECIDED.
|
|
2185
|
+
|
|
2186
|
+
Re-run after any engine change:
|
|
2187
|
+
|
|
2188
|
+
```bash
|
|
2189
|
+
npx tsx packages/mcp/skill/reference/probes/opp-conditional.mts # ~2.7M matches, ~31s
|
|
2190
|
+
SEED_NAMESPACE=confirm-2 npx tsx … # a fresh disjoint sample of the same fixtures. Knockout should reproduce;
|
|
2191
|
+
# League's T1 is not expected to — it flips between samples at this N, which is the League finding itself
|
|
2192
|
+
N=4000 npx tsx … # smoke; the tables are shapes, and the pinned-rate control cannot apply at any default-namespace
|
|
2193
|
+
# N other than 10,000 — those streams overlap the pins, so the two rates are paired rather than comparable
|
|
2194
|
+
# 132 `zero-ctl:` control fixtures (22 same-label, 110 pair-key) sit at a family-adjusted bar, which
|
|
2195
|
+
# bounds a valid run's abort probability at 5% — they are NOT independent (the seed key omits the mode,
|
|
2196
|
+
# so each pair's League and Knockout controls share a stream), so 5% is a ceiling, not a rate. The
|
|
2197
|
+
# reported diagonals are NOT checked. Investigate an
|
|
2198
|
+
# exceedance first; only then re-run under another SEED_NAMESPACE — never widen the bar to pass
|
|
2199
|
+
```
|
|
2200
|
+
|
|
2201
|
+
---
|
|
2202
|
+
|
|
1706
2203
|
## Note for the maintainers
|
|
1707
2204
|
|
|
1708
2205
|
> **2026-08-19: this note was rewritten twice in one day.** First it said the
|