pog-mcp 0.9.15 → 0.9.17

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pog-mcp",
3
- "version": "0.9.15",
3
+ "version": "0.9.17",
4
4
  "type": "module",
5
5
  "description": "MCP server that lets an AI agent play Proof of Goal — wallet, sign-in, squad building, and matches as typed tools.",
6
6
  "license": "MIT",
@@ -664,18 +664,18 @@ sets it — there is no `fair:` outside the engine's own tests, and every engine
664
664
  input is built from an explicit field list — so it is a global constant in
665
665
  practice, and setting it on both squads IS the rule change.
666
666
 
667
- **Not swept: the zone-press coefficient.** At the time its selector was broken —
668
- every candidate weight was `0.0`, so `weightedPick` fell back to "the first non-GK
669
- slot" — and a coefficient in front of an undefined selection measures nothing. The
670
- selector was fixed in #730 (`zp-selector.mts` below); the coefficient is sweepable
671
- now but has not been swept.
667
+ **The zone-press coefficient waited for a selector.** While every candidate
668
+ weight at that scene was `0.0`, `weightedPick` fell back to "the first non-GK
669
+ slot" and a coefficient in front of an undefined selection measured nothing.
670
+ #730 fixed the selector (`zp-selector.mts` below), so the axis became sweepable
671
+ and it is swept here, over 1.5..4.5.
672
672
 
673
673
  ### The statistic, and why not the obvious one
674
674
 
675
675
  Rows below are in **win-equivalent share**: a win is 1, a draw is 0.5, a loss is
676
676
  0, and every match counts. The obvious alternative — win rate among DECIDED
677
677
  matches — conditions on something the swept constants themselves move. Free-kick
678
- occurrence takes the regulation draw rate from 57.9% to 38.1% across its range,
678
+ occurrence takes the regulation draw rate from 57.7% to 38.0% across its range,
679
679
  so a decided-only percentage mixes "who wins more" with "which matches got
680
680
  decided at all". It also inflates the number the activation gate is read
681
681
  against: an 11/80/9 record reads 55% decided, clearing a +5pp gate, while the
@@ -722,7 +722,7 @@ to count rather than a failure:
722
722
 
723
723
  | Measured every run | Result |
724
724
  | --- | --- |
725
- | Knockout draws, across both baskets | **0 in 5,440,000** knockout matches |
725
+ | Knockout draws, across both baskets | **0 in 10,880,000** knockout matches |
726
726
 
727
727
  That draw count is a claim this section originally got wrong in the other
728
728
  direction. It said knockout mode "has no draws by construction". It nearly does:
@@ -741,9 +741,10 @@ without — a paired test cannot detect misaligned columns from its own output,
741
741
  a freshly patched shipped-value engine is required to reproduce the default
742
742
  column match for match, in order.
743
743
 
744
- 3,400,000 matches in the narrow basket (5 archetypes, 10 fixtures, 10,000
745
- matches each, 17 conditions, both modes) plus 7,480,000 in the wide one.
746
- **Bonferroni over m = 8,320 pre-registered comparisons, two-sided: |z| >= 4.53.**
744
+ 6,800,000 matches in the narrow basket (5 archetypes, 10 fixtures, 10,000
745
+ matches each, 34 conditions, both modes) plus 14,960,000 in the wide one
746
+ (11 archetypes, 55 fixtures, 4,000 matches each, the same 34 conditions).
747
+ **Bonferroni over m = 17,160 pre-registered comparisons, two-sided: |z| >= 4.68.**
747
748
  That m prices the SEARCHES, not just the tests. A gate cell does not spend one
748
749
  hypothesis, and BOTH of its sides are data-chosen. The baseline side picks a
749
750
  leader out of K and builds a band around it — noisy standings could have named
@@ -770,10 +771,10 @@ power, not a minimum detectable effect. The 80%-power figure adds z(0.80):
770
771
 
771
772
  | median / worst | significance threshold | 80%-power effect |
772
773
  | --- | --- | --- |
773
- | narrow, regulation | 1.0 / 2.2pp | 1.2 / 2.6pp |
774
- | narrow, knockouts | 1.8 / 3.0pp | 2.1 / 3.6pp |
775
- | wide, regulation | 1.8 / 3.8pp | 2.1 / 4.5pp |
776
- | wide, knockouts | 2.9 / 4.8pp | 3.4 / **5.7pp** |
774
+ | narrow, regulation | 0.8 / 2.2pp | 0.9 / 2.6pp |
775
+ | narrow, knockouts | 1.5 / 3.1pp | 1.8 / 3.6pp |
776
+ | wide, regulation | 1.5 / 4.0pp | 1.7 / 4.7pp |
777
+ | wide, knockouts | 2.5 / 4.9pp | 3.0 / **5.7pp** |
777
778
 
778
779
  Read the worst column, not the median, since "every test is powered enough" is a
779
780
  claim about the tail — and read it honestly: **the worst wide-knockout fixture
@@ -796,10 +797,12 @@ gate, and says so.
796
797
  That only settles anything if the band is SHARP — noisy standings widen it, and
797
798
  a wide enough band absorbs every new leader, which would be a power limitation
798
799
  wearing the costume of a result. So the probe reports the band's own
799
- discrimination: it holds **2 of the 5** narrow squads and **3 of the 11** wide
800
- ones, with its edge **1.45pp** and **1.60pp** from the leader. A band that tight
801
- is not swallowing the basket. That is the number a smaller rerun moves, and the
802
- one to check before reusing this conclusion.
800
+ discrimination at its WORST over the searched weights: it holds **2 of the 5**
801
+ narrow squads and **3 of the 11** wide ones, with its edge **1.51pp** and
802
+ **2.29pp** from the leader. Two of five is not a band that has resolved the
803
+ basket — it is why this section's verdict is UNRESOLVED against a wide band
804
+ rather than "no contest", and it is the number a rerun with a further-apart
805
+ basket would move.
803
806
 
804
807
  ### The table
805
808
 
@@ -807,8 +810,11 @@ Basket: `role-433`, `def-433`, `def-532`, `shootonly-433`, `flat-433` — five
807
810
  outfield templates, the shared `squad-lib.mts` builders, so a row here and a row
808
811
  in the round-robin above are the same squads.
809
812
 
810
- - **`sig`** — fixture deltas clearing the corrected bar, out of 40 in the narrow
811
- basket (10 fixtures x 4 swept values) and out of 220 in the wide one (55 x 4).
813
+ - **`sig`** — fixture deltas clearing the corrected bar, out of **10 fixtures x
814
+ that axis's swept values** in the narrow basket and **55 x the same** in the
815
+ wide one. Most axes sweep four, so most rows read out of 40 and 220; set-piece
816
+ `dBonus` sweeps five to cover 3..8 end to end, so its rows read out of 50 and
817
+ 275.
812
818
  Both are printed, because an axis can be classified on the wide basket's
813
819
  evidence alone and a table that hid it would publish a verdict whose decisive
814
820
  measurement is invisible.
@@ -821,28 +827,93 @@ in the round-robin above are the same squads.
821
827
 
822
828
  | Axis (shipped) | Mode | sig (narrow) | sig (wide) | rev | Largest narrow delta | gate | What that row leaves FREE |
823
829
  | --- | --- | --- | --- | --- | --- | --- | --- |
824
- | **FK occurrence** `rand(10)<1` | reg | 21/40 | 102/220 | 0 | **+11.8pp** `role-433` vs `flat-433` @x4 | none | the FK taker: every basket squad kicks from slot 9, so this is "more free kicks", never "more free kicks AND a better taker" |
825
- | | ko | 18/40 | 100/220 | 0 | +8.8pp `def-532` vs `flat-433` @x4 | none — +7.5pp against `role-433` alone, 0.0pp against the band it belongs to | shootout exposure — it is a property of the PAIRING, and the basket fixes both halves |
826
- | **Post play** `rand(10)<1.5` | reg | 18/40 | 50/220 | 0 | −4.8pp `role-433` vs `flat-433` @x4 | none | the PP2 shooter is drawn by `getPlayer`, so the row confounds "more post plays" with "who the draw lands on" |
827
- | | ko | 5/40 | 40/220 | 0 | −3.4pp `role-433` vs `shootonly-433` @x4 | none | same, plus FW count: 433 and 532 field three and two forwards, and `fwIdx.length` gates the branch |
828
- | **Counter** `THRESHOLD=5` | reg | 3/40 | 5/220 | 0 | −2.0pp `shootonly-433` vs `flat-433` @0% | none | the four sampled thresholds span 0% to 100%, but changing this constant reroutes RNG consumption, so outcomes need not interpolate between them |
829
- | | ko | **0/40** | 7/220 | 0 | −1.2pp (below bar) | none | as above. This is the control: it never produced a gate in either basket, which is the only behaviour that matters here |
830
- | **`fair`** default 5 | reg | 13/40 | 37/220 | **1** | **+10.9pp** `shootonly-433` vs `flat-433` @22 | none | `fair` is uniform across all 22 players here; per-player or per-position `fair` is a different (and unmeasured) axis |
831
- | | ko | 7/40 | 23/220 | 0 | +10.9pp `shootonly-433` vs `flat-433` @22 | none | keeper `defense`, which every extra penalty routes through — held fixed inside each template |
832
-
833
- **Nothing here becomes a candidate.** All eight rows move margins somewhere — the
834
- control included, once the wide basket is counted — and
835
- the one pairing that re-prices anything does so against a squad that was already
836
- tied for best. The pass set at the plan's own activation gate is empty over both
837
- baskets, and the answer took a day rather than a release.
838
-
839
- **Empty for these four axes, which is not the same as empty for the pool.** The
840
- plan's weekly-rule pool also names set-piece `dBonus`, penalty attack noise, and
841
- base `dBonus`; none of them is swept here, and this run says nothing about them.
842
- The axes measured are the ones the issue behind this section listed — the two
843
- hardcoded rates, `fair`, and the counter control. Anyone reading this as
844
- "schedule nothing" should read it as "schedule none of these four, and go
845
- measure the other three before concluding anything about the pool."
830
+ | **FK occurrence** `rand(10)<1` | reg | 23/40 | 98/220 | 0 | **+11.8pp** `def-532` vs `flat-433` @x4 | none | the FK taker: every basket squad kicks from slot 9, so this is "more free kicks", never "more free kicks AND a better taker" |
831
+ | | ko | 19/40 | 92/220 | 0 | +8.5pp `def-532` vs `flat-433` @x4 | none | shootout exposure — it is a property of the PAIRING, and the basket fixes both halves |
832
+ | **Post play** `rand(10)<1.5` | reg | 16/40 | 52/220 | 0 | −5.0pp `role-433` vs `flat-433` @x4 | none | the PP2 shooter is drawn by `getPlayer`, so the row confounds "more post plays" with "who the draw lands on" |
833
+ | | ko | 7/40 | 33/220 | 0 | −4.1pp `def-532` vs `shootonly-433` @x4 | none | same, plus FW count: 433 and 532 field three and two forwards, and `fwIdx.length` gates the branch |
834
+ | **Counter** `THRESHOLD=5` | reg | 2/40 | 5/220 | 0 | −1.8pp `shootonly-433` vs `flat-433` @0% | none | the four sampled thresholds span 0% to 100%, but changing this constant reroutes RNG consumption, so outcomes need not interpolate between them |
835
+ | | ko | **0/40** | 3/220 | 0 | +1.5pp (below bar) | none | as above. This is the control: it never produced a gate in either basket, which is the only behaviour that matters here |
836
+ | **`fair`** default 5 | reg | 13/40 | 33/220 | **1** | **+11.4pp** `shootonly-433` vs `flat-433` @22 | none | `fair` is uniform across all 22 players here; per-player or per-position `fair` is a different (and unmeasured) axis |
837
+ | | ko | 7/40 | 28/220 | 0 | +10.7pp `shootonly-433` vs `flat-433` @22 | none | keeper `defense`, which every extra penalty routes through — held fixed inside each template |
838
+ | **Set-piece `dBonus`** 5 | reg | 15/50 | 56/275 | 0 | −1.7pp `def-532` vs `flat-433` @8 | none | it widens the DEFENDER's roll at FK1/FK2 only, and the taker is slot 9 in every basket squad. Five swept values, not four: 3..8 end to end with the default column carrying 5 |
839
+ | | ko | 10/50 | 74/275 | 0 | +1.4pp `role-433` vs `shootonly-433` @8 | none | same, plus the shootout, which this axis does not touch (`takePkShot` has its own width) |
840
+ | **Base `dBonus`** 8 | reg | 25/40 | 43/220 | 0 | +2.9pp `def-532` vs `flat-433` @6 | none | the widest-reaching of the seven — it is the defender roll for every non-set-piece contest — and still nothing re-prices |
841
+ | | ko | 14/40 | 41/220 | 0 | +3.7pp `def-532` vs `shootonly-433` @6 | none | as above |
842
+ | **PK attack noise** `rand(16)` | reg | 9/40 | 25/220 | 0 | +0.9pp `role-433` vs `def-532` @8 | none | in regulation it reaches only awarded penalties, which is why the row is nearly flat |
843
+ | | ko | 25/40 | 102/220 | 0 | **+9.7pp** `role-433` vs `def-532` @8 | none | the shootout routes every kick through it, so knockouts feel it and regulation barely does — the largest mode split in the table, and still no gate |
844
+ | **Zone press** `ratio * 3` | reg | **0/40** | **0/220** | 0 | −0.8pp `def-433` vs `flat-433` @4.5 | none | the axis #730 unblocked, and the only one in the file that moves NOTHING: zero significant deltas in either basket, in either mode |
845
+ | | ko | **0/40** | **0/220** | 0 | +1.0pp `def-433` vs `def-532` @1.5 | none | as above. INERT rather than MARGINS — even the counter control clears the bar somewhere |
846
+
847
+ **Nothing here becomes a candidate.** Fourteen of the sixteen rows move margins
848
+ somewhere — the control included, once the wide basket is counted — and the one
849
+ pairing that re-prices anything does so against a squad that was already tied for
850
+ best. The two that move nothing at all are zone press, in both modes. The pass
851
+ set at the plan's own activation gate is empty over both baskets.
852
+
853
+ **The pool is now measured end to end, and nothing in it clears the gate.** An
854
+ earlier run covered only four
855
+ axes and said so: schedule none of those four, and go measure the other three
856
+ before concluding anything about the pool. Those three — set-piece `dBonus`,
857
+ base `dBonus` and penalty attack noise — are in the table above and land where
858
+ the first four did.
859
+
860
+ And the pool is now COMPLETE: the zone-press coefficient, the one parameter the
861
+ plan named but forbade "until #730 decides the selector", is swept here too — the
862
+ selector was decided, so the prohibition lapsed. It comes back INERT, the only
863
+ axis in the file that moves no fixture in either basket in either mode.
864
+
865
+ Two axes are covered END TO END, because the plan gives them ranges and the sweep
866
+ spans them: set-piece `dBonus` over 3..8 (five swept values plus the default 5)
867
+ and base `dBonus` over 6..10 (four plus the default 8). **Penalty attack noise
868
+ and zone press are not**: §8 3.1 names their constants — `rand(16)` and
869
+ `ratio * 3` — and no range, so this run picked ±25% and ±50% around each and the
870
+ conclusions are bounded to 8..24 and 1.5..4.5. A rule proposing a value outside
871
+ those spans is unmeasured, and these do not interpolate — changing them reroutes
872
+ the RNG stream. **No axis produced a candidate that clears the gate: the largest
873
+ gain reachable at ANY exposure weight, across all eight axes, both baskets and
874
+ 20,146 solved interior weights, is 0.00pp against a 5pp threshold.**
875
+
876
+ **And that is a failure to distinguish, not a demonstration that nothing is
877
+ there.** The distinction is the whole reading of this section, so it is worth
878
+ being exact about. A gate cell is scored by pairing the swept optimum against the
879
+ stale one; every one of the 99 cells per basket formed NO PAIR, because the
880
+ swept optimum was already inside the default band. A cell that forms no pair
881
+ contributes a zero gain, and 99 zeros is what produces the 0.00pp above — so the
882
+ headline number is the ABSENCE of a resolvable contest, not the presence of a
883
+ measured null. Whether an axis re-prices the decision is therefore **UNRESOLVED**
884
+ on this evidence.
885
+
886
+ What would resolve it is a sharper band. At its worst the default band holds 2 of
887
+ the 5 narrow squads, with the edge 1.51pp from the leader, and 3 of 11 in the wide
888
+ basket at 2.29pp: a band that wide absorbs any new leader before the gate can
889
+ weigh it. A basket whose members sit further apart would shrink it. A lower bar
890
+ would not — that would only relabel the same inability.
891
+
892
+ One thing this is definitely NOT: "these constants do nothing". PK attack noise
893
+ moves a knockout fixture by 9.7pp and base `dBonus` moves a regulation one by
894
+ 2.9pp; margins move everywhere, and the wide basket below changes the top build
895
+ in 12 of 33 knockout cells. What none of that reaches is the gate.
896
+
897
+ **Scope, stated once and plainly.** Every row above is measured at `cond` 5 on
898
+ both sides, and production has not been that since #794 turned the form system
899
+ on. `getPoint` reads condition in both its individual and its team term, so a
900
+ swept threshold can interact with it; this run cannot see that interaction. So
901
+ the claim this section supports is **no pool axis clears the gate under neutral
902
+ conditions**, not "under production conditions" — closing that gap means
903
+ re-running the pool over representative frozen condition states, which is a
904
+ separate measurement and is not done here.
905
+
906
+ **What that means for 3.1.** The pool produced no schedulable candidate, and
907
+ this run cannot say whether that is because none exists or because its basket
908
+ cannot resolve one. So 3.1's premise — that a weekly rule rotates which build is
909
+ worth fielding — is **not supported by anything measured here, and not refuted
910
+ either.** Two measurements would move it: a basket whose members sit further
911
+ apart, which is what the unformed gates are asking for, and a rerun over
912
+ production's condition states. Building the feature on this evidence would be
913
+ building on an unresolved gate; shelving it permanently on this evidence would
914
+ be over-reading the same gate in the other direction. That is a product call
915
+ rather than a
916
+ measurement one, and it is recorded on #733.
846
917
 
847
918
  The verdict is deliberately absent from the rows above: it belongs to the AXIS,
848
919
  not to a mode. The next section is where it is issued.
@@ -899,7 +970,7 @@ nothing, and no amount of simulation shrinks that. Nor can the difference be
899
970
  paired away: the leader and its weakest opponent are argmax and argmin picks
900
971
  that can change IDENTITY between the two conditions, so there is no fixed pair
901
972
  to difference. What this probe does instead is a 2-sigma DIRECTION screen at
902
- +3.0. It refuses far more than a corrected test would — that would sit near 6.4
973
+ +3.0. It refuses far more than a corrected test would — that would sit near 6.6
903
974
  — which is the safe direction for a gate whose job is to refuse. It is a
904
975
  heuristic, it says so in the probe's own legend, and it is not the plan's
905
976
  sentence. The plan does not put numbers on the feel bands, so the ones
@@ -917,10 +988,14 @@ new leader drawn FROM the band scores zero by construction.
917
988
 
918
989
  | Axis | Best value | Gate (ko weight 0.081) | dom | domAbs | draw | goals | Blocked by |
919
990
  | --- | --- | --- | --- | --- | --- | --- | --- |
920
- | FK occurrence | x0.25 | **0.00pp** | +0.4 | 0.4 | **+5.1pp** | −16% | gain — and it would also fail the draw band |
921
- | Post play | x0.25 | **0.00pp** | −0.1 | 0.0 | −1.0pp | +3% | gain |
922
- | Counter (control) | 0% | **0.00pp** | +0.7 | 0.9 | +1.4pp | −6% | gain |
923
- | `fair` | 1 | **0.00pp** | +0.5 | 0.6 | +3.6pp | −12% | gain |
991
+ | FK occurrence | x0.25 | **0.00pp** | +0.1 | 0.3 | **+5.0pp** | −16% | gain — and it would also fail the draw band |
992
+ | Post play | x0.25 | **0.00pp** | +0.7 | 1.1 | −1.1pp | +3% | gain |
993
+ | Counter (control) | 0% | **0.00pp** | +0.4 | 0.8 | +1.4pp | −6% | gain |
994
+ | `fair` | 1 | **0.00pp** | −0.2 | 0.2 | +3.7pp | −12% | gain |
995
+ | Set-piece `dBonus` | 3 | **0.00pp** | +0.6 | 0.2 | −2.4pp | +8% | gain |
996
+ | Base `dBonus` | 6 | **0.00pp** | −0.2 | 0.2 | **−5.1pp** | +22% | gain |
997
+ | PK attack noise | 8 | **0.00pp** | +1.1 | 0.1 | +1.1pp | −4% | gain |
998
+ | Zone press | x0.5 | **0.00pp** | +1.1 | 0.2 | −1.3pp | +5% | gain — and it is INERT, not MARGINS |
924
999
 
925
1000
  `dom` is the CHANGE in the leader's weakest margin; `domAbs` is that margin
926
1001
  itself, in the eleven-archetype basket. Both block, and the two are not
@@ -928,7 +1003,7 @@ interchangeable: a rule that leaves an already-dominant squad exactly where it
928
1003
  stood scores a `dom` near zero while failing the requirement outright, which is
929
1004
  why the absolute column exists at all. `domAbs` is the one carrying the plan's
930
1005
  sentence — at or past the corrected bar, the same |z| the rest of this section
931
- uses, 4.53 at these sample sizes, one squad significantly beats every other,
1006
+ uses, 4.68 at these sample sizes, one squad significantly beats every other,
932
1007
  which is what a dominant build is. `dom` is the direction screen described
933
1008
  above, at a threshold this probe declares rather than quotes.
934
1009
 
@@ -957,6 +1032,10 @@ table is not four verdicts.)
957
1032
  | Post play | 0.00pp | 0.00pp | 0.00pp |
958
1033
  | Counter (control) | 0.00pp | 0.00pp | 0.00pp |
959
1034
  | `fair` | 0.00pp | 0.00pp | 0.00pp |
1035
+ | Set-piece `dBonus` | 0.00pp | 0.00pp | 0.00pp |
1036
+ | Base `dBonus` | 0.00pp | 0.00pp | 0.00pp |
1037
+ | PK attack noise | 0.00pp | 0.00pp | 0.00pp |
1038
+ | Zone press | 0.00pp | 0.00pp | 0.00pp |
960
1039
 
961
1040
  Those three points do not settle the interior on their own: each squad's
962
1041
  weighted score is linear in the weight, the leader is their upper envelope, and
@@ -979,11 +1058,11 @@ on its own grid answers a different question. Evaluating every boundary, and
979
1058
  inside every resulting interval its midpoint plus both ONE-SIDED limits (the
980
1059
  gain jumps where the leader changes, so a shared endpoint reports the wrong
981
1060
  squad), plus the weights where the two baskets' gain lines cross — that is where
982
- `min(narrow, wide)` peaks when their slopes oppose — comes to 9,816 probe
1061
+ `min(narrow, wide)` peaks when their slopes oppose — comes to 20,146 probe
983
1062
  weights across every value:
984
1063
 
985
1064
  > **The largest gain reachable at ANY exposure weight is 0.00pp**, against a 5pp
986
- > threshold. At all 9,816 probe weights the swept optimum was already inside the
1065
+ > threshold. At all 20,146 probe weights the swept optimum was already inside the
987
1066
  > default band in at least one basket.
988
1067
 
989
1068
  That statement deliberately carries no significance claim, and does not need
@@ -1017,8 +1096,8 @@ on every run instead of implying it did.
1017
1096
 
1018
1097
  **And the near miss is not a near miss.** FK x0.25 fails the gain gate outright —
1019
1098
  0.00pp, because the build it promotes was already inside the default band — and
1020
- it would fail the draw band too, at +5.1pp against a +/-5pp limit. Its dominance
1021
- is fine on the weighted season (+0.4 change, 0.4 absolute), which is worth
1099
+ it would fail the draw band too, at +5.0pp against a +/-5pp limit. Its dominance
1100
+ is fine on the weighted season (+0.1 change, 0.3 absolute), which is worth
1022
1101
  stating precisely: an earlier version of this section read dominance off the two
1023
1102
  pure modes and reported +4.4, and that was an artefact of not weighting. The
1024
1103
  candidate that looked closest to shippable still fails, but on the gain and the
@@ -1027,8 +1106,9 @@ feel band, not on concentration.
1027
1106
  ### What the control calibrates
1028
1107
 
1029
1108
  The counter trigger is in this sweep to check that the bar rejects things, and
1030
- under a paired test it is **not** inert: 3 of 40 narrow fixture deltas clear the
1031
- bar, and 12 of 440 across the wide basket's two modes. They are small — and no
1109
+ under a paired test it is **not** inert: 2 of its 80 narrow fixture deltas clear
1110
+ the bar (both in regulation; knockouts give 0 of 40), and 8 of 440 across the
1111
+ wide basket's two modes. They are small — and no
1032
1112
  value of it ever produced a gate, in either basket, at any exposure weight. That
1033
1113
  is the useful result:
1034
1114
 
@@ -1038,22 +1118,34 @@ is the useful result:
1038
1118
 
1039
1119
  It also supplies a floor. In the NARROW basket — same N, same ten fixtures on
1040
1120
  both sides, the only comparison here that is apples to apples — the largest
1041
- movement this known non-lever produces is **2.0pp**. Anything a candidate does
1121
+ movement this known non-lever produces is **1.8pp**. Anything a candidate does
1042
1122
  that is not comfortably above that is not distinguishable from what a knob
1043
1123
  nobody would ship already does:
1044
1124
 
1045
- | Axis | Mode | Largest fixture move | vs the control's 2.0pp |
1125
+ | Axis | Mode | Largest fixture move | vs the control's 1.8pp |
1046
1126
  | --- | --- | --- | --- |
1047
- | FK occurrence | reg | 11.8pp | **6.0x** |
1048
- | FK occurrence | ko | 8.8pp | **4.4x** |
1049
- | `fair` | reg | 10.9pp | **5.5x** |
1050
- | `fair` | ko | 10.9pp | **5.5x** |
1051
- | Post play | reg | 4.8pp | 2.4x |
1052
- | Post play | ko | 3.4pp | **1.7x** — the closest any candidate comes to the floor |
1053
-
1054
- The probe flags a row as indistinguishable from the control below **1.5x**, so
1055
- post play in knockouts clears that line — but only just, and by less than any
1056
- other candidate-mode pair in the table.
1127
+ | FK occurrence | reg | 11.8pp | **6.6x** |
1128
+ | FK occurrence | ko | 8.5pp | **4.8x** |
1129
+ | `fair` | reg | 11.4pp | **6.4x** |
1130
+ | `fair` | ko | 10.7pp | **6.0x** |
1131
+ | PK attack noise | ko | 9.7pp | **5.5x** |
1132
+ | Post play | reg | 5.0pp | 2.8x |
1133
+ | Post play | ko | 4.1pp | 2.3x |
1134
+ | Base `dBonus` | ko | 3.7pp | 2.1x |
1135
+ | Base `dBonus` | reg | 2.9pp | 1.6x |
1136
+ | Set-piece `dBonus` | reg | 1.7pp | **0.9x — inside the control's own noise** |
1137
+ | Set-piece `dBonus` | ko | 1.4pp | **0.8x — inside the control's own noise** |
1138
+ | PK attack noise | reg | 0.9pp | **0.5x — inside the control's own noise** |
1139
+ | Zone press | ko | 1.0pp | **0.6x — inside the control's own noise** |
1140
+ | Zone press | reg | 0.8pp | **0.5x — inside the control's own noise** |
1141
+
1142
+ The probe flags a row as indistinguishable from the control below **1.5x**.
1143
+ Five of the fourteen candidate rows are below it — the control is the floor, so
1144
+ it is not one of them — and all five belong to the axes this run added: set-piece
1145
+ `dBonus` in both modes, PK noise in regulation, and zone press in both modes move
1146
+ a fixture LESS than a knob nobody would ship. PK noise in knockouts is the
1147
+ opposite case — 5.5x, fourth largest behind FK regulation at 6.6x and `fair` at
1148
+ 6.4x and 6.0x — which is the shootout, and still no gate.
1057
1149
 
1058
1150
  The control is swept in BOTH baskets — that is what the gate needs — but this
1059
1151
  ratio is deliberately narrow-only on both sides. It compares a MAXIMUM over a
@@ -1065,43 +1157,53 @@ these ratios by a factor of two between runs that changed no measurement.
1065
1157
  ### Per axis, what actually moved
1066
1158
 
1067
1159
  **FK occurrence is the strongest axis, and what it does is CONCENTRATE.** From
1068
- 2.5% to 40% the regulation draw rate runs 57.9% -> 52.8% (shipped) -> 38.1% and
1069
- goals per match 0.653 -> 0.779 -> 1.288. Every large delta at x4 has the same
1160
+ 2.5% to 40% the regulation draw rate runs 57.7% -> 52.6% (shipped) -> 38.0% and
1161
+ goals per match 0.658 -> 0.787 -> 1.290. Every large delta at x4 has the same
1070
1162
  shape: the differentiated squads pull away from `flat-433` (`role-433` vs
1071
- `flat-433` 68.4% -> 80.2% of the win-equivalent share). Turning it DOWN
1072
- compresses the field instead — `flat-433`'s points per match rises 0.769 ->
1073
- 0.842 while every other squad's falls. So the axis is a skill-expression dial:
1163
+ `flat-433` 68.1% of the win-equivalent share at the shipped rate, +11.4pp of it
1164
+ at x4). Turning it DOWN
1165
+ compresses the field instead — `flat-433`'s own share rises 0.335 -> 0.368 while
1166
+ every other squad's falls. So the axis is a skill-expression dial:
1074
1167
  more free kicks means more of the match decided by who invested, less means more
1075
1168
  of it decided by nothing. Turning it down therefore costs on both feel gates at
1076
1169
  once — more draws AND fewer goals — and in knockouts alone it also raises the
1077
- leader's weakest margin from z 6.3 to 8.9, which points at concentration rather
1170
+ leader's weakest margin from z 5.2 to 6.6, which points at concentration rather
1078
1171
  than rotation. Read that as the knockout-only reading it is: on the
1079
1172
  exposure-weighted season the same change is +0.4, well inside the direction
1080
1173
  screen, and the activation table above reports the weighted number.
1081
1174
 
1082
1175
  **`fair` is the only axis that ROTATES demand rather than amplifying it.**
1083
- Raising it from 4.5% to 100% moves `shootonly-433` from 1.127 to 1.302 points
1084
- per match in regulation and 1.370 to 1.518 in knockouts, while `role-433` and
1085
- `def-433` stand still — more fouls means more free kicks and penalties, and
1086
- those are converted by `shoot`. It carries the run's single reversal: `role-433`
1087
- vs `def-433` goes from **47.5%** at the shipped default to **50.8%** at
1088
- `fair 22`. Read that precisely — the DELTA is significant, the new level is a
1176
+ Raising it from the shipped 22.7% to 100% moves `shootonly-433`'s
1177
+ win-equivalent share from 0.451 to 0.498 in regulation and 0.456 to 0.505 in
1178
+ knockouts, while `role-433` and `def-433` stand still or fall — more fouls means
1179
+ more free kicks and penalties, and those are converted by `shoot`. It carries the
1180
+ run's single reversal: `role-433` vs `def-433` goes from **47.2%** at the shipped
1181
+ default to **50.3%** at `fair 22`. Read that precisely — the DELTA is significant, the new level is a
1089
1182
  coin flip. The rule turned "the defensive lean is better" into "there is no
1090
1183
  difference", which is a re-pricing; it did not turn it into "the attacking lean
1091
1184
  is better".
1092
1185
 
1093
1186
  `fair` is also the only axis whose regulation dominance column does not move
1094
1187
  (+0.0). At the midpoint (`fair 11`, 50%) it still moves four fixtures
1095
- significantly while taking the draw rate DOWN, 52.8% -> 48.6%, and goals up,
1096
- 0.779 -> 0.899 — the one cell in the whole sweep whose feel side-effects point
1188
+ significantly while taking the draw rate DOWN, 52.6% -> 48.4%, and goals up,
1189
+ 0.787 -> 0.903 — the one cell in the whole sweep whose feel side-effects point
1097
1190
  in a direction anyone would ask for.
1098
1191
 
1099
- **Post play is the weakest candidate.** Its largest move is 2.4x the control's
1100
- in regulation and 1.7x in knockouts — the closest any candidate comes to a knob
1101
- that gates 3.9% of goals. It needs x4 (60% of box entries) to reach even that,
1102
- the feel barely registers (draw rate +4.3pp, goals −0.095), and across all eight
1103
- of its wide-basket cells — four swept values in each mode — the top of the
1104
- eleven-archetype table never changes.
1192
+ **Set-piece `dBonus` is the weakest axis that moves anything at all.** Zone press
1193
+ is weaker and moves nothing: it is the only INERT row here, 0 significant deltas
1194
+ in either basket. Set-piece `dBonus` does clear the bar in 15 narrow and 56 wide
1195
+ regulation cells, but its largest move is 0.9x the control's in regulation and
1196
+ 0.8x in knockouts — a knob that gates 3.9% of goals moves a fixture MORE than it
1197
+ does. The feel barely registers either —
1198
+ its regulation draw rate spans 50.2% to 55.9% across the whole 3..8 range — yet
1199
+ it still changes the top of the
1200
+ eleven-archetype knockout table in two of its ten wide-basket cells, which is
1201
+ the whole lesson of the wide basket restated: moving a table's top is cheap,
1202
+ and moving the decision is what the gate asks about.
1203
+
1204
+ **Post play needs its whole range to reach 2.8x.** It gets there only at x4
1205
+ (60% of box entries), and the feel barely registers on the way (draw rate
1206
+ +4.4pp, goals −0.096).
1105
1207
 
1106
1208
  **The counter trigger never gates.** Swept from "counters never happen" to
1107
1209
  "every eligible turnover becomes one", goals per match moved 0.036, the basket's
@@ -1117,8 +1219,7 @@ The narrow basket's biggest weakness is that five archetypes may simply not
1117
1219
  contain the alternative a rule rewards. So EVERY swept value of every axis — the
1118
1220
  control included, so it is checked on the same surface as the candidates — is
1119
1221
  re-asked over the ELEVEN archetypes the round-robin publishes, 55 fixtures,
1120
- 4,000 matches each. That is 34 cells; the probe prints them all, and the shape
1121
- is simple enough to state:
1222
+ 4,000 matches each. That is 68 cells, 34 per mode; the probe prints them all.
1122
1223
 
1123
1224
  Standings here are the **win-equivalent share** — the same 1 / 0.5 / 0 statistic
1124
1225
  the gate measures, so the leader they name is the squad the gate is then
@@ -1126,26 +1227,47 @@ evaluated on. Ranking by league points instead is not a monotonic
1126
1227
  transformation of it when draw rates differ, and these constants move draw rates
1127
1228
  by fifteen points.
1128
1229
 
1129
- | Mode | Shipped rule's top four | Last | Cells whose top changed |
1230
+ | Mode | Shipped rule's top four | Last | Swept cells whose top changed |
1130
1231
  | --- | --- | --- | --- |
1131
- | reg | `gkmin-433` .603, `def-433` .586, `def-532` .585, `role-433` .564 | `flat-433` .365 | **none of 16** |
1132
- | ko | `role-433` .615, `gkheavy-433` .599, `def-433` .591, `def-532` .573 | `flat-433` .259 | **2 of 16**, both FK down |
1133
-
1134
- In regulation, `gkmin-433` leads every single cell — nothing moves the answer at
1135
- any value of any axis. The two knockout exceptions are FK x0.25 and FK x0.5,
1136
- which both promote `gkheavy-433` (+7.5pp z 4.7 and +6.0pp z 3.8 against
1137
- `role-433`; 0.0pp against the tied band it belongs to). `flat-433` is last in 32
1138
- of the 34 cells; the two it is not are FK x0.25 and x0.5 in regulation, where
1139
- `stars-433` drops below it — that build wins and loses far more than it draws,
1140
- so a metric that counts draws at half costs it more than a points table does.
1141
-
1142
- **That one axis-and-direction is the most interesting number in this section.**
1143
- Read as a pair in both rules:
1144
-
1145
- > `gkheavy-433` vs `role-433`, knockout: **49.8% (z −0.3) at the shipped FK rate
1146
- > -> 53.7% (z +4.7) at FK x0.25.** A coin flip becomes a significant win. Those
1147
- > are shares; the +7.5pp above is the same reading as a win-rate advantage, which
1148
- > is exactly twice the share advantage over 50%.
1232
+ | reg | `gkmin-433` .601, `def-532` .586, `def-433` .586, `role-433` .557 | `flat-433` .365 | **1 of 33** (`fair 22`, and not significantly) |
1233
+ | ko | `role-433` .603, `gkheavy-433` .597, `def-433` .595, `def-532` .572 | `flat-433` .258 | **12 of 33**, two of them significant |
1234
+
1235
+ Regulation is almost immovable: `gkmin-433` leads 32 of the 33 swept cells, and
1236
+ the one exception (`fair 22`, where `def-532` reads .592 against `gkmin-433`'s
1237
+ .590) is not
1238
+ significant. `flat-433` is last in all 68 cells.
1239
+
1240
+ **Knockouts are where the newly swept axes show up.** Twelve cells change the
1241
+ top build, and the two that do so significantly are `fk x0.25` (+9.8pp, z 6.2)
1242
+ and `pknoise 8` (+11.4pp, z 7.2), both promoting `gkheavy-433` over `role-433`;
1243
+ `pknoise 12` (+7.3pp, z 4.7) misses the corrected bar. Base `dBonus` and set-piece `dBonus` change the
1244
+ top in six more cells, none significantly.
1245
+
1246
+ **Read the width axes as mean AND variance, not variance alone.** `randFloat(w)`
1247
+ is uniform on [0, w), so its mean is w/2: taking base `dBonus` from 8 to 10 does
1248
+ not just widen the defender's roll, it raises that defender's expected score from
1249
+ 4 to 5, and narrowing it to 6 lowers the mean to 3. The kicker's PK roll moves the
1250
+ same way in the other direction. So the pattern — `fk 0.25`/`0.5`, `pknoise 8`/`12`
1251
+ and `dbonus 9`/`10` promoting the 29-point keeper, while `dbonus 6`/`7` and
1252
+ `pknoise 24` hand `def-433` the lead — is consistent with "whose contest is made
1253
+ decisive" AND with a straight shift in who is favoured on average, and this run
1254
+ cannot separate the two. A rule designer reading a mechanism off these rows would
1255
+ be guessing; separating them needs rolls centred on their shipped mean, which is
1256
+ a different experiment from the one the plan's parameter names.
1257
+
1258
+ **None of it produces a gate.** The promoted build was already inside the
1259
+ default band every time — that is why the gate column reads 0.00pp for all seven
1260
+ axes while this table moves as much as it does. Changing which build is at the
1261
+ TOP of a table is not the same as re-pricing the DECISION, and this section is
1262
+ the clearest demonstration of the difference in the file.
1263
+
1264
+ **The FK direction is the pair worth reading in full**, because the same
1265
+ promotion arrives three separate ways and this is the one with a shipped-rule
1266
+ baseline to compare against:
1267
+
1268
+ > `gkheavy-433` vs `role-433`, knockout: **50.9% (z +1.1) at the shipped FK rate
1269
+ > -> 54.9% (z +6.2) at FK x0.25.** A coin flip becomes a significant win. Those
1270
+ > are shares; a win-rate advantage is exactly twice the share advantage over 50%.
1149
1271
 
1150
1272
  Mechanically that is coherent: fewer free kicks means fewer set-piece goals,
1151
1273
  more matches level at full time, more shootouts, and the 29-point keeper the
@@ -1169,31 +1291,32 @@ since none of these constants is scheduled.
1169
1291
 
1170
1292
  And it is worth **0.0pp at the activation gate** — in the mode with the LEAST
1171
1293
  exposure, since knockouts are cup-only and the cup runs once a day for the
1172
- ladder's top 48. Two things reduce it from +7.5pp to nothing, and both are in
1294
+ ladder's top 48. Two things reduce it from +9.8pp to nothing, and both are in
1173
1295
  the activation section above. `gkheavy-433` is inside the shipped rule's
1174
1296
  unresolved top band, so a manager could already have been on it for free. And
1175
- the same cell moves the draw rate +5.1pp, outside the +/-5pp band a rule change
1176
- is allowed to move it, so even a real gain there would have been refused.
1297
+ the same cell moves the draw rate +5.0pp, at the edge of the +/-5pp band a rule
1298
+ change is allowed to move it, so even a real gain there would have been
1299
+ argued about before it shipped.
1177
1300
 
1178
1301
  Incidentally, the shipped rows above are an 11-archetype round-robin at 4,000
1179
1302
  matches per fixture — roughly seven times the 300 seeds behind the published
1180
1303
  table — and they agree with it where that table says it is decidable:
1181
1304
  `gkmin-433` clear at the top of regulation, `role-433` at the top of knockouts,
1182
- `flat-433` at the bottom of all but two. The middle band still reorders between
1305
+ `flat-433` at the bottom of every one of the 68 cells. The middle band still reorders between
1183
1306
  the two runs, as that section says it must — and note that the bottom is where
1184
1307
  the metric matters: on a points table `stars-433` is above `flat-433`, and on
1185
1308
  the win-equivalent share the two swap in a couple of cells. The shipped `role-433` vs `def-433`
1186
1309
  fixture also independently reproduces the Skill's "solidity beats aggression"
1187
- pair at 10,000 matches: the attacking lean takes 47.5% of the win-equivalent
1188
- share in regulation and 53.1% in knockouts.
1310
+ pair at 10,000 matches: the attacking lean takes 47.2% of the win-equivalent
1311
+ share in regulation and 52.6% in knockouts.
1189
1312
 
1190
- **Do not read that 47.5% against the Skill's 56.5%** — they are different
1313
+ **Do not read that 47.2% against the Skill's 56.5%** — they are different
1191
1314
  statistics on the same fact. The Skill quotes DECIDED matches, this section
1192
1315
  quotes the win-equivalent share, and half of all regulation matches are draws,
1193
1316
  so the share is compressed toward 50 exactly as the round-robin correction above
1194
- warns. The knockout figures (53.1% here, 53.6% there) are directly comparable,
1195
- because knockout draws are vanishingly rare — zero in the 5,440,000 knockout matches
1196
- this run played, though not structurally impossible.
1317
+ warns. The knockout figures (52.6% here, 53.6% there) are directly comparable,
1318
+ because knockout draws are vanishingly rare — zero in the 10,880,000 knockout
1319
+ matches this run played, though not structurally impossible.
1197
1320
 
1198
1321
  ### Degrees of freedom this sweep LEFT free
1199
1322
 
@@ -1211,12 +1334,18 @@ uncontrolled degree of freedom each time. So, explicitly:
1211
1334
  and the zone press, the team-contribution terms and possession all read it.
1212
1335
  - **Set-piece slots.** `squad-lib` pins the FK taker to slot 9 and the penalty
1213
1336
  taker to slot 10 for every candidate. The FK rows are the RATE axis only.
1214
- - **Slot order.** Identical across the basket, which pins the broken zone-press
1215
- fallback to the same position in all of them. Held constant, not measured.
1337
+ - **Slot order.** Identical across the basket. It used to carry a confound —
1338
+ the zone-press fallback made the first non-GK slot a hidden lever — and #730
1339
+ removed it, so the zone-press rows here are not contaminated by it. Held
1340
+ constant either way, and not measured.
1216
1341
  - **Interactions.** One axis moves at a time. Nothing here says what FK x4 does
1217
1342
  while `fair` is 22, and a pool built from two axes at once is untested.
1218
- - **`cond`.** Fixed at 5 everywhere, as production does today. A form system
1219
- would invalidate every row.
1343
+ - **`cond`.** Fixed at 5 everywhere, and production no longer is: #794 turned
1344
+ the form system on, so competitive matches now arrive with per-player
1345
+ conditions derived and frozen at claim time. `getPoint` reads `cond` in both
1346
+ the individual and the team term, so a swept threshold can interact with it,
1347
+ and **every row here is a neutral-condition row.** That is the single largest
1348
+ scope limit on this section's conclusion — see the verdict, which states it.
1220
1349
  - **Opponent population.** Archetypes, not the live ladder — and production bots
1221
1350
  are a single build, which is neither.
1222
1351
 
@@ -1703,6 +1832,374 @@ N=2000 npx tsx … # ~50s; the bar is fixed, the detectable EFFECT moves
1703
1832
 
1704
1833
  ---
1705
1834
 
1835
+ ## Does the best BUILD depend on who you play? (`opp-conditional.mts`)
1836
+
1837
+ Recorded 2026-09-09 — docs/tactics/01 track 0.2 ②, #729. Track 2.4 shipped a
1838
+ kickoff lock and a scouting surface that shows the opponent's last fielded
1839
+ eleven (#726). The plan's own condition for that surface having a DECISION
1840
+ value rather than only an information value is written in 01 §5: across a grid
1841
+ of opponent builds, take my best response to each — *if the distinct argmax is
1842
+ 1, scouting's decision value is 0 and 2.4's reason to exist is gone.* Until this
1843
+ run that measurement had not been made; every reversal in this file is
1844
+ conditional on the mode or on my own attack strength, not on the opponent.
1845
+
1846
+ The section above asked the same question of a different axis — which SHEET is
1847
+ best against each opponent, with my 212 points fixed — and found one sheet in
1848
+ every cell's band. This one asks it of the axis a manager actually controls
1849
+ today, the 212-point build.
1850
+
1851
+ ### The statistic, the three tests, and the controls
1852
+
1853
+ The eleven published archetypes from `strategy-probe.mts`, built by the shared
1854
+ `squad-lib.mts` so they are the round-robin's squads by construction, play each
1855
+ other in both modes at 10,000 matches per fixture (5,000 seeds, home and away,
1856
+ SHA-256 keyed on both labels and the side). Each unordered pair is simulated
1857
+ once and read from both sides, so the matrix is antisymmetric by construction.
1858
+ The diagonal is played too: a manager who knows the opponent's archetype can
1859
+ field the same one, so that fixture is a real response with a real answer (50
1860
+ by symmetry). Its 22 cells have a known truth, but they are REPORTED data — the
1861
+ same-build response in their own column, feeding the tests — so nothing aborts on
1862
+ them. The controls are separate fixtures on their own seed keys (below).
1863
+
1864
+ For every opponent and mode, my candidates are all eleven builds. The LEADER is
1865
+ the candidate with the highest win-equivalent share (draws half); the BAND is
1866
+ every candidate the sample cannot separate from the leader under the corrected
1867
+ bar; a cell is RESOLVED when its band has one member. Two candidates against the
1868
+ same opponent play different seed columns, so every comparison is UNPAIRED.
1869
+
1870
+ Bonferroni is priced before the run over the search a band performs: one
1871
+ comparison per unordered pair per (opponent, mode) cell, 11·10/2 = 55 pairs over
1872
+ 22 cells — **m = 1,210, two-sided, |z| ≥ 4.10**. Enumerating every pair already
1873
+ covers whichever pair the data selects as leader and runner-up, so that selection
1874
+ costs nothing extra, and the two-sided tail already covers both directions; an
1875
+ earlier version charged 2 × 55 AND halved α, pricing direction twice for a bar of
1876
+ 4.26. A too-wide bar is not the safe side here — it widens every band and every
1877
+ equivalence bound, and this grid has boundary-sensitive readings. Worst standard
1878
+ error anywhere in the run 0.71 points of share, so the significance threshold is
1879
+ **2.90 points of share (5.8pp of win rate)** and the 80%-power effect
1880
+ **3.49 (7.0pp)**.
1881
+
1882
+ Three tests are reported, and their provenance is not the same. **T1 was
1883
+ declared before the run.** **T2 and T3 were not**: T2 was added after the first
1884
+ run had been read, because "in every band" is a non-rejection and a
1885
+ non-rejection is not an equivalence; T3 was added after T2's result had been
1886
+ read, because "T1 fires and T2 fails" is not a magnitude claim either — T2
1887
+ failing only says no build was SHOWN within the gate everywhere, and every true
1888
+ effect could still be under it. On the run the tables below come from, T2 and
1889
+ T3 are therefore **post-hoc readings of intervals that were already priced**
1890
+ (the Bonferroni family covers every pairwise contrast they use), and the honest
1891
+ name for that is exploratory. So all three were then re-evaluated on a
1892
+ **confirmatory run** — the same fixtures under a fresh seed namespace
1893
+ (`SEED_NAMESPACE=confirm-1`, disjoint seeds, same N), with T2 and T3 fixed in
1894
+ the code before it was started. Its verdicts are reported in their own table
1895
+ below; the numbers in the sections in between are the first run's. The tests
1896
+ are reported **independently**; a difference can be statistically resolved and
1897
+ still smaller than the practical margin, so T1 and T2 can hold at once:
1898
+
1899
+ - **T1 — conditionality DETECTED**: no build is in every opponent's band; each
1900
+ is significantly beaten as a best response by some build against some
1901
+ opponent. A rejection-based positive claim.
1902
+ - **T2 — equivalence ESTABLISHED**: some build is non-inferior to the TRUE best
1903
+ response against every opponent — for that build, the corrected upper bound
1904
+ of (candidate − it) over *every other candidate*, not only the sample leader
1905
+ (the two differ: a bound against the leader alone is trivially satisfied by
1906
+ the leader itself, while one noisy non-leader can deny the simultaneous
1907
+ bound — Knockout's `role-433` and `def-433` columns qualify nobody for that
1908
+ reason),
1909
+ is under the margin in all eleven columns. A positive equivalence claim.
1910
+ - **T3 — conditionality BEYOND THE MARGIN**: every build is beaten, against some
1911
+ opponent, by more than the margin with the corrected LOWER bound of (leader −
1912
+ it). The positive magnitude claim, read from the same priced contrasts at the
1913
+ other end.
1914
+ - The readings: T1 with T3 → **conditional beyond the margin**; T1 with T2 →
1915
+ conditional but within it; T1 alone → conditional, magnitude
1916
+ unresolved; T2 without T1 → **equivalent**; neither T1 nor T2 →
1917
+ **UNDECIDED** at this sample.
1918
+
1919
+ **The margin is 5pp of win rate WITHIN a mode, and that is NOT the repository's
1920
+ gate.** It is deliberately the same number — below it, re-solving stops being
1921
+ worth an agent's trouble — but the gate proper (`01-engine-expansion.md` §5,
1922
+ implemented by `constant-sweep.mts`) is specified on the **exposure-weighted
1923
+ mode mix**, and the modes are nowhere near equally exposed. A qualifier's day is
1924
+ about 1.33 knockout matches against 12 ladder plus the cup's own 3 group matches,
1925
+ which are played in regulation: **8.1% knockout**, and 0% outside the season's
1926
+ top 48. `constant-sweep.mts` says in as many words that issuing a verdict per
1927
+ mode would call cleared what the mix has not.
1928
+
1929
+ **Which way that cuts is not decided here, and the arithmetic must not pretend
1930
+ otherwise.** T3 is EXISTENTIAL: each build meets SOME opponent that beats it by
1931
+ more than the margin. Multiplying that by the 8.1% mode exposure would assume the
1932
+ adverse opponent is met in every knockout fixture — a bracket may hold it rarely
1933
+ or not at all — so the product is not a lower bound on the mix, and the weighted
1934
+ regret can be anywhere from zero upward. This grid also never plays one locked
1935
+ build across the weighted opponent-and-mode mix, which is what the gate actually
1936
+ asks. So clearing the repository gate is **not established**
1937
+ here, and **not ruled out** either. What is established is that the best response
1938
+ is opponent-dependent *within a mode* by more than a within-mode 5pp — the
1939
+ prerequisite for scouting to be worth anything.
1940
+
1941
+ Controls, every one of which aborts the run: each of the eleven builds must
1942
+ reproduce a pinned signature over every field the engine reads from a player —
1943
+ `name`, position, the four attributes, `total`, `cond`, `fair` and both kicker
1944
+ flags — for all eleven players (slot totals alone would let ratios or a scarcity repair
1945
+ drift under the same labels); no two builds byte-identical; **132 true-zero cells**
1946
+ within a family-adjusted |z| < 3.55 of 50 (worst 2.3, `stars-433` on the
1947
+ `stars-433|bal-442` League control key) —
1948
+ correcting matters here: at an uncorrected two-sided 3σ each true null exceeds
1949
+ with probability ≈0.27%, so across 132 streams a valid rerun would fail about
1950
+ three times in ten (1 − 0.9973^132 ≈ 30%). All 132 are CONTROL FIXTURES of their own,
1951
+ never cells read back out of the grid: 22 on a same-label key `a|a` and 110 on a
1952
+ two-distinct-label key, the shape every published cell uses. Both halves matter.
1953
+ A seeding defect keyed on the pair of labels — which the retired FNV one was —
1954
+ passes a same-label-only control while biasing every cell that control exists to
1955
+ validate. And every one carries a `zero-ctl:` prefix, so no control shares a
1956
+ stream with a reported cell: the controls abort at their α, and a rerun under a
1957
+ fresh namespace selected on a control that shared those streams would be picked
1958
+ partly on the estimates themselves. And a
1959
+ published fact held — `cycle-probe.mts` reports `role-433` beating each of
1960
+ its seven other candidates in Knockout, and here the same eight squads say so
1961
+ under the same statistic, smallest decided z 3.8 — played on a `drift-ctl:` seed
1962
+ family disjoint from every reported cell, because a guard that aborts on the
1963
+ reported cells' own estimates would accept only runs reproducing them. That last control is a
1964
+ **compatibility** test, not a demand for renewed significance: at
1965
+ `cycle-probe`'s N=4000 the weakest of the seven pairs would replicate |z| ≥ 3.32
1966
+ with roughly even odds even if nothing had changed, so a rerun that required it
1967
+ would be aborting on its own sampling noise. Instead `role-433`'s decided win
1968
+ rate against each of the seven from the default-namespace run is pinned in the
1969
+ probe, and a run fails only when its own rate is incompatible with the pin at
1970
+ the Bonferroni-7 bound |z| ≤ 2.69 — what "the squads changed or the engine
1971
+ moved" looks like at any N. Worst |z| against the pins: 0.0 in the default
1972
+ namespace (the pins are that family's own default values) and 0.8 in
1973
+ `confirm-1`.
1974
+ The Knockout column also reproduces the cheap keeper's published collapse
1975
+ (`gkmin-433` at 27.8% against `def-532`, inside the 27–42% this file already
1976
+ states), and the League column reproduces the two pairings `cycle-probe.mts`
1977
+ could not separate (`gkmin-433` 50.6% vs `def-433` and 49.2% vs `def-532`,
1978
+ neither significant here either).
1979
+
1980
+ ### League — UNDECIDED: T1 does not fire, T2 does not pass
1981
+
1982
+ | Opponent | Leader | Share | Runner-up | Gap (z) | Band | Within 5pp of the true best |
1983
+ | --- | --- | --- | --- | --- | --- | --- |
1984
+ | `flat-433` | **`gkmin-433`** | 72.0 | `def-433` | +2.5 (5.7) | `gkmin-433` | `gkmin-433` |
1985
+ | `role-433` | `gkmin-433` | 54.0 | `def-532` | +1.0 (2.2) | `def-433` / `gkmin-433` / `def-532` | `gkmin-433` |
1986
+ | `def-433` | `gkmin-433` | 50.6 | `def-433` | +0.2 (0.5) | `def-433` / `gkmin-433` / `def-532` | `def-433` / `gkmin-433` / `def-532` |
1987
+ | `stars-433` | **`gkmin-433`** | 68.1 | `def-532` | +4.0 (7.3) | `gkmin-433` | `gkmin-433` |
1988
+ | `gkheavy-433` | `gkmin-433` | 60.8 | `def-433` | +1.9 (4.1) | `def-433` / `gkmin-433` | `gkmin-433` |
1989
+ | `gkmin-433` | `def-532` | 50.8 | `gkmin-433` | +0.8 (1.8) | `def-433` / `gkmin-433` / `def-532` | `def-532` |
1990
+ | `shootonly-433` | **`gkmin-433`** | 64.3 | `def-433` | +3.7 (7.2) | `gkmin-433` | `gkmin-433` |
1991
+ | `passonly-433` | `def-433` | 64.7 | `gkmin-433` | +1.8 (3.9) | `def-433` / `gkmin-433` | `def-433` |
1992
+ | `def-532` | `def-532` | 50.1 | `def-433` | +0.1 (0.3) | `def-433` / `gkmin-433` / `def-532` | `def-433` / `gkmin-433` / `def-532` |
1993
+ | `bal-442` | `gkmin-433` | 58.8 | `def-433` | +1.2 (2.6) | `def-433` / `gkmin-433` | `gkmin-433` |
1994
+ | `atk-352` | `gkmin-433` | 63.0 | `def-433` | +1.4 (3.0) | `def-433` / `gkmin-433` | `gkmin-433` |
1995
+
1996
+ > **Measured on** — every one of the eleven published archetypes as the
1997
+ > opponent, with all eleven as my candidates (the opponent's own build among
1998
+ > them: its EXPECTED share is 50 by symmetry, and the sampled cell — 50.1 for
1999
+ > `def-532` here — enters the tests like any other, a known truth on reported
2000
+ > data rather than a control). Bold
2001
+ > leaders are RESOLVED cells: the band holds one build. The last column is T2's
2002
+ > per-cell reading — which builds the sample shows within 2.5 points of share of
2003
+ > the TRUE best response: the corrected upper bound of (candidate − it) is under
2004
+ > the margin for EVERY other candidate in the column, not just for the sample
2005
+ > leader. That is why a column can list nobody even though its leader is
2006
+ > trivially zero behind itself — one noisy candidate denies the simultaneous
2007
+ > bound.
2008
+ > **Sample** — 10,000 matches per fixture, home and away, SHA-256 seeds; 66
2009
+ > fixtures in this mode.
2010
+ > **Mode** — League (`gameFlg 0`, draws count half).
2011
+ > **Effect** — the raw leader is `gkmin-433` in 8 of 11 columns and 3 cells
2012
+ > resolve, all three to `gkmin-433`. **T1 does not fire**: `gkmin-433` is in the
2013
+ > band of all eleven opponents, so conditionality is NOT detected — but see the
2014
+ > confirmation below, where a disjoint sample of the same fixtures puts it
2015
+ > outside one band and T1 does fire. That instability is itself the finding
2016
+ > here. **T2 does not pass**: no build — `gkmin-433` included — is shown within
2017
+ > 5pp of every other candidate in every column; against its own archetype the
2018
+ > best sample response is `def-532` and the corrected upper bound of that
2019
+ > 0.8-point gap exceeds the margin, and against `passonly-433` the leader is
2020
+ > `def-433`. So
2021
+ > the League reading is **UNDECIDED**: this sample did not detect an opponent
2022
+ > that re-prices the build, and it cannot rule out re-pricing of up to its own
2023
+ > power — 7.0pp of win rate on the worst cell — either. One build,
2024
+ > `gkmin-433`, is never beaten by MORE than the gate in any column (T3's
2025
+ > per-build reading) — which is consistent with equivalence and is not it. "Field `gkmin-433` regardless" is consistent with the grid; the grid
2026
+ > does not establish it.
2027
+ > **Source** — `opp-conditional.mts`, sections (2) and (3), `REG`.
2028
+
2029
+ ### Knockout — CONDITIONAL BEYOND THE WITHIN-MODE MARGIN: T1 fires, T3 holds
2030
+
2031
+ | Opponent | Leader | Share | Runner-up | Gap (z) | Shootout reach | Band | Within 5pp of the true best |
2032
+ | --- | --- | --- | --- | --- | --- | --- | --- |
2033
+ | `flat-433` | `def-433` | 82.8 | `def-532` | +0.3 (0.6) | 21% | `role-433` / `def-433` / `def-532` | `def-433` |
2034
+ | `role-433` | `gkheavy-433` | 50.6 | `role-433` | +0.3 (0.5) | 26% | `role-433` / `gkheavy-433` | — |
2035
+ | `def-433` | `role-433` | 52.4 | `gkheavy-433` | +0.5 (0.7) | 32% | `role-433` / `def-433` / `gkheavy-433` | `role-433` |
2036
+ | `stars-433` | **`gkmin-433`** | 69.9 | `def-532` | +4.8 (7.3) | 5% | `gkmin-433` | `gkmin-433` |
2037
+ | `gkheavy-433` | `gkmin-433` | 52.0 | `gkheavy-433` | +1.6 (2.2) | 28% | `role-433` / `gkheavy-433` / `gkmin-433` | `gkmin-433` |
2038
+ | `gkmin-433` | **`def-532`** | 72.2 | `def-433` | +8.7 (13.3) | 24% | `def-532` | `def-532` |
2039
+ | `shootonly-433` | `role-433` | 63.9 | `def-433` | +2.3 (3.3) | 18% | `role-433` / `def-433` | `role-433` |
2040
+ | `passonly-433` | `def-433` | 67.8 | `gkheavy-433` | +0.5 (0.7) | 25% | `role-433` / `def-433` / `gkheavy-433` | `def-433` |
2041
+ | `def-532` | **`gkheavy-433`** | 59.9 | `role-433` | +3.0 (4.3) | 47% | `gkheavy-433` | `gkheavy-433` |
2042
+ | `bal-442` | **`gkheavy-433`** | 58.9 | `role-433` | +4.0 (5.7) | 31% | `gkheavy-433` | `gkheavy-433` |
2043
+ | `atk-352` | `gkheavy-433` | 62.3 | `role-433` | +1.7 (2.5) | 26% | `role-433` / `gkheavy-433` | `gkheavy-433` |
2044
+
2045
+ > **Measured on** — the same eleven archetypes, the same construction. Shootout
2046
+ > reach is the share of that opponent's matches, averaged over all eleven
2047
+ > candidates, that went to penalties — DESCRIPTIVE, see below.
2048
+ > **Sample** — 10,000 matches per fixture, home and away, SHA-256 seeds; 66
2049
+ > fixtures in this mode.
2050
+ > **Mode** — Knockout (`gameFlg 1`: extra time and shootout, every match
2051
+ > decided).
2052
+ > **Effect** — five distinct raw leaders and **4 of 11 cells resolve, to three
2053
+ > different builds**: `gkmin-433` against `stars-433` (z 7.3 over the
2054
+ > runner-up), `def-532` against `gkmin-433` (z 13.3 — the cheap keeper's
2055
+ > published collapse, seen from the other side), and `gkheavy-433` against both
2056
+ > `def-532` (z 4.3) and `bal-442` (z 5.7). **T1 fires**: no build is in every
2057
+ > band. `role-433` and `gkheavy-433` share the most bands at seven each, and
2058
+ > neither survives: `role-433` is significantly excluded from all four resolved
2059
+ > cells, `gkheavy-433` from two.
2060
+ > The exclusions are carried by resolved cells at z 4.3 to 13.3, not by noise.
2061
+ > **T2 does not pass**: no build is within 5pp of every other candidate in every
2062
+ > column — and the `role-433` column names nobody at all once the bound is taken
2063
+ > over every candidate rather than the sample leader.
2064
+ > **T3 holds**: every one of the eleven builds is beaten by MORE than 5pp of win
2065
+ > rate, with the corrected lower bound, against at least one opponent. The
2066
+ > `gkmin-433` column alone does it for the other ten — `def-532`'s 72.2 there
2067
+ > clears the margin against every one of them, the closest being `def-433` at
2068
+ > 63.5 — and `def-532` itself is beaten beyond the margin elsewhere, most
2069
+ > clearly by `gkheavy-433` in its own column (59.9 against `def-532`'s 49.5).
2070
+ > So the Knockout
2071
+ > reading is **CONDITIONAL BEYOND THE WITHIN-MODE MARGIN**: the best build is
2072
+ > opponent-conditional, and by more than 5pp of win rate *inside Knockout* — a
2073
+ > claim T3 makes, not one inferred from T2 failing. Whether that clears the
2074
+ > repository's exposure-weighted gate is neither established nor ruled out here:
2075
+ > 5pp is a lower bound, and this grid never plays a locked build across the mix.
2076
+ > **Source** — `opp-conditional.mts`, sections (2) and (3), `KO`.
2077
+
2078
+ ### Confirmation on a fresh seed namespace
2079
+
2080
+ Same eleven squads, same N, seeds prefixed `confirm-1|` so no fixture shares a
2081
+ match with the run above; T2 and T3 were in the code before it started. Every
2082
+ control holds (worst true-zero |z| 2.4, `def-433` on the
2083
+ `def-433|gkheavy-433` League control key; pinned-rate guard worst 1.2), and
2084
+ Knockout reproduces exactly — **League does not**:
2085
+
2086
+ | run | mode | T1 | T2 | T3 | reading | resolved leaders | in every band | never beaten beyond the margin |
2087
+ | --- | --- | --- | --- | --- | --- | --- | --- | --- |
2088
+ | tables above | League | no | no | no | UNDECIDED | `gkmin-433` | `gkmin-433` | `gkmin-433` |
2089
+ | `confirm-1` | League | **fires** | no | no | CONDITIONAL, MAGNITUDE UNRESOLVED | `gkmin-433`, `def-433` | none | `gkmin-433`, `def-532` |
2090
+ | tables above | Knockout | **fires** | no | **holds** | CONDITIONAL BEYOND THE MARGIN (within mode) | `gkmin-433`, `def-532`, `gkheavy-433` | none | none |
2091
+ | `confirm-1` | Knockout | **fires** | no | **holds** | CONDITIONAL BEYOND THE MARGIN (within mode) | `gkmin-433`, `def-532`, `gkheavy-433` | none | none |
2092
+
2093
+ **League's T1 is not reproducible at this sample size, and that is the League
2094
+ result.** T1 asks whether any build sits in every opponent's band; `gkmin-433`
2095
+ does in the reported run and in `confirm-2`, and does not in `confirm-1` — three
2096
+ disjoint samples of the same eleven fixtures, two landing one way and one the
2097
+ other. The `passonly-433` column is where it turns: `gkmin-433` is inside that
2098
+ band at z 3.9 in the reported run and in `confirm-2`, and outside it at z 4.3 in
2099
+ `confirm-1`. Which side of T1 a sample lands on is decided by noise at N=10,000,
2100
+ so nothing about League should be read as "conditionality was detected" or "was
2101
+ not"; the honest statement is UNDECIDED, and a larger N is what would settle it.
2102
+ The Knockout verdict carries none of this: T1 fires and T3 holds in all three,
2103
+ with the same three resolved leaders.
2104
+
2105
+ The confirmatory runs are what the T2/T3 verdicts rest on; the first run's
2106
+ tables remain the published numbers because they are the run every figure in
2107
+ this section was read from, and mixing them would leave no run owning them.
2108
+
2109
+ ### What the grid does and does not say about 2.4
2110
+
2111
+ **What it says.** In Knockout the answer to "which build should I field" changes
2112
+ with the opponent, by amounts the sample resolves (3.0 to 8.7 points of share
2113
+ between the leader and the runner-up in the four resolved cells: 4.8, 8.7, 3.0
2114
+ and 4.0). In League
2115
+ the same grid neither finds such an opponent nor rules one out. The plan's
2116
+ sentence in 01 §5 was written mode-blind; measured, the axis is a decision in
2117
+ one mode and undecided in the other.
2118
+
2119
+ **What it does not say — and an earlier version of this section said it.** It
2120
+ does not say that the cup lock is where that decision lives. The lock (#726)
2121
+ freezes one eleven when the cup opens and reuses it for every fixture of the
2122
+ day, group stage (`gameFlg 0`) and knockout rounds (`gameFlg 1`) alike. A
2123
+ manager cannot learn a knockout opponent and then switch to the build this grid
2124
+ says beats it; the decision the lock actually forces is *one build against a
2125
+ bracket of opponents in sequence*, and a per-opponent best response is not that
2126
+ object. What this grid establishes is the prerequisite: that in the mode the
2127
+ knockout rounds are played in, a per-opponent best response EXISTS to be
2128
+ informed by. Whether a single locked build should be chosen differently given
2129
+ the bracket — and whether the scouting surface changes that choice — needs a
2130
+ bracket-level experiment, and this file does not contain one. As #729 itself
2131
+ notes, a distinct-argmax count of 1 would not have meant "remove 2.4" either:
2132
+ the lock is a commitment device on its own.
2133
+
2134
+ **What produced the conditional responses is not identified.** The three builds
2135
+ that are unique best responses somewhere — `gkmin-433`, `gkheavy-433`,
2136
+ `def-532` — differ from the rest of the grid in keeper budget, but not only in
2137
+ that: `def-532` carries an ordinary keeper and changes formation and back-line
2138
+ allocation, and `GK_MIN`/`GK_HEAVY` redistribute the keeper's points across the
2139
+ whole outfield. The keeper section of this file supplies a mechanism that is
2140
+ consistent with the pattern — in Knockout the keeper's total is
2141
+ outfield-conditional through shootout exposure, and the table's shootout reach
2142
+ runs from 5% (`stars-433`) to 47% (`def-532`) across opponents — but consistency
2143
+ is not attribution. Isolating the keeper would take a keeper-only,
2144
+ fixed-outfield contrast, which this grid does not run.
2145
+
2146
+ Two secondary readings survive at the descriptive level. In League the
2147
+ cheap-keeper build is the raw leader in 8 of 11 columns, which is this file's
2148
+ "smaller keeper total is better in matches that allow draws" seen across the
2149
+ whole grid at once. And `role-433` — the Knockout dominant in the 8-candidate
2150
+ round robin above — is the raw leader in only 2 of 11 Knockout columns once the
2151
+ keeper-budget archetypes are in the grid, and is the unique best response to
2152
+ nobody. Dominance in a round robin and being the best response to each opponent
2153
+ are different questions, and the second is the one scouting asks.
2154
+
2155
+ ### What this grid CANNOT decide
2156
+
2157
+ - **The build space.** Eleven hand-built archetypes, the same set on both axes.
2158
+ That is a sample, not an argmax over 212-point squads. A conditional best
2159
+ response can only appear among builds that exist in the grid, and a build
2160
+ that beats every leader here could exist outside it. Every claim above is a
2161
+ claim about these eleven.
2162
+ - **Nobody re-optimises against a KNOWN opponent.** Every candidate is a
2163
+ pre-built archetype. A squad rebuilt with the opponent's eleven in hand is a
2164
+ different 212-point problem — and it is the search the scouting surface would
2165
+ actually enable. This grid measures whether the choice AMONG fixed builds is
2166
+ conditional, which is the weaker, prerequisite question.
2167
+ - **The bracket.** The lock forces one build against a sequence of opponents;
2168
+ the grid asks about one opponent at a time. The object the lock decides is
2169
+ not measured here.
2170
+ - **Mechanism.** Formation, keeper budget and outfield allocation move together
2171
+ across these archetypes. Which of them re-prices the build is not separated.
2172
+ - **`cond` fixed at 5, tenure 0.** As everywhere in this file. Production
2173
+ applies growth before the engine and, since #794, an involvement-driven
2174
+ condition; neither is in these squads.
2175
+ - **Sheets at the default.** The section above showed the sheet axis moves more
2176
+ than any build change does; this grid holds it at the shipped weights on both
2177
+ sides.
2178
+ - **Power.** 8 of 11 League cells and 7 of 11 Knockout cells are unresolved,
2179
+ and League's T1 flips between two disjoint samples at this N.
2180
+ T2's margin is 5pp within the mode; the worst cell's 80%-power effect is 7.0pp
2181
+ of win rate, so this sample could not have established equivalence at that margin
2182
+ even where it holds — and T2's bound is taken over every candidate, so a
2183
+ weak candidate with a large standard error can deny it on its own. A larger
2184
+ N is the way to move League out of UNDECIDED.
2185
+
2186
+ Re-run after any engine change:
2187
+
2188
+ ```bash
2189
+ npx tsx packages/mcp/skill/reference/probes/opp-conditional.mts # ~2.7M matches, ~31s
2190
+ SEED_NAMESPACE=confirm-2 npx tsx … # a fresh disjoint sample of the same fixtures. Knockout should reproduce;
2191
+ # League's T1 is not expected to — it flips between samples at this N, which is the League finding itself
2192
+ N=4000 npx tsx … # smoke; the tables are shapes, and the pinned-rate control cannot apply at any default-namespace
2193
+ # N other than 10,000 — those streams overlap the pins, so the two rates are paired rather than comparable
2194
+ # 132 `zero-ctl:` control fixtures (22 same-label, 110 pair-key) sit at a family-adjusted bar, which
2195
+ # bounds a valid run's abort probability at 5% — they are NOT independent (the seed key omits the mode,
2196
+ # so each pair's League and Knockout controls share a stream), so 5% is a ceiling, not a rate. The
2197
+ # reported diagonals are NOT checked. Investigate an
2198
+ # exceedance first; only then re-run under another SEED_NAMESPACE — never widen the bar to pass
2199
+ ```
2200
+
2201
+ ---
2202
+
1706
2203
  ## Note for the maintainers
1707
2204
 
1708
2205
  > **2026-08-19: this note was rewritten twice in one day.** First it said the