pog-mcp 0.9.14 → 0.9.16

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pog-mcp",
3
- "version": "0.9.14",
3
+ "version": "0.9.16",
4
4
  "type": "module",
5
5
  "description": "MCP server that lets an AI agent play Proof of Goal — wallet, sign-in, squad building, and matches as typed tools.",
6
6
  "license": "MIT",
@@ -664,10 +664,11 @@ sets it — there is no `fair:` outside the engine's own tests, and every engine
664
664
  input is built from an explicit field list — so it is a global constant in
665
665
  practice, and setting it on both squads IS the rule change.
666
666
 
667
- **Not swept: the zone-press coefficient.** Its selector is broken — every
668
- candidate weight is `0.0`, so `weightedPick` falls back to "the first non-GK
669
- slot" — and a coefficient in front of an undefined selection measures nothing.
670
- Forbidden until the selector is decided.
667
+ **Not swept: the zone-press coefficient.** At the time its selector was broken —
668
+ every candidate weight was `0.0`, so `weightedPick` fell back to "the first non-GK
669
+ slot" — and a coefficient in front of an undefined selection measures nothing. The
670
+ selector was fixed in #730 (`zp-selector.mts` below); the coefficient is sweepable
671
+ now but has not been swept.
671
672
 
672
673
  ### The statistic, and why not the obvious one
673
674
 
@@ -1702,6 +1703,374 @@ N=2000 npx tsx … # ~50s; the bar is fixed, the detectable EFFECT moves
1702
1703
 
1703
1704
  ---
1704
1705
 
1706
+ ## Does the best BUILD depend on who you play? (`opp-conditional.mts`)
1707
+
1708
+ Recorded 2026-09-09 — docs/tactics/01 track 0.2 ②, #729. Track 2.4 shipped a
1709
+ kickoff lock and a scouting surface that shows the opponent's last fielded
1710
+ eleven (#726). The plan's own condition for that surface having a DECISION
1711
+ value rather than only an information value is written in 01 §5: across a grid
1712
+ of opponent builds, take my best response to each — *if the distinct argmax is
1713
+ 1, scouting's decision value is 0 and 2.4's reason to exist is gone.* Until this
1714
+ run that measurement had not been made; every reversal in this file is
1715
+ conditional on the mode or on my own attack strength, not on the opponent.
1716
+
1717
+ The section above asked the same question of a different axis — which SHEET is
1718
+ best against each opponent, with my 212 points fixed — and found one sheet in
1719
+ every cell's band. This one asks it of the axis a manager actually controls
1720
+ today, the 212-point build.
1721
+
1722
+ ### The statistic, the three tests, and the controls
1723
+
1724
+ The eleven published archetypes from `strategy-probe.mts`, built by the shared
1725
+ `squad-lib.mts` so they are the round-robin's squads by construction, play each
1726
+ other in both modes at 10,000 matches per fixture (5,000 seeds, home and away,
1727
+ SHA-256 keyed on both labels and the side). Each unordered pair is simulated
1728
+ once and read from both sides, so the matrix is antisymmetric by construction.
1729
+ The diagonal is played too: a manager who knows the opponent's archetype can
1730
+ field the same one, so that fixture is a real response with a real answer (50
1731
+ by symmetry). Its 22 cells have a known truth, but they are REPORTED data — the
1732
+ same-build response in their own column, feeding the tests — so nothing aborts on
1733
+ them. The controls are separate fixtures on their own seed keys (below).
1734
+
1735
+ For every opponent and mode, my candidates are all eleven builds. The LEADER is
1736
+ the candidate with the highest win-equivalent share (draws half); the BAND is
1737
+ every candidate the sample cannot separate from the leader under the corrected
1738
+ bar; a cell is RESOLVED when its band has one member. Two candidates against the
1739
+ same opponent play different seed columns, so every comparison is UNPAIRED.
1740
+
1741
+ Bonferroni is priced before the run over the search a band performs: one
1742
+ comparison per unordered pair per (opponent, mode) cell, 11·10/2 = 55 pairs over
1743
+ 22 cells — **m = 1,210, two-sided, |z| ≥ 4.10**. Enumerating every pair already
1744
+ covers whichever pair the data selects as leader and runner-up, so that selection
1745
+ costs nothing extra, and the two-sided tail already covers both directions; an
1746
+ earlier version charged 2 × 55 AND halved α, pricing direction twice for a bar of
1747
+ 4.26. A too-wide bar is not the safe side here — it widens every band and every
1748
+ equivalence bound, and this grid has boundary-sensitive readings. Worst standard
1749
+ error anywhere in the run 0.71 points of share, so the significance threshold is
1750
+ **2.90 points of share (5.8pp of win rate)** and the 80%-power effect
1751
+ **3.49 (7.0pp)**.
1752
+
1753
+ Three tests are reported, and their provenance is not the same. **T1 was
1754
+ declared before the run.** **T2 and T3 were not**: T2 was added after the first
1755
+ run had been read, because "in every band" is a non-rejection and a
1756
+ non-rejection is not an equivalence; T3 was added after T2's result had been
1757
+ read, because "T1 fires and T2 fails" is not a magnitude claim either — T2
1758
+ failing only says no build was SHOWN within the gate everywhere, and every true
1759
+ effect could still be under it. On the run the tables below come from, T2 and
1760
+ T3 are therefore **post-hoc readings of intervals that were already priced**
1761
+ (the Bonferroni family covers every pairwise contrast they use), and the honest
1762
+ name for that is exploratory. So all three were then re-evaluated on a
1763
+ **confirmatory run** — the same fixtures under a fresh seed namespace
1764
+ (`SEED_NAMESPACE=confirm-1`, disjoint seeds, same N), with T2 and T3 fixed in
1765
+ the code before it was started. Its verdicts are reported in their own table
1766
+ below; the numbers in the sections in between are the first run's. The tests
1767
+ are reported **independently**; a difference can be statistically resolved and
1768
+ still smaller than the practical margin, so T1 and T2 can hold at once:
1769
+
1770
+ - **T1 — conditionality DETECTED**: no build is in every opponent's band; each
1771
+ is significantly beaten as a best response by some build against some
1772
+ opponent. A rejection-based positive claim.
1773
+ - **T2 — equivalence ESTABLISHED**: some build is non-inferior to the TRUE best
1774
+ response against every opponent — for that build, the corrected upper bound
1775
+ of (candidate − it) over *every other candidate*, not only the sample leader
1776
+ (the two differ: a bound against the leader alone is trivially satisfied by
1777
+ the leader itself, while one noisy non-leader can deny the simultaneous
1778
+ bound — Knockout's `role-433` and `def-433` columns qualify nobody for that
1779
+ reason),
1780
+ is under the margin in all eleven columns. A positive equivalence claim.
1781
+ - **T3 — conditionality BEYOND THE MARGIN**: every build is beaten, against some
1782
+ opponent, by more than the margin with the corrected LOWER bound of (leader −
1783
+ it). The positive magnitude claim, read from the same priced contrasts at the
1784
+ other end.
1785
+ - The readings: T1 with T3 → **conditional beyond the margin**; T1 with T2 →
1786
+ conditional but within it; T1 alone → conditional, magnitude
1787
+ unresolved; T2 without T1 → **equivalent**; neither T1 nor T2 →
1788
+ **UNDECIDED** at this sample.
1789
+
1790
+ **The margin is 5pp of win rate WITHIN a mode, and that is NOT the repository's
1791
+ gate.** It is deliberately the same number — below it, re-solving stops being
1792
+ worth an agent's trouble — but the gate proper (`01-engine-expansion.md` §5,
1793
+ implemented by `constant-sweep.mts`) is specified on the **exposure-weighted
1794
+ mode mix**, and the modes are nowhere near equally exposed. A qualifier's day is
1795
+ about 1.33 knockout matches against 12 ladder plus the cup's own 3 group matches,
1796
+ which are played in regulation: **8.1% knockout**, and 0% outside the season's
1797
+ top 48. `constant-sweep.mts` says in as many words that issuing a verdict per
1798
+ mode would call cleared what the mix has not.
1799
+
1800
+ **Which way that cuts is not decided here, and the arithmetic must not pretend
1801
+ otherwise.** T3 is EXISTENTIAL: each build meets SOME opponent that beats it by
1802
+ more than the margin. Multiplying that by the 8.1% mode exposure would assume the
1803
+ adverse opponent is met in every knockout fixture — a bracket may hold it rarely
1804
+ or not at all — so the product is not a lower bound on the mix, and the weighted
1805
+ regret can be anywhere from zero upward. This grid also never plays one locked
1806
+ build across the weighted opponent-and-mode mix, which is what the gate actually
1807
+ asks. So clearing the repository gate is **not established**
1808
+ here, and **not ruled out** either. What is established is that the best response
1809
+ is opponent-dependent *within a mode* by more than a within-mode 5pp — the
1810
+ prerequisite for scouting to be worth anything.
1811
+
1812
+ Controls, every one of which aborts the run: each of the eleven builds must
1813
+ reproduce a pinned signature over every field the engine reads from a player —
1814
+ `name`, position, the four attributes, `total`, `cond`, `fair` and both kicker
1815
+ flags — for all eleven players (slot totals alone would let ratios or a scarcity repair
1816
+ drift under the same labels); no two builds byte-identical; **132 true-zero cells**
1817
+ within a family-adjusted |z| < 3.55 of 50 (worst 2.3, `stars-433` on the
1818
+ `stars-433|bal-442` League control key) —
1819
+ correcting matters here: at an uncorrected two-sided 3σ each true null exceeds
1820
+ with probability ≈0.27%, so across 132 streams a valid rerun would fail about
1821
+ three times in ten (1 − 0.9973^132 ≈ 30%). All 132 are CONTROL FIXTURES of their own,
1822
+ never cells read back out of the grid: 22 on a same-label key `a|a` and 110 on a
1823
+ two-distinct-label key, the shape every published cell uses. Both halves matter.
1824
+ A seeding defect keyed on the pair of labels — which the retired FNV one was —
1825
+ passes a same-label-only control while biasing every cell that control exists to
1826
+ validate. And every one carries a `zero-ctl:` prefix, so no control shares a
1827
+ stream with a reported cell: the controls abort at their α, and a rerun under a
1828
+ fresh namespace selected on a control that shared those streams would be picked
1829
+ partly on the estimates themselves. And a
1830
+ published fact held — `cycle-probe.mts` reports `role-433` beating each of
1831
+ its seven other candidates in Knockout, and here the same eight squads say so
1832
+ under the same statistic, smallest decided z 3.8 — played on a `drift-ctl:` seed
1833
+ family disjoint from every reported cell, because a guard that aborts on the
1834
+ reported cells' own estimates would accept only runs reproducing them. That last control is a
1835
+ **compatibility** test, not a demand for renewed significance: at
1836
+ `cycle-probe`'s N=4000 the weakest of the seven pairs would replicate |z| ≥ 3.32
1837
+ with roughly even odds even if nothing had changed, so a rerun that required it
1838
+ would be aborting on its own sampling noise. Instead `role-433`'s decided win
1839
+ rate against each of the seven from the default-namespace run is pinned in the
1840
+ probe, and a run fails only when its own rate is incompatible with the pin at
1841
+ the Bonferroni-7 bound |z| ≤ 2.69 — what "the squads changed or the engine
1842
+ moved" looks like at any N. Worst |z| against the pins: 0.0 in the default
1843
+ namespace (the pins are that family's own default values) and 0.8 in
1844
+ `confirm-1`.
1845
+ The Knockout column also reproduces the cheap keeper's published collapse
1846
+ (`gkmin-433` at 27.8% against `def-532`, inside the 27–42% this file already
1847
+ states), and the League column reproduces the two pairings `cycle-probe.mts`
1848
+ could not separate (`gkmin-433` 50.6% vs `def-433` and 49.2% vs `def-532`,
1849
+ neither significant here either).
1850
+
1851
+ ### League — UNDECIDED: T1 does not fire, T2 does not pass
1852
+
1853
+ | Opponent | Leader | Share | Runner-up | Gap (z) | Band | Within 5pp of the true best |
1854
+ | --- | --- | --- | --- | --- | --- | --- |
1855
+ | `flat-433` | **`gkmin-433`** | 72.0 | `def-433` | +2.5 (5.7) | `gkmin-433` | `gkmin-433` |
1856
+ | `role-433` | `gkmin-433` | 54.0 | `def-532` | +1.0 (2.2) | `def-433` / `gkmin-433` / `def-532` | `gkmin-433` |
1857
+ | `def-433` | `gkmin-433` | 50.6 | `def-433` | +0.2 (0.5) | `def-433` / `gkmin-433` / `def-532` | `def-433` / `gkmin-433` / `def-532` |
1858
+ | `stars-433` | **`gkmin-433`** | 68.1 | `def-532` | +4.0 (7.3) | `gkmin-433` | `gkmin-433` |
1859
+ | `gkheavy-433` | `gkmin-433` | 60.8 | `def-433` | +1.9 (4.1) | `def-433` / `gkmin-433` | `gkmin-433` |
1860
+ | `gkmin-433` | `def-532` | 50.8 | `gkmin-433` | +0.8 (1.8) | `def-433` / `gkmin-433` / `def-532` | `def-532` |
1861
+ | `shootonly-433` | **`gkmin-433`** | 64.3 | `def-433` | +3.7 (7.2) | `gkmin-433` | `gkmin-433` |
1862
+ | `passonly-433` | `def-433` | 64.7 | `gkmin-433` | +1.8 (3.9) | `def-433` / `gkmin-433` | `def-433` |
1863
+ | `def-532` | `def-532` | 50.1 | `def-433` | +0.1 (0.3) | `def-433` / `gkmin-433` / `def-532` | `def-433` / `gkmin-433` / `def-532` |
1864
+ | `bal-442` | `gkmin-433` | 58.8 | `def-433` | +1.2 (2.6) | `def-433` / `gkmin-433` | `gkmin-433` |
1865
+ | `atk-352` | `gkmin-433` | 63.0 | `def-433` | +1.4 (3.0) | `def-433` / `gkmin-433` | `gkmin-433` |
1866
+
1867
+ > **Measured on** — every one of the eleven published archetypes as the
1868
+ > opponent, with all eleven as my candidates (the opponent's own build among
1869
+ > them: its EXPECTED share is 50 by symmetry, and the sampled cell — 50.1 for
1870
+ > `def-532` here — enters the tests like any other, a known truth on reported
1871
+ > data rather than a control). Bold
1872
+ > leaders are RESOLVED cells: the band holds one build. The last column is T2's
1873
+ > per-cell reading — which builds the sample shows within 2.5 points of share of
1874
+ > the TRUE best response: the corrected upper bound of (candidate − it) is under
1875
+ > the margin for EVERY other candidate in the column, not just for the sample
1876
+ > leader. That is why a column can list nobody even though its leader is
1877
+ > trivially zero behind itself — one noisy candidate denies the simultaneous
1878
+ > bound.
1879
+ > **Sample** — 10,000 matches per fixture, home and away, SHA-256 seeds; 66
1880
+ > fixtures in this mode.
1881
+ > **Mode** — League (`gameFlg 0`, draws count half).
1882
+ > **Effect** — the raw leader is `gkmin-433` in 8 of 11 columns and 3 cells
1883
+ > resolve, all three to `gkmin-433`. **T1 does not fire**: `gkmin-433` is in the
1884
+ > band of all eleven opponents, so conditionality is NOT detected — but see the
1885
+ > confirmation below, where a disjoint sample of the same fixtures puts it
1886
+ > outside one band and T1 does fire. That instability is itself the finding
1887
+ > here. **T2 does not pass**: no build — `gkmin-433` included — is shown within
1888
+ > 5pp of every other candidate in every column; against its own archetype the
1889
+ > best sample response is `def-532` and the corrected upper bound of that
1890
+ > 0.8-point gap exceeds the margin, and against `passonly-433` the leader is
1891
+ > `def-433`. So
1892
+ > the League reading is **UNDECIDED**: this sample did not detect an opponent
1893
+ > that re-prices the build, and it cannot rule out re-pricing of up to its own
1894
+ > power — 7.0pp of win rate on the worst cell — either. One build,
1895
+ > `gkmin-433`, is never beaten by MORE than the gate in any column (T3's
1896
+ > per-build reading) — which is consistent with equivalence and is not it. "Field `gkmin-433` regardless" is consistent with the grid; the grid
1897
+ > does not establish it.
1898
+ > **Source** — `opp-conditional.mts`, sections (2) and (3), `REG`.
1899
+
1900
+ ### Knockout — CONDITIONAL BEYOND THE WITHIN-MODE MARGIN: T1 fires, T3 holds
1901
+
1902
+ | Opponent | Leader | Share | Runner-up | Gap (z) | Shootout reach | Band | Within 5pp of the true best |
1903
+ | --- | --- | --- | --- | --- | --- | --- | --- |
1904
+ | `flat-433` | `def-433` | 82.8 | `def-532` | +0.3 (0.6) | 21% | `role-433` / `def-433` / `def-532` | `def-433` |
1905
+ | `role-433` | `gkheavy-433` | 50.6 | `role-433` | +0.3 (0.5) | 26% | `role-433` / `gkheavy-433` | — |
1906
+ | `def-433` | `role-433` | 52.4 | `gkheavy-433` | +0.5 (0.7) | 32% | `role-433` / `def-433` / `gkheavy-433` | `role-433` |
1907
+ | `stars-433` | **`gkmin-433`** | 69.9 | `def-532` | +4.8 (7.3) | 5% | `gkmin-433` | `gkmin-433` |
1908
+ | `gkheavy-433` | `gkmin-433` | 52.0 | `gkheavy-433` | +1.6 (2.2) | 28% | `role-433` / `gkheavy-433` / `gkmin-433` | `gkmin-433` |
1909
+ | `gkmin-433` | **`def-532`** | 72.2 | `def-433` | +8.7 (13.3) | 24% | `def-532` | `def-532` |
1910
+ | `shootonly-433` | `role-433` | 63.9 | `def-433` | +2.3 (3.3) | 18% | `role-433` / `def-433` | `role-433` |
1911
+ | `passonly-433` | `def-433` | 67.8 | `gkheavy-433` | +0.5 (0.7) | 25% | `role-433` / `def-433` / `gkheavy-433` | `def-433` |
1912
+ | `def-532` | **`gkheavy-433`** | 59.9 | `role-433` | +3.0 (4.3) | 47% | `gkheavy-433` | `gkheavy-433` |
1913
+ | `bal-442` | **`gkheavy-433`** | 58.9 | `role-433` | +4.0 (5.7) | 31% | `gkheavy-433` | `gkheavy-433` |
1914
+ | `atk-352` | `gkheavy-433` | 62.3 | `role-433` | +1.7 (2.5) | 26% | `role-433` / `gkheavy-433` | `gkheavy-433` |
1915
+
1916
+ > **Measured on** — the same eleven archetypes, the same construction. Shootout
1917
+ > reach is the share of that opponent's matches, averaged over all eleven
1918
+ > candidates, that went to penalties — DESCRIPTIVE, see below.
1919
+ > **Sample** — 10,000 matches per fixture, home and away, SHA-256 seeds; 66
1920
+ > fixtures in this mode.
1921
+ > **Mode** — Knockout (`gameFlg 1`: extra time and shootout, every match
1922
+ > decided).
1923
+ > **Effect** — five distinct raw leaders and **4 of 11 cells resolve, to three
1924
+ > different builds**: `gkmin-433` against `stars-433` (z 7.3 over the
1925
+ > runner-up), `def-532` against `gkmin-433` (z 13.3 — the cheap keeper's
1926
+ > published collapse, seen from the other side), and `gkheavy-433` against both
1927
+ > `def-532` (z 4.3) and `bal-442` (z 5.7). **T1 fires**: no build is in every
1928
+ > band. `role-433` and `gkheavy-433` share the most bands at seven each, and
1929
+ > neither survives: `role-433` is significantly excluded from all four resolved
1930
+ > cells, `gkheavy-433` from two.
1931
+ > The exclusions are carried by resolved cells at z 4.3 to 13.3, not by noise.
1932
+ > **T2 does not pass**: no build is within 5pp of every other candidate in every
1933
+ > column — and the `role-433` column names nobody at all once the bound is taken
1934
+ > over every candidate rather than the sample leader.
1935
+ > **T3 holds**: every one of the eleven builds is beaten by MORE than 5pp of win
1936
+ > rate, with the corrected lower bound, against at least one opponent. The
1937
+ > `gkmin-433` column alone does it for the other ten — `def-532`'s 72.2 there
1938
+ > clears the margin against every one of them, the closest being `def-433` at
1939
+ > 63.5 — and `def-532` itself is beaten beyond the margin elsewhere, most
1940
+ > clearly by `gkheavy-433` in its own column (59.9 against `def-532`'s 49.5).
1941
+ > So the Knockout
1942
+ > reading is **CONDITIONAL BEYOND THE WITHIN-MODE MARGIN**: the best build is
1943
+ > opponent-conditional, and by more than 5pp of win rate *inside Knockout* — a
1944
+ > claim T3 makes, not one inferred from T2 failing. Whether that clears the
1945
+ > repository's exposure-weighted gate is neither established nor ruled out here:
1946
+ > 5pp is a lower bound, and this grid never plays a locked build across the mix.
1947
+ > **Source** — `opp-conditional.mts`, sections (2) and (3), `KO`.
1948
+
1949
+ ### Confirmation on a fresh seed namespace
1950
+
1951
+ Same eleven squads, same N, seeds prefixed `confirm-1|` so no fixture shares a
1952
+ match with the run above; T2 and T3 were in the code before it started. Every
1953
+ control holds (worst true-zero |z| 2.4, `def-433` on the
1954
+ `def-433|gkheavy-433` League control key; pinned-rate guard worst 1.2), and
1955
+ Knockout reproduces exactly — **League does not**:
1956
+
1957
+ | run | mode | T1 | T2 | T3 | reading | resolved leaders | in every band | never beaten beyond the margin |
1958
+ | --- | --- | --- | --- | --- | --- | --- | --- | --- |
1959
+ | tables above | League | no | no | no | UNDECIDED | `gkmin-433` | `gkmin-433` | `gkmin-433` |
1960
+ | `confirm-1` | League | **fires** | no | no | CONDITIONAL, MAGNITUDE UNRESOLVED | `gkmin-433`, `def-433` | none | `gkmin-433`, `def-532` |
1961
+ | tables above | Knockout | **fires** | no | **holds** | CONDITIONAL BEYOND THE MARGIN (within mode) | `gkmin-433`, `def-532`, `gkheavy-433` | none | none |
1962
+ | `confirm-1` | Knockout | **fires** | no | **holds** | CONDITIONAL BEYOND THE MARGIN (within mode) | `gkmin-433`, `def-532`, `gkheavy-433` | none | none |
1963
+
1964
+ **League's T1 is not reproducible at this sample size, and that is the League
1965
+ result.** T1 asks whether any build sits in every opponent's band; `gkmin-433`
1966
+ does in the reported run and in `confirm-2`, and does not in `confirm-1` — three
1967
+ disjoint samples of the same eleven fixtures, two landing one way and one the
1968
+ other. The `passonly-433` column is where it turns: `gkmin-433` is inside that
1969
+ band at z 3.9 in the reported run and in `confirm-2`, and outside it at z 4.3 in
1970
+ `confirm-1`. Which side of T1 a sample lands on is decided by noise at N=10,000,
1971
+ so nothing about League should be read as "conditionality was detected" or "was
1972
+ not"; the honest statement is UNDECIDED, and a larger N is what would settle it.
1973
+ The Knockout verdict carries none of this: T1 fires and T3 holds in all three,
1974
+ with the same three resolved leaders.
1975
+
1976
+ The confirmatory runs are what the T2/T3 verdicts rest on; the first run's
1977
+ tables remain the published numbers because they are the run every figure in
1978
+ this section was read from, and mixing them would leave no run owning them.
1979
+
1980
+ ### What the grid does and does not say about 2.4
1981
+
1982
+ **What it says.** In Knockout the answer to "which build should I field" changes
1983
+ with the opponent, by amounts the sample resolves (3.0 to 8.7 points of share
1984
+ between the leader and the runner-up in the four resolved cells: 4.8, 8.7, 3.0
1985
+ and 4.0). In League
1986
+ the same grid neither finds such an opponent nor rules one out. The plan's
1987
+ sentence in 01 §5 was written mode-blind; measured, the axis is a decision in
1988
+ one mode and undecided in the other.
1989
+
1990
+ **What it does not say — and an earlier version of this section said it.** It
1991
+ does not say that the cup lock is where that decision lives. The lock (#726)
1992
+ freezes one eleven when the cup opens and reuses it for every fixture of the
1993
+ day, group stage (`gameFlg 0`) and knockout rounds (`gameFlg 1`) alike. A
1994
+ manager cannot learn a knockout opponent and then switch to the build this grid
1995
+ says beats it; the decision the lock actually forces is *one build against a
1996
+ bracket of opponents in sequence*, and a per-opponent best response is not that
1997
+ object. What this grid establishes is the prerequisite: that in the mode the
1998
+ knockout rounds are played in, a per-opponent best response EXISTS to be
1999
+ informed by. Whether a single locked build should be chosen differently given
2000
+ the bracket — and whether the scouting surface changes that choice — needs a
2001
+ bracket-level experiment, and this file does not contain one. As #729 itself
2002
+ notes, a distinct-argmax count of 1 would not have meant "remove 2.4" either:
2003
+ the lock is a commitment device on its own.
2004
+
2005
+ **What produced the conditional responses is not identified.** The three builds
2006
+ that are unique best responses somewhere — `gkmin-433`, `gkheavy-433`,
2007
+ `def-532` — differ from the rest of the grid in keeper budget, but not only in
2008
+ that: `def-532` carries an ordinary keeper and changes formation and back-line
2009
+ allocation, and `GK_MIN`/`GK_HEAVY` redistribute the keeper's points across the
2010
+ whole outfield. The keeper section of this file supplies a mechanism that is
2011
+ consistent with the pattern — in Knockout the keeper's total is
2012
+ outfield-conditional through shootout exposure, and the table's shootout reach
2013
+ runs from 5% (`stars-433`) to 47% (`def-532`) across opponents — but consistency
2014
+ is not attribution. Isolating the keeper would take a keeper-only,
2015
+ fixed-outfield contrast, which this grid does not run.
2016
+
2017
+ Two secondary readings survive at the descriptive level. In League the
2018
+ cheap-keeper build is the raw leader in 8 of 11 columns, which is this file's
2019
+ "smaller keeper total is better in matches that allow draws" seen across the
2020
+ whole grid at once. And `role-433` — the Knockout dominant in the 8-candidate
2021
+ round robin above — is the raw leader in only 2 of 11 Knockout columns once the
2022
+ keeper-budget archetypes are in the grid, and is the unique best response to
2023
+ nobody. Dominance in a round robin and being the best response to each opponent
2024
+ are different questions, and the second is the one scouting asks.
2025
+
2026
+ ### What this grid CANNOT decide
2027
+
2028
+ - **The build space.** Eleven hand-built archetypes, the same set on both axes.
2029
+ That is a sample, not an argmax over 212-point squads. A conditional best
2030
+ response can only appear among builds that exist in the grid, and a build
2031
+ that beats every leader here could exist outside it. Every claim above is a
2032
+ claim about these eleven.
2033
+ - **Nobody re-optimises against a KNOWN opponent.** Every candidate is a
2034
+ pre-built archetype. A squad rebuilt with the opponent's eleven in hand is a
2035
+ different 212-point problem — and it is the search the scouting surface would
2036
+ actually enable. This grid measures whether the choice AMONG fixed builds is
2037
+ conditional, which is the weaker, prerequisite question.
2038
+ - **The bracket.** The lock forces one build against a sequence of opponents;
2039
+ the grid asks about one opponent at a time. The object the lock decides is
2040
+ not measured here.
2041
+ - **Mechanism.** Formation, keeper budget and outfield allocation move together
2042
+ across these archetypes. Which of them re-prices the build is not separated.
2043
+ - **`cond` fixed at 5, tenure 0.** As everywhere in this file. Production
2044
+ applies growth before the engine and, since #794, an involvement-driven
2045
+ condition; neither is in these squads.
2046
+ - **Sheets at the default.** The section above showed the sheet axis moves more
2047
+ than any build change does; this grid holds it at the shipped weights on both
2048
+ sides.
2049
+ - **Power.** 8 of 11 League cells and 7 of 11 Knockout cells are unresolved,
2050
+ and League's T1 flips between two disjoint samples at this N.
2051
+ T2's margin is 5pp within the mode; the worst cell's 80%-power effect is 7.0pp
2052
+ of win rate, so this sample could not have established equivalence at that margin
2053
+ even where it holds — and T2's bound is taken over every candidate, so a
2054
+ weak candidate with a large standard error can deny it on its own. A larger
2055
+ N is the way to move League out of UNDECIDED.
2056
+
2057
+ Re-run after any engine change:
2058
+
2059
+ ```bash
2060
+ npx tsx packages/mcp/skill/reference/probes/opp-conditional.mts # ~2.7M matches, ~31s
2061
+ SEED_NAMESPACE=confirm-2 npx tsx … # a fresh disjoint sample of the same fixtures. Knockout should reproduce;
2062
+ # League's T1 is not expected to — it flips between samples at this N, which is the League finding itself
2063
+ N=4000 npx tsx … # smoke; the tables are shapes, and the pinned-rate control cannot apply at any default-namespace
2064
+ # N other than 10,000 — those streams overlap the pins, so the two rates are paired rather than comparable
2065
+ # 132 `zero-ctl:` control fixtures (22 same-label, 110 pair-key) sit at a family-adjusted bar, which
2066
+ # bounds a valid run's abort probability at 5% — they are NOT independent (the seed key omits the mode,
2067
+ # so each pair's League and Knockout controls share a stream), so 5% is a ceiling, not a rate. The
2068
+ # reported diagonals are NOT checked. Investigate an
2069
+ # exceedance first; only then re-run under another SEED_NAMESPACE — never widen the bar to pass
2070
+ ```
2071
+
2072
+ ---
2073
+
1705
2074
  ## Note for the maintainers
1706
2075
 
1707
2076
  > **2026-08-19: this note was rewritten twice in one day.** First it said the
@@ -1810,7 +2179,9 @@ another is classified as a `swap` in `FormationPitch`, which calls `handleDrop`
1810
2179
  so the drop handler wired there never fires. One drag reaches this at team creation; an
1811
2180
  existing team cannot be reordered that way.
1812
2181
 
1813
- **The mechanism is unidentified, and both candidates are back open.**
2182
+ **The mechanism is unidentified, and both candidates are back open.** *(Superseded:
2183
+ `zp-selector.mts` below identified it — the zone-press fallback, removed by #730. The
2184
+ paragraphs that follow are kept as the record of why the first ablation missed it.)*
1814
2185
 
1815
2186
  | candidate | test | result |
1816
2187
  | --- | --- | --- |
@@ -4279,3 +4650,183 @@ literal early broke one seven times; the last was a markdown table's `` `7036372
4279
4650
  the edit script itself. No tsconfig includes those. The remedy is asserting the match count on
4280
4651
  every replacement, so a broken script dies before it touches the file — which is exactly what
4281
4652
  happened. A smoke run is a behavioural check, not the parse gate.
4653
+
4654
+ ## Is the slot-order lever the zone-press fallback? (`zp-selector.mts`)
4655
+
4656
+ Yes — on the one core where a lever exists, and a two-line engine change removes
4657
+ it. This settles the question the depth-cap note above left open ("the mechanism
4658
+ is unidentified, and both candidates are back open"): the candidate it withdrew
4659
+ for the wrong reason was the right one.
4660
+
4661
+ ### What was broken, exactly
4662
+
4663
+ `tryZonePress` drew the presser with the DEFENSIVE appearance table, whose
4664
+ `SCENE.ZP` column is all-zero for every position (`D_*_APPR[ZP] = 0/0/0/0`), so
4665
+ `weightedPick` had no candidate and fell through to `players[first non-GK]` on
4666
+ every zone press. `getPoint` then drew the possessing side's responder on the
4667
+ same scene with the same table — same fallback, same slot. Both sides of every
4668
+ zone press were therefore the first non-GK slot of their team, a slot the manager
4669
+ chooses freely at team creation. The offence column `O_*_APPR[ZP]` (FW 1.0, OMF
4670
+ 0.9, DMF 0.8, DF 0.7) was populated and unreachable.
4671
+
4672
+ The fix draws both with the offence table — the side `getPoint` already scores
4673
+ the presser as (`O_*_RATE[ZP]`). It invents no table; it revives one.
4674
+
4675
+ It also had to define one case the old code left undefined. `getPoint`'s DMF
4676
+ branch carries a ported defect: on the home side it subtracts the ball carrier
4677
+ `this.kiten` rather than the acting player. For a zone press `this.kiten` indexes
4678
+ the POSSESSING team while the calculation is the PRESSING team's, so it names an
4679
+ arbitrary same-numbered player — the keeper included — and makes the press depend
4680
+ on the OPPONENT's slot order. The fallback hid this: the presser was always the
4681
+ first non-GK slot, so a DMF pressed only when a manager happened to put one
4682
+ there. Drawing by weight makes DMF pressers ~19% of presses, so `SCENE.ZP` now
4683
+ takes the acting player, like every other position in that switch. The defect is
4684
+ untouched everywhere it was measured. That this path really was unreachable is
4685
+ now asserted, not assumed: reverting the two selector lines in a source copy
4686
+ reproduces the pre-#725 corpus digest exactly.
4687
+
4688
+ ### Method
4689
+
4690
+ Four cores (`flat`, `shaped`, `gkheavy`, `gkmin`), each against a basket of the
4691
+ other three, both modes exposure-weighted (knockout 8.1%). Units are pp of WIN
4692
+ RATE on that mix. Two engines run the same matches from the same seeds: the
4693
+ shipped engine and a copy with the two selector lines patched in a `mkdtemp`
4694
+ directory (the `appr-role.mts` pattern; nothing is written inside the
4695
+ repository). Three readings per core:
4696
+
4697
+ 1. **Sweep**: swap each slot j into the first non-GK slot, paired against the
4698
+ base order, on both engines, plus the paired **shipped − fixed contrast** on
4699
+ the same matches. If the lever is the fallback, the shipped column moves and
4700
+ the fixed column does not.
4701
+ 2. **Best permutation, selected TWICE — once per engine**: 120 permutations
4702
+ scored at 200 seeds, the winner published at 1500 fresh seeds on both
4703
+ engines (selection column and publication column are different seeds, so the
4704
+ selection optimism is not in the published number). The search runs once
4705
+ against the shipped engine and once, independently, against the fixed one.
4706
+ Optimising only against the shipped engine can show the old exploit is gone;
4707
+ it cannot show the fixed engine has no lever of its own, because a
4708
+ permutation the fixed engine would favour is never tried.
4709
+ 3. **Base shift**: fixed − shipped for the base order itself, descriptive only
4710
+ and outside the family, because it answers a different question (does the fix
4711
+ move the level, not the lever).
4712
+
4713
+ Pre-registered family m=132, Bonferroni at α=0.05 → |t| ≥ 3.56 (df 1499). The
4714
+ family is every comparison the run prints with a significance mark — each swap
4715
+ on both engines (72), every swap's own engine contrast (36, not a preselected
4716
+ "widest": nothing picks one and the readings below use several), and the two
4717
+ selected permutations on both engines with their contrasts (16 + 4). Every
4718
+ interval is a paired per-seed t interval.
4719
+
4720
+ Controls, every one of which aborts the run if it fails: the import-rewritten
4721
+ engine with no selector change reproduces 1000 stock matches byte-for-byte
4722
+ (whole `MatchResult`); a fallback census over 400 matches counts 1056 zero-weight
4723
+ fallbacks on the shipped engine (one per zone press) and 0 on the fixed one;
4724
+ mirror cells (true z = 0) read z = 0.3/0.2 shipped and 1.2/1.3 fixed; and the
4725
+ base order replayed reproduces itself exactly on both engines.
4726
+
4727
+ ### Result
4728
+
4729
+ | core | shipped-selected perm, on shipped | the same perm, on fixed | shipped − fixed (paired) | **fixed-selected perm, on fixed** | sweep swaps significant, shipped / fixed |
4730
+ | --- | --- | --- | --- | --- | --- |
4731
+ | `shaped` | **+5.20pp** [+0.45, +9.95], t 3.9 | +0.76pp [−3.97, +5.50], t 0.6 | **+4.44pp** [+0.13, +8.74], t 3.7 | **+0.02pp** [−4.60, +4.65], t 0.0 | **4** / 0 |
4732
+ | `flat` | −1.57pp [−4.44, +1.30], t −1.9 | −0.79pp [−3.71, +2.13] | −0.78pp [−3.46, +1.90], t −1.0 | −0.21pp [−3.19, +2.76], t −0.3 | 1 / 0 |
4733
+ | `gkheavy` | +1.10pp [−2.20, +4.41], t 1.2 | +1.18pp [−2.16, +4.52] | −0.08pp [−3.00, +2.84], t −0.1 | +0.82pp [−2.50, +4.13], t 0.9 | 0 / 0 |
4734
+ | `gkmin` | −0.45pp [−3.80, +2.89], t −0.5 | −1.65pp [−5.12, +1.82] | +1.19pp [−1.87, +4.25], t 1.4 | +0.39pp [−3.06, +3.85], t 0.4 | 0 / 0 |
4735
+
4736
+ `shaped` is the only core with a lever on the shipped engine, and the fix is what
4737
+ takes it away — with one honest limit on how far that can be pushed.
4738
+
4739
+ What is RESOLVED: the paired engine contrast. The best permutation's +5.20pp,
4740
+ which put a FW (t28) in the slot a DF (t12) held, reads +0.75pp on the fixed
4741
+ engine, and the shipped−fixed difference over the same permutation and the same
4742
+ matches is **+4.44pp [+0.13, +8.74], t 3.7** — a resolved, corrected difference,
4743
+ so the drop is the fix and not noise. The sweep says the same thing with nine
4744
+ paired swaps instead of one permutation: the four significant ones (slots 7–10,
4745
+ +5.19 to +7.02pp, an OMF or FW into the first slot) read +0.46 to −0.31pp fixed,
4746
+ every one of their contrasts significant (t 4.3 to 6.5), and the range of the
4747
+ whole sweep collapses from 7.02pp to 1.71pp — 0 of 9 significant.
4748
+
4749
+ What that one column does NOT resolve: the residual on the SHIPPED-selected
4750
+ permutation. +0.76pp carries [−3.97, +5.50], whose upper bound still reaches
4751
+ past the 5pp gate, so that reading alone cannot certify the leftover as *below*
4752
+ it — non-significance is not equivalence, and the probe says `UNDECIDED vs the
4753
+ gate` there rather than `REMOVED`.
4754
+
4755
+ What answers the question it was actually asked: give the FIXED engine its own
4756
+ search. A permutation chosen against the shipped engine is not what a manager
4757
+ optimising against the fixed one would ever pick, so the shipped-selected winner
4758
+ cannot speak to whether a lever remains. Under an independent 120-permutation
4759
+ search on the fixed engine, **every core comes in below the gate**: `shaped`
4760
+ +0.02pp [−4.60, +4.65], `flat` −0.21pp, `gkheavy` +0.82pp, `gkmin` +0.39pp, none
4761
+ significant (|t| ≤ 0.9). That is an equivalence result at the gate, on all four
4762
+ cores, for the engine that ships after this fix. The sweep says the same with
4763
+ nine paired swaps: 0 of 9 significant on every core.
4764
+
4765
+ `flat` shows the lever's other face: moving its OMF into slot 1 is a LOSS of
4766
+ 3.48pp (t −4.5) on the shipped engine and +0.30pp on the fixed one, contrast
4767
+ −3.78pp (t −4.9). The fallback was hurting managers who put the wrong player
4768
+ first as much as it rewarded the right one.
4769
+
4770
+ The fixed engine does not create a lever on any core: 0 of 36 sweep swaps are
4771
+ significant across the four cores, no fixed-engine gain clears the gate in
4772
+ either direction, and its own independent search finds nothing (above).
4773
+ One isolated contrast stays marked: `gkheavy` slot 10, −2.77pp (t −3.6), with
4774
+ neither engine's own swap significant (−1.49pp t −1.8 shipped, +1.28pp t 1.7
4775
+ fixed) — a resolved difference between two unresolved readings, reported and not
4776
+ read as a lever. `gkmin` slot 5 (−2.27pp, t −3.5) cleared the old m=88 bar and
4777
+ does not clear the m=132 one; it is no longer marked.
4778
+
4779
+ Who presses under the fix, from the revived weights, on a 4-3-3: FW 36%, DMF
4780
+ 19%, OMF 11%, DF 34% (on `gkmin`'s 4-4-2: 24/20/22/34). Under the shipped
4781
+ engine: the first slot, a DF in all four cores, 100% of the time.
4782
+
4783
+ Base shift for the record: `shaped` +2.75pp [−0.32, +5.81], `flat` −1.20, `gkheavy`
4784
+ −0.57, `gkmin` −2.64pp [−4.94, −0.34]. The fix moves levels between cores by a
4785
+ few points, as any change to who presses must.
4786
+
4787
+ ### What it means
4788
+
4789
+ - **The +5.7pp of the depth-cap note was the fallback.** Its clone ablation
4790
+ withdrew the candidate because it held POSITION fixed and varied only
4791
+ attributes — but what the fallback delivers into the first slot is a player
4792
+ whose `getPoint` branch depends on position, so a same-position swap could
4793
+ not see it. Disabling the path and re-running the same permutations, which
4794
+ that note said was needed, is what this probe did.
4795
+ - **The lever is gone once the fix ships**, so the Skill records nothing about
4796
+ slot order: a lever that no longer exists must not be documented as one.
4797
+ - **Determinism** (P5): completed matches are persisted as events and do not
4798
+ change; the live stream and the replay both read those persisted events, so
4799
+ nothing about a match is observable before it is persisted. Every future
4800
+ match changes — a zone press happens about once a match (1.07 on the test
4801
+ teams, up to 7), and the real draw consumes RNG the fallback did not, so every
4802
+ later draw in that match shifts. The scheduler is one Vercel Cron invocation
4803
+ a minute, not a fleet of replicas: a deploy lands between ticks, and the only
4804
+ match that can straddle it is one whose invocation died mid-tick and whose
4805
+ 5-minute lease is then reclaimed by a tick running the new build. That match
4806
+ is simulated once, by the new build, from the stored snapshot and seed — the
4807
+ old build never produced a result for it. Determinism is per build; every
4808
+ engine change (#737, #786, #794 before this one) has made such a boundary, and
4809
+ an engine-replay verification whose window spans a deploy sees it as an
4810
+ ENGINE boundary.
4811
+ - **Workload tally** (F-1/F-2): predicate D — actor plus responder minus
4812
+ zero-weight fallbacks — is unchanged in meaning, and on the shipped engine
4813
+ every fallback was a zone-press engagement, so D excluded zone presses
4814
+ entirely. Now it counts them: on the test teams that is **+0.10 engagements
4815
+ per player-match, +6.1%** of the tally (5.75% of the fixed engine's D). The
4816
+ F-2 curve is `cond = 5 + a − b·L/(L+L₀)` with `b = 0.4914`, `L₀ = 59.03`; a
4817
+ 6.1% inflation of the load L moves condition by at most **0.0075** points
4818
+ (at L = L₀; 0.0006 at the neutral load 1.146, 0.0047 at the p95 14.76) — an
4819
+ order of magnitude below the 0.5 band step and far below anything
4820
+ `cond-effect.mts` could detect. `PLAYER_WORKLOAD_VERSION` is therefore NOT
4821
+ bumped: the version is a hard filter in the form store, so a bump erases
4822
+ every accumulated load league-wide — a reset three orders of magnitude larger
4823
+ than the drift it would prevent. The boundary is recorded here instead, and
4824
+ `L_neutral`/`L₀` are re-measured by `engagement-load.mts` on the first
4825
+ production window after the deploy (C-2 follow-up).
4826
+ - **Open sub-decision, recorded**: the responder. The original appears to have
4827
+ skipped the defensive role at ZP entirely (port comment at `game.pl:1924`);
4828
+ drawing it from the same offence column is the smallest change that invents
4829
+ nothing, and "no responder, team term only" was not adopted because the exact
4830
+ original expression cannot be recovered.
4831
+
4832
+ 5,648,400 matches in about a minute. Whole run: `zp-selector.mts`.