pog-mcp 0.9.15 → 0.9.16

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pog-mcp",
3
- "version": "0.9.15",
3
+ "version": "0.9.16",
4
4
  "type": "module",
5
5
  "description": "MCP server that lets an AI agent play Proof of Goal — wallet, sign-in, squad building, and matches as typed tools.",
6
6
  "license": "MIT",
@@ -1703,6 +1703,374 @@ N=2000 npx tsx … # ~50s; the bar is fixed, the detectable EFFECT moves
1703
1703
 
1704
1704
  ---
1705
1705
 
1706
+ ## Does the best BUILD depend on who you play? (`opp-conditional.mts`)
1707
+
1708
+ Recorded 2026-09-09 — docs/tactics/01 track 0.2 ②, #729. Track 2.4 shipped a
1709
+ kickoff lock and a scouting surface that shows the opponent's last fielded
1710
+ eleven (#726). The plan's own condition for that surface having a DECISION
1711
+ value rather than only an information value is written in 01 §5: across a grid
1712
+ of opponent builds, take my best response to each — *if the distinct argmax is
1713
+ 1, scouting's decision value is 0 and 2.4's reason to exist is gone.* Until this
1714
+ run that measurement had not been made; every reversal in this file is
1715
+ conditional on the mode or on my own attack strength, not on the opponent.
1716
+
1717
+ The section above asked the same question of a different axis — which SHEET is
1718
+ best against each opponent, with my 212 points fixed — and found one sheet in
1719
+ every cell's band. This one asks it of the axis a manager actually controls
1720
+ today, the 212-point build.
1721
+
1722
+ ### The statistic, the three tests, and the controls
1723
+
1724
+ The eleven published archetypes from `strategy-probe.mts`, built by the shared
1725
+ `squad-lib.mts` so they are the round-robin's squads by construction, play each
1726
+ other in both modes at 10,000 matches per fixture (5,000 seeds, home and away,
1727
+ SHA-256 keyed on both labels and the side). Each unordered pair is simulated
1728
+ once and read from both sides, so the matrix is antisymmetric by construction.
1729
+ The diagonal is played too: a manager who knows the opponent's archetype can
1730
+ field the same one, so that fixture is a real response with a real answer (50
1731
+ by symmetry). Its 22 cells have a known truth, but they are REPORTED data — the
1732
+ same-build response in their own column, feeding the tests — so nothing aborts on
1733
+ them. The controls are separate fixtures on their own seed keys (below).
1734
+
1735
+ For every opponent and mode, my candidates are all eleven builds. The LEADER is
1736
+ the candidate with the highest win-equivalent share (draws half); the BAND is
1737
+ every candidate the sample cannot separate from the leader under the corrected
1738
+ bar; a cell is RESOLVED when its band has one member. Two candidates against the
1739
+ same opponent play different seed columns, so every comparison is UNPAIRED.
1740
+
1741
+ Bonferroni is priced before the run over the search a band performs: one
1742
+ comparison per unordered pair per (opponent, mode) cell, 11·10/2 = 55 pairs over
1743
+ 22 cells — **m = 1,210, two-sided, |z| ≥ 4.10**. Enumerating every pair already
1744
+ covers whichever pair the data selects as leader and runner-up, so that selection
1745
+ costs nothing extra, and the two-sided tail already covers both directions; an
1746
+ earlier version charged 2 × 55 AND halved α, pricing direction twice for a bar of
1747
+ 4.26. A too-wide bar is not the safe side here — it widens every band and every
1748
+ equivalence bound, and this grid has boundary-sensitive readings. Worst standard
1749
+ error anywhere in the run 0.71 points of share, so the significance threshold is
1750
+ **2.90 points of share (5.8pp of win rate)** and the 80%-power effect
1751
+ **3.49 (7.0pp)**.
1752
+
1753
+ Three tests are reported, and their provenance is not the same. **T1 was
1754
+ declared before the run.** **T2 and T3 were not**: T2 was added after the first
1755
+ run had been read, because "in every band" is a non-rejection and a
1756
+ non-rejection is not an equivalence; T3 was added after T2's result had been
1757
+ read, because "T1 fires and T2 fails" is not a magnitude claim either — T2
1758
+ failing only says no build was SHOWN within the gate everywhere, and every true
1759
+ effect could still be under it. On the run the tables below come from, T2 and
1760
+ T3 are therefore **post-hoc readings of intervals that were already priced**
1761
+ (the Bonferroni family covers every pairwise contrast they use), and the honest
1762
+ name for that is exploratory. So all three were then re-evaluated on a
1763
+ **confirmatory run** — the same fixtures under a fresh seed namespace
1764
+ (`SEED_NAMESPACE=confirm-1`, disjoint seeds, same N), with T2 and T3 fixed in
1765
+ the code before it was started. Its verdicts are reported in their own table
1766
+ below; the numbers in the sections in between are the first run's. The tests
1767
+ are reported **independently**; a difference can be statistically resolved and
1768
+ still smaller than the practical margin, so T1 and T2 can hold at once:
1769
+
1770
+ - **T1 — conditionality DETECTED**: no build is in every opponent's band; each
1771
+ is significantly beaten as a best response by some build against some
1772
+ opponent. A rejection-based positive claim.
1773
+ - **T2 — equivalence ESTABLISHED**: some build is non-inferior to the TRUE best
1774
+ response against every opponent — for that build, the corrected upper bound
1775
+ of (candidate − it) over *every other candidate*, not only the sample leader
1776
+ (the two differ: a bound against the leader alone is trivially satisfied by
1777
+ the leader itself, while one noisy non-leader can deny the simultaneous
1778
+ bound — Knockout's `role-433` and `def-433` columns qualify nobody for that
1779
+ reason),
1780
+ is under the margin in all eleven columns. A positive equivalence claim.
1781
+ - **T3 — conditionality BEYOND THE MARGIN**: every build is beaten, against some
1782
+ opponent, by more than the margin with the corrected LOWER bound of (leader −
1783
+ it). The positive magnitude claim, read from the same priced contrasts at the
1784
+ other end.
1785
+ - The readings: T1 with T3 → **conditional beyond the margin**; T1 with T2 →
1786
+ conditional but within it; T1 alone → conditional, magnitude
1787
+ unresolved; T2 without T1 → **equivalent**; neither T1 nor T2 →
1788
+ **UNDECIDED** at this sample.
1789
+
1790
+ **The margin is 5pp of win rate WITHIN a mode, and that is NOT the repository's
1791
+ gate.** It is deliberately the same number — below it, re-solving stops being
1792
+ worth an agent's trouble — but the gate proper (`01-engine-expansion.md` §5,
1793
+ implemented by `constant-sweep.mts`) is specified on the **exposure-weighted
1794
+ mode mix**, and the modes are nowhere near equally exposed. A qualifier's day is
1795
+ about 1.33 knockout matches against 12 ladder plus the cup's own 3 group matches,
1796
+ which are played in regulation: **8.1% knockout**, and 0% outside the season's
1797
+ top 48. `constant-sweep.mts` says in as many words that issuing a verdict per
1798
+ mode would call cleared what the mix has not.
1799
+
1800
+ **Which way that cuts is not decided here, and the arithmetic must not pretend
1801
+ otherwise.** T3 is EXISTENTIAL: each build meets SOME opponent that beats it by
1802
+ more than the margin. Multiplying that by the 8.1% mode exposure would assume the
1803
+ adverse opponent is met in every knockout fixture — a bracket may hold it rarely
1804
+ or not at all — so the product is not a lower bound on the mix, and the weighted
1805
+ regret can be anywhere from zero upward. This grid also never plays one locked
1806
+ build across the weighted opponent-and-mode mix, which is what the gate actually
1807
+ asks. So clearing the repository gate is **not established**
1808
+ here, and **not ruled out** either. What is established is that the best response
1809
+ is opponent-dependent *within a mode* by more than a within-mode 5pp — the
1810
+ prerequisite for scouting to be worth anything.
1811
+
1812
+ Controls, every one of which aborts the run: each of the eleven builds must
1813
+ reproduce a pinned signature over every field the engine reads from a player —
1814
+ `name`, position, the four attributes, `total`, `cond`, `fair` and both kicker
1815
+ flags — for all eleven players (slot totals alone would let ratios or a scarcity repair
1816
+ drift under the same labels); no two builds byte-identical; **132 true-zero cells**
1817
+ within a family-adjusted |z| < 3.55 of 50 (worst 2.3, `stars-433` on the
1818
+ `stars-433|bal-442` League control key) —
1819
+ correcting matters here: at an uncorrected two-sided 3σ each true null exceeds
1820
+ with probability ≈0.27%, so across 132 streams a valid rerun would fail about
1821
+ three times in ten (1 − 0.9973^132 ≈ 30%). All 132 are CONTROL FIXTURES of their own,
1822
+ never cells read back out of the grid: 22 on a same-label key `a|a` and 110 on a
1823
+ two-distinct-label key, the shape every published cell uses. Both halves matter.
1824
+ A seeding defect keyed on the pair of labels — which the retired FNV one was —
1825
+ passes a same-label-only control while biasing every cell that control exists to
1826
+ validate. And every one carries a `zero-ctl:` prefix, so no control shares a
1827
+ stream with a reported cell: the controls abort at their α, and a rerun under a
1828
+ fresh namespace selected on a control that shared those streams would be picked
1829
+ partly on the estimates themselves. And a
1830
+ published fact held — `cycle-probe.mts` reports `role-433` beating each of
1831
+ its seven other candidates in Knockout, and here the same eight squads say so
1832
+ under the same statistic, smallest decided z 3.8 — played on a `drift-ctl:` seed
1833
+ family disjoint from every reported cell, because a guard that aborts on the
1834
+ reported cells' own estimates would accept only runs reproducing them. That last control is a
1835
+ **compatibility** test, not a demand for renewed significance: at
1836
+ `cycle-probe`'s N=4000 the weakest of the seven pairs would replicate |z| ≥ 3.32
1837
+ with roughly even odds even if nothing had changed, so a rerun that required it
1838
+ would be aborting on its own sampling noise. Instead `role-433`'s decided win
1839
+ rate against each of the seven from the default-namespace run is pinned in the
1840
+ probe, and a run fails only when its own rate is incompatible with the pin at
1841
+ the Bonferroni-7 bound |z| ≤ 2.69 — what "the squads changed or the engine
1842
+ moved" looks like at any N. Worst |z| against the pins: 0.0 in the default
1843
+ namespace (the pins are that family's own default values) and 0.8 in
1844
+ `confirm-1`.
1845
+ The Knockout column also reproduces the cheap keeper's published collapse
1846
+ (`gkmin-433` at 27.8% against `def-532`, inside the 27–42% this file already
1847
+ states), and the League column reproduces the two pairings `cycle-probe.mts`
1848
+ could not separate (`gkmin-433` 50.6% vs `def-433` and 49.2% vs `def-532`,
1849
+ neither significant here either).
1850
+
1851
+ ### League — UNDECIDED: T1 does not fire, T2 does not pass
1852
+
1853
+ | Opponent | Leader | Share | Runner-up | Gap (z) | Band | Within 5pp of the true best |
1854
+ | --- | --- | --- | --- | --- | --- | --- |
1855
+ | `flat-433` | **`gkmin-433`** | 72.0 | `def-433` | +2.5 (5.7) | `gkmin-433` | `gkmin-433` |
1856
+ | `role-433` | `gkmin-433` | 54.0 | `def-532` | +1.0 (2.2) | `def-433` / `gkmin-433` / `def-532` | `gkmin-433` |
1857
+ | `def-433` | `gkmin-433` | 50.6 | `def-433` | +0.2 (0.5) | `def-433` / `gkmin-433` / `def-532` | `def-433` / `gkmin-433` / `def-532` |
1858
+ | `stars-433` | **`gkmin-433`** | 68.1 | `def-532` | +4.0 (7.3) | `gkmin-433` | `gkmin-433` |
1859
+ | `gkheavy-433` | `gkmin-433` | 60.8 | `def-433` | +1.9 (4.1) | `def-433` / `gkmin-433` | `gkmin-433` |
1860
+ | `gkmin-433` | `def-532` | 50.8 | `gkmin-433` | +0.8 (1.8) | `def-433` / `gkmin-433` / `def-532` | `def-532` |
1861
+ | `shootonly-433` | **`gkmin-433`** | 64.3 | `def-433` | +3.7 (7.2) | `gkmin-433` | `gkmin-433` |
1862
+ | `passonly-433` | `def-433` | 64.7 | `gkmin-433` | +1.8 (3.9) | `def-433` / `gkmin-433` | `def-433` |
1863
+ | `def-532` | `def-532` | 50.1 | `def-433` | +0.1 (0.3) | `def-433` / `gkmin-433` / `def-532` | `def-433` / `gkmin-433` / `def-532` |
1864
+ | `bal-442` | `gkmin-433` | 58.8 | `def-433` | +1.2 (2.6) | `def-433` / `gkmin-433` | `gkmin-433` |
1865
+ | `atk-352` | `gkmin-433` | 63.0 | `def-433` | +1.4 (3.0) | `def-433` / `gkmin-433` | `gkmin-433` |
1866
+
1867
+ > **Measured on** — every one of the eleven published archetypes as the
1868
+ > opponent, with all eleven as my candidates (the opponent's own build among
1869
+ > them: its EXPECTED share is 50 by symmetry, and the sampled cell — 50.1 for
1870
+ > `def-532` here — enters the tests like any other, a known truth on reported
1871
+ > data rather than a control). Bold
1872
+ > leaders are RESOLVED cells: the band holds one build. The last column is T2's
1873
+ > per-cell reading — which builds the sample shows within 2.5 points of share of
1874
+ > the TRUE best response: the corrected upper bound of (candidate − it) is under
1875
+ > the margin for EVERY other candidate in the column, not just for the sample
1876
+ > leader. That is why a column can list nobody even though its leader is
1877
+ > trivially zero behind itself — one noisy candidate denies the simultaneous
1878
+ > bound.
1879
+ > **Sample** — 10,000 matches per fixture, home and away, SHA-256 seeds; 66
1880
+ > fixtures in this mode.
1881
+ > **Mode** — League (`gameFlg 0`, draws count half).
1882
+ > **Effect** — the raw leader is `gkmin-433` in 8 of 11 columns and 3 cells
1883
+ > resolve, all three to `gkmin-433`. **T1 does not fire**: `gkmin-433` is in the
1884
+ > band of all eleven opponents, so conditionality is NOT detected — but see the
1885
+ > confirmation below, where a disjoint sample of the same fixtures puts it
1886
+ > outside one band and T1 does fire. That instability is itself the finding
1887
+ > here. **T2 does not pass**: no build — `gkmin-433` included — is shown within
1888
+ > 5pp of every other candidate in every column; against its own archetype the
1889
+ > best sample response is `def-532` and the corrected upper bound of that
1890
+ > 0.8-point gap exceeds the margin, and against `passonly-433` the leader is
1891
+ > `def-433`. So
1892
+ > the League reading is **UNDECIDED**: this sample did not detect an opponent
1893
+ > that re-prices the build, and it cannot rule out re-pricing of up to its own
1894
+ > power — 7.0pp of win rate on the worst cell — either. One build,
1895
+ > `gkmin-433`, is never beaten by MORE than the gate in any column (T3's
1896
+ > per-build reading) — which is consistent with equivalence and is not it. "Field `gkmin-433` regardless" is consistent with the grid; the grid
1897
+ > does not establish it.
1898
+ > **Source** — `opp-conditional.mts`, sections (2) and (3), `REG`.
1899
+
1900
+ ### Knockout — CONDITIONAL BEYOND THE WITHIN-MODE MARGIN: T1 fires, T3 holds
1901
+
1902
+ | Opponent | Leader | Share | Runner-up | Gap (z) | Shootout reach | Band | Within 5pp of the true best |
1903
+ | --- | --- | --- | --- | --- | --- | --- | --- |
1904
+ | `flat-433` | `def-433` | 82.8 | `def-532` | +0.3 (0.6) | 21% | `role-433` / `def-433` / `def-532` | `def-433` |
1905
+ | `role-433` | `gkheavy-433` | 50.6 | `role-433` | +0.3 (0.5) | 26% | `role-433` / `gkheavy-433` | — |
1906
+ | `def-433` | `role-433` | 52.4 | `gkheavy-433` | +0.5 (0.7) | 32% | `role-433` / `def-433` / `gkheavy-433` | `role-433` |
1907
+ | `stars-433` | **`gkmin-433`** | 69.9 | `def-532` | +4.8 (7.3) | 5% | `gkmin-433` | `gkmin-433` |
1908
+ | `gkheavy-433` | `gkmin-433` | 52.0 | `gkheavy-433` | +1.6 (2.2) | 28% | `role-433` / `gkheavy-433` / `gkmin-433` | `gkmin-433` |
1909
+ | `gkmin-433` | **`def-532`** | 72.2 | `def-433` | +8.7 (13.3) | 24% | `def-532` | `def-532` |
1910
+ | `shootonly-433` | `role-433` | 63.9 | `def-433` | +2.3 (3.3) | 18% | `role-433` / `def-433` | `role-433` |
1911
+ | `passonly-433` | `def-433` | 67.8 | `gkheavy-433` | +0.5 (0.7) | 25% | `role-433` / `def-433` / `gkheavy-433` | `def-433` |
1912
+ | `def-532` | **`gkheavy-433`** | 59.9 | `role-433` | +3.0 (4.3) | 47% | `gkheavy-433` | `gkheavy-433` |
1913
+ | `bal-442` | **`gkheavy-433`** | 58.9 | `role-433` | +4.0 (5.7) | 31% | `gkheavy-433` | `gkheavy-433` |
1914
+ | `atk-352` | `gkheavy-433` | 62.3 | `role-433` | +1.7 (2.5) | 26% | `role-433` / `gkheavy-433` | `gkheavy-433` |
1915
+
1916
+ > **Measured on** — the same eleven archetypes, the same construction. Shootout
1917
+ > reach is the share of that opponent's matches, averaged over all eleven
1918
+ > candidates, that went to penalties — DESCRIPTIVE, see below.
1919
+ > **Sample** — 10,000 matches per fixture, home and away, SHA-256 seeds; 66
1920
+ > fixtures in this mode.
1921
+ > **Mode** — Knockout (`gameFlg 1`: extra time and shootout, every match
1922
+ > decided).
1923
+ > **Effect** — five distinct raw leaders and **4 of 11 cells resolve, to three
1924
+ > different builds**: `gkmin-433` against `stars-433` (z 7.3 over the
1925
+ > runner-up), `def-532` against `gkmin-433` (z 13.3 — the cheap keeper's
1926
+ > published collapse, seen from the other side), and `gkheavy-433` against both
1927
+ > `def-532` (z 4.3) and `bal-442` (z 5.7). **T1 fires**: no build is in every
1928
+ > band. `role-433` and `gkheavy-433` share the most bands at seven each, and
1929
+ > neither survives: `role-433` is significantly excluded from all four resolved
1930
+ > cells, `gkheavy-433` from two.
1931
+ > The exclusions are carried by resolved cells at z 4.3 to 13.3, not by noise.
1932
+ > **T2 does not pass**: no build is within 5pp of every other candidate in every
1933
+ > column — and the `role-433` column names nobody at all once the bound is taken
1934
+ > over every candidate rather than the sample leader.
1935
+ > **T3 holds**: every one of the eleven builds is beaten by MORE than 5pp of win
1936
+ > rate, with the corrected lower bound, against at least one opponent. The
1937
+ > `gkmin-433` column alone does it for the other ten — `def-532`'s 72.2 there
1938
+ > clears the margin against every one of them, the closest being `def-433` at
1939
+ > 63.5 — and `def-532` itself is beaten beyond the margin elsewhere, most
1940
+ > clearly by `gkheavy-433` in its own column (59.9 against `def-532`'s 49.5).
1941
+ > So the Knockout
1942
+ > reading is **CONDITIONAL BEYOND THE WITHIN-MODE MARGIN**: the best build is
1943
+ > opponent-conditional, and by more than 5pp of win rate *inside Knockout* — a
1944
+ > claim T3 makes, not one inferred from T2 failing. Whether that clears the
1945
+ > repository's exposure-weighted gate is neither established nor ruled out here:
1946
+ > 5pp is a lower bound, and this grid never plays a locked build across the mix.
1947
+ > **Source** — `opp-conditional.mts`, sections (2) and (3), `KO`.
1948
+
1949
+ ### Confirmation on a fresh seed namespace
1950
+
1951
+ Same eleven squads, same N, seeds prefixed `confirm-1|` so no fixture shares a
1952
+ match with the run above; T2 and T3 were in the code before it started. Every
1953
+ control holds (worst true-zero |z| 2.4, `def-433` on the
1954
+ `def-433|gkheavy-433` League control key; pinned-rate guard worst 1.2), and
1955
+ Knockout reproduces exactly — **League does not**:
1956
+
1957
+ | run | mode | T1 | T2 | T3 | reading | resolved leaders | in every band | never beaten beyond the margin |
1958
+ | --- | --- | --- | --- | --- | --- | --- | --- | --- |
1959
+ | tables above | League | no | no | no | UNDECIDED | `gkmin-433` | `gkmin-433` | `gkmin-433` |
1960
+ | `confirm-1` | League | **fires** | no | no | CONDITIONAL, MAGNITUDE UNRESOLVED | `gkmin-433`, `def-433` | none | `gkmin-433`, `def-532` |
1961
+ | tables above | Knockout | **fires** | no | **holds** | CONDITIONAL BEYOND THE MARGIN (within mode) | `gkmin-433`, `def-532`, `gkheavy-433` | none | none |
1962
+ | `confirm-1` | Knockout | **fires** | no | **holds** | CONDITIONAL BEYOND THE MARGIN (within mode) | `gkmin-433`, `def-532`, `gkheavy-433` | none | none |
1963
+
1964
+ **League's T1 is not reproducible at this sample size, and that is the League
1965
+ result.** T1 asks whether any build sits in every opponent's band; `gkmin-433`
1966
+ does in the reported run and in `confirm-2`, and does not in `confirm-1` — three
1967
+ disjoint samples of the same eleven fixtures, two landing one way and one the
1968
+ other. The `passonly-433` column is where it turns: `gkmin-433` is inside that
1969
+ band at z 3.9 in the reported run and in `confirm-2`, and outside it at z 4.3 in
1970
+ `confirm-1`. Which side of T1 a sample lands on is decided by noise at N=10,000,
1971
+ so nothing about League should be read as "conditionality was detected" or "was
1972
+ not"; the honest statement is UNDECIDED, and a larger N is what would settle it.
1973
+ The Knockout verdict carries none of this: T1 fires and T3 holds in all three,
1974
+ with the same three resolved leaders.
1975
+
1976
+ The confirmatory runs are what the T2/T3 verdicts rest on; the first run's
1977
+ tables remain the published numbers because they are the run every figure in
1978
+ this section was read from, and mixing them would leave no run owning them.
1979
+
1980
+ ### What the grid does and does not say about 2.4
1981
+
1982
+ **What it says.** In Knockout the answer to "which build should I field" changes
1983
+ with the opponent, by amounts the sample resolves (3.0 to 8.7 points of share
1984
+ between the leader and the runner-up in the four resolved cells: 4.8, 8.7, 3.0
1985
+ and 4.0). In League
1986
+ the same grid neither finds such an opponent nor rules one out. The plan's
1987
+ sentence in 01 §5 was written mode-blind; measured, the axis is a decision in
1988
+ one mode and undecided in the other.
1989
+
1990
+ **What it does not say — and an earlier version of this section said it.** It
1991
+ does not say that the cup lock is where that decision lives. The lock (#726)
1992
+ freezes one eleven when the cup opens and reuses it for every fixture of the
1993
+ day, group stage (`gameFlg 0`) and knockout rounds (`gameFlg 1`) alike. A
1994
+ manager cannot learn a knockout opponent and then switch to the build this grid
1995
+ says beats it; the decision the lock actually forces is *one build against a
1996
+ bracket of opponents in sequence*, and a per-opponent best response is not that
1997
+ object. What this grid establishes is the prerequisite: that in the mode the
1998
+ knockout rounds are played in, a per-opponent best response EXISTS to be
1999
+ informed by. Whether a single locked build should be chosen differently given
2000
+ the bracket — and whether the scouting surface changes that choice — needs a
2001
+ bracket-level experiment, and this file does not contain one. As #729 itself
2002
+ notes, a distinct-argmax count of 1 would not have meant "remove 2.4" either:
2003
+ the lock is a commitment device on its own.
2004
+
2005
+ **What produced the conditional responses is not identified.** The three builds
2006
+ that are unique best responses somewhere — `gkmin-433`, `gkheavy-433`,
2007
+ `def-532` — differ from the rest of the grid in keeper budget, but not only in
2008
+ that: `def-532` carries an ordinary keeper and changes formation and back-line
2009
+ allocation, and `GK_MIN`/`GK_HEAVY` redistribute the keeper's points across the
2010
+ whole outfield. The keeper section of this file supplies a mechanism that is
2011
+ consistent with the pattern — in Knockout the keeper's total is
2012
+ outfield-conditional through shootout exposure, and the table's shootout reach
2013
+ runs from 5% (`stars-433`) to 47% (`def-532`) across opponents — but consistency
2014
+ is not attribution. Isolating the keeper would take a keeper-only,
2015
+ fixed-outfield contrast, which this grid does not run.
2016
+
2017
+ Two secondary readings survive at the descriptive level. In League the
2018
+ cheap-keeper build is the raw leader in 8 of 11 columns, which is this file's
2019
+ "smaller keeper total is better in matches that allow draws" seen across the
2020
+ whole grid at once. And `role-433` — the Knockout dominant in the 8-candidate
2021
+ round robin above — is the raw leader in only 2 of 11 Knockout columns once the
2022
+ keeper-budget archetypes are in the grid, and is the unique best response to
2023
+ nobody. Dominance in a round robin and being the best response to each opponent
2024
+ are different questions, and the second is the one scouting asks.
2025
+
2026
+ ### What this grid CANNOT decide
2027
+
2028
+ - **The build space.** Eleven hand-built archetypes, the same set on both axes.
2029
+ That is a sample, not an argmax over 212-point squads. A conditional best
2030
+ response can only appear among builds that exist in the grid, and a build
2031
+ that beats every leader here could exist outside it. Every claim above is a
2032
+ claim about these eleven.
2033
+ - **Nobody re-optimises against a KNOWN opponent.** Every candidate is a
2034
+ pre-built archetype. A squad rebuilt with the opponent's eleven in hand is a
2035
+ different 212-point problem — and it is the search the scouting surface would
2036
+ actually enable. This grid measures whether the choice AMONG fixed builds is
2037
+ conditional, which is the weaker, prerequisite question.
2038
+ - **The bracket.** The lock forces one build against a sequence of opponents;
2039
+ the grid asks about one opponent at a time. The object the lock decides is
2040
+ not measured here.
2041
+ - **Mechanism.** Formation, keeper budget and outfield allocation move together
2042
+ across these archetypes. Which of them re-prices the build is not separated.
2043
+ - **`cond` fixed at 5, tenure 0.** As everywhere in this file. Production
2044
+ applies growth before the engine and, since #794, an involvement-driven
2045
+ condition; neither is in these squads.
2046
+ - **Sheets at the default.** The section above showed the sheet axis moves more
2047
+ than any build change does; this grid holds it at the shipped weights on both
2048
+ sides.
2049
+ - **Power.** 8 of 11 League cells and 7 of 11 Knockout cells are unresolved,
2050
+ and League's T1 flips between two disjoint samples at this N.
2051
+ T2's margin is 5pp within the mode; the worst cell's 80%-power effect is 7.0pp
2052
+ of win rate, so this sample could not have established equivalence at that margin
2053
+ even where it holds — and T2's bound is taken over every candidate, so a
2054
+ weak candidate with a large standard error can deny it on its own. A larger
2055
+ N is the way to move League out of UNDECIDED.
2056
+
2057
+ Re-run after any engine change:
2058
+
2059
+ ```bash
2060
+ npx tsx packages/mcp/skill/reference/probes/opp-conditional.mts # ~2.7M matches, ~31s
2061
+ SEED_NAMESPACE=confirm-2 npx tsx … # a fresh disjoint sample of the same fixtures. Knockout should reproduce;
2062
+ # League's T1 is not expected to — it flips between samples at this N, which is the League finding itself
2063
+ N=4000 npx tsx … # smoke; the tables are shapes, and the pinned-rate control cannot apply at any default-namespace
2064
+ # N other than 10,000 — those streams overlap the pins, so the two rates are paired rather than comparable
2065
+ # 132 `zero-ctl:` control fixtures (22 same-label, 110 pair-key) sit at a family-adjusted bar, which
2066
+ # bounds a valid run's abort probability at 5% — they are NOT independent (the seed key omits the mode,
2067
+ # so each pair's League and Knockout controls share a stream), so 5% is a ceiling, not a rate. The
2068
+ # reported diagonals are NOT checked. Investigate an
2069
+ # exceedance first; only then re-run under another SEED_NAMESPACE — never widen the bar to pass
2070
+ ```
2071
+
2072
+ ---
2073
+
1706
2074
  ## Note for the maintainers
1707
2075
 
1708
2076
  > **2026-08-19: this note was rewritten twice in one day.** First it said the