pog-mcp 0.9.17 → 0.9.19
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/skill/reference/measurements.md +366 -0
package/package.json
CHANGED
|
@@ -4959,3 +4959,369 @@ few points, as any change to who presses must.
|
|
|
4959
4959
|
original expression cannot be recovered.
|
|
4960
4960
|
|
|
4961
4961
|
5,648,400 matches in about a minute. Whole run: `zp-selector.mts`.
|
|
4962
|
+
|
|
4963
|
+
## Do the six pog axis candidates move the outcome, and are any of them a decision? (`pog-axes.mts`)
|
|
4964
|
+
|
|
4965
|
+
Six candidates for the pog playbook (issue #809, gate G0): `balance`, `line`,
|
|
4966
|
+
`press`, `build`, `keeper`, `set_piece`. Every one ships as a **selection**
|
|
4967
|
+
preference — abilities stay frozen from creation (F-6) and a playbook only
|
|
4968
|
+
chooses who starts, at which position, and who takes kicks. This probe
|
|
4969
|
+
measures by sweeping 212-point TEMPLATES: the question is whether the engine
|
|
4970
|
+
is sensitive to the profile difference at all; selection realizes it
|
|
4971
|
+
discretely, checked separately in a 17-player-pool companion measurement.
|
|
4972
|
+
|
|
4973
|
+
Five of the six realizations change `squad-lib.mts`'s `build()` inputs;
|
|
4974
|
+
`set_piece` instead reassigns the already-built team's `fkKicker`/`pkKicker`
|
|
4975
|
+
flags directly (`build()` itself hardcodes those flags onto slots 9/10, so
|
|
4976
|
+
this is still a 212-template property, just set via a different code path).
|
|
4977
|
+
`D_*_APPR` (issue #745) is the one axis that patches the engine, via the
|
|
4978
|
+
`appr-role.mts`/#728 exact-match-count-asserted source-replace pattern.
|
|
4979
|
+
|
|
4980
|
+
Eight Codex review rounds (41 findings total, 1 deferred to a follow-up
|
|
4981
|
+
issue rather than fixed inline — see "What Codex found") progressively
|
|
4982
|
+
hardened the isolation guarantees. Highlights load-bearing for the tables
|
|
4983
|
+
below: `balance`/`press` pin the goalkeeper BYTE-IDENTICAL to the base
|
|
4984
|
+
squad, including for `mid`. `line` (round 8) instead clones all 11 existing
|
|
4985
|
+
players byte-for-byte and relabels only the `position` field to the new
|
|
4986
|
+
formation's slot — no attribute rebuild at all, so no GK-pinning pipeline
|
|
4987
|
+
is even needed for this axis. `D_*_APPR`'s focus policies override the
|
|
4988
|
+
argmax player's selection weight to 1,000,000 (not 1024 — see "What Codex
|
|
4989
|
+
found"). `set_piece` ranks FK/PK candidates independently by the actual
|
|
4990
|
+
engine formula: `getPoint`/`takePkShot` show both use `atkSkill + cond/2 +
|
|
4991
|
+
total/3 + roll`, and FK's single `fkKicker` flag triggers a 50/50 mix of a
|
|
4992
|
+
shoot-based and a pass-based branch, so the expected-value-optimal ranking
|
|
4993
|
+
for BOTH roles is `(pass+shoot)/2 + total/3` — with a tie-break that keeps
|
|
4994
|
+
the two roles on different players. This omitted `ofTeamContrib`, a
|
|
4995
|
+
team-aggregate term `getPoint` also adds — round 8's Finding 2, deferred at
|
|
4996
|
+
the time (see "What Codex found") and later closed via issue #852: PK needs
|
|
4997
|
+
no correction (`takePkShot` bypasses `getPoint` entirely), and FK's ranking
|
|
4998
|
+
now adds `0.5*ofTeamContribDirect` (only the direct/shoot branch reaches a
|
|
4999
|
+
nonzero team-aggregate term; the crossed/pass branch's rate is 0 for every
|
|
5000
|
+
position). Closing it flipped `set_piece`'s own monotone classification —
|
|
5001
|
+
see Result and "What Codex found" below.
|
|
5002
|
+
|
|
5003
|
+
### Method
|
|
5004
|
+
|
|
5005
|
+
Four cores (`flat`, `shaped`, `gkheavy`, `gkmin`). Each axis is a `low`/`mid`/
|
|
5006
|
+
`high` ladder; `low`/`high` change exactly one dimension of `build()`'s
|
|
5007
|
+
inputs (see `docs/pog/02-axes-measurement.md` §2). Opponent grid = the same
|
|
5008
|
+
four cores, both modes, `N=20,000` per cell, seeds keyed on
|
|
5009
|
+
`(axis, myCoreKey, oppCoreKey, mode)`. Bonferroni family size (m=706)
|
|
5010
|
+
computed analytically before any comparison runs. Non-monotone requires BOTH
|
|
5011
|
+
ladder legs to individually clear the bar with opposite signs, on the
|
|
5012
|
+
self-play cell. Opponent-conditional requires a candidate to beat both
|
|
5013
|
+
others at the same bar, for each opponent.
|
|
5014
|
+
|
|
5015
|
+
Adoption is judged on "significant against any of the 4 opponents", which
|
|
5016
|
+
does not always match the displayed self-play row alone — both counts are
|
|
5017
|
+
now reported side by side for reconciliation (see Result below).
|
|
5018
|
+
|
|
5019
|
+
Controls: patch-identity; mirror cells; a `D_*_APPR` policy mirror; a live
|
|
5020
|
+
assertion that no knockout match returns a null winner; every 17-player
|
|
5021
|
+
pool's full 2^17=131,072 possible 11-subsets checked and validated; and the
|
|
5022
|
+
reported grand total checked against a value computed analytically from the
|
|
5023
|
+
loop structure.
|
|
5024
|
+
|
|
5025
|
+
### Result
|
|
5026
|
+
|
|
5027
|
+
| axis | adopted | sig (self-play only) | sig (any of 4 opp) | max \|Δ\| pp | non-monotone | opponent-conditional |
|
|
5028
|
+
| --- | --- | --- | --- | --- | --- | --- |
|
|
5029
|
+
| `balance` | yes | 7/8 | 8/8 | 33.89 | **0/8** | **0/8** |
|
|
5030
|
+
| `line` | yes | 8/8 | 8/8 | 22.87 | 1/8 | 3/8 |
|
|
5031
|
+
| `press` | yes | 7/8 | 8/8 | 39.88 | 2/8 | **7/8** |
|
|
5032
|
+
| `build` | yes | 7/8 | 8/8 | 12.52 | 1/8 | 0/8 |
|
|
5033
|
+
| `keeper` | yes | 8/8 | 8/8 | 19.02 | **0/8** | **0/8** |
|
|
5034
|
+
| `set_piece` | yes | 8/8 | 8/8 | 11.38 | **0/8** | 0/8 |
|
|
5035
|
+
|
|
5036
|
+
All six adopted (judged on the any-opponent column, matching the code's
|
|
5037
|
+
adoption rule); `line`/`press`/`build` non-monotone, `balance`/`keeper`/
|
|
5038
|
+
`set_piece` cleanly monotone. G0 passes (3 of 6 non-monotone, above the
|
|
5039
|
+
2-required bar). `press` remains the strongest trade-off: 7 of 8 core×mode
|
|
5040
|
+
cells have a supported leader that flips across the 4-opponent grid.
|
|
5041
|
+
`set_piece`'s non-monotone count in this table (0/8) reflects the #852 fix
|
|
5042
|
+
below, not the original round-8 measurement (4/8, `flat`×2 + `gkmin`×2) —
|
|
5043
|
+
closing the `ofTeamContrib` gap moved `high`'s FK pick on those two cores
|
|
5044
|
+
onto a candidate matching `mid`'s own default FK kicker (same position and
|
|
5045
|
+
attributes); the accompanying PK-tiebreak shift measures indistinguishably
|
|
5046
|
+
from `mid` too, though for a more specific reason than "identical kickers"
|
|
5047
|
+
(see the round-3 correction below) — together removing the dip that made
|
|
5048
|
+
those two cores' ladders non-monotone. `shaped`/`gkheavy` are byte-identical to the
|
|
5049
|
+
pre-#852 numbers — their position groups aren't attribute-uniform, so the
|
|
5050
|
+
added team-aggregate term didn't move the top-ranked candidate. Round 6
|
|
5051
|
+
additionally found FK's ranking was missing half its own mechanic (see
|
|
5052
|
+
"What Codex found").
|
|
5053
|
+
|
|
5054
|
+
`D_*_APPR` (#745's three questions): `focus-total` is now restricted to
|
|
5055
|
+
FORWARDS specifically (round 7 — searching all outfielders let a DF win the
|
|
5056
|
+
argmax whenever rounding gave one the squad's top total, e.g. `gkheavy`'s
|
|
5057
|
+
19-point DF against 18-point forwards, silently testing "focus on whoever
|
|
5058
|
+
has the most points" rather than #745's actual forward-specific claim). The
|
|
5059
|
+
picture sharpened considerably: only `shaped` (STARS weighting, clearly
|
|
5060
|
+
uneven totals with forwards on top) benefits from `focus-total`
|
|
5061
|
+
(+16.65/+19.80pp); the other three cores, where totals are flatter or a
|
|
5062
|
+
forward isn't the max, take a LARGE loss from forcing defensive duty onto a
|
|
5063
|
+
forward (−30 to −46pp) — a much more precise reproduction of #745's own
|
|
5064
|
+
conditional claim ("only works when totals are uneven") than the
|
|
5065
|
+
all-position version showed. The reversal boundary is confirmed in both
|
|
5066
|
+
modes at outfield total-spread 17 (`focus-total` wins, z=36.61 REG / 30.03
|
|
5067
|
+
KO); at every lower spread tested (1, 6, 11) `focus-defense` wins
|
|
5068
|
+
confirmed, now by far larger margins (up to −47.6pp) since the two policies
|
|
5069
|
+
target genuinely different players. Opponent-conditional: 0/8.
|
|
5070
|
+
|
|
5071
|
+
**17-pool realizability (exploratory, not gating)**: same construction as
|
|
5072
|
+
before (three pools of 17, all 2^17 subsets checked). Two limitations now
|
|
5073
|
+
documented, pulling in OPPOSITE directions: (1, round 5) the pool assumes
|
|
5074
|
+
each of the 17 individuals' POSITION and kicker flags are fixed at their
|
|
5075
|
+
build-time template, but `AGENTS.md` states F-6 freezes ABILITY VALUES
|
|
5076
|
+
only, not "pool membership, mutable unminted positions/names, or FK/PK
|
|
5077
|
+
choices" — the real product may allow repositioning or kicker reassignment
|
|
5078
|
+
this check cannot see, understating realizability; (2, round 7) an
|
|
5079
|
+
alternate's total is copied from the slot it replaces, which squad-lib's
|
|
5080
|
+
own `slotTotals()` clamps to [10,29] — wider than the production FA
|
|
5081
|
+
generator's actual reachable range of 16-28 (`AGENTS.md` §3-16). `pool-shaped`/
|
|
5082
|
+
`pool-gkheavy` do produce alternates outside that window (as low as 12, as
|
|
5083
|
+
high as 29 — a keeper on `pool-gkheavy`), i.e. synthetic reserves the real
|
|
5084
|
+
generator could never mint, which OVERSTATES realizability for any verdict
|
|
5085
|
+
depending on one. This exploratory check does not net the two limitations
|
|
5086
|
+
against each other — the true direction is unresolved. Full table and
|
|
5087
|
+
caveats: `docs/pog/02-axes-measurement.md` §6.
|
|
5088
|
+
|
|
5089
|
+
### What Codex found, and what changed (eight review rounds, 41 findings)
|
|
5090
|
+
|
|
5091
|
+
Rounds 1–4 are summarized in `docs/pog/02-axes-measurement.md` §7 (GK
|
|
5092
|
+
weight-share leaks across three different axes and two different root
|
|
5093
|
+
causes, cross-slot scarcity-repair leaks, unvalidated/undersized/
|
|
5094
|
+
misnamed pool XIs, unsupported opponent-conditional argmax, combined then
|
|
5095
|
+
still-colliding FK/PK selection, proportional instead of concentrated
|
|
5096
|
+
`D_*_APPR` policies, `mid` built through a different algorithm than
|
|
5097
|
+
`low`/`high`, several doc-accuracy corrections).
|
|
5098
|
+
|
|
5099
|
+
**Round 5 (2 P1, 2 P2)** — new defects inside rounds 1–4's fixes:
|
|
5100
|
+
|
|
5101
|
+
- `D_*_APPR`'s focus-weight override of 1024 (round 3's fix) wasn't large
|
|
5102
|
+
enough: in the highest-competition scenes (A0/A1, where up to 4 DF plus
|
|
5103
|
+
the GK each carry ~100 stock weight, summing to ~500), 1024 only bought
|
|
5104
|
+
the focal player 62-72% of the draw rather than near-total concentration
|
|
5105
|
+
(measured on `shaped`/A0). `appr-role.mts`'s 1024 was calibrated for the
|
|
5106
|
+
offence table's much smaller competing sums, not this one. Raised to
|
|
5107
|
+
1,000,000 — a three-orders-of-magnitude safety margin that guarantees
|
|
5108
|
+
>99.9% share regardless of competing weight.
|
|
5109
|
+
- `set_piece`'s FK/PK ranking (round 4's independent-selection fix) still
|
|
5110
|
+
ranked by the bare `shoot` / `(pass+shoot)/2` stat alone, omitting the
|
|
5111
|
+
`total/3` term `getPoint` and `takePkShot` actually add to the kicker's
|
|
5112
|
+
point. On `flat` (every outfielder has identical `shoot`), ties were
|
|
5113
|
+
broken by JS's stable-sort array order — which happened to correlate with
|
|
5114
|
+
`slotTotals()`'s rounding drift (a couple of slots get one extra total
|
|
5115
|
+
point) — landing the "worst" AND "best" ranking on the same higher-total
|
|
5116
|
+
DF/DMF slot while `mid` uses the FW slots, manufacturing a discontinuity
|
|
5117
|
+
unrelated to kicker quality. Including `total/3` makes ties genuine
|
|
5118
|
+
engine-level ties rather than rounding artifacts wearing a tie's mask.
|
|
5119
|
+
- Two P2s: the 17-pool check doesn't model that unminted players' positions
|
|
5120
|
+
and kicker flags are mutable in the real product (`AGENTS.md`) — documented
|
|
5121
|
+
as a lower-bound limitation rather than modeled (modeling it properly needs
|
|
5122
|
+
a lineup solver, P1-2's job); and the "significant" cell counts didn't
|
|
5123
|
+
match the displayed self-play rows (adoption is judged on "any of 4
|
|
5124
|
+
opponents", which a reader can't verify from the self-play table alone) —
|
|
5125
|
+
now both counts are reported side by side.
|
|
5126
|
+
|
|
5127
|
+
**Round 6 (1 P1)** — a new defect inside round 5's own fix:
|
|
5128
|
+
|
|
5129
|
+
- `set_piece`'s FK ranking (round 5's `total/3` fix) still measured only
|
|
5130
|
+
HALF of the `fkKicker` mechanic. `doFk` sends every free kick through one
|
|
5131
|
+
50/50 coin flip between `doFkDirect` (`getPoint(..., 'shoot', SCENE.FK1)`)
|
|
5132
|
+
and `doFkCross` (`getPoint(..., 'pass', SCENE.FK2)`) — the same single
|
|
5133
|
+
flag governs both branches. Ranking by `shoot + total/3` alone matched
|
|
5134
|
+
only the direct-kick half in isolation (presumably what
|
|
5135
|
+
`docs/tactics/00-plan.md`'s narrower "FK=shoot" finding itself measured),
|
|
5136
|
+
not the blended mechanic the flag actually controls. Reading `getPoint`
|
|
5137
|
+
directly confirmed both branches use the identical
|
|
5138
|
+
`atkSkill + cond/2 + total/3 + roll` shape with only `atkSkill` swapped
|
|
5139
|
+
between `shoot` and `pass`, so the expected-value-optimal FK ranking is
|
|
5140
|
+
`(shoot+pass)/2 + total/3` — the same formula PK already uses. The two
|
|
5141
|
+
roles now genuinely converge on the same "best all-round kicker" ranking,
|
|
5142
|
+
which is the correct consequence of the engine's real mechanics, not a
|
|
5143
|
+
bug — it just means round 4's "PK falls to the next-best candidate when
|
|
5144
|
+
it ties with FK's pick" tie-break now fires on nearly every ladder cell
|
|
5145
|
+
instead of occasionally. **This held only through round 8** — issue #852
|
|
5146
|
+
(below) later added a term to FK's formula alone, so FK and PK diverge
|
|
5147
|
+
again and the tiebreak fires only when they happen to coincide (on
|
|
5148
|
+
`flat`/`gkmin` specifically, it no longer fires at all — see #852's own
|
|
5149
|
+
entry for the mechanics).
|
|
5150
|
+
|
|
5151
|
+
**Round 7 (1 P1, 4 P2)** — new defects inside earlier rounds' fixes:
|
|
5152
|
+
|
|
5153
|
+
- `D_*_APPR`'s `focus-total` searched the argmax total across ALL outfield
|
|
5154
|
+
positions, not specifically forwards — #745's own claim names "your
|
|
5155
|
+
biggest FORWARD" explicitly. On `gkheavy` (19-point DF vs 18-point
|
|
5156
|
+
forwards) and the B3 boundary's alpha=0 squad, rounding made a DF the
|
|
5157
|
+
squad-wide max, so `focus-total` was silently testing "focus on whoever
|
|
5158
|
+
has the most points" rather than #745's forward-specific policy.
|
|
5159
|
+
Restricting the search to `position === 'FW'` sharpened the picture
|
|
5160
|
+
considerably (see Result above): only `shaped` (clearly uneven totals,
|
|
5161
|
+
forwards on top) benefits, and the other three cores take a large loss
|
|
5162
|
+
(−30 to −46pp) from forcing defensive duty onto a forward — a much more
|
|
5163
|
+
precise reproduction of #745's conditional claim than the all-position
|
|
5164
|
+
version showed.
|
|
5165
|
+
- Four P2s: both documents still described FK's ranking as `shoot + total/3`
|
|
5166
|
+
even after round 6 corrected the code to `(pass+shoot)/2 + total/3` —
|
|
5167
|
+
fixed to match the committed treatment; the probe's own header comment
|
|
5168
|
+
claimed all six axes are template-input changes via `build()`, when
|
|
5169
|
+
`set_piece` actually reassigns already-built kicker flags directly — now
|
|
5170
|
+
distinguished; the header also still claimed `mid` is "always the bare
|
|
5171
|
+
core", which stopped being true for `balance`/`press`/`line` once round 4
|
|
5172
|
+
rebuilt `mid` through the same pinned pipeline as `low`/`high` (the exact
|
|
5173
|
+
difference round 4 used to find `balance`'s spurious reversal) — corrected
|
|
5174
|
+
to state the real per-axis contract; and the 17-pool's bench alternates can
|
|
5175
|
+
carry totals outside the production FA generator's actual reachable range
|
|
5176
|
+
of 16-28 (`AGENTS.md` §3-16) — as low as 12, as high as 29 for a keeper —
|
|
5177
|
+
documented as a SECOND limitation pulling in the OPPOSITE direction from
|
|
5178
|
+
round 5's position/kicker-flexibility one (this one makes some "가능"/
|
|
5179
|
+
"부분" verdicts optimistic rather than conservative; the two are not
|
|
5180
|
+
netted against each other).
|
|
5181
|
+
|
|
5182
|
+
**Round 8 (2 P1)** — new defects inside earlier rounds' fixes:
|
|
5183
|
+
|
|
5184
|
+
- **[Fixed]** `line`'s realization conflated formation change with attribute
|
|
5185
|
+
reallocation. `AGENTS.md` states F-6 freezes ABILITY VALUES only —
|
|
5186
|
+
unminted players' position labels and FK/PK assignments stay mutable — but
|
|
5187
|
+
the pre-round-8 `line` implementation rebuilt the outfield through
|
|
5188
|
+
`buildGkPinnedVariant` using the new formation's mix ratios, so a position
|
|
5189
|
+
change also dragged along an attribute redistribution the real product
|
|
5190
|
+
doesn't require. Rewritten to clone all 11 existing players (GK included)
|
|
5191
|
+
byte-for-byte and relabel only `position` to the new formation's slot —
|
|
5192
|
+
`buildGkPinnedVariant` is no longer needed for this axis at all, since no
|
|
5193
|
+
attribute is touched. The now-dead `avgOutfieldWeightByPosition()` helper
|
|
5194
|
+
was deleted. Re-measured: `flat` (EVEN weighting, where the old and new
|
|
5195
|
+
approaches coincide by construction) is byte-identical; the other three
|
|
5196
|
+
cores show smaller, more isolated formation-only effects (e.g. `shaped`
|
|
5197
|
+
−8.43→−5.00pp) — the axis's own max \|Δ\| drops from 32.20pp to 22.87pp.
|
|
5198
|
+
Non-monotone/opponent-conditional classification is unchanged (1/8, 3/8) —
|
|
5199
|
+
only the magnitude was wrong, not the qualitative call.
|
|
5200
|
+
- **[Deferred to a follow-up issue, later closed — see below]** `set_piece`'s
|
|
5201
|
+
FK/PK ranking (`(shoot+pass)/2 + total/3`, fixed in rounds 5–6) omitted
|
|
5202
|
+
`ofTeamContrib`, the team-aggregate term `getPoint`'s full formula also
|
|
5203
|
+
adds: `ofPoint = ofPlayerP + ofTeamContrib`, where `ofTeamContrib =
|
|
5204
|
+
(atkCondSum*atkRate*0.075 + atkAttrSum*atkRate*0.1) * (teamPow/100)`.
|
|
5205
|
+
Codex supplied a concrete counterexample on the `flat` core where including
|
|
5206
|
+
this term could flip the top-ranked candidate. Closing this gap requires
|
|
5207
|
+
extracting the engine's `O_*_RATE` tables and `teamPow` computation — new
|
|
5208
|
+
engine-internals research outside this issue's (#809) measurement scope.
|
|
5209
|
+
Per the coordinator's round-7 scope-management guidance, this was left as a
|
|
5210
|
+
documented "KNOWN LIMITATION" code comment rather than extending the review
|
|
5211
|
+
loop further, and tracked in a separate GitHub issue (linked from PR #840)
|
|
5212
|
+
instead. **At the time, this was believed not to affect `set_piece`'s own
|
|
5213
|
+
adopted/non-monotone verdict — the current ranking already captures the
|
|
5214
|
+
dominant term of the real mechanism; only a secondary correction term is
|
|
5215
|
+
missing.** That belief turned out to be wrong.
|
|
5216
|
+
|
|
5217
|
+
**Issue #852 (post-#809) — closed the deferred gap, and it was not merely
|
|
5218
|
+
"secondary":** reading `engine.ts` directly settled the exact formula each
|
|
5219
|
+
role needs. `takePkShot` bypasses `getPoint` entirely (its own comment says
|
|
5220
|
+
so); its formula is `(pass+shoot)/2 + cond/2 + total/3 + roll`, which
|
|
5221
|
+
`pkRanking` (`(pass+shoot)/2 + total/3`) matches in RELATIVE ORDER — not
|
|
5222
|
+
term-for-term, since it omits `cond/2` and the roll — under this probe's own
|
|
5223
|
+
uniform-`cond:5` invariant (now runtime-asserted, not just assumed). PK
|
|
5224
|
+
needs no further correction. FK's two branches reach `ofTeamContrib` asymmetrically:
|
|
5225
|
+
`getPoint`'s `effScene = (skillType==='shoot' || scene===SCENE.A0) ? 0 :
|
|
5226
|
+
scene` forces `doFkDirect` (`skillType:'shoot'`) to always read the A0
|
|
5227
|
+
column of `O_*_RATE` (FW 1.0 / OMF 0.5 / DMF 0.2 / DF 0.1, transcribed
|
|
5228
|
+
verbatim from the engine source) regardless of `SCENE.FK1`'s own numeric
|
|
5229
|
+
value, while `doFkCross` (`skillType:'pass'`, `scene: SCENE.FK2`) reads
|
|
5230
|
+
`O_*_RATE[9]` — 0.0 for every position, confirmed by reading all four
|
|
5231
|
+
arrays. `reassignKicker()` now adds `0.5*ofTeamContribDirect(p)` (half-
|
|
5232
|
+
weighted, matching the 50/50 direct/cross split) to the FK score only.
|
|
5233
|
+
Every squad this probe builds has uniform `cond:5` and sums to exactly 212
|
|
5234
|
+
(`squad-lib.mts`'s `TOTAL`), so `ofTeamContribDirect` is derived from those
|
|
5235
|
+
values rather than hardcoded, keeping the formula self-evidently correct.
|
|
5236
|
+
Re-running `N=20,000` end to end (14,066,000 matches, same loop structure,
|
|
5237
|
+
self-check passed) confirmed Codex's counterexample was real: on `flat`
|
|
5238
|
+
and `gkmin`, the top **FK** candidate actually changes from a marginally
|
|
5239
|
+
higher-total DF/OMF to a lower-total FW, because `O_FW_RATE[0]` (1.0) is
|
|
5240
|
+
10× `O_DF_RATE[0]` (0.1) — enough to overturn the old ranking's tiny
|
|
5241
|
+
`total/3`-driven tiebreak. That FW matches `mid`'s own default FK kicker
|
|
5242
|
+
(slot 9) in both position and attributes.
|
|
5243
|
+
|
|
5244
|
+
**PK is a separate story — `pkRanking` itself is unchanged by this fix**
|
|
5245
|
+
[round-3 review correction]. `ofTeamContribDirect` was added to `fkScore`
|
|
5246
|
+
only; `pkRanking` still ranks by the bare `(pass+shoot)/2+total/3`. What
|
|
5247
|
+
changes is round 4's tiebreak condition (`pkRanking[0] === fkTarget`): under
|
|
5248
|
+
the OLD formula, `fkTarget` coincided with `pkRanking`'s own top pick (both
|
|
5249
|
+
picked the same highest-`total` candidate), so the tiebreak fired and PK
|
|
5250
|
+
fell to `pkRanking[1]`. Under the NEW formula `fkTarget` moves to a
|
|
5251
|
+
different candidate (the promoted FW), so the tiebreak no longer fires and
|
|
5252
|
+
PK now gets `pkRanking[0]` directly — a candidate that was previously
|
|
5253
|
+
pushed down to second place. The two test cores land differently: on `flat`,
|
|
5254
|
+
`pkRanking[0]` is the 20-total DF, so `high`'s PK kicker is now a 20-total
|
|
5255
|
+
DF — genuinely DIFFERENT from `mid`'s 19-total FW PK kicker in both
|
|
5256
|
+
position and total (not "identical" — the 1-point total gap is just small
|
|
5257
|
+
enough that `high%`≈`mid%` at this sample size). On `gkmin`, `pkRanking[0]`
|
|
5258
|
+
is an OMF whose `(shoot+pass)/2` (6, from a 6/6 split) happens to exactly
|
|
5259
|
+
equal `mid`'s FW PK kicker's `(shoot+pass)/2` (also 6, from a 10/2 split) —
|
|
5260
|
+
different position, same average, so `pkScore` itself is identical between
|
|
5261
|
+
them, which is the real reason `high` and `mid` measure indistinguishably
|
|
5262
|
+
there. `shaped`/`gkheavy` are byte-identical before/after: their position
|
|
5263
|
+
groups are non-uniform, so the added term didn't reach far enough to flip
|
|
5264
|
+
the top FK candidate there (and the PK tiebreak dynamics above never
|
|
5265
|
+
trigger). Net
|
|
5266
|
+
effect: `set_piece`'s non-monotone count dropped from 4/8 to 0/8 —
|
|
5267
|
+
`set_piece` reclassifies from "adopted, non-monotone" to "adopted,
|
|
5268
|
+
monotone," joining `balance`/`keeper`. This does NOT change G0's own
|
|
5269
|
+
pass/fail verdict (non-monotone axis count: 4/6→3/6, still above the
|
|
5270
|
+
2-required bar) — but it does mean the deferral's own "does not affect the
|
|
5271
|
+
verdict" reasoning was incomplete: a secondary correction term changed a
|
|
5272
|
+
qualitative classification, not just a number.
|
|
5273
|
+
|
|
5274
|
+
**Remaining limitation (PR #858 review, tracked in issue #859) — the cross
|
|
5275
|
+
branch still scores the wrong player.** `fkScore`'s 0.5-weighted cross
|
|
5276
|
+
contribution still credits the KICKER's own `shoot`/`ofTeamContribDirect`,
|
|
5277
|
+
but `doFkCross`'s `crossPoints` (`getPoint(..., 'pass', SCENE.FK2, ...)`) is
|
|
5278
|
+
only a binary gate — it never reads `shoot`, and its `ofTeamContrib` is
|
|
5279
|
+
provably zero (`O_*_RATE[9]` is 0.0 for every position). If the gate
|
|
5280
|
+
succeeds, the actual shot is taken by a SEPARATELY weight-drawn receiver
|
|
5281
|
+
(`getPlayer(atk, SCENE.A1, true, kickerIdx)` — excludes the kicker, weighted
|
|
5282
|
+
by `O_*_APPR[SCENE.A1]`: OMF 100/FW 50/DMF 30/DF 20), scored by their own
|
|
5283
|
+
stats. Promoting a candidate to `fkKicker` therefore has an unmodeled
|
|
5284
|
+
second-order cost: removing them from the cross branch's receiver pool —
|
|
5285
|
+
exactly what the `flat`/`gkmin` FW promotions above do. Closing this needs
|
|
5286
|
+
the receiver draw's expected value (tractable — a weighted average, reusing
|
|
5287
|
+
`ofTeamContribDirect`) AND the gate's success probability, which is NOT
|
|
5288
|
+
context-free: it depends on which specific opponent defender contests
|
|
5289
|
+
`SCENE.FK2` that match, a value this ranking (evaluated once per `Candidate`,
|
|
5290
|
+
before any opponent is chosen) cannot know — unlike the 4-opponent grid this
|
|
5291
|
+
probe deliberately measures against elsewhere. A genuinely deeper,
|
|
5292
|
+
opponent-dependent piece of engine-mechanics research than #852's own gap,
|
|
5293
|
+
tracked separately rather than folded into this fix. The measured numbers
|
|
5294
|
+
above are still valid (they come from real simulations); the claim that this
|
|
5295
|
+
formula's picks are the true optimum is exactly as caveated as this gap.
|
|
5296
|
+
|
|
5297
|
+
### What it means
|
|
5298
|
+
|
|
5299
|
+
- **G0 passes.** 6/6 adopted, 3/6 non-monotone (originally reported as 4/6 —
|
|
5300
|
+
see issue #852 above: closing `set_piece`'s `ofTeamContrib` gap
|
|
5301
|
+
reclassified it from non-monotone to monotone; the pass/fail verdict
|
|
5302
|
+
itself is unaffected, still comfortably above the 2-required bar).
|
|
5303
|
+
`docs/pog/02-axes-measurement.md` has the full write-up and all eight
|
|
5304
|
+
review rounds' changelogs.
|
|
5305
|
+
- **The realizability column's direction is explicitly unresolved, not
|
|
5306
|
+
presented as a clean lower bound** — round 7 found a second limitation
|
|
5307
|
+
(out-of-range synthetic reserves) pulling the opposite way from round 5's
|
|
5308
|
+
(position/kicker flexibility), and this exploratory check does not net
|
|
5309
|
+
them. Important for W1's planning: some "불가능" verdicts could turn
|
|
5310
|
+
"가능" once P1-2 builds a real fixture, but some "가능" verdicts here
|
|
5311
|
+
could also turn out to be unreachable.
|
|
5312
|
+
- **Every finding across all eight rounds lived inside a FIX, not the
|
|
5313
|
+
original measurement design** — the generalizable lesson: a control that
|
|
5314
|
+
checks the harness is unbiased (mirror, patch-identity) says nothing about
|
|
5315
|
+
whether a specific numeric choice (a weight constant, a ranking formula, an
|
|
5316
|
+
allocator) actually matches the real mechanism it's meant to approximate.
|
|
5317
|
+
Each of those needed independent verification against the actual engine
|
|
5318
|
+
formula or actual competing magnitudes, not just "does the effect look
|
|
5319
|
+
directionally right."
|
|
5320
|
+
|
|
5321
|
+
14,066,000 matches in 182.5s (N=20,000, round-8 re-run — the loop structure
|
|
5322
|
+
is unchanged so the match count is identical, only wall-clock varies between
|
|
5323
|
+
runs; self-checked against the loop structure). The default `N=6,000`
|
|
5324
|
+
reaches the same G0 PASS verdict in about a third of the time. Whole run:
|
|
5325
|
+
`pog-axes.mts`. Re-run post-#852 (same `N=20,000`, same loop structure):
|
|
5326
|
+
identical 14,066,000 matches in 151.6s — wall-clock varies, the match count
|
|
5327
|
+
does not.
|