pog-mcp 0.9.7 → 0.9.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pog-mcp",
3
- "version": "0.9.7",
3
+ "version": "0.9.8",
4
4
  "type": "module",
5
5
  "description": "MCP server that lets an AI agent play Proof of Goal — wallet, sign-in, squad building, and matches as typed tools.",
6
6
  "license": "MIT",
@@ -2544,18 +2544,24 @@ that was an artefact of an unattainable +4 buff and is withdrawn.
2544
2544
  the short side drops by the same Δ. A real fatigued squad is uneven.
2545
2545
 
2546
2546
 
2547
- ## What growth-neutral load would a player accrue? (`engagement-load.mts` + `c2-team-load.ts`)
2548
-
2549
- **This section does NOT close C-2.** An earlier version said it did; that was wrong. `τ` was
2550
- never measured the probe's 24h default was simply used — and the neutral point was left
2551
- unchosen between its two branches. An F-1 implementer following that decision would have made
2552
- a probe default the production recovery rate. And `τ` cannot be measured from this data: at the
2553
- time of this measurement F-1 did not exist, so production had no fatigue and there was no observed
2554
- recovery curve to fit. What measurement CAN do is price the choices, which is what the τ sweep
2555
- does. Predicate C versus D was the third product choice at publication; it was subsequently
2556
- resolved as **D** on 2026-09-04. The published values remain conditional on C; switching to D
2557
- moves τ=24h median/p95 `L` by −5.3%/−10.4%. That measures the price of the choice, and selecting D
2558
- does not retroactively relabel C-conditioned calibration as D-conditioned evidence.
2547
+ ## What growth-neutral load would a player accrue? measurement and product decision
2548
+
2549
+ The first #783/#784 measurement did **not** close C-2. `τ` was not measurable the probe's 24h
2550
+ default had merely been used — and the neutral point was left between two branches. Production
2551
+ had no fatigue curve to fit before F-1; measurement could price the choices but could not make the
2552
+ product choice. Predicate C versus D was also open at publication and F-1 subsequently persisted
2553
+ **D** on 2026-09-04. The C-conditioned tables and correction trail below remain historical evidence;
2554
+ they are not relabelled as D.
2555
+
2556
+ **Product decision, 2026-09-06 F-2 reference curve.** Re-integrating the same owner-only export
2557
+ and stored seed column under D, then independently validating only the selected curve, chooses
2558
+ **`τ=24h`**, active-appearance neutral load **`L_neutral=8.076776`**,
2559
+ **`L₀=90.717176`** (4× D p95 `22.679294`), **`b=0.5399`**, and therefore
2560
+ **`a=b·L_neutral/(L_neutral+L₀)=0.044139`**. The versioned equation is
2561
+ `cond=5+a−bL/(L+L₀)`. This does not activate fatigue: F-2 first lands behind a default-OFF flag,
2562
+ and G-3 can flip only after the maximum legal role×condition interaction, 8–11-match cup depletion,
2563
+ the launch-amplitude prize/cost ordering, H ceiling, and I-c CAP clear against this same curve.
2564
+ The D product rerun is recorded below; the older C run remains the audit trail.
2559
2565
 
2560
2566
  ### F-1 implementation status — predicate D is now durable
2561
2567
 
@@ -3418,14 +3424,64 @@ rule-of-three bound was bolted on for it. Under empirical Bernstein the range te
3418
3424
  `3R ln(3/δ)/n` survives at SD = 0, so the interval never collapses in the first place: no
3419
3425
  guard, no special case, and one fewer argument resting on its own iid assumption.
3420
3426
 
3427
+ ### Post-F-1 product rerun — certify only the selected D · 4× curve
3428
+
3429
+ Results 6 through the safe-amplitude section above preserve the historical predicate-C run and
3430
+ its joint 1×/2×/4× candidate family. The product values were not obtained by scaling that table.
3431
+ All four τ values were re-integrated and tested from scratch with
3432
+ **`LOAD_PRED=D L0_SWEEP=4`**, using the same stored seed column and the same owner-only export,
3433
+ SHA-256 `1c7335ba9fa11388363de57628207ecc4820dc9bb02285066bda849adf571b6f`.
3434
+ Letting discarded 1×/2× curves choose the worst cell or enlarge the family would charge the
3435
+ shipped curve for hypotheses the product does not ship. The 4× choice was fixed from input-load
3436
+ geometry, not outcome seeds, before the result bounds; its family is
3437
+ `(26+26+26+24)×6 opponents = 612`.
3438
+
3439
+ | τ(h) | D `L` median | D `L` p95 | 0.2-derived `b` (4×) | one match (pp)\* | cells over | **product-curve safe `b`** | worst upper at low end |
3440
+ | --- | --- | --- | --- | --- | --- | --- | --- |
3441
+ | 6 | 2.51 | 10.52 | 3.548 | 0.2948 | **14/16** | 0.4296–0.4331 | 3.74pp |
3442
+ | 12 | 4.07 | 14.29 | 3.009 | 0.1876 | **7/16** | 0.4672–0.4701 | 3.44pp |
3443
+ | **24** | **8.08** | **22.68** | **2.446** | **0.0980** | **0/18** | **0.5399–0.5423** | **3.69pp** |
3444
+ | 48 | 16.69 | 40.03 | 2.119 | 0.0488 | **0/17** | 0.5546–0.5567 | 3.70pp |
3445
+
3446
+ \* “One match” is the secant conversion at each row's `b` derived from E's 0.2 median spread,
3447
+ kept only for comparison. It is not the one-match cost at the selected safe `b`. The safe range
3448
+ is not a confidence interval either: its left endpoint is the largest grid point independently
3449
+ established below for every actual cell, while the right endpoint is the smallest grid point not
3450
+ established below. No row established a lower bound above 5pp, so none gets an upper threshold.
3451
+
3452
+ The product choices are:
3453
+
3454
+ - **`τ=24h`.** The 6h and 12h 0.2-derived curves put 14 and 7 cells over; 24h is the first with
3455
+ zero. 48h is also zero but halves the one-match secant again, 0.0980→0.0488pp. This is a product
3456
+ compromise between daily rotation and fairness, not a recovery estimate fitted to observations.
3457
+ - **`L₀=90.717176`.** This is 4× the D/τ24 pre-appearance p95 `22.679294`, and 2.12× the observed
3458
+ maximum `42.87`. At that maximum the marginal slope retained relative to zero load is
3459
+ 12.0%/26.4%/46.1% for 1×/2×/4×. Only 4× meaningfully follows §3-1's prescription to keep the
3460
+ curvature scale above the realistic range so marginal cost does not die early.
3461
+ - **Neutral `L=8.076776`, `b=0.5399`, `a=0.044139`.** The median active appearance maps to
3462
+ condition 5. Zero load maps to 5.044, D p95 to 4.936, this window's observed maximum to 4.871,
3463
+ and the asymptote to 4.504. At `b=0.5399`, the independent worst upper bound is 3.69pp across
3464
+ all 18 actual cells.
3465
+ - **Prize versus cost remains unresolved.** At the safe amplitude's maximum recoverable deficit
3466
+ 0.173, the prize is 0.19pp [−0.07, 0.45], the selected cheapest one-point replacement costs
3467
+ 0.16pp, and paired cost−prize is −0.03pp [−0.50, 0.43]. Together with the unmeasured legal
3468
+ role×condition and cup extremes, that keeps F-2 default OFF until the follow-up gates and G-3.
3469
+
3470
+ The canonical constants are therefore `τ=24h` and
3471
+ `cond = 5 + 0.044139 − 0.5399·L/(L + 90.717176)`. Implementation derives `a` from the other
3472
+ three constants so documentation rounding cannot move the exact neutral point.
3473
+
3421
3474
  ### Stage 3 — a multi-wallet feeder is exempt, not merely advantaged
3422
3475
 
3423
3476
  F-1's side rule writes nothing for a ladder away side that is human; `playoff-finalize.ts`
3424
3477
  calls `applyResult` for both; `selectOpponentByStrength` draws from the five strength-nearest
3425
3478
  in-division; the cooldown is initiator-only. Observed at-appearance `L` reaches 55 for an
3426
3479
  honest player while a team taking the same match count as the away side accumulates zero. That
3427
- is exemption, not a ratio — bounded only by what bounds the feature, since the prize is 0.59pp at g=0.4 (0.16pp at the τ=24h safe amplitude)
3428
- and the team-level gap is 3.43pp at the safe amplitude, under the gate.
3480
+ is exemption, not a ratio — bounded only by what bounds the feature. The historical `g=0.4`
3481
+ single-player prize is 0.59pp [0.25, 0.93]; at the selected D product amplitude the maximum
3482
+ recoverable deficit gives 0.19pp [−0.07, 0.45], and the independently certified team-level worst
3483
+ upper bound is 3.69pp across the 18 observed cells. The unmeasured legal role×condition extreme
3484
+ is why that bound is not yet an activation verdict.
3429
3485
 
3430
3486
  ### What this does not settle
3431
3487
 
@@ -3459,10 +3515,10 @@ states while `allCells` did not. A deficit of exactly zero is not evidence of cl
3459
3515
  gate either, so such teams are excluded outright and the output says how many. None occur in
3460
3516
  this window.
3461
3517
 
3462
- The prize-versus-cost ordering AT THE AMPLITUDE THAT WOULD SHIP it splits at `g=0.4`
3463
- (−0.43pp [−0.76, 0.10]) but not at the safe amplitude's own deficit (0.133; prize 0.16pp
3464
- [−0.09, 0.40], cheapest cost 0.16pp, paired difference 0.00pp [−0.47, 0.47]), and that second
3465
- one is the decision. The neutral-point fork. Maximum LEGAL role concentration (the
3518
+ The prize-versus-cost ordering at the D product amplitude remains unresolved: at its maximum
3519
+ recoverable deficit 0.173, prize is 0.19pp [−0.07, 0.45], cheapest selected cost is 0.16pp, and
3520
+ paired cost−prize is −0.03pp [−0.50, 0.43]. The neutral-point fork is now closed at the D active-appearance median;
3521
+ what remains is maximum LEGAL role concentration (the
3466
3522
  replay is what happened, not what the rules permit). Player identity is the snapshot's `actorIds` playerId, keyed GLOBALLY rather than per team —
3467
3523
  a minted player sold mid-window would otherwise get two chains starting from zero, understating
3468
3524
  the buyer's vector. A missing actor map or even one player that fails to join is rejected by the
@@ -3486,9 +3542,15 @@ itself. The cost is that nothing with a period longer than 11 days is visible.
3486
3542
  C2_ENV_FILE=<main checkout>/.env.deploy DAYS=30 WINDOW_END=2026-09-01 \
3487
3543
  EXPORT_SNAPSHOTS=/tmp/prod-tl.json \
3488
3544
  pnpm --filter @sws26/api exec tsx scripts/c2-team-load.ts
3545
+ # Post-F-1 product curve: predicate D, testing only the selected L₀=4×p95
3546
+ MATCHES=3000 SNAPSHOTS=/tmp/prod-tl.json REPLAYS=5 TAU_SWEEP=6,12,24,48 \
3547
+ LOAD_PRED=D L0_SWEEP=4 LAMBDA_MED=6 LAMBDA_MAX=21 \
3548
+ pnpm --filter pog-mcp exec tsx skill/reference/probes/engagement-load.mts
3549
+
3550
+ # Historical C-conditioned 1×/2×/4× table only
3489
3551
  MATCHES=3000 SNAPSHOTS=/tmp/prod-tl.json REPLAYS=5 TAU_SWEEP=6,12,24,48 \
3490
- LAMBDA_MED=6 LAMBDA_MAX=21 \
3491
- npx tsx packages/mcp/skill/reference/probes/engagement-load.mts
3552
+ LOAD_PRED=C L0_SWEEP=1,2,4 LAMBDA_MED=6 LAMBDA_MAX=21 \
3553
+ pnpm --filter pog-mcp exec tsx skill/reference/probes/engagement-load.mts
3492
3554
  ```
3493
3555
 
3494
3556
  **The window's end is now stated explicitly, so these figures can be rebuilt.** The default is
@@ -3567,9 +3629,12 @@ from an artificially low `L` and `Lmed`, the gate vectors and the safe amplitude
3567
3629
  understated. Engine changes are input-dependent, so interspersed failures are a normal
3568
3630
  possibility rather than a clean release boundary.
3569
3631
 
3570
- `LOAD_PRED` selects the integration predicate (default `C`). C is published, and the probe also
3571
- integrates at D at every τ and prints the difference. F-1 subsequently selected D, but the
3572
- published figures remain C-conditioned: calibration must always carry the predicate it used.
3632
+ `LOAD_PRED` selects the integration predicate (default `D`). D is the F-1 product predicate, and
3633
+ the probe also integrates the opposite predicate C at every τ and prints the difference. Set
3634
+ `LOAD_PRED=C` explicitly to reproduce the historical C-conditioned figures: calibration must
3635
+ always carry the predicate it used. `L0_SWEEP` selects the p95 multiples under test. Its default
3636
+ `1,2,4` preserves the historical comparison; the product amplitude is certified separately with
3637
+ `L0_SWEEP=4`, so an unselected curve neither enlarges the family nor chooses the worst cell.
3573
3638
 
3574
3639
  Without `TAU_SWEEP` only `TAU_H` runs. **Calling a single-τ run a "decision" makes a probe
3575
3640
  default the production recovery rate** — that is why the sweep exists.