@coinrithm/mcp-trading 0.7.12 → 0.7.14

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,546 +1,613 @@
1
- # Changelog
2
-
3
- All notable changes to `@coinrithm/mcp-trading` are documented here. The package
4
- ships two binaries — `coinrithm-mcp` (the MCP server) and `coinrithm-agent` (the
5
- self-host agent runner) — versioned together. The CoinRithm **API contract** is
6
- versioned separately (see `openapi.yaml` `info.version`, currently `1.7.0`).
7
-
8
- ## 0.7.12 — 2026-09-15
9
-
10
- Clarify `whoami`, `cancel_spot_order` and `report_pm_opportunity` descriptions,
11
- removing execution-cost prose unrelated to these operations. Document actual
12
- authentication, side effects, result fields and retry behavior.
13
-
14
- Spot cancellation now advertises its existing idempotent behavior. Opportunity
15
- reporting no longer advertises unconditional idempotency: duplicate protection
16
- requires `decisionId` (or `agentTrace.decisionId`) under the same API key; the
17
- first stored record wins. Reporting remains a write despite requiring only the
18
- `read` scope, and its evidence remains explicitly self-reported.
19
-
20
- Tool names, accepted inputs and execution behavior are unchanged. This source
21
- entry does not establish registry publication, deployment or a new Glama score.
22
-
23
- ## 0.7.11 — 2026-09-15
24
-
25
- Fix opportunity reporting that previously treated resolved API failures as
26
- successful submissions. Only `ok: true` confirms a report. HTTP errors remain
27
- unconfirmed; transport failures and exceptions have an unknown delivery outcome.
28
-
29
- `CycleResult.opportunity` now contains confirmed reports only. The additive
30
- `opportunityReport` field preserves the attempted payload, outcome and status,
31
- without API error bodies or exception details. Existing consumers should use
32
- this field when they need attempted rather than confirmed evidence.
33
-
34
- The reporter retains its one-invocation-per-cycle latch and adds no retries.
35
- Focused regressions compare successful and failed reporting in skip/act cycles
36
- and verify unchanged trading results and runner state. SDK versions are unchanged.
37
- This source entry does not establish registry publication or hosted deployment.
38
-
39
- ## 0.7.10 — 2026-09-15
40
-
41
- Source changes following the published 0.7.9 release. The TypeScript SDK remains
42
- 0.3.1 and Python remains 1.8.1; their runtime source is unchanged.
43
-
44
- - Persist file-backed run identity before execution so a first-cycle process
45
- crash cannot discard the idempotency identity. Transport uncertainty is still
46
- not automatically replayed or recorded as a completed trade.
47
- - Fix runner startup on Node 18: use the imported Node crypto API instead of
48
- depending on a global crypto object. The installed-package matrix reproduced
49
- this failure on Linux, Windows and macOS.
50
- - Add the supported `@coinrithm/mcp-trading/engine` entry point, preserving
51
- existing deep imports, and separate observation accounting and opportunity
52
- reporting from cycle ordering.
53
- - Add opt-in, machine-checked crypto return predicates and visible warnings for
54
- inactive/reserved configuration. Existing agents are not automatically opted in.
55
- - Serialize scheduler migrations in a bounded transaction; add an offline
56
- credential-rotation helper and interruption/recovery rehearsal.
57
- - Pin workflow actions and verify release-tool checksums. Add installed-package
58
- compatibility and restart smoke checks across operating systems and runtimes.
59
-
60
- 0.7.10 was published and deployed on 2026-09-15. Registry and GitHub downloads
61
- matched the CI-tested archive; hosted MCP and scheduler deployments finished.
62
- See the [release record](https://github.com/CoinRithm/coinrithm-agent-trading/releases/tag/mcp-trading-v0.7.10).
63
-
64
- ## 0.7.9 - 2026-09-15
65
-
66
- Published on npm and verified on 2026-09-15: the registry archive matches the
67
- reviewed artifact and passes fresh-install checks. See the
68
- [combined release notes](https://github.com/CoinRithm/coinrithm-agent-trading/releases/tag/mcp-trading-v0.7.9)
69
- for source provenance, verification and community acknowledgments.
70
-
71
- Public market-data fidelity and runner reliability release. No MCP tool was
72
- renamed or removed, and the API **contract stays 1.7.0**.
73
-
74
- **PM evaluation budget.** Event-driven periodic prediction-market evaluations
75
- now respect `maxLlmCallsPerHour` after their cooldown elapses. Budget skips make
76
- no provider call and consume no call allowance. PM keeps its own cooldown;
77
- open-position management and explicit always-on behavior retain their existing
78
- exemptions. This runner gate is separate from hosted provider-capacity admission.
79
-
80
- **Retry-After parsing.** Missing, blank or malformed headers no longer become
81
- zero-delay retries. The runner API client uses its existing five-second fallback;
82
- explicit zero, numeric seconds and HTTP dates remain supported. Model-provider
83
- cooldowns share the parser and retain their existing one-hour cap.
84
-
85
- **API request deadlines.** Each runner API operation now has a 30-second total
86
- deadline covering response headers, body reads and all 429 retry waits. The
87
- same client serves the hosted scheduler. Embedded callers can set a finite
88
- `requestTimeoutMs` and supply an `AbortSignal`. A timeout or cancellation returns
89
- an uncertain transport result without automatically replaying a trading write.
90
- Timers and listeners are removed when the operation finishes.
91
-
92
- **State persistence.** Self-host state is serialized to a private temporary file
93
- and atomically renamed over the previous state. A failed serialization or rename
94
- leaves the prior state intact. This is atomic replacement, not a claim of durable
95
- storage across power loss.
96
-
97
- **Agent conversion.** `coinrithm-agent eject` preserves explicit `triggerPolicy`
98
- and `capitalSizing` blocks. Previously conversion could restore default hourly
99
- budgets and drop equity sizing.
100
-
101
- **Release verification.** All-source coverage gates, mandatory PostgreSQL CI,
102
- dependency updates and corrected client setup docs are included. See the
103
- [reliability record](https://github.com/CoinRithm/coinrithm-agent-trading/blob/main/docs/RELIABILITY.md).
104
-
105
- **Confirmed-action journal.** Completed-action memory now requires an action
106
- to be both accepted and executed. Failed writes and uncertain transport results
107
- retain their attempt evidence without becoming completed moves in the next
108
- decision prompt. Dry-run proposals remain unexecuted.
109
-
110
- **Direct NVIDIA retry.** One complete HTTP 500/502/503/504 response can be
111
- retried once on the identical direct NVIDIA route within the original deadline.
112
- Both attempts are retained. This does not retry trading writes or change the
113
- hosted shared-pool routing policy.
114
-
115
- **Capital reconciliation.** Frozen-balance rounding residue down to -1e-8 is
116
- normalized only in the sizing calculation after the independent reads agree.
117
- Negative spendable cash still fails closed; wallet balances are not changed.
118
-
119
- **Private decision input evidence.** The bounded numeric projection includes
120
- nested indicator inputs and context movers, with legacy v1 records still readable.
121
- It does not retain hidden reasoning or raw model output.
122
-
123
- **Compact prediction-market evidence.** Discovery and compact event-detail
124
- responses now retain the API's `source.quoteScale`, `source.methodology` and
125
- `source.supportsMarketMetrics`, plus `spreadPoints`, `probabilityBook` and
126
- each retained outcome's `normalizedProbability`. Venue-native bid/ask quotes
127
- are never rescaled or interpreted from magnitude. Normalization remains the
128
- API's calculation over the original full book, not the truncated top-five
129
- outcome list. Existing payload bounds and explicit `detail: full` behavior
130
- are unchanged.
131
-
132
- **Settlement-time provenance.** Compact events retain `resolvedAtBasis` and
133
- `settlementWindowClosedAt`, keeping provider expiration distinct from an
134
- announced settlement time. Null and absent upstream evidence stay null and
135
- absent; the MCP does not infer missing values.
136
-
137
- **Candle semantics.** The `get_candles` description now states that these are
138
- sampled composite-price bars. Each bar's `v` is a mean rolling 24-hour
139
- quote-volume observation in USD, not volume traded during the candle, and
140
- must not be summed across bars.
141
-
142
- **HTTP completion diagnostics.** The hosted HTTP
143
- entry now has a bounded, stderr-only completion observer with final SDK-result
144
- and finish/abort accounting. Initialization, discovery, tool failures and
145
- successful delivery are distinct; unknown tool names are normalized. Records
146
- contain no arguments, bodies, credentials, caller/RPC IDs or caller-origin labels.
147
- Credential presence is not authentication. Durations describe the HTTP request,
148
- shared by batch members; server finish does not prove client receipt or use.
149
- Stdio, tools, authentication and dependency versions are unchanged. See
150
- `DEPLOY.md` for the measurement and retention limits. Hosted source/image
151
- `18a0bb6a8a0665e91cebc10225fec6f7ebcdaaf7` passed a bounded anonymous smoke on
152
- 2026-09-13. Hosted verification and npm publication are separate release steps.
153
-
154
- **Deployment boundaries.** Hosted scheduler admission reasons are private
155
- scheduler telemetry, not a new SDK or MCP response field. The API's corrected
156
- comparison probabilities and enriched spread names use the existing response
157
- shape and reach current clients through fresh API reads. Outcome display names
158
- may change; use source/event/outcome identifiers for identity, never summed
159
- prices or matching labels alone. These fixes do not establish trading returns.
160
-
161
- ## 0.7.8
162
-
163
- Runner decision-quality, evidence and paper-capital release. Additive: no MCP
164
- tool was renamed or removed, and the API **contract stays 1.7.0**. This release
165
- contains all package changes since published 0.7.7 (`gitHead` `80d0cae`), not
166
- just the previously listed thesis work.
167
-
168
- **Thesis exits.** Every opening action (`futures_open`, `spot_order`,
169
- `pm_open`) now carries a `thesis`: a one-sentence summary plus an
170
- `invalidation` with at least one machine-checkable condition (`priceBelow` /
171
- `priceAbove` for coins, `probabilityBelow` / `probabilityAbove` for prediction
172
- markets, a `maxHoldMinutes` time stop, and a free-text `catalyst` the model
173
- re-judges itself). The runner binds the thesis to the position the server
174
- returns, sanitized side-aware (a rising price never invalidates a long; a
175
- wrong-side level is dropped rather than re-signed; the time stop is clamped to
176
- 60 minutes .. 30 days), persists it in the run state (`RunState.theses`, the
177
- same state file / `agent_state` JSON as before, no schema change) and
178
- re-evaluates it every cycle. A futures position whose price level or time stop
179
- is breached is closed by the runner before the model is asked anything, logged
180
- as a `thesis_invalidated` exit with its own idempotency key, after the
181
- kill-switch and drawdown checks and never instead of them. Prediction-market
182
- positions have no close endpoint, so a broken PM thesis is surfaced to the
183
- model instead (do not add, let it settle). The parser is tolerant (a malformed
184
- thesis never fails the open; a thesis copied onto a close is ignored) and the
185
- structured-output schema requires it, so schema-enforced hosted models always
186
- emit one.
187
-
188
- **Fundamentals in the observation.** Each watch entry now carries
189
- `fundamentals` sourced only from calls the runner already makes: `categories`,
190
- `marketCapRank` and `marketCapUsd` from the market context; `volume24hUsd` from
191
- the candles the `indicators` capability already fetches (live-probed
192
- 2026-09-02: each bar's `v` is a rolling 24h volume, so the latest bar is the
193
- 24h figure, never the sum); and up to three `headlines` with `publishedAt`
194
- timestamps from the one `news` call, attributed through the curated coin-news
195
- graph. Discovered PM markets carry `endDate` and `liquidityUsd`; open PM
196
- positions carry their title, side, entry and current probability and
197
- `openedAt`; open futures positions carry `openedAt`. The system prompt states
198
- the thesis contract, the runner-enforced exit and how to grade a trade on the
199
- fundamentals. Not carried, because no agent endpoint serves them: an "about"
200
- text per coin, a 24h probability change and a cross-venue divergence per PM
201
- market.
202
-
203
- **Fix:** the public movers feed serializes `change24h` / `currentPrice` as
204
- decimal strings; the universe-scan context rows read them strictly as numbers
205
- and shipped `undefined` for every mover.
206
-
207
- **Opt-in equity-based paper sizing.** A runner can size entries from a
208
- conservative fraction of its independently attributed paper book instead of a
209
- fixed stake/margin. The book is accepted only when wallet identity, cash
210
- partitions, held-position attribution and spot-mark coverage reconcile. Quotes
211
- then enforce per-entry, per-symbol, deployed-capital and daily-entry limits;
212
- fee buffers and the API's fee-inclusive quote evidence are included. Any
213
- missing or inconsistent evidence fails closed. Legacy positions on a different
214
- book remain visible for management but never inflate the current book's buying
215
- power.
216
-
217
- **Prediction-market decisions use executable economics.** PM opens now reject
218
- an invalid raw probability and a model forecast that does not clear the quoted
219
- entry price. Forecast edge is measured against the actual fee/slippage-adjusted
220
- fill, not the headline market probability. Quote-expiry outcomes are recorded
221
- separately from risk/balance rejection, and futures risk/reward validation uses
222
- fee-inclusive entry and stop economics.
223
-
224
- **Decision evidence is structured and bounded.** Cycles can expose a sanitized,
225
- partial private decision-input record: configuration and observation
226
- fingerprints, daily budget and guard state, plus bounded observation rows with
227
- explicit omission counts. It is not a prompt, transcript, raw model output or
228
- hidden reasoning record. The runner also reports quote/validation evidence for
229
- abstained, forecast-only and quote-expired PM opportunities. Hosted persistence
230
- and retention remain the caller's responsibility.
231
-
232
- **Runtime controls are more faithful.** The model sees the remaining daily
233
- entry/add budget rather than only static maxima. Entry caps still block new
234
- risk, while closes and other risk-reducing actions remain available. Direct
235
- provider HTTP 429 responses are capacity skips rather than model failures, so
236
- BYO agents do not build a failure streak during ordinary quota pressure.
237
- Structured-tool decisions remain required where the provider supports that
238
- contract.
239
-
240
- **Scorecard fix.** Maximum drawdown now measures decline from starting equity,
241
- so an immediate loss is no longer hidden by treating the first post-trade point
242
- as the high-water mark.
243
-
244
- ## 0.7.7
245
-
246
- Reliability release. Every change here came from a live production failure, not
247
- from a roadmap. Additive: no tool renamed or removed, and the API **contract
248
- stays 1.7.0** because nothing on the documented surface changed.
249
-
250
- **Model requests are now built from a declared capability table, not
251
- assumptions.** `providerCapabilities.ts` states, per model family, which
252
- parameter carries the completion budget, whether a non-default temperature is
253
- allowed, and what extra body fields the family needs. Two failures this fixes:
254
-
255
- - **OpenAI's current models rejected our requests outright.** `gpt-5*` and
256
- `o*` refuse `max_tokens` and any non-default `temperature`; they take
257
- `max_completion_tokens`. The family is detected by MODEL id, not just the
258
- provider name, so an OpenAI-compatible gateway serving `gpt-5` gets the same
259
- shape. If you brought your own OpenAI key, this is why it now works.
260
- - **NVIDIA Nemotron models emitted a think-chain where the JSON decision
261
- belonged**, which failed every cycle. The `chat_template_kwargs.enable_thinking=false`
262
- switch and the "detailed thinking off" system hint are now encoded as data
263
- rather than re-learned by failing.
264
-
265
- **New: `probeDecisionContract()`.** An HTTP 200 is not proof a route can run an
266
- agent. Both production failure modes returned 200s: a think-chain in the JSON
267
- slot, and an empty completion because a reasoning model spent its whole budget
268
- before answering. The probe sends a canned mini-observation through the REAL
269
- decision parser at a >=1024 completion allowance and classifies the result as
270
- `http`, `empty` or `parse`. Use it before adopting any model id; provider
271
- catalogs list ids that 404 on invoke.
272
-
273
- **Provider trouble no longer disables an agent.** A permanent-looking model
274
- error (404/410/decommissioned) used to disable the agent after a threshold. On
275
- 2026-08-26 NVIDIA end-of-lifed an entire model line and 35 agents died on that
276
- path. The runner now reports a hold and keeps retrying each cadence, recovering
277
- by itself when the provider does. Disables remain for what deserves them:
278
- revoked credentials, drawdown, kill-switch, user action.
279
-
280
- **Failures carry structured metadata.** A failed `decide()` now returns
281
- `status` and, when the provider sends one, `retryAfterMs` (parsed from
282
- `Retry-After` in both delta-seconds and HTTP-date form, capped at an hour), so
283
- a caller can tell a 429 from a 5xx without parsing strings. Error text is
284
- unchanged.
285
-
286
- **`ClientConfig.extraHeaders`.** Headers attached to every request, spread
287
- before auth so they can never clobber it. Self-host has nothing to put here;
288
- it exists so CoinRithm's own hosted scheduler can present its attestation
289
- channel.
290
-
291
- **Model names corrected throughout.** The retired Llama 3.x line is gone from
292
- the README, the runtime defaults and the `quant-reference` example, which is
293
- relocked onto `nvidia/nemotron-3-nano-30b-a3b`.
294
-
295
- ## 0.7.6
296
-
297
- Agent capability release: universe discovery, first-class behavioral guards,
298
- and the hosted prose budget made visible. Additive — no tool renamed or
299
- removed. Contract moves to **1.7.0** (two keyless paths declared).
300
-
301
- **New: agents can look beyond their own watchlist.**
302
-
303
- - **`get_crypto_movers` tool.** Keyless scan of the tracked coin universe for
304
- the biggest 24h gainers or losers. Rows carry `coinId`, `symbol`, `name`,
305
- `slug`, `change24hPct`, `priceUsd`.
306
- - **`universe_scan` capability** for the self-host runner. Each cycle it pulls
307
- the top movers, promotes the strongest few into full watch entries marked
308
- `discovered: true`, and passes the remainder as compact context. Watchlist
309
- and blocklist symbols are excluded up front, so a discovered row can never
310
- duplicate a configured pair or bypass the deny list.
311
- - **Both now carry the coinId through.** The movers row's `ucid` IS the
312
- `coinId` that `get_candles` / `get_market_context` / the futures quote path
313
- take. It was previously stripped from the tool response and re-derived from
314
- the SYMBOL via a resolve round-trip — a wasted call per discovered mover and
315
- a real correctness hazard, because symbols collide across listings and the
316
- resolver could return a different coin than the one that actually moved.
317
-
318
- **New: contract declares the endpoints the tools call.**
319
-
320
- - `/api/coins/top-gainers` and `/api/coins/top-losers` are now in
321
- `openapi.yaml` (tag `public-crypto-data`), so both SDKs can reach the
322
- surface `get_crypto_movers` uses. Probe-verified against prod: bare array,
323
- no envelope; `change24h` / `currentPrice` are decimal STRINGS; default
324
- `limit` is 3 and out-of-range values return 400 rather than clamping.
325
-
326
- **New: personality and boundaries are configurable, and documented.**
327
-
328
- - **`character/guards.md`** — first-class hard behavioral guards, merged into
329
- the strategy prose as a distinct section rather than buried in the thesis.
330
- - **`examples/agents/pia-pump-fader`** — a full bundle demonstrating
331
- capabilities plus boundary configuration (watchlist/blocklist interaction,
332
- the five-point risk gate, re-entry discipline).
333
- - **`examples/agents/FORKING.md`** — a file-by-file map of what is strategy
334
- and what is plumbing, so a fork knows what it is allowed to change.
335
- - **QUICKSTART** documents capabilities, and a docs-drift tripwire fails the
336
- suite when a capability ships undocumented (`universe_scan` shipped
337
- invisible in every user surface once; that cannot recur silently).
338
-
339
- **Fixed.**
340
-
341
- - **Hosted prose budget is validated, not discovered at deploy.**
342
- `coinrithm-agent validate --hosted` now checks the 8,000-character merged
343
- prose budget and reports the exact overage. A bundle could previously
344
- validate clean and still be undeployable. YAML frontmatter is stripped
345
- before the count (and before the model sees it — it was being fed in as if
346
- it were strategy). `pia-pump-fader` was rebuilt to fit at 7,932.
347
- - **Permanent failures stop being revived.** A disabled agent whose model is
348
- gone or whose key is invalid is no longer resurrected by the scheduler's
349
- revive pass; only transient failures are retried.
350
- - **Fresh scaffolds are no longer bricked** by the capabilities field, and
351
- action-confidence tolerance was widened to match what models actually emit.
352
- - **False market-data licensing assertion corrected** in both READMEs.
353
-
354
- ⚠ Publishing to npm remains a **manual** step — `publish-mcp.yml` pushes
355
- `server.json` to the MCP registry only.
356
-
357
- ## 0.7.5
358
-
359
- **Release-hygiene bump. Everything below was already merged but never
360
- reached npm** — the 0.7.4 tarball was published 2026-07-26T23:01:11Z and
361
- five commits landed after that instant without a version bump, so the
362
- repository's 0.7.4 and the published 0.7.4 were different code under one
363
- version number. This release makes the published artifact match the
364
- source again.
365
-
366
- - **Security.** MCP dependency audit fixes, 6 findings to 0 (`da65e9e`).
367
- Anyone on published 0.7.4 is running the pre-audit dependency set.
368
- - **`pm_data` Gemini exposure** for agents (`3ab04ae`).
369
- - **Venue methodology and health** exposed as tools (`74cc495`).
370
- - **Reproducible decision receipts** persisted by the agent runner
371
- (`8bebc9e`).
372
- - **Docs.** Contract version drift corrected and the placeholder SDK
373
- README replaced, so the docs stop advertising an install that 404s
374
- (`6d61b92`).
375
-
376
- No tool was renamed or removed; this is additive plus a dependency
377
- refresh.
378
-
379
- ⚠ Publishing this package is a **manual** step — `publish-mcp.yml` only
380
- pushes `server.json` to the MCP registry, it does not run `npm publish`.
381
- That asymmetry is exactly how the drift above accumulated unnoticed.
382
-
383
- ## 0.7.4
384
-
385
- Docs-only. No tool behavior change, no API-surface change.
386
-
387
- - **Acceptable Use of Market Data.** The README (root and this package) and
388
- `openapi.yaml` (`info.termsOfService`, `info.description`, and the
389
- `public-pm-data` tag) now reference and summarize CoinRithm's licensing
390
- flow-down restriction on Market Data from third-party prediction-market
391
- venues: read-only use for paper-trading context and settled-outcome
392
- scoring only — no model training/fine-tuning/benchmarking, no
393
- redistribution or bulk-extraction, no use to build a competing product.
394
- Full terms: <https://www.coinrithm.com/en/terms-of-use>
395
-
396
- ## 0.7.3
397
-
398
- Quality-engine surfaces + independent forecasts. Additive; no breaking change.
399
-
400
- - **Quality verdicts in tool responses.** `discover_pm_markets`, `pm_quote`, and
401
- the `pm_data_*` tools now surface the persisted truth-engine `quality` object
402
- (`decisionEligible`, warning/block reason codes, `policyVersion`, `assessedAt`).
403
- Markets with critical failures stay visible but cannot drive paper opens or
404
- alerts.
405
- - **`openBlocked` preview on `pm_quote`.** Quotes preview the open-time quality
406
- gate (`openBlocked` + `openBlockReasons`), so an agent can skip a market that
407
- would 422 before burning the open attempt. The self-host runner
408
- (`coinrithm-agent`) does this skip automatically.
409
- - **Independent forecasts in the runner.** The self-host agent runner elicits the
410
- model's OWN probability (judged from the question/resolution criteria/deadline,
411
- never anchored to the market price) and submits it as `forecastProbability` on
412
- PM opens — feeding the public calibration dataset with proper-scoring-rule
413
- forecasts. Clamped to [1,99]; omitted (never faked) when the model does not
414
- produce one; `HOUSE_AGENT_FORECAST_ENABLED=false` disables.
415
- - **`crossPlatform` on event lists** documented in the API contract: sibling
416
- venues pricing the same question, on list rows.
417
- - **ForecastEx venue truth.** Public MCP discovery copy and registry metadata
418
- now describe all 11 live venues, including ForecastEx.
419
- - **Contract synchronization.** Runner templates and example bundles pin the
420
- served OpenAPI 1.6.0 contract; canonical scorecard paths are unambiguous.
421
-
422
- ## 0.7.2
423
-
424
- Docs-truth + privacy release. No tool behavior change, no API-surface change.
425
-
426
- - **Ten venues in the public listing.** `pm_data_*` tool copy, the README, and
427
- `server.json` now name all ten venues (adds Futuur and Myriad). npm `0.7.1` was
428
- published before those landed, so the registry listing still advertised "eight
429
- venues"; npm versions are immutable, so correcting the public listing required
430
- a new release.
431
- - **`source` parameter description** on `pm_data_events` / `pm_data_event_detail`
432
- now enumerates all ten venue slugs. Accepted values are unchanged — this is
433
- description text only, which is why it is a patch and not a minor.
434
- - **Privacy.** Raw model output is no longer persisted, enforcing the package's
435
- no-chain-of-thought promise.
436
- - **New tripwire.** `server.json` (the MCP-registry listing) is now guarded
437
- against version and venue-count drift; it had no guard, which is how it went
438
- stale in the first place.
439
- - Refreshed stale Arena-gate example copy.
440
-
441
- ## 0.7.1
442
-
443
- Docs + registry-metadata release; no tool behavior changes.
444
-
445
- - **README refresh**: the keyless `pm_data_*` data surface is now front and
446
- center — 8 venues (Polymarket, Kalshi, Smarkets, Limitless, Manifold,
447
- Metaculus, PredictIt, Rothera — the "seven venues" line predated Rothera),
448
- the anonymous hosted-endpoint path, and the `referenceProbability` /
449
- `volumeHistory` fields the data tools return.
450
- - **`server.json`**: hosted endpoint's `Authorization` header marked optional
451
- (the `pm_data_*` tools work anonymously — verified live) and the server
452
- description now leads with the keyless data surface.
453
- - Ships the post-0.7.0 commits: `pm_data_event` advertises `volumeHistory`,
454
- `pm_data_events` advertises `referenceProbability` on list items, and the
455
- hosted MCP root (`GET /`) serves a self-describing JSON landing (with a
456
- 405 + hint on `GET /mcp`).
457
-
458
- ## 0.7.0
459
-
460
- (Retroactive entry — released 2026-07-05 without a changelog note.)
461
-
462
- - **Four keyless `pm_data_*` tools** — CoinRithm's free public cross-venue
463
- prediction-market dataset over MCP, no API key required and yours is never
464
- attached: `pm_data_overview` (market-wide stats), `pm_data_events`
465
- (cross-venue event list), `pm_data_event` (detail incl.
466
- `crossSourceMatches` + resolution evidence), `pm_data_whales`
467
- (large-trade tape).
468
-
469
- ## 0.5.0
470
-
471
- Agent-runner quality + reliability release. `coinrithm-agent` got materially
472
- smarter and less noisy; `coinrithm-mcp` is unchanged in shape. Bundles the work
473
- since 0.4.0.
474
-
475
- - **Prediction markets are a first-class venue** in the decide prompt: a short
476
- `pmN` ref so small models trade PM reliably, eligible-outcome filtering (only
477
- backend-openable outcomes reach the model), and futures-capped agents steered
478
- to PM (a separate budget) instead of re-rejecting.
479
- - **PM anti-churn — now actually effective.** The candidate list is pre-filtered
480
- to exclude markets the agent already holds, and the runner + server both block
481
- re-betting a held market+outcome (no more one-agent, 25-identical-bets churn).
482
- An earlier version read the wrong `/positions/pm` fields and was silently dead;
483
- fixed.
484
- - **Settlement-feedback learning loop.** The agent sees how its own recent bets
485
- actually resolved (win/loss/void + realized PnL) as reflective context, so it
486
- learns from outcomes across cycles.
487
- - **Per-trade reasoning stays honest about the market.** A multi-action decision
488
- no longer stamps its primary rationale onto a secondary trade about a different
489
- market — the trade's public Arena "why" always matches the market it's on.
490
- - **Futures reliability.** The model is unblinded to per-position mark /
491
- liquidation / stop / take-profit prices; take-profit is auto-clamped to a valid
492
- R:R target off the stop (kills the `take_profit_not_*_mark` reject waves); and a
493
- marking-down PM book now trips the equity-drawdown kill-switch too.
494
- - **News capability.** Recent high-importance news for the watchlist coins is fed
495
- into the decide context as a market-catalyst layer.
496
- - **Robustness + contract accuracy.** Scheduler/runner hardening, flat-state
497
- prompt steers (weak models stop hallucinating closes), manage-enum
498
- normalization, and the PM contract now documents real entry friction rather
499
- than a disclose-only stance.
500
- - **Security.** `hono` bumped to 4.12.27 (high-severity advisories: serve-static
501
- path traversal, CORS wildcard-with-credentials, body-limit bypass).
502
- - **Docs.** The npm README leads with value / free / OKF / Studio; stale
503
- scheduler and `minDecidedTrades` claims corrected.
504
-
505
- ## 0.4.0
506
-
507
- - **Deterministic scorecard engine (`computeScorecard`).** The reproducible-
508
- evaluation engine for `coinrithm.agent.scorecard.v1` — pure math over an
509
- agent's realized track record (no network, no model): realized PnL, win rate,
510
- expectancy, profit factor, reward-to-risk, Sharpe, Sortino, deflated /
511
- probabilistic Sharpe (Bailey & López de Prado — skill vs luck with a multiple-
512
- testing penalty), max drawdown, and Brier + ECE calibration for probabilistic
513
- calls. Same inputs → identical metrics **and** a sha256 `contentHash` of the
514
- canonicalized result, so a scorecard whose hash doesn't reproduce isn't
515
- trusted. Metrics are computed AFTER the run from immutable evidence (leakage-
516
- separation), so tuning-to-the-metric is structurally impossible. Returns
517
- `null` for thin records — never a fabricated number.
518
- - **Resolver: committable file metadata + functionality pin.** The OKF resolver
519
- now carries per-file metadata and pins functionality through resolution, so a
520
- bundle's behavior is reproducible from its committed files.
521
-
522
- ## 0.3.0
523
-
524
- - **Agent risk config: coin deny-list (`blocklist`).** `risk.blocklist` lets an
525
- agent name symbols it must never open, even if they are on the watchlist —
526
- deny wins over allow. Enforced in the runner's decision validator (rejects
527
- `futures_open` / `spot_order` on a denied symbol) and surfaced in the system
528
- prompt, so the model is told the boundary and the runner re-checks it.
529
- - **Docs: Open Knowledge Format positioning.** Clarified that a CoinRithm agent
530
- is an OKF bundle — a portable directory of markdown + frontmatter that is
531
- model-agnostic (run the same definition on any model). Develop and prove it
532
- free on paper, then run it anywhere.
533
- - npm keywords refreshed (`open-knowledge-format`, `okf`, `model-agnostic`,
534
- `gemini`) for registry discovery.
535
-
536
- ## 0.2.0
537
-
538
- - Added the **`coinrithm-agent`** self-host runner binary alongside the MCP
539
- server: a folder-as-architecture (OKF) agent you bring your own model key to,
540
- with caps enforced by the runner (not the model), dry-run by default,
541
- paper-only.
542
-
543
- ## 0.1.x
544
-
545
- - Initial `coinrithm-mcp` MCP server: reads, quotes, scoped spot/futures/PM
546
- writes, ledger export, and Agent Arena integration over a user-minted API key.
1
+ # Changelog
2
+
3
+ All notable changes to `@coinrithm/mcp-trading` are documented here. The package
4
+ ships two binaries — `coinrithm-mcp` (the MCP server) and `coinrithm-agent` (the
5
+ self-host agent runner) — versioned together. The CoinRithm **API contract** is
6
+ versioned separately (see `openapi.yaml` `info.version`, currently `1.7.0`).
7
+
8
+ ## 0.7.14
9
+
10
+ - Scope permanent model-error streaks to the attempted provider/model. Discard
11
+ legacy unattributed streaks and reset availability failures on a successful
12
+ provider response, including malformed decisions. Old-route failures cannot
13
+ request an early hold on a replacement model; agent risk limits are unchanged.
14
+ - Attribute a routed permanent-error hold to the attempt that produced the error,
15
+ even when a later fallback is rate-limited; keep actual-call metering unchanged.
16
+ - Preflight futures protection updates against observed side, mark and liquidation
17
+ prices, including retained triggers. Reject known invalid end-states with an
18
+ actionable reason before a write; the API remains authoritative as prices move.
19
+ - Add the optional `risk.pmMinEntryProbabilityPct` policy (0..100 points): the
20
+ runner rejects `pm_open` when the chosen outcome's raw market probability at
21
+ entry is below the floor, with a reason that names both numbers. Fees stay
22
+ in the forecast-edge check, so a 19-point outcome fails a 20-point floor even
23
+ when fees lift its cost above 20. Absent keeps today's behaviour; a set floor
24
+ with no quoted probability fails closed. The floor is rendered in the prompt's
25
+ hard caps.
26
+ - Retain candle timestamps and compact coverage/interval evidence alongside
27
+ indicators, separately from market-price freshness. Missing timestamps stay
28
+ unknown; nominal five-minute cadence does not imply current or regular bars.
29
+ Model input and private decision evidence retain the same context.
30
+ - Load optional `character/entries.md`, `exits.md`, `sizing.md` and `research.md`
31
+ from local bundles so Studio strategy sections reach the model, manifest and
32
+ compiled definition. Existing bundles without these files are unchanged.
33
+ - Include the exact compiled strategy definition and its digest in local
34
+ inspection; expose the same snapshot builder to engine consumers.
35
+ - Add `run --expect-definition` to reject a changed baseline before model or
36
+ account access. The digest is not a full replay record or broker adapter.
37
+ - Hosted validation accepts `limits.maxDailyLossMusd: 0` as no daily loss cap,
38
+ matching the runner and the hosted API; negative and non-finite values are
39
+ still rejected.
40
+ - Tactic cap merges treat 0 as unlimited for `maxTradesPerDay` and
41
+ `maxDailyLossMusd`: a tactic may tighten 0 to a positive cap but can no
42
+ longer turn a positive cap into 0. Writes and margins keep lower-is-tighter.
43
+
44
+ Publication and hosted deployment are verified separately. See the repository's
45
+ [release status](https://github.com/CoinRithm/coinrithm-agent-trading#version-clarity)
46
+ for registry availability and the hosted runtime evidence.
47
+
48
+ ## 0.7.13
49
+
50
+ - Preserve optional per-outcome venue terms in compact prediction-market output,
51
+ including explicit unavailable values and valid false/zero facts.
52
+ - Include candle volume coverage in agent fundamentals: positive counts identify
53
+ missing expected venues; unknown coverage stays unknown. Preserve measured zero
54
+ volume and keep each volume paired with its own coverage.
55
+ - Classify provider capacity failures consistently, including HTTP 503
56
+ resource exhaustion, in runner results and configured same-model retries.
57
+ - Export `classifyProviderFailure` from the supported engine entry point.
58
+ Retry counts and trading behavior are unchanged.
59
+ - Surface signed cumulative futures funding and its applied-through timestamp to
60
+ the agent observation so funding already reflected in balances is not counted twice.
61
+ - Add public whale-wallet context to the 40-tool surface: bounded 7-day or
62
+ 30-day wallet activity and source/wallet movement detail, with provenance and
63
+ provider-reported context kept distinct from CoinRithm's observed trade flow.
64
+ - Clarify `pm_data_calibration` as market-price calibration: the primary lane
65
+ uses one complete-book snapshot in the inclusive 20–28 hour pre-resolution
66
+ window, with event-weighted ECE and separate `finalPrice`/`ownCapture` lanes.
67
+
68
+ Publication and hosted deployment are verified separately. See the repository's
69
+ [release status](https://github.com/CoinRithm/coinrithm-agent-trading#version-clarity)
70
+ for registry availability and the hosted runtime evidence.
71
+
72
+ ## 0.7.12 — 2026-09-15
73
+
74
+ Clarify `whoami`, `cancel_spot_order` and `report_pm_opportunity` descriptions,
75
+ removing execution-cost prose unrelated to these operations. Document actual
76
+ authentication, side effects, result fields and retry behavior.
77
+
78
+ Spot cancellation now advertises its existing idempotent behavior. Opportunity
79
+ reporting no longer advertises unconditional idempotency: duplicate protection
80
+ requires `decisionId` (or `agentTrace.decisionId`) under the same API key; the
81
+ first stored record wins. Reporting remains a write despite requiring only the
82
+ `read` scope, and its evidence remains explicitly self-reported.
83
+
84
+ Tool names, accepted inputs and execution behavior are unchanged. Published on
85
+ npm and deployed to hosted MCP on 2026-09-15; registry and GitHub downloads match
86
+ the CI archive. A clean registry installation passed startup and tool metadata
87
+ checks. See the [release record](https://github.com/CoinRithm/coinrithm-agent-trading/releases/tag/mcp-trading-v0.7.12).
88
+ Glama's evaluation score is separate from package delivery.
89
+
90
+ ## 0.7.11 — 2026-09-15
91
+
92
+ Fix opportunity reporting that previously treated resolved API failures as
93
+ successful submissions. Only `ok: true` confirms a report. HTTP errors remain
94
+ unconfirmed; transport failures and exceptions have an unknown delivery outcome.
95
+
96
+ `CycleResult.opportunity` now contains confirmed reports only. The additive
97
+ `opportunityReport` field preserves the attempted payload, outcome and status,
98
+ without API error bodies or exception details. Existing consumers should use
99
+ this field when they need attempted rather than confirmed evidence.
100
+
101
+ The reporter retains its one-invocation-per-cycle latch and adds no retries.
102
+ Focused regressions compare successful and failed reporting in skip/act cycles
103
+ and verify unchanged trading results and runner state. SDK versions are unchanged.
104
+ This source entry does not establish registry publication or hosted deployment.
105
+
106
+ ## 0.7.10 — 2026-09-15
107
+
108
+ Source changes following the published 0.7.9 release. The TypeScript SDK remains
109
+ 0.3.1 and Python remains 1.8.1; their runtime source is unchanged.
110
+
111
+ - Persist file-backed run identity before execution so a first-cycle process
112
+ crash cannot discard the idempotency identity. Transport uncertainty is still
113
+ not automatically replayed or recorded as a completed trade.
114
+ - Fix runner startup on Node 18: use the imported Node crypto API instead of
115
+ depending on a global crypto object. The installed-package matrix reproduced
116
+ this failure on Linux, Windows and macOS.
117
+ - Add the supported `@coinrithm/mcp-trading/engine` entry point, preserving
118
+ existing deep imports, and separate observation accounting and opportunity
119
+ reporting from cycle ordering.
120
+ - Add opt-in, machine-checked crypto return predicates and visible warnings for
121
+ inactive/reserved configuration. Existing agents are not automatically opted in.
122
+ - Serialize scheduler migrations in a bounded transaction; add an offline
123
+ credential-rotation helper and interruption/recovery rehearsal.
124
+ - Pin workflow actions and verify release-tool checksums. Add installed-package
125
+ compatibility and restart smoke checks across operating systems and runtimes.
126
+
127
+ 0.7.10 was published and deployed on 2026-09-15. Registry and GitHub downloads
128
+ matched the CI-tested archive; hosted MCP and scheduler deployments finished.
129
+ See the [release record](https://github.com/CoinRithm/coinrithm-agent-trading/releases/tag/mcp-trading-v0.7.10).
130
+
131
+ ## 0.7.9 - 2026-09-15
132
+
133
+ Published on npm and verified on 2026-09-15: the registry archive matches the
134
+ reviewed artifact and passes fresh-install checks. See the
135
+ [combined release notes](https://github.com/CoinRithm/coinrithm-agent-trading/releases/tag/mcp-trading-v0.7.9)
136
+ for source provenance, verification and community acknowledgments.
137
+
138
+ Public market-data fidelity and runner reliability release. No MCP tool was
139
+ renamed or removed, and the API **contract stays 1.7.0**.
140
+
141
+ **PM evaluation budget.** Event-driven periodic prediction-market evaluations
142
+ now respect `maxLlmCallsPerHour` after their cooldown elapses. Budget skips make
143
+ no provider call and consume no call allowance. PM keeps its own cooldown;
144
+ open-position management and explicit always-on behavior retain their existing
145
+ exemptions. This runner gate is separate from hosted provider-capacity admission.
146
+
147
+ **Retry-After parsing.** Missing, blank or malformed headers no longer become
148
+ zero-delay retries. The runner API client uses its existing five-second fallback;
149
+ explicit zero, numeric seconds and HTTP dates remain supported. Model-provider
150
+ cooldowns share the parser and retain their existing one-hour cap.
151
+
152
+ **API request deadlines.** Each runner API operation now has a 30-second total
153
+ deadline covering response headers, body reads and all 429 retry waits. The
154
+ same client serves the hosted scheduler. Embedded callers can set a finite
155
+ `requestTimeoutMs` and supply an `AbortSignal`. A timeout or cancellation returns
156
+ an uncertain transport result without automatically replaying a trading write.
157
+ Timers and listeners are removed when the operation finishes.
158
+
159
+ **State persistence.** Self-host state is serialized to a private temporary file
160
+ and atomically renamed over the previous state. A failed serialization or rename
161
+ leaves the prior state intact. This is atomic replacement, not a claim of durable
162
+ storage across power loss.
163
+
164
+ **Agent conversion.** `coinrithm-agent eject` preserves explicit `triggerPolicy`
165
+ and `capitalSizing` blocks. Previously conversion could restore default hourly
166
+ budgets and drop equity sizing.
167
+
168
+ **Release verification.** All-source coverage gates, mandatory PostgreSQL CI,
169
+ dependency updates and corrected client setup docs are included. See the
170
+ [reliability record](https://github.com/CoinRithm/coinrithm-agent-trading/blob/main/docs/RELIABILITY.md).
171
+
172
+ **Confirmed-action journal.** Completed-action memory now requires an action
173
+ to be both accepted and executed. Failed writes and uncertain transport results
174
+ retain their attempt evidence without becoming completed moves in the next
175
+ decision prompt. Dry-run proposals remain unexecuted.
176
+
177
+ **Direct NVIDIA retry.** One complete HTTP 500/502/503/504 response can be
178
+ retried once on the identical direct NVIDIA route within the original deadline.
179
+ Both attempts are retained. This does not retry trading writes or change the
180
+ hosted shared-pool routing policy.
181
+
182
+ **Capital reconciliation.** Frozen-balance rounding residue down to -1e-8 is
183
+ normalized only in the sizing calculation after the independent reads agree.
184
+ Negative spendable cash still fails closed; wallet balances are not changed.
185
+
186
+ **Private decision input evidence.** The bounded numeric projection includes
187
+ nested indicator inputs and context movers, with legacy v1 records still readable.
188
+ It does not retain hidden reasoning or raw model output.
189
+
190
+ **Compact prediction-market evidence.** Discovery and compact event-detail
191
+ responses now retain the API's `source.quoteScale`, `source.methodology` and
192
+ `source.supportsMarketMetrics`, plus `spreadPoints`, `probabilityBook` and
193
+ each retained outcome's `normalizedProbability`. Venue-native bid/ask quotes
194
+ are never rescaled or interpreted from magnitude. Normalization remains the
195
+ API's calculation over the original full book, not the truncated top-five
196
+ outcome list. Existing payload bounds and explicit `detail: full` behavior
197
+ are unchanged.
198
+
199
+ **Settlement-time provenance.** Compact events retain `resolvedAtBasis` and
200
+ `settlementWindowClosedAt`, keeping provider expiration distinct from an
201
+ announced settlement time. Null and absent upstream evidence stay null and
202
+ absent; the MCP does not infer missing values.
203
+
204
+ **Candle semantics.** The `get_candles` description now states that these are
205
+ sampled composite-price bars. Each bar's `v` is a mean rolling 24-hour
206
+ quote-volume observation in USD, not volume traded during the candle, and
207
+ must not be summed across bars.
208
+
209
+ **HTTP completion diagnostics.** The hosted HTTP
210
+ entry now has a bounded, stderr-only completion observer with final SDK-result
211
+ and finish/abort accounting. Initialization, discovery, tool failures and
212
+ successful delivery are distinct; unknown tool names are normalized. Records
213
+ contain no arguments, bodies, credentials, caller/RPC IDs or caller-origin labels.
214
+ Credential presence is not authentication. Durations describe the HTTP request,
215
+ shared by batch members; server finish does not prove client receipt or use.
216
+ Stdio, tools, authentication and dependency versions are unchanged. See
217
+ `DEPLOY.md` for the measurement and retention limits. Hosted source/image
218
+ `18a0bb6a8a0665e91cebc10225fec6f7ebcdaaf7` passed a bounded anonymous smoke on
219
+ 2026-09-13. Hosted verification and npm publication are separate release steps.
220
+
221
+ **Deployment boundaries.** Hosted scheduler admission reasons are private
222
+ scheduler telemetry, not a new SDK or MCP response field. The API's corrected
223
+ comparison probabilities and enriched spread names use the existing response
224
+ shape and reach current clients through fresh API reads. Outcome display names
225
+ may change; use source/event/outcome identifiers for identity, never summed
226
+ prices or matching labels alone. These fixes do not establish trading returns.
227
+
228
+ ## 0.7.8
229
+
230
+ Runner decision-quality, evidence and paper-capital release. Additive: no MCP
231
+ tool was renamed or removed, and the API **contract stays 1.7.0**. This release
232
+ contains all package changes since published 0.7.7 (`gitHead` `80d0cae`), not
233
+ just the previously listed thesis work.
234
+
235
+ **Thesis exits.** Every opening action (`futures_open`, `spot_order`,
236
+ `pm_open`) now carries a `thesis`: a one-sentence summary plus an
237
+ `invalidation` with at least one machine-checkable condition (`priceBelow` /
238
+ `priceAbove` for coins, `probabilityBelow` / `probabilityAbove` for prediction
239
+ markets, a `maxHoldMinutes` time stop, and a free-text `catalyst` the model
240
+ re-judges itself). The runner binds the thesis to the position the server
241
+ returns, sanitized side-aware (a rising price never invalidates a long; a
242
+ wrong-side level is dropped rather than re-signed; the time stop is clamped to
243
+ 60 minutes .. 30 days), persists it in the run state (`RunState.theses`, the
244
+ same state file / `agent_state` JSON as before, no schema change) and
245
+ re-evaluates it every cycle. A futures position whose price level or time stop
246
+ is breached is closed by the runner before the model is asked anything, logged
247
+ as a `thesis_invalidated` exit with its own idempotency key, after the
248
+ kill-switch and drawdown checks and never instead of them. Prediction-market
249
+ positions have no close endpoint, so a broken PM thesis is surfaced to the
250
+ model instead (do not add, let it settle). The parser is tolerant (a malformed
251
+ thesis never fails the open; a thesis copied onto a close is ignored) and the
252
+ structured-output schema requires it, so schema-enforced hosted models always
253
+ emit one.
254
+
255
+ **Fundamentals in the observation.** Each watch entry now carries
256
+ `fundamentals` sourced only from calls the runner already makes: `categories`,
257
+ `marketCapRank` and `marketCapUsd` from the market context; `volume24hUsd` from
258
+ the candles the `indicators` capability already fetches (live-probed
259
+ 2026-09-02: each bar's `v` is a rolling 24h volume, so the latest bar is the
260
+ 24h figure, never the sum); and up to three `headlines` with `publishedAt`
261
+ timestamps from the one `news` call, attributed through the curated coin-news
262
+ graph. Discovered PM markets carry `endDate` and `liquidityUsd`; open PM
263
+ positions carry their title, side, entry and current probability and
264
+ `openedAt`; open futures positions carry `openedAt`. The system prompt states
265
+ the thesis contract, the runner-enforced exit and how to grade a trade on the
266
+ fundamentals. Not carried, because no agent endpoint serves them: an "about"
267
+ text per coin, a 24h probability change and a cross-venue divergence per PM
268
+ market.
269
+
270
+ **Fix:** the public movers feed serializes `change24h` / `currentPrice` as
271
+ decimal strings; the universe-scan context rows read them strictly as numbers
272
+ and shipped `undefined` for every mover.
273
+
274
+ **Opt-in equity-based paper sizing.** A runner can size entries from a
275
+ conservative fraction of its independently attributed paper book instead of a
276
+ fixed stake/margin. The book is accepted only when wallet identity, cash
277
+ partitions, held-position attribution and spot-mark coverage reconcile. Quotes
278
+ then enforce per-entry, per-symbol, deployed-capital and daily-entry limits;
279
+ fee buffers and the API's fee-inclusive quote evidence are included. Any
280
+ missing or inconsistent evidence fails closed. Legacy positions on a different
281
+ book remain visible for management but never inflate the current book's buying
282
+ power.
283
+
284
+ **Prediction-market decisions use executable economics.** PM opens now reject
285
+ an invalid raw probability and a model forecast that does not clear the quoted
286
+ entry price. Forecast edge is measured against the actual fee/slippage-adjusted
287
+ fill, not the headline market probability. Quote-expiry outcomes are recorded
288
+ separately from risk/balance rejection, and futures risk/reward validation uses
289
+ fee-inclusive entry and stop economics.
290
+
291
+ **Decision evidence is structured and bounded.** Cycles can expose a sanitized,
292
+ partial private decision-input record: configuration and observation
293
+ fingerprints, daily budget and guard state, plus bounded observation rows with
294
+ explicit omission counts. It is not a prompt, transcript, raw model output or
295
+ hidden reasoning record. The runner also reports quote/validation evidence for
296
+ abstained, forecast-only and quote-expired PM opportunities. Hosted persistence
297
+ and retention remain the caller's responsibility.
298
+
299
+ **Runtime controls are more faithful.** The model sees the remaining daily
300
+ entry/add budget rather than only static maxima. Entry caps still block new
301
+ risk, while closes and other risk-reducing actions remain available. Direct
302
+ provider HTTP 429 responses are capacity skips rather than model failures, so
303
+ BYO agents do not build a failure streak during ordinary quota pressure.
304
+ Structured-tool decisions remain required where the provider supports that
305
+ contract.
306
+
307
+ **Scorecard fix.** Maximum drawdown now measures decline from starting equity,
308
+ so an immediate loss is no longer hidden by treating the first post-trade point
309
+ as the high-water mark.
310
+
311
+ ## 0.7.7
312
+
313
+ Reliability release. Every change here came from a live production failure, not
314
+ from a roadmap. Additive: no tool renamed or removed, and the API **contract
315
+ stays 1.7.0** because nothing on the documented surface changed.
316
+
317
+ **Model requests are now built from a declared capability table, not
318
+ assumptions.** `providerCapabilities.ts` states, per model family, which
319
+ parameter carries the completion budget, whether a non-default temperature is
320
+ allowed, and what extra body fields the family needs. Two failures this fixes:
321
+
322
+ - **OpenAI's current models rejected our requests outright.** `gpt-5*` and
323
+ `o*` refuse `max_tokens` and any non-default `temperature`; they take
324
+ `max_completion_tokens`. The family is detected by MODEL id, not just the
325
+ provider name, so an OpenAI-compatible gateway serving `gpt-5` gets the same
326
+ shape. If you brought your own OpenAI key, this is why it now works.
327
+ - **NVIDIA Nemotron models emitted a think-chain where the JSON decision
328
+ belonged**, which failed every cycle. The `chat_template_kwargs.enable_thinking=false`
329
+ switch and the "detailed thinking off" system hint are now encoded as data
330
+ rather than re-learned by failing.
331
+
332
+ **New: `probeDecisionContract()`.** An HTTP 200 is not proof a route can run an
333
+ agent. Both production failure modes returned 200s: a think-chain in the JSON
334
+ slot, and an empty completion because a reasoning model spent its whole budget
335
+ before answering. The probe sends a canned mini-observation through the REAL
336
+ decision parser at a >=1024 completion allowance and classifies the result as
337
+ `http`, `empty` or `parse`. Use it before adopting any model id; provider
338
+ catalogs list ids that 404 on invoke.
339
+
340
+ **Provider trouble no longer disables an agent.** A permanent-looking model
341
+ error (404/410/decommissioned) used to disable the agent after a threshold. On
342
+ 2026-08-26 NVIDIA end-of-lifed an entire model line and 35 agents died on that
343
+ path. The runner now reports a hold and keeps retrying each cadence, recovering
344
+ by itself when the provider does. Disables remain for what deserves them:
345
+ revoked credentials, drawdown, kill-switch, user action.
346
+
347
+ **Failures carry structured metadata.** A failed `decide()` now returns
348
+ `status` and, when the provider sends one, `retryAfterMs` (parsed from
349
+ `Retry-After` in both delta-seconds and HTTP-date form, capped at an hour), so
350
+ a caller can tell a 429 from a 5xx without parsing strings. Error text is
351
+ unchanged.
352
+
353
+ **`ClientConfig.extraHeaders`.** Headers attached to every request, spread
354
+ before auth so they can never clobber it. Self-host has nothing to put here;
355
+ it exists so CoinRithm's own hosted scheduler can present its attestation
356
+ channel.
357
+
358
+ **Model names corrected throughout.** The retired Llama 3.x line is gone from
359
+ the README, the runtime defaults and the `quant-reference` example, which is
360
+ relocked onto `nvidia/nemotron-3-nano-30b-a3b`.
361
+
362
+ ## 0.7.6
363
+
364
+ Agent capability release: universe discovery, first-class behavioral guards,
365
+ and the hosted prose budget made visible. Additive — no tool renamed or
366
+ removed. Contract moves to **1.7.0** (two keyless paths declared).
367
+
368
+ **New: agents can look beyond their own watchlist.**
369
+
370
+ - **`get_crypto_movers` tool.** Keyless scan of the tracked coin universe for
371
+ the biggest 24h gainers or losers. Rows carry `coinId`, `symbol`, `name`,
372
+ `slug`, `change24hPct`, `priceUsd`.
373
+ - **`universe_scan` capability** for the self-host runner. Each cycle it pulls
374
+ the top movers, promotes the strongest few into full watch entries marked
375
+ `discovered: true`, and passes the remainder as compact context. Watchlist
376
+ and blocklist symbols are excluded up front, so a discovered row can never
377
+ duplicate a configured pair or bypass the deny list.
378
+ - **Both now carry the coinId through.** The movers row's `ucid` IS the
379
+ `coinId` that `get_candles` / `get_market_context` / the futures quote path
380
+ take. It was previously stripped from the tool response and re-derived from
381
+ the SYMBOL via a resolve round-trip — a wasted call per discovered mover and
382
+ a real correctness hazard, because symbols collide across listings and the
383
+ resolver could return a different coin than the one that actually moved.
384
+
385
+ **New: contract declares the endpoints the tools call.**
386
+
387
+ - `/api/coins/top-gainers` and `/api/coins/top-losers` are now in
388
+ `openapi.yaml` (tag `public-crypto-data`), so both SDKs can reach the
389
+ surface `get_crypto_movers` uses. Probe-verified against prod: bare array,
390
+ no envelope; `change24h` / `currentPrice` are decimal STRINGS; default
391
+ `limit` is 3 and out-of-range values return 400 rather than clamping.
392
+
393
+ **New: personality and boundaries are configurable, and documented.**
394
+
395
+ - **`character/guards.md`** — first-class hard behavioral guards, merged into
396
+ the strategy prose as a distinct section rather than buried in the thesis.
397
+ - **`examples/agents/pia-pump-fader`** — a full bundle demonstrating
398
+ capabilities plus boundary configuration (watchlist/blocklist interaction,
399
+ the five-point risk gate, re-entry discipline).
400
+ - **`examples/agents/FORKING.md`** — a file-by-file map of what is strategy
401
+ and what is plumbing, so a fork knows what it is allowed to change.
402
+ - **QUICKSTART** documents capabilities, and a docs-drift tripwire fails the
403
+ suite when a capability ships undocumented (`universe_scan` shipped
404
+ invisible in every user surface once; that cannot recur silently).
405
+
406
+ **Fixed.**
407
+
408
+ - **Hosted prose budget is validated, not discovered at deploy.**
409
+ `coinrithm-agent validate --hosted` now checks the 8,000-character merged
410
+ prose budget and reports the exact overage. A bundle could previously
411
+ validate clean and still be undeployable. YAML frontmatter is stripped
412
+ before the count (and before the model sees it — it was being fed in as if
413
+ it were strategy). `pia-pump-fader` was rebuilt to fit at 7,932.
414
+ - **Permanent failures stop being revived.** A disabled agent whose model is
415
+ gone or whose key is invalid is no longer resurrected by the scheduler's
416
+ revive pass; only transient failures are retried.
417
+ - **Fresh scaffolds are no longer bricked** by the capabilities field, and
418
+ action-confidence tolerance was widened to match what models actually emit.
419
+ - **False market-data licensing assertion corrected** in both READMEs.
420
+
421
+ ⚠ Publishing to npm remains a **manual** step — `publish-mcp.yml` pushes
422
+ `server.json` to the MCP registry only.
423
+
424
+ ## 0.7.5
425
+
426
+ **Release-hygiene bump. Everything below was already merged but never
427
+ reached npm** — the 0.7.4 tarball was published 2026-07-26T23:01:11Z and
428
+ five commits landed after that instant without a version bump, so the
429
+ repository's 0.7.4 and the published 0.7.4 were different code under one
430
+ version number. This release makes the published artifact match the
431
+ source again.
432
+
433
+ - **Security.** MCP dependency audit fixes, 6 findings to 0 (`da65e9e`).
434
+ Anyone on published 0.7.4 is running the pre-audit dependency set.
435
+ - **`pm_data` Gemini exposure** for agents (`3ab04ae`).
436
+ - **Venue methodology and health** exposed as tools (`74cc495`).
437
+ - **Reproducible decision receipts** persisted by the agent runner
438
+ (`8bebc9e`).
439
+ - **Docs.** Contract version drift corrected and the placeholder SDK
440
+ README replaced, so the docs stop advertising an install that 404s
441
+ (`6d61b92`).
442
+
443
+ No tool was renamed or removed; this is additive plus a dependency
444
+ refresh.
445
+
446
+ ⚠ Publishing this package is a **manual** step — `publish-mcp.yml` only
447
+ pushes `server.json` to the MCP registry, it does not run `npm publish`.
448
+ That asymmetry is exactly how the drift above accumulated unnoticed.
449
+
450
+ ## 0.7.4
451
+
452
+ Docs-only. No tool behavior change, no API-surface change.
453
+
454
+ - **Acceptable Use of Market Data.** The README (root and this package) and
455
+ `openapi.yaml` (`info.termsOfService`, `info.description`, and the
456
+ `public-pm-data` tag) now reference and summarize CoinRithm's licensing
457
+ flow-down restriction on Market Data from third-party prediction-market
458
+ venues: read-only use for paper-trading context and settled-outcome
459
+ scoring only — no model training/fine-tuning/benchmarking, no
460
+ redistribution or bulk-extraction, no use to build a competing product.
461
+ Full terms: <https://www.coinrithm.com/en/terms-of-use>
462
+
463
+ ## 0.7.3
464
+
465
+ Quality-engine surfaces + independent forecasts. Additive; no breaking change.
466
+
467
+ - **Quality verdicts in tool responses.** `discover_pm_markets`, `pm_quote`, and
468
+ the `pm_data_*` tools now surface the persisted truth-engine `quality` object
469
+ (`decisionEligible`, warning/block reason codes, `policyVersion`, `assessedAt`).
470
+ Markets with critical failures stay visible but cannot drive paper opens or
471
+ alerts.
472
+ - **`openBlocked` preview on `pm_quote`.** Quotes preview the open-time quality
473
+ gate (`openBlocked` + `openBlockReasons`), so an agent can skip a market that
474
+ would 422 before burning the open attempt. The self-host runner
475
+ (`coinrithm-agent`) does this skip automatically.
476
+ - **Independent forecasts in the runner.** The self-host agent runner elicits the
477
+ model's OWN probability (judged from the question/resolution criteria/deadline,
478
+ never anchored to the market price) and submits it as `forecastProbability` on
479
+ PM opens — feeding the public calibration dataset with proper-scoring-rule
480
+ forecasts. Clamped to [1,99]; omitted (never faked) when the model does not
481
+ produce one; `HOUSE_AGENT_FORECAST_ENABLED=false` disables.
482
+ - **`crossPlatform` on event lists** documented in the API contract: sibling
483
+ venues pricing the same question, on list rows.
484
+ - **ForecastEx venue truth.** Public MCP discovery copy and registry metadata
485
+ now describe all 11 live venues, including ForecastEx.
486
+ - **Contract synchronization.** Runner templates and example bundles pin the
487
+ served OpenAPI 1.6.0 contract; canonical scorecard paths are unambiguous.
488
+
489
+ ## 0.7.2
490
+
491
+ Docs-truth + privacy release. No tool behavior change, no API-surface change.
492
+
493
+ - **Ten venues in the public listing.** `pm_data_*` tool copy, the README, and
494
+ `server.json` now name all ten venues (adds Futuur and Myriad). npm `0.7.1` was
495
+ published before those landed, so the registry listing still advertised "eight
496
+ venues"; npm versions are immutable, so correcting the public listing required
497
+ a new release.
498
+ - **`source` parameter description** on `pm_data_events` / `pm_data_event_detail`
499
+ now enumerates all ten venue slugs. Accepted values are unchanged — this is
500
+ description text only, which is why it is a patch and not a minor.
501
+ - **Privacy.** Raw model output is no longer persisted, enforcing the package's
502
+ no-chain-of-thought promise.
503
+ - **New tripwire.** `server.json` (the MCP-registry listing) is now guarded
504
+ against version and venue-count drift; it had no guard, which is how it went
505
+ stale in the first place.
506
+ - Refreshed stale Arena-gate example copy.
507
+
508
+ ## 0.7.1
509
+
510
+ Docs + registry-metadata release; no tool behavior changes.
511
+
512
+ - **README refresh**: the keyless `pm_data_*` data surface is now front and
513
+ center — 8 venues (Polymarket, Kalshi, Smarkets, Limitless, Manifold,
514
+ Metaculus, PredictIt, Rothera — the "seven venues" line predated Rothera),
515
+ the anonymous hosted-endpoint path, and the `referenceProbability` /
516
+ `volumeHistory` fields the data tools return.
517
+ - **`server.json`**: hosted endpoint's `Authorization` header marked optional
518
+ (the `pm_data_*` tools work anonymously — verified live) and the server
519
+ description now leads with the keyless data surface.
520
+ - Ships the post-0.7.0 commits: `pm_data_event` advertises `volumeHistory`,
521
+ `pm_data_events` advertises `referenceProbability` on list items, and the
522
+ hosted MCP root (`GET /`) serves a self-describing JSON landing (with a
523
+ 405 + hint on `GET /mcp`).
524
+
525
+ ## 0.7.0
526
+
527
+ (Retroactive entry — released 2026-07-05 without a changelog note.)
528
+
529
+ - **Four keyless `pm_data_*` tools** — CoinRithm's free public cross-venue
530
+ prediction-market dataset over MCP, no API key required and yours is never
531
+ attached: `pm_data_overview` (market-wide stats), `pm_data_events`
532
+ (cross-venue event list), `pm_data_event` (detail incl.
533
+ `crossSourceMatches` + resolution evidence), `pm_data_whales`
534
+ (large-trade tape).
535
+
536
+ ## 0.5.0
537
+
538
+ Agent-runner quality + reliability release. `coinrithm-agent` got materially
539
+ smarter and less noisy; `coinrithm-mcp` is unchanged in shape. Bundles the work
540
+ since 0.4.0.
541
+
542
+ - **Prediction markets are a first-class venue** in the decide prompt: a short
543
+ `pmN` ref so small models trade PM reliably, eligible-outcome filtering (only
544
+ backend-openable outcomes reach the model), and futures-capped agents steered
545
+ to PM (a separate budget) instead of re-rejecting.
546
+ - **PM anti-churn — now actually effective.** The candidate list is pre-filtered
547
+ to exclude markets the agent already holds, and the runner + server both block
548
+ re-betting a held market+outcome (no more one-agent, 25-identical-bets churn).
549
+ An earlier version read the wrong `/positions/pm` fields and was silently dead;
550
+ fixed.
551
+ - **Settlement-feedback learning loop.** The agent sees how its own recent bets
552
+ actually resolved (win/loss/void + realized PnL) as reflective context, so it
553
+ learns from outcomes across cycles.
554
+ - **Per-trade reasoning stays honest about the market.** A multi-action decision
555
+ no longer stamps its primary rationale onto a secondary trade about a different
556
+ market — the trade's public Arena "why" always matches the market it's on.
557
+ - **Futures reliability.** The model is unblinded to per-position mark /
558
+ liquidation / stop / take-profit prices; take-profit is auto-clamped to a valid
559
+ R:R target off the stop (kills the `take_profit_not_*_mark` reject waves); and a
560
+ marking-down PM book now trips the equity-drawdown kill-switch too.
561
+ - **News capability.** Recent high-importance news for the watchlist coins is fed
562
+ into the decide context as a market-catalyst layer.
563
+ - **Robustness + contract accuracy.** Scheduler/runner hardening, flat-state
564
+ prompt steers (weak models stop hallucinating closes), manage-enum
565
+ normalization, and the PM contract now documents real entry friction rather
566
+ than a disclose-only stance.
567
+ - **Security.** `hono` bumped to 4.12.27 (high-severity advisories: serve-static
568
+ path traversal, CORS wildcard-with-credentials, body-limit bypass).
569
+ - **Docs.** The npm README leads with value / free / OKF / Studio; stale
570
+ scheduler and `minDecidedTrades` claims corrected.
571
+
572
+ ## 0.4.0
573
+
574
+ - **Deterministic scorecard engine (`computeScorecard`).** The reproducible-
575
+ evaluation engine for `coinrithm.agent.scorecard.v1` — pure math over an
576
+ agent's realized track record (no network, no model): realized PnL, win rate,
577
+ expectancy, profit factor, reward-to-risk, Sharpe, Sortino, deflated /
578
+ probabilistic Sharpe (Bailey & López de Prado — skill vs luck with a multiple-
579
+ testing penalty), max drawdown, and Brier + ECE calibration for probabilistic
580
+ calls. Same inputs → identical metrics **and** a sha256 `contentHash` of the
581
+ canonicalized result, so a scorecard whose hash doesn't reproduce isn't
582
+ trusted. Metrics are computed AFTER the run from immutable evidence (leakage-
583
+ separation), so tuning-to-the-metric is structurally impossible. Returns
584
+ `null` for thin records — never a fabricated number.
585
+ - **Resolver: committable file metadata + functionality pin.** The OKF resolver
586
+ now carries per-file metadata and pins functionality through resolution, so a
587
+ bundle's behavior is reproducible from its committed files.
588
+
589
+ ## 0.3.0
590
+
591
+ - **Agent risk config: coin deny-list (`blocklist`).** `risk.blocklist` lets an
592
+ agent name symbols it must never open, even if they are on the watchlist —
593
+ deny wins over allow. Enforced in the runner's decision validator (rejects
594
+ `futures_open` / `spot_order` on a denied symbol) and surfaced in the system
595
+ prompt, so the model is told the boundary and the runner re-checks it.
596
+ - **Docs: Open Knowledge Format positioning.** Clarified that a CoinRithm agent
597
+ is an OKF bundle — a portable directory of markdown + frontmatter that is
598
+ model-agnostic (run the same definition on any model). Develop and prove it
599
+ free on paper, then run it anywhere.
600
+ - npm keywords refreshed (`open-knowledge-format`, `okf`, `model-agnostic`,
601
+ `gemini`) for registry discovery.
602
+
603
+ ## 0.2.0
604
+
605
+ - Added the **`coinrithm-agent`** self-host runner binary alongside the MCP
606
+ server: a folder-as-architecture (OKF) agent you bring your own model key to,
607
+ with caps enforced by the runner (not the model), dry-run by default,
608
+ paper-only.
609
+
610
+ ## 0.1.x
611
+
612
+ - Initial `coinrithm-mcp` MCP server: reads, quotes, scoped spot/futures/PM
613
+ writes, ledger export, and Agent Arena integration over a user-minted API key.