@coinrithm/mcp-trading 0.7.12 → 0.7.13

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,546 +1,573 @@
1
- # Changelog
2
-
3
- All notable changes to `@coinrithm/mcp-trading` are documented here. The package
4
- ships two binaries — `coinrithm-mcp` (the MCP server) and `coinrithm-agent` (the
5
- self-host agent runner) — versioned together. The CoinRithm **API contract** is
6
- versioned separately (see `openapi.yaml` `info.version`, currently `1.7.0`).
7
-
8
- ## 0.7.12 — 2026-09-15
9
-
10
- Clarify `whoami`, `cancel_spot_order` and `report_pm_opportunity` descriptions,
11
- removing execution-cost prose unrelated to these operations. Document actual
12
- authentication, side effects, result fields and retry behavior.
13
-
14
- Spot cancellation now advertises its existing idempotent behavior. Opportunity
15
- reporting no longer advertises unconditional idempotency: duplicate protection
16
- requires `decisionId` (or `agentTrace.decisionId`) under the same API key; the
17
- first stored record wins. Reporting remains a write despite requiring only the
18
- `read` scope, and its evidence remains explicitly self-reported.
19
-
20
- Tool names, accepted inputs and execution behavior are unchanged. This source
21
- entry does not establish registry publication, deployment or a new Glama score.
22
-
23
- ## 0.7.11 — 2026-09-15
24
-
25
- Fix opportunity reporting that previously treated resolved API failures as
26
- successful submissions. Only `ok: true` confirms a report. HTTP errors remain
27
- unconfirmed; transport failures and exceptions have an unknown delivery outcome.
28
-
29
- `CycleResult.opportunity` now contains confirmed reports only. The additive
30
- `opportunityReport` field preserves the attempted payload, outcome and status,
31
- without API error bodies or exception details. Existing consumers should use
32
- this field when they need attempted rather than confirmed evidence.
33
-
34
- The reporter retains its one-invocation-per-cycle latch and adds no retries.
35
- Focused regressions compare successful and failed reporting in skip/act cycles
36
- and verify unchanged trading results and runner state. SDK versions are unchanged.
37
- This source entry does not establish registry publication or hosted deployment.
38
-
39
- ## 0.7.10 — 2026-09-15
40
-
41
- Source changes following the published 0.7.9 release. The TypeScript SDK remains
42
- 0.3.1 and Python remains 1.8.1; their runtime source is unchanged.
43
-
44
- - Persist file-backed run identity before execution so a first-cycle process
45
- crash cannot discard the idempotency identity. Transport uncertainty is still
46
- not automatically replayed or recorded as a completed trade.
47
- - Fix runner startup on Node 18: use the imported Node crypto API instead of
48
- depending on a global crypto object. The installed-package matrix reproduced
49
- this failure on Linux, Windows and macOS.
50
- - Add the supported `@coinrithm/mcp-trading/engine` entry point, preserving
51
- existing deep imports, and separate observation accounting and opportunity
52
- reporting from cycle ordering.
53
- - Add opt-in, machine-checked crypto return predicates and visible warnings for
54
- inactive/reserved configuration. Existing agents are not automatically opted in.
55
- - Serialize scheduler migrations in a bounded transaction; add an offline
56
- credential-rotation helper and interruption/recovery rehearsal.
57
- - Pin workflow actions and verify release-tool checksums. Add installed-package
58
- compatibility and restart smoke checks across operating systems and runtimes.
59
-
60
- 0.7.10 was published and deployed on 2026-09-15. Registry and GitHub downloads
61
- matched the CI-tested archive; hosted MCP and scheduler deployments finished.
62
- See the [release record](https://github.com/CoinRithm/coinrithm-agent-trading/releases/tag/mcp-trading-v0.7.10).
63
-
64
- ## 0.7.9 - 2026-09-15
65
-
66
- Published on npm and verified on 2026-09-15: the registry archive matches the
67
- reviewed artifact and passes fresh-install checks. See the
68
- [combined release notes](https://github.com/CoinRithm/coinrithm-agent-trading/releases/tag/mcp-trading-v0.7.9)
69
- for source provenance, verification and community acknowledgments.
70
-
71
- Public market-data fidelity and runner reliability release. No MCP tool was
72
- renamed or removed, and the API **contract stays 1.7.0**.
73
-
74
- **PM evaluation budget.** Event-driven periodic prediction-market evaluations
75
- now respect `maxLlmCallsPerHour` after their cooldown elapses. Budget skips make
76
- no provider call and consume no call allowance. PM keeps its own cooldown;
77
- open-position management and explicit always-on behavior retain their existing
78
- exemptions. This runner gate is separate from hosted provider-capacity admission.
79
-
80
- **Retry-After parsing.** Missing, blank or malformed headers no longer become
81
- zero-delay retries. The runner API client uses its existing five-second fallback;
82
- explicit zero, numeric seconds and HTTP dates remain supported. Model-provider
83
- cooldowns share the parser and retain their existing one-hour cap.
84
-
85
- **API request deadlines.** Each runner API operation now has a 30-second total
86
- deadline covering response headers, body reads and all 429 retry waits. The
87
- same client serves the hosted scheduler. Embedded callers can set a finite
88
- `requestTimeoutMs` and supply an `AbortSignal`. A timeout or cancellation returns
89
- an uncertain transport result without automatically replaying a trading write.
90
- Timers and listeners are removed when the operation finishes.
91
-
92
- **State persistence.** Self-host state is serialized to a private temporary file
93
- and atomically renamed over the previous state. A failed serialization or rename
94
- leaves the prior state intact. This is atomic replacement, not a claim of durable
95
- storage across power loss.
96
-
97
- **Agent conversion.** `coinrithm-agent eject` preserves explicit `triggerPolicy`
98
- and `capitalSizing` blocks. Previously conversion could restore default hourly
99
- budgets and drop equity sizing.
100
-
101
- **Release verification.** All-source coverage gates, mandatory PostgreSQL CI,
102
- dependency updates and corrected client setup docs are included. See the
103
- [reliability record](https://github.com/CoinRithm/coinrithm-agent-trading/blob/main/docs/RELIABILITY.md).
104
-
105
- **Confirmed-action journal.** Completed-action memory now requires an action
106
- to be both accepted and executed. Failed writes and uncertain transport results
107
- retain their attempt evidence without becoming completed moves in the next
108
- decision prompt. Dry-run proposals remain unexecuted.
109
-
110
- **Direct NVIDIA retry.** One complete HTTP 500/502/503/504 response can be
111
- retried once on the identical direct NVIDIA route within the original deadline.
112
- Both attempts are retained. This does not retry trading writes or change the
113
- hosted shared-pool routing policy.
114
-
115
- **Capital reconciliation.** Frozen-balance rounding residue down to -1e-8 is
116
- normalized only in the sizing calculation after the independent reads agree.
117
- Negative spendable cash still fails closed; wallet balances are not changed.
118
-
119
- **Private decision input evidence.** The bounded numeric projection includes
120
- nested indicator inputs and context movers, with legacy v1 records still readable.
121
- It does not retain hidden reasoning or raw model output.
122
-
123
- **Compact prediction-market evidence.** Discovery and compact event-detail
124
- responses now retain the API's `source.quoteScale`, `source.methodology` and
125
- `source.supportsMarketMetrics`, plus `spreadPoints`, `probabilityBook` and
126
- each retained outcome's `normalizedProbability`. Venue-native bid/ask quotes
127
- are never rescaled or interpreted from magnitude. Normalization remains the
128
- API's calculation over the original full book, not the truncated top-five
129
- outcome list. Existing payload bounds and explicit `detail: full` behavior
130
- are unchanged.
131
-
132
- **Settlement-time provenance.** Compact events retain `resolvedAtBasis` and
133
- `settlementWindowClosedAt`, keeping provider expiration distinct from an
134
- announced settlement time. Null and absent upstream evidence stay null and
135
- absent; the MCP does not infer missing values.
136
-
137
- **Candle semantics.** The `get_candles` description now states that these are
138
- sampled composite-price bars. Each bar's `v` is a mean rolling 24-hour
139
- quote-volume observation in USD, not volume traded during the candle, and
140
- must not be summed across bars.
141
-
142
- **HTTP completion diagnostics.** The hosted HTTP
143
- entry now has a bounded, stderr-only completion observer with final SDK-result
144
- and finish/abort accounting. Initialization, discovery, tool failures and
145
- successful delivery are distinct; unknown tool names are normalized. Records
146
- contain no arguments, bodies, credentials, caller/RPC IDs or caller-origin labels.
147
- Credential presence is not authentication. Durations describe the HTTP request,
148
- shared by batch members; server finish does not prove client receipt or use.
149
- Stdio, tools, authentication and dependency versions are unchanged. See
150
- `DEPLOY.md` for the measurement and retention limits. Hosted source/image
151
- `18a0bb6a8a0665e91cebc10225fec6f7ebcdaaf7` passed a bounded anonymous smoke on
152
- 2026-09-13. Hosted verification and npm publication are separate release steps.
153
-
154
- **Deployment boundaries.** Hosted scheduler admission reasons are private
155
- scheduler telemetry, not a new SDK or MCP response field. The API's corrected
156
- comparison probabilities and enriched spread names use the existing response
157
- shape and reach current clients through fresh API reads. Outcome display names
158
- may change; use source/event/outcome identifiers for identity, never summed
159
- prices or matching labels alone. These fixes do not establish trading returns.
160
-
161
- ## 0.7.8
162
-
163
- Runner decision-quality, evidence and paper-capital release. Additive: no MCP
164
- tool was renamed or removed, and the API **contract stays 1.7.0**. This release
165
- contains all package changes since published 0.7.7 (`gitHead` `80d0cae`), not
166
- just the previously listed thesis work.
167
-
168
- **Thesis exits.** Every opening action (`futures_open`, `spot_order`,
169
- `pm_open`) now carries a `thesis`: a one-sentence summary plus an
170
- `invalidation` with at least one machine-checkable condition (`priceBelow` /
171
- `priceAbove` for coins, `probabilityBelow` / `probabilityAbove` for prediction
172
- markets, a `maxHoldMinutes` time stop, and a free-text `catalyst` the model
173
- re-judges itself). The runner binds the thesis to the position the server
174
- returns, sanitized side-aware (a rising price never invalidates a long; a
175
- wrong-side level is dropped rather than re-signed; the time stop is clamped to
176
- 60 minutes .. 30 days), persists it in the run state (`RunState.theses`, the
177
- same state file / `agent_state` JSON as before, no schema change) and
178
- re-evaluates it every cycle. A futures position whose price level or time stop
179
- is breached is closed by the runner before the model is asked anything, logged
180
- as a `thesis_invalidated` exit with its own idempotency key, after the
181
- kill-switch and drawdown checks and never instead of them. Prediction-market
182
- positions have no close endpoint, so a broken PM thesis is surfaced to the
183
- model instead (do not add, let it settle). The parser is tolerant (a malformed
184
- thesis never fails the open; a thesis copied onto a close is ignored) and the
185
- structured-output schema requires it, so schema-enforced hosted models always
186
- emit one.
187
-
188
- **Fundamentals in the observation.** Each watch entry now carries
189
- `fundamentals` sourced only from calls the runner already makes: `categories`,
190
- `marketCapRank` and `marketCapUsd` from the market context; `volume24hUsd` from
191
- the candles the `indicators` capability already fetches (live-probed
192
- 2026-09-02: each bar's `v` is a rolling 24h volume, so the latest bar is the
193
- 24h figure, never the sum); and up to three `headlines` with `publishedAt`
194
- timestamps from the one `news` call, attributed through the curated coin-news
195
- graph. Discovered PM markets carry `endDate` and `liquidityUsd`; open PM
196
- positions carry their title, side, entry and current probability and
197
- `openedAt`; open futures positions carry `openedAt`. The system prompt states
198
- the thesis contract, the runner-enforced exit and how to grade a trade on the
199
- fundamentals. Not carried, because no agent endpoint serves them: an "about"
200
- text per coin, a 24h probability change and a cross-venue divergence per PM
201
- market.
202
-
203
- **Fix:** the public movers feed serializes `change24h` / `currentPrice` as
204
- decimal strings; the universe-scan context rows read them strictly as numbers
205
- and shipped `undefined` for every mover.
206
-
207
- **Opt-in equity-based paper sizing.** A runner can size entries from a
208
- conservative fraction of its independently attributed paper book instead of a
209
- fixed stake/margin. The book is accepted only when wallet identity, cash
210
- partitions, held-position attribution and spot-mark coverage reconcile. Quotes
211
- then enforce per-entry, per-symbol, deployed-capital and daily-entry limits;
212
- fee buffers and the API's fee-inclusive quote evidence are included. Any
213
- missing or inconsistent evidence fails closed. Legacy positions on a different
214
- book remain visible for management but never inflate the current book's buying
215
- power.
216
-
217
- **Prediction-market decisions use executable economics.** PM opens now reject
218
- an invalid raw probability and a model forecast that does not clear the quoted
219
- entry price. Forecast edge is measured against the actual fee/slippage-adjusted
220
- fill, not the headline market probability. Quote-expiry outcomes are recorded
221
- separately from risk/balance rejection, and futures risk/reward validation uses
222
- fee-inclusive entry and stop economics.
223
-
224
- **Decision evidence is structured and bounded.** Cycles can expose a sanitized,
225
- partial private decision-input record: configuration and observation
226
- fingerprints, daily budget and guard state, plus bounded observation rows with
227
- explicit omission counts. It is not a prompt, transcript, raw model output or
228
- hidden reasoning record. The runner also reports quote/validation evidence for
229
- abstained, forecast-only and quote-expired PM opportunities. Hosted persistence
230
- and retention remain the caller's responsibility.
231
-
232
- **Runtime controls are more faithful.** The model sees the remaining daily
233
- entry/add budget rather than only static maxima. Entry caps still block new
234
- risk, while closes and other risk-reducing actions remain available. Direct
235
- provider HTTP 429 responses are capacity skips rather than model failures, so
236
- BYO agents do not build a failure streak during ordinary quota pressure.
237
- Structured-tool decisions remain required where the provider supports that
238
- contract.
239
-
240
- **Scorecard fix.** Maximum drawdown now measures decline from starting equity,
241
- so an immediate loss is no longer hidden by treating the first post-trade point
242
- as the high-water mark.
243
-
244
- ## 0.7.7
245
-
246
- Reliability release. Every change here came from a live production failure, not
247
- from a roadmap. Additive: no tool renamed or removed, and the API **contract
248
- stays 1.7.0** because nothing on the documented surface changed.
249
-
250
- **Model requests are now built from a declared capability table, not
251
- assumptions.** `providerCapabilities.ts` states, per model family, which
252
- parameter carries the completion budget, whether a non-default temperature is
253
- allowed, and what extra body fields the family needs. Two failures this fixes:
254
-
255
- - **OpenAI's current models rejected our requests outright.** `gpt-5*` and
256
- `o*` refuse `max_tokens` and any non-default `temperature`; they take
257
- `max_completion_tokens`. The family is detected by MODEL id, not just the
258
- provider name, so an OpenAI-compatible gateway serving `gpt-5` gets the same
259
- shape. If you brought your own OpenAI key, this is why it now works.
260
- - **NVIDIA Nemotron models emitted a think-chain where the JSON decision
261
- belonged**, which failed every cycle. The `chat_template_kwargs.enable_thinking=false`
262
- switch and the "detailed thinking off" system hint are now encoded as data
263
- rather than re-learned by failing.
264
-
265
- **New: `probeDecisionContract()`.** An HTTP 200 is not proof a route can run an
266
- agent. Both production failure modes returned 200s: a think-chain in the JSON
267
- slot, and an empty completion because a reasoning model spent its whole budget
268
- before answering. The probe sends a canned mini-observation through the REAL
269
- decision parser at a >=1024 completion allowance and classifies the result as
270
- `http`, `empty` or `parse`. Use it before adopting any model id; provider
271
- catalogs list ids that 404 on invoke.
272
-
273
- **Provider trouble no longer disables an agent.** A permanent-looking model
274
- error (404/410/decommissioned) used to disable the agent after a threshold. On
275
- 2026-08-26 NVIDIA end-of-lifed an entire model line and 35 agents died on that
276
- path. The runner now reports a hold and keeps retrying each cadence, recovering
277
- by itself when the provider does. Disables remain for what deserves them:
278
- revoked credentials, drawdown, kill-switch, user action.
279
-
280
- **Failures carry structured metadata.** A failed `decide()` now returns
281
- `status` and, when the provider sends one, `retryAfterMs` (parsed from
282
- `Retry-After` in both delta-seconds and HTTP-date form, capped at an hour), so
283
- a caller can tell a 429 from a 5xx without parsing strings. Error text is
284
- unchanged.
285
-
286
- **`ClientConfig.extraHeaders`.** Headers attached to every request, spread
287
- before auth so they can never clobber it. Self-host has nothing to put here;
288
- it exists so CoinRithm's own hosted scheduler can present its attestation
289
- channel.
290
-
291
- **Model names corrected throughout.** The retired Llama 3.x line is gone from
292
- the README, the runtime defaults and the `quant-reference` example, which is
293
- relocked onto `nvidia/nemotron-3-nano-30b-a3b`.
294
-
295
- ## 0.7.6
296
-
297
- Agent capability release: universe discovery, first-class behavioral guards,
298
- and the hosted prose budget made visible. Additive — no tool renamed or
299
- removed. Contract moves to **1.7.0** (two keyless paths declared).
300
-
301
- **New: agents can look beyond their own watchlist.**
302
-
303
- - **`get_crypto_movers` tool.** Keyless scan of the tracked coin universe for
304
- the biggest 24h gainers or losers. Rows carry `coinId`, `symbol`, `name`,
305
- `slug`, `change24hPct`, `priceUsd`.
306
- - **`universe_scan` capability** for the self-host runner. Each cycle it pulls
307
- the top movers, promotes the strongest few into full watch entries marked
308
- `discovered: true`, and passes the remainder as compact context. Watchlist
309
- and blocklist symbols are excluded up front, so a discovered row can never
310
- duplicate a configured pair or bypass the deny list.
311
- - **Both now carry the coinId through.** The movers row's `ucid` IS the
312
- `coinId` that `get_candles` / `get_market_context` / the futures quote path
313
- take. It was previously stripped from the tool response and re-derived from
314
- the SYMBOL via a resolve round-trip — a wasted call per discovered mover and
315
- a real correctness hazard, because symbols collide across listings and the
316
- resolver could return a different coin than the one that actually moved.
317
-
318
- **New: contract declares the endpoints the tools call.**
319
-
320
- - `/api/coins/top-gainers` and `/api/coins/top-losers` are now in
321
- `openapi.yaml` (tag `public-crypto-data`), so both SDKs can reach the
322
- surface `get_crypto_movers` uses. Probe-verified against prod: bare array,
323
- no envelope; `change24h` / `currentPrice` are decimal STRINGS; default
324
- `limit` is 3 and out-of-range values return 400 rather than clamping.
325
-
326
- **New: personality and boundaries are configurable, and documented.**
327
-
328
- - **`character/guards.md`** — first-class hard behavioral guards, merged into
329
- the strategy prose as a distinct section rather than buried in the thesis.
330
- - **`examples/agents/pia-pump-fader`** — a full bundle demonstrating
331
- capabilities plus boundary configuration (watchlist/blocklist interaction,
332
- the five-point risk gate, re-entry discipline).
333
- - **`examples/agents/FORKING.md`** — a file-by-file map of what is strategy
334
- and what is plumbing, so a fork knows what it is allowed to change.
335
- - **QUICKSTART** documents capabilities, and a docs-drift tripwire fails the
336
- suite when a capability ships undocumented (`universe_scan` shipped
337
- invisible in every user surface once; that cannot recur silently).
338
-
339
- **Fixed.**
340
-
341
- - **Hosted prose budget is validated, not discovered at deploy.**
342
- `coinrithm-agent validate --hosted` now checks the 8,000-character merged
343
- prose budget and reports the exact overage. A bundle could previously
344
- validate clean and still be undeployable. YAML frontmatter is stripped
345
- before the count (and before the model sees it — it was being fed in as if
346
- it were strategy). `pia-pump-fader` was rebuilt to fit at 7,932.
347
- - **Permanent failures stop being revived.** A disabled agent whose model is
348
- gone or whose key is invalid is no longer resurrected by the scheduler's
349
- revive pass; only transient failures are retried.
350
- - **Fresh scaffolds are no longer bricked** by the capabilities field, and
351
- action-confidence tolerance was widened to match what models actually emit.
352
- - **False market-data licensing assertion corrected** in both READMEs.
353
-
354
- ⚠ Publishing to npm remains a **manual** step — `publish-mcp.yml` pushes
355
- `server.json` to the MCP registry only.
356
-
357
- ## 0.7.5
358
-
359
- **Release-hygiene bump. Everything below was already merged but never
360
- reached npm** — the 0.7.4 tarball was published 2026-07-26T23:01:11Z and
361
- five commits landed after that instant without a version bump, so the
362
- repository's 0.7.4 and the published 0.7.4 were different code under one
363
- version number. This release makes the published artifact match the
364
- source again.
365
-
366
- - **Security.** MCP dependency audit fixes, 6 findings to 0 (`da65e9e`).
367
- Anyone on published 0.7.4 is running the pre-audit dependency set.
368
- - **`pm_data` Gemini exposure** for agents (`3ab04ae`).
369
- - **Venue methodology and health** exposed as tools (`74cc495`).
370
- - **Reproducible decision receipts** persisted by the agent runner
371
- (`8bebc9e`).
372
- - **Docs.** Contract version drift corrected and the placeholder SDK
373
- README replaced, so the docs stop advertising an install that 404s
374
- (`6d61b92`).
375
-
376
- No tool was renamed or removed; this is additive plus a dependency
377
- refresh.
378
-
379
- ⚠ Publishing this package is a **manual** step — `publish-mcp.yml` only
380
- pushes `server.json` to the MCP registry, it does not run `npm publish`.
381
- That asymmetry is exactly how the drift above accumulated unnoticed.
382
-
383
- ## 0.7.4
384
-
385
- Docs-only. No tool behavior change, no API-surface change.
386
-
387
- - **Acceptable Use of Market Data.** The README (root and this package) and
388
- `openapi.yaml` (`info.termsOfService`, `info.description`, and the
389
- `public-pm-data` tag) now reference and summarize CoinRithm's licensing
390
- flow-down restriction on Market Data from third-party prediction-market
391
- venues: read-only use for paper-trading context and settled-outcome
392
- scoring only — no model training/fine-tuning/benchmarking, no
393
- redistribution or bulk-extraction, no use to build a competing product.
394
- Full terms: <https://www.coinrithm.com/en/terms-of-use>
395
-
396
- ## 0.7.3
397
-
398
- Quality-engine surfaces + independent forecasts. Additive; no breaking change.
399
-
400
- - **Quality verdicts in tool responses.** `discover_pm_markets`, `pm_quote`, and
401
- the `pm_data_*` tools now surface the persisted truth-engine `quality` object
402
- (`decisionEligible`, warning/block reason codes, `policyVersion`, `assessedAt`).
403
- Markets with critical failures stay visible but cannot drive paper opens or
404
- alerts.
405
- - **`openBlocked` preview on `pm_quote`.** Quotes preview the open-time quality
406
- gate (`openBlocked` + `openBlockReasons`), so an agent can skip a market that
407
- would 422 before burning the open attempt. The self-host runner
408
- (`coinrithm-agent`) does this skip automatically.
409
- - **Independent forecasts in the runner.** The self-host agent runner elicits the
410
- model's OWN probability (judged from the question/resolution criteria/deadline,
411
- never anchored to the market price) and submits it as `forecastProbability` on
412
- PM opens — feeding the public calibration dataset with proper-scoring-rule
413
- forecasts. Clamped to [1,99]; omitted (never faked) when the model does not
414
- produce one; `HOUSE_AGENT_FORECAST_ENABLED=false` disables.
415
- - **`crossPlatform` on event lists** documented in the API contract: sibling
416
- venues pricing the same question, on list rows.
417
- - **ForecastEx venue truth.** Public MCP discovery copy and registry metadata
418
- now describe all 11 live venues, including ForecastEx.
419
- - **Contract synchronization.** Runner templates and example bundles pin the
420
- served OpenAPI 1.6.0 contract; canonical scorecard paths are unambiguous.
421
-
422
- ## 0.7.2
423
-
424
- Docs-truth + privacy release. No tool behavior change, no API-surface change.
425
-
426
- - **Ten venues in the public listing.** `pm_data_*` tool copy, the README, and
427
- `server.json` now name all ten venues (adds Futuur and Myriad). npm `0.7.1` was
428
- published before those landed, so the registry listing still advertised "eight
429
- venues"; npm versions are immutable, so correcting the public listing required
430
- a new release.
431
- - **`source` parameter description** on `pm_data_events` / `pm_data_event_detail`
432
- now enumerates all ten venue slugs. Accepted values are unchanged — this is
433
- description text only, which is why it is a patch and not a minor.
434
- - **Privacy.** Raw model output is no longer persisted, enforcing the package's
435
- no-chain-of-thought promise.
436
- - **New tripwire.** `server.json` (the MCP-registry listing) is now guarded
437
- against version and venue-count drift; it had no guard, which is how it went
438
- stale in the first place.
439
- - Refreshed stale Arena-gate example copy.
440
-
441
- ## 0.7.1
442
-
443
- Docs + registry-metadata release; no tool behavior changes.
444
-
445
- - **README refresh**: the keyless `pm_data_*` data surface is now front and
446
- center — 8 venues (Polymarket, Kalshi, Smarkets, Limitless, Manifold,
447
- Metaculus, PredictIt, Rothera — the "seven venues" line predated Rothera),
448
- the anonymous hosted-endpoint path, and the `referenceProbability` /
449
- `volumeHistory` fields the data tools return.
450
- - **`server.json`**: hosted endpoint's `Authorization` header marked optional
451
- (the `pm_data_*` tools work anonymously — verified live) and the server
452
- description now leads with the keyless data surface.
453
- - Ships the post-0.7.0 commits: `pm_data_event` advertises `volumeHistory`,
454
- `pm_data_events` advertises `referenceProbability` on list items, and the
455
- hosted MCP root (`GET /`) serves a self-describing JSON landing (with a
456
- 405 + hint on `GET /mcp`).
457
-
458
- ## 0.7.0
459
-
460
- (Retroactive entry — released 2026-07-05 without a changelog note.)
461
-
462
- - **Four keyless `pm_data_*` tools** — CoinRithm's free public cross-venue
463
- prediction-market dataset over MCP, no API key required and yours is never
464
- attached: `pm_data_overview` (market-wide stats), `pm_data_events`
465
- (cross-venue event list), `pm_data_event` (detail incl.
466
- `crossSourceMatches` + resolution evidence), `pm_data_whales`
467
- (large-trade tape).
468
-
469
- ## 0.5.0
470
-
471
- Agent-runner quality + reliability release. `coinrithm-agent` got materially
472
- smarter and less noisy; `coinrithm-mcp` is unchanged in shape. Bundles the work
473
- since 0.4.0.
474
-
475
- - **Prediction markets are a first-class venue** in the decide prompt: a short
476
- `pmN` ref so small models trade PM reliably, eligible-outcome filtering (only
477
- backend-openable outcomes reach the model), and futures-capped agents steered
478
- to PM (a separate budget) instead of re-rejecting.
479
- - **PM anti-churn — now actually effective.** The candidate list is pre-filtered
480
- to exclude markets the agent already holds, and the runner + server both block
481
- re-betting a held market+outcome (no more one-agent, 25-identical-bets churn).
482
- An earlier version read the wrong `/positions/pm` fields and was silently dead;
483
- fixed.
484
- - **Settlement-feedback learning loop.** The agent sees how its own recent bets
485
- actually resolved (win/loss/void + realized PnL) as reflective context, so it
486
- learns from outcomes across cycles.
487
- - **Per-trade reasoning stays honest about the market.** A multi-action decision
488
- no longer stamps its primary rationale onto a secondary trade about a different
489
- market — the trade's public Arena "why" always matches the market it's on.
490
- - **Futures reliability.** The model is unblinded to per-position mark /
491
- liquidation / stop / take-profit prices; take-profit is auto-clamped to a valid
492
- R:R target off the stop (kills the `take_profit_not_*_mark` reject waves); and a
493
- marking-down PM book now trips the equity-drawdown kill-switch too.
494
- - **News capability.** Recent high-importance news for the watchlist coins is fed
495
- into the decide context as a market-catalyst layer.
496
- - **Robustness + contract accuracy.** Scheduler/runner hardening, flat-state
497
- prompt steers (weak models stop hallucinating closes), manage-enum
498
- normalization, and the PM contract now documents real entry friction rather
499
- than a disclose-only stance.
500
- - **Security.** `hono` bumped to 4.12.27 (high-severity advisories: serve-static
501
- path traversal, CORS wildcard-with-credentials, body-limit bypass).
502
- - **Docs.** The npm README leads with value / free / OKF / Studio; stale
503
- scheduler and `minDecidedTrades` claims corrected.
504
-
505
- ## 0.4.0
506
-
507
- - **Deterministic scorecard engine (`computeScorecard`).** The reproducible-
508
- evaluation engine for `coinrithm.agent.scorecard.v1` — pure math over an
509
- agent's realized track record (no network, no model): realized PnL, win rate,
510
- expectancy, profit factor, reward-to-risk, Sharpe, Sortino, deflated /
511
- probabilistic Sharpe (Bailey & López de Prado — skill vs luck with a multiple-
512
- testing penalty), max drawdown, and Brier + ECE calibration for probabilistic
513
- calls. Same inputs → identical metrics **and** a sha256 `contentHash` of the
514
- canonicalized result, so a scorecard whose hash doesn't reproduce isn't
515
- trusted. Metrics are computed AFTER the run from immutable evidence (leakage-
516
- separation), so tuning-to-the-metric is structurally impossible. Returns
517
- `null` for thin records — never a fabricated number.
518
- - **Resolver: committable file metadata + functionality pin.** The OKF resolver
519
- now carries per-file metadata and pins functionality through resolution, so a
520
- bundle's behavior is reproducible from its committed files.
521
-
522
- ## 0.3.0
523
-
524
- - **Agent risk config: coin deny-list (`blocklist`).** `risk.blocklist` lets an
525
- agent name symbols it must never open, even if they are on the watchlist —
526
- deny wins over allow. Enforced in the runner's decision validator (rejects
527
- `futures_open` / `spot_order` on a denied symbol) and surfaced in the system
528
- prompt, so the model is told the boundary and the runner re-checks it.
529
- - **Docs: Open Knowledge Format positioning.** Clarified that a CoinRithm agent
530
- is an OKF bundle — a portable directory of markdown + frontmatter that is
531
- model-agnostic (run the same definition on any model). Develop and prove it
532
- free on paper, then run it anywhere.
533
- - npm keywords refreshed (`open-knowledge-format`, `okf`, `model-agnostic`,
534
- `gemini`) for registry discovery.
535
-
536
- ## 0.2.0
537
-
538
- - Added the **`coinrithm-agent`** self-host runner binary alongside the MCP
539
- server: a folder-as-architecture (OKF) agent you bring your own model key to,
540
- with caps enforced by the runner (not the model), dry-run by default,
541
- paper-only.
542
-
543
- ## 0.1.x
544
-
545
- - Initial `coinrithm-mcp` MCP server: reads, quotes, scoped spot/futures/PM
546
- writes, ledger export, and Agent Arena integration over a user-minted API key.
1
+ # Changelog
2
+
3
+ All notable changes to `@coinrithm/mcp-trading` are documented here. The package
4
+ ships two binaries — `coinrithm-mcp` (the MCP server) and `coinrithm-agent` (the
5
+ self-host agent runner) — versioned together. The CoinRithm **API contract** is
6
+ versioned separately (see `openapi.yaml` `info.version`, currently `1.7.0`).
7
+
8
+ ## 0.7.13
9
+
10
+ - Preserve optional per-outcome venue terms in compact prediction-market output,
11
+ including explicit unavailable values and valid false/zero facts.
12
+ - Include candle volume coverage in agent fundamentals: positive counts identify
13
+ missing expected venues; unknown coverage stays unknown. Preserve measured zero
14
+ volume and keep each volume paired with its own coverage.
15
+ - Classify provider capacity failures consistently, including HTTP 503
16
+ resource exhaustion, in runner results and configured same-model retries.
17
+ - Export `classifyProviderFailure` from the supported engine entry point.
18
+ Retry counts and trading behavior are unchanged.
19
+ - Surface signed cumulative futures funding and its applied-through timestamp to
20
+ the agent observation so funding already reflected in balances is not counted twice.
21
+ - Add public whale-wallet context to the 40-tool surface: bounded 7-day or
22
+ 30-day wallet activity and source/wallet movement detail, with provenance and
23
+ provider-reported context kept distinct from CoinRithm's observed trade flow.
24
+ - Clarify `pm_data_calibration` as market-price calibration: the primary lane
25
+ uses one complete-book snapshot in the inclusive 20–28 hour pre-resolution
26
+ window, with event-weighted ECE and separate `finalPrice`/`ownCapture` lanes.
27
+
28
+ Publication and hosted deployment are verified separately. See the repository's
29
+ [release status](https://github.com/CoinRithm/coinrithm-agent-trading#version-clarity)
30
+ for registry availability and the hosted runtime evidence.
31
+
32
+ ## 0.7.12 — 2026-09-15
33
+
34
+ Clarify `whoami`, `cancel_spot_order` and `report_pm_opportunity` descriptions,
35
+ removing execution-cost prose unrelated to these operations. Document actual
36
+ authentication, side effects, result fields and retry behavior.
37
+
38
+ Spot cancellation now advertises its existing idempotent behavior. Opportunity
39
+ reporting no longer advertises unconditional idempotency: duplicate protection
40
+ requires `decisionId` (or `agentTrace.decisionId`) under the same API key; the
41
+ first stored record wins. Reporting remains a write despite requiring only the
42
+ `read` scope, and its evidence remains explicitly self-reported.
43
+
44
+ Tool names, accepted inputs and execution behavior are unchanged. Published on
45
+ npm and deployed to hosted MCP on 2026-09-15; registry and GitHub downloads match
46
+ the CI archive. A clean registry installation passed startup and tool metadata
47
+ checks. See the [release record](https://github.com/CoinRithm/coinrithm-agent-trading/releases/tag/mcp-trading-v0.7.12).
48
+ Glama's evaluation score is separate from package delivery.
49
+
50
+ ## 0.7.11 — 2026-09-15
51
+
52
+ Fix opportunity reporting that previously treated resolved API failures as
53
+ successful submissions. Only `ok: true` confirms a report. HTTP errors remain
54
+ unconfirmed; transport failures and exceptions have an unknown delivery outcome.
55
+
56
+ `CycleResult.opportunity` now contains confirmed reports only. The additive
57
+ `opportunityReport` field preserves the attempted payload, outcome and status,
58
+ without API error bodies or exception details. Existing consumers should use
59
+ this field when they need attempted rather than confirmed evidence.
60
+
61
+ The reporter retains its one-invocation-per-cycle latch and adds no retries.
62
+ Focused regressions compare successful and failed reporting in skip/act cycles
63
+ and verify unchanged trading results and runner state. SDK versions are unchanged.
64
+ This source entry does not establish registry publication or hosted deployment.
65
+
66
+ ## 0.7.10 — 2026-09-15
67
+
68
+ Source changes following the published 0.7.9 release. The TypeScript SDK remains
69
+ 0.3.1 and Python remains 1.8.1; their runtime source is unchanged.
70
+
71
+ - Persist file-backed run identity before execution so a first-cycle process
72
+ crash cannot discard the idempotency identity. Transport uncertainty is still
73
+ not automatically replayed or recorded as a completed trade.
74
+ - Fix runner startup on Node 18: use the imported Node crypto API instead of
75
+ depending on a global crypto object. The installed-package matrix reproduced
76
+ this failure on Linux, Windows and macOS.
77
+ - Add the supported `@coinrithm/mcp-trading/engine` entry point, preserving
78
+ existing deep imports, and separate observation accounting and opportunity
79
+ reporting from cycle ordering.
80
+ - Add opt-in, machine-checked crypto return predicates and visible warnings for
81
+ inactive/reserved configuration. Existing agents are not automatically opted in.
82
+ - Serialize scheduler migrations in a bounded transaction; add an offline
83
+ credential-rotation helper and interruption/recovery rehearsal.
84
+ - Pin workflow actions and verify release-tool checksums. Add installed-package
85
+ compatibility and restart smoke checks across operating systems and runtimes.
86
+
87
+ 0.7.10 was published and deployed on 2026-09-15. Registry and GitHub downloads
88
+ matched the CI-tested archive; hosted MCP and scheduler deployments finished.
89
+ See the [release record](https://github.com/CoinRithm/coinrithm-agent-trading/releases/tag/mcp-trading-v0.7.10).
90
+
91
+ ## 0.7.9 - 2026-09-15
92
+
93
+ Published on npm and verified on 2026-09-15: the registry archive matches the
94
+ reviewed artifact and passes fresh-install checks. See the
95
+ [combined release notes](https://github.com/CoinRithm/coinrithm-agent-trading/releases/tag/mcp-trading-v0.7.9)
96
+ for source provenance, verification and community acknowledgments.
97
+
98
+ Public market-data fidelity and runner reliability release. No MCP tool was
99
+ renamed or removed, and the API **contract stays 1.7.0**.
100
+
101
+ **PM evaluation budget.** Event-driven periodic prediction-market evaluations
102
+ now respect `maxLlmCallsPerHour` after their cooldown elapses. Budget skips make
103
+ no provider call and consume no call allowance. PM keeps its own cooldown;
104
+ open-position management and explicit always-on behavior retain their existing
105
+ exemptions. This runner gate is separate from hosted provider-capacity admission.
106
+
107
+ **Retry-After parsing.** Missing, blank or malformed headers no longer become
108
+ zero-delay retries. The runner API client uses its existing five-second fallback;
109
+ explicit zero, numeric seconds and HTTP dates remain supported. Model-provider
110
+ cooldowns share the parser and retain their existing one-hour cap.
111
+
112
+ **API request deadlines.** Each runner API operation now has a 30-second total
113
+ deadline covering response headers, body reads and all 429 retry waits. The
114
+ same client serves the hosted scheduler. Embedded callers can set a finite
115
+ `requestTimeoutMs` and supply an `AbortSignal`. A timeout or cancellation returns
116
+ an uncertain transport result without automatically replaying a trading write.
117
+ Timers and listeners are removed when the operation finishes.
118
+
119
+ **State persistence.** Self-host state is serialized to a private temporary file
120
+ and atomically renamed over the previous state. A failed serialization or rename
121
+ leaves the prior state intact. This is atomic replacement, not a claim of durable
122
+ storage across power loss.
123
+
124
+ **Agent conversion.** `coinrithm-agent eject` preserves explicit `triggerPolicy`
125
+ and `capitalSizing` blocks. Previously conversion could restore default hourly
126
+ budgets and drop equity sizing.
127
+
128
+ **Release verification.** All-source coverage gates, mandatory PostgreSQL CI,
129
+ dependency updates and corrected client setup docs are included. See the
130
+ [reliability record](https://github.com/CoinRithm/coinrithm-agent-trading/blob/main/docs/RELIABILITY.md).
131
+
132
+ **Confirmed-action journal.** Completed-action memory now requires an action
133
+ to be both accepted and executed. Failed writes and uncertain transport results
134
+ retain their attempt evidence without becoming completed moves in the next
135
+ decision prompt. Dry-run proposals remain unexecuted.
136
+
137
+ **Direct NVIDIA retry.** One complete HTTP 500/502/503/504 response can be
138
+ retried once on the identical direct NVIDIA route within the original deadline.
139
+ Both attempts are retained. This does not retry trading writes or change the
140
+ hosted shared-pool routing policy.
141
+
142
+ **Capital reconciliation.** Frozen-balance rounding residue down to -1e-8 is
143
+ normalized only in the sizing calculation after the independent reads agree.
144
+ Negative spendable cash still fails closed; wallet balances are not changed.
145
+
146
+ **Private decision input evidence.** The bounded numeric projection includes
147
+ nested indicator inputs and context movers, with legacy v1 records still readable.
148
+ It does not retain hidden reasoning or raw model output.
149
+
150
+ **Compact prediction-market evidence.** Discovery and compact event-detail
151
+ responses now retain the API's `source.quoteScale`, `source.methodology` and
152
+ `source.supportsMarketMetrics`, plus `spreadPoints`, `probabilityBook` and
153
+ each retained outcome's `normalizedProbability`. Venue-native bid/ask quotes
154
+ are never rescaled or interpreted from magnitude. Normalization remains the
155
+ API's calculation over the original full book, not the truncated top-five
156
+ outcome list. Existing payload bounds and explicit `detail: full` behavior
157
+ are unchanged.
158
+
159
+ **Settlement-time provenance.** Compact events retain `resolvedAtBasis` and
160
+ `settlementWindowClosedAt`, keeping provider expiration distinct from an
161
+ announced settlement time. Null and absent upstream evidence stay null and
162
+ absent; the MCP does not infer missing values.
163
+
164
+ **Candle semantics.** The `get_candles` description now states that these are
165
+ sampled composite-price bars. Each bar's `v` is a mean rolling 24-hour
166
+ quote-volume observation in USD, not volume traded during the candle, and
167
+ must not be summed across bars.
168
+
169
+ **HTTP completion diagnostics.** The hosted HTTP
170
+ entry now has a bounded, stderr-only completion observer with final SDK-result
171
+ and finish/abort accounting. Initialization, discovery, tool failures and
172
+ successful delivery are distinct; unknown tool names are normalized. Records
173
+ contain no arguments, bodies, credentials, caller/RPC IDs or caller-origin labels.
174
+ Credential presence is not authentication. Durations describe the HTTP request,
175
+ shared by batch members; server finish does not prove client receipt or use.
176
+ Stdio, tools, authentication and dependency versions are unchanged. See
177
+ `DEPLOY.md` for the measurement and retention limits. Hosted source/image
178
+ `18a0bb6a8a0665e91cebc10225fec6f7ebcdaaf7` passed a bounded anonymous smoke on
179
+ 2026-09-13. Hosted verification and npm publication are separate release steps.
180
+
181
+ **Deployment boundaries.** Hosted scheduler admission reasons are private
182
+ scheduler telemetry, not a new SDK or MCP response field. The API's corrected
183
+ comparison probabilities and enriched spread names use the existing response
184
+ shape and reach current clients through fresh API reads. Outcome display names
185
+ may change; use source/event/outcome identifiers for identity, never summed
186
+ prices or matching labels alone. These fixes do not establish trading returns.
187
+
188
+ ## 0.7.8
189
+
190
+ Runner decision-quality, evidence and paper-capital release. Additive: no MCP
191
+ tool was renamed or removed, and the API **contract stays 1.7.0**. This release
192
+ contains all package changes since published 0.7.7 (`gitHead` `80d0cae`), not
193
+ just the previously listed thesis work.
194
+
195
+ **Thesis exits.** Every opening action (`futures_open`, `spot_order`,
196
+ `pm_open`) now carries a `thesis`: a one-sentence summary plus an
197
+ `invalidation` with at least one machine-checkable condition (`priceBelow` /
198
+ `priceAbove` for coins, `probabilityBelow` / `probabilityAbove` for prediction
199
+ markets, a `maxHoldMinutes` time stop, and a free-text `catalyst` the model
200
+ re-judges itself). The runner binds the thesis to the position the server
201
+ returns, sanitized side-aware (a rising price never invalidates a long; a
202
+ wrong-side level is dropped rather than re-signed; the time stop is clamped to
203
+ 60 minutes .. 30 days), persists it in the run state (`RunState.theses`, the
204
+ same state file / `agent_state` JSON as before, no schema change) and
205
+ re-evaluates it every cycle. A futures position whose price level or time stop
206
+ is breached is closed by the runner before the model is asked anything, logged
207
+ as a `thesis_invalidated` exit with its own idempotency key, after the
208
+ kill-switch and drawdown checks and never instead of them. Prediction-market
209
+ positions have no close endpoint, so a broken PM thesis is surfaced to the
210
+ model instead (do not add, let it settle). The parser is tolerant (a malformed
211
+ thesis never fails the open; a thesis copied onto a close is ignored) and the
212
+ structured-output schema requires it, so schema-enforced hosted models always
213
+ emit one.
214
+
215
+ **Fundamentals in the observation.** Each watch entry now carries
216
+ `fundamentals` sourced only from calls the runner already makes: `categories`,
217
+ `marketCapRank` and `marketCapUsd` from the market context; `volume24hUsd` from
218
+ the candles the `indicators` capability already fetches (live-probed
219
+ 2026-09-02: each bar's `v` is a rolling 24h volume, so the latest bar is the
220
+ 24h figure, never the sum); and up to three `headlines` with `publishedAt`
221
+ timestamps from the one `news` call, attributed through the curated coin-news
222
+ graph. Discovered PM markets carry `endDate` and `liquidityUsd`; open PM
223
+ positions carry their title, side, entry and current probability and
224
+ `openedAt`; open futures positions carry `openedAt`. The system prompt states
225
+ the thesis contract, the runner-enforced exit and how to grade a trade on the
226
+ fundamentals. Not carried, because no agent endpoint serves them: an "about"
227
+ text per coin, a 24h probability change and a cross-venue divergence per PM
228
+ market.
229
+
230
+ **Fix:** the public movers feed serializes `change24h` / `currentPrice` as
231
+ decimal strings; the universe-scan context rows read them strictly as numbers
232
+ and shipped `undefined` for every mover.
233
+
234
+ **Opt-in equity-based paper sizing.** A runner can size entries from a
235
+ conservative fraction of its independently attributed paper book instead of a
236
+ fixed stake/margin. The book is accepted only when wallet identity, cash
237
+ partitions, held-position attribution and spot-mark coverage reconcile. Quotes
238
+ then enforce per-entry, per-symbol, deployed-capital and daily-entry limits;
239
+ fee buffers and the API's fee-inclusive quote evidence are included. Any
240
+ missing or inconsistent evidence fails closed. Legacy positions on a different
241
+ book remain visible for management but never inflate the current book's buying
242
+ power.
243
+
244
+ **Prediction-market decisions use executable economics.** PM opens now reject
245
+ an invalid raw probability and a model forecast that does not clear the quoted
246
+ entry price. Forecast edge is measured against the actual fee/slippage-adjusted
247
+ fill, not the headline market probability. Quote-expiry outcomes are recorded
248
+ separately from risk/balance rejection, and futures risk/reward validation uses
249
+ fee-inclusive entry and stop economics.
250
+
251
+ **Decision evidence is structured and bounded.** Cycles can expose a sanitized,
252
+ partial private decision-input record: configuration and observation
253
+ fingerprints, daily budget and guard state, plus bounded observation rows with
254
+ explicit omission counts. It is not a prompt, transcript, raw model output or
255
+ hidden reasoning record. The runner also reports quote/validation evidence for
256
+ abstained, forecast-only and quote-expired PM opportunities. Hosted persistence
257
+ and retention remain the caller's responsibility.
258
+
259
+ **Runtime controls are more faithful.** The model sees the remaining daily
260
+ entry/add budget rather than only static maxima. Entry caps still block new
261
+ risk, while closes and other risk-reducing actions remain available. Direct
262
+ provider HTTP 429 responses are capacity skips rather than model failures, so
263
+ BYO agents do not build a failure streak during ordinary quota pressure.
264
+ Structured-tool decisions remain required where the provider supports that
265
+ contract.
266
+
267
+ **Scorecard fix.** Maximum drawdown now measures decline from starting equity,
268
+ so an immediate loss is no longer hidden by treating the first post-trade point
269
+ as the high-water mark.
270
+
271
+ ## 0.7.7
272
+
273
+ Reliability release. Every change here came from a live production failure, not
274
+ from a roadmap. Additive: no tool renamed or removed, and the API **contract
275
+ stays 1.7.0** because nothing on the documented surface changed.
276
+
277
+ **Model requests are now built from a declared capability table, not
278
+ assumptions.** `providerCapabilities.ts` states, per model family, which
279
+ parameter carries the completion budget, whether a non-default temperature is
280
+ allowed, and what extra body fields the family needs. Two failures this fixes:
281
+
282
+ - **OpenAI's current models rejected our requests outright.** `gpt-5*` and
283
+ `o*` refuse `max_tokens` and any non-default `temperature`; they take
284
+ `max_completion_tokens`. The family is detected by MODEL id, not just the
285
+ provider name, so an OpenAI-compatible gateway serving `gpt-5` gets the same
286
+ shape. If you brought your own OpenAI key, this is why it now works.
287
+ - **NVIDIA Nemotron models emitted a think-chain where the JSON decision
288
+ belonged**, which failed every cycle. The `chat_template_kwargs.enable_thinking=false`
289
+ switch and the "detailed thinking off" system hint are now encoded as data
290
+ rather than re-learned by failing.
291
+
292
+ **New: `probeDecisionContract()`.** An HTTP 200 is not proof a route can run an
293
+ agent. Both production failure modes returned 200s: a think-chain in the JSON
294
+ slot, and an empty completion because a reasoning model spent its whole budget
295
+ before answering. The probe sends a canned mini-observation through the REAL
296
+ decision parser at a >=1024 completion allowance and classifies the result as
297
+ `http`, `empty` or `parse`. Use it before adopting any model id; provider
298
+ catalogs list ids that 404 on invoke.
299
+
300
+ **Provider trouble no longer disables an agent.** A permanent-looking model
301
+ error (404/410/decommissioned) used to disable the agent after a threshold. On
302
+ 2026-08-26 NVIDIA end-of-lifed an entire model line and 35 agents died on that
303
+ path. The runner now reports a hold and keeps retrying each cadence, recovering
304
+ by itself when the provider does. Disables remain for what deserves them:
305
+ revoked credentials, drawdown, kill-switch, user action.
306
+
307
+ **Failures carry structured metadata.** A failed `decide()` now returns
308
+ `status` and, when the provider sends one, `retryAfterMs` (parsed from
309
+ `Retry-After` in both delta-seconds and HTTP-date form, capped at an hour), so
310
+ a caller can tell a 429 from a 5xx without parsing strings. Error text is
311
+ unchanged.
312
+
313
+ **`ClientConfig.extraHeaders`.** Headers attached to every request, spread
314
+ before auth so they can never clobber it. Self-host has nothing to put here;
315
+ it exists so CoinRithm's own hosted scheduler can present its attestation
316
+ channel.
317
+
318
+ **Model names corrected throughout.** The retired Llama 3.x line is gone from
319
+ the README, the runtime defaults and the `quant-reference` example, which is
320
+ relocked onto `nvidia/nemotron-3-nano-30b-a3b`.
321
+
322
+ ## 0.7.6
323
+
324
+ Agent capability release: universe discovery, first-class behavioral guards,
325
+ and the hosted prose budget made visible. Additive — no tool renamed or
326
+ removed. Contract moves to **1.7.0** (two keyless paths declared).
327
+
328
+ **New: agents can look beyond their own watchlist.**
329
+
330
+ - **`get_crypto_movers` tool.** Keyless scan of the tracked coin universe for
331
+ the biggest 24h gainers or losers. Rows carry `coinId`, `symbol`, `name`,
332
+ `slug`, `change24hPct`, `priceUsd`.
333
+ - **`universe_scan` capability** for the self-host runner. Each cycle it pulls
334
+ the top movers, promotes the strongest few into full watch entries marked
335
+ `discovered: true`, and passes the remainder as compact context. Watchlist
336
+ and blocklist symbols are excluded up front, so a discovered row can never
337
+ duplicate a configured pair or bypass the deny list.
338
+ - **Both now carry the coinId through.** The movers row's `ucid` IS the
339
+ `coinId` that `get_candles` / `get_market_context` / the futures quote path
340
+ take. It was previously stripped from the tool response and re-derived from
341
+ the SYMBOL via a resolve round-trip — a wasted call per discovered mover and
342
+ a real correctness hazard, because symbols collide across listings and the
343
+ resolver could return a different coin than the one that actually moved.
344
+
345
+ **New: contract declares the endpoints the tools call.**
346
+
347
+ - `/api/coins/top-gainers` and `/api/coins/top-losers` are now in
348
+ `openapi.yaml` (tag `public-crypto-data`), so both SDKs can reach the
349
+ surface `get_crypto_movers` uses. Probe-verified against prod: bare array,
350
+ no envelope; `change24h` / `currentPrice` are decimal STRINGS; default
351
+ `limit` is 3 and out-of-range values return 400 rather than clamping.
352
+
353
+ **New: personality and boundaries are configurable, and documented.**
354
+
355
+ - **`character/guards.md`** — first-class hard behavioral guards, merged into
356
+ the strategy prose as a distinct section rather than buried in the thesis.
357
+ - **`examples/agents/pia-pump-fader`** — a full bundle demonstrating
358
+ capabilities plus boundary configuration (watchlist/blocklist interaction,
359
+ the five-point risk gate, re-entry discipline).
360
+ - **`examples/agents/FORKING.md`** — a file-by-file map of what is strategy
361
+ and what is plumbing, so a fork knows what it is allowed to change.
362
+ - **QUICKSTART** documents capabilities, and a docs-drift tripwire fails the
363
+ suite when a capability ships undocumented (`universe_scan` shipped
364
+ invisible in every user surface once; that cannot recur silently).
365
+
366
+ **Fixed.**
367
+
368
+ - **Hosted prose budget is validated, not discovered at deploy.**
369
+ `coinrithm-agent validate --hosted` now checks the 8,000-character merged
370
+ prose budget and reports the exact overage. A bundle could previously
371
+ validate clean and still be undeployable. YAML frontmatter is stripped
372
+ before the count (and before the model sees it — it was being fed in as if
373
+ it were strategy). `pia-pump-fader` was rebuilt to fit at 7,932.
374
+ - **Permanent failures stop being revived.** A disabled agent whose model is
375
+ gone or whose key is invalid is no longer resurrected by the scheduler's
376
+ revive pass; only transient failures are retried.
377
+ - **Fresh scaffolds are no longer bricked** by the capabilities field, and
378
+ action-confidence tolerance was widened to match what models actually emit.
379
+ - **False market-data licensing assertion corrected** in both READMEs.
380
+
381
+ ⚠ Publishing to npm remains a **manual** step — `publish-mcp.yml` pushes
382
+ `server.json` to the MCP registry only.
383
+
384
+ ## 0.7.5
385
+
386
+ **Release-hygiene bump. Everything below was already merged but never
387
+ reached npm** — the 0.7.4 tarball was published 2026-07-26T23:01:11Z and
388
+ five commits landed after that instant without a version bump, so the
389
+ repository's 0.7.4 and the published 0.7.4 were different code under one
390
+ version number. This release makes the published artifact match the
391
+ source again.
392
+
393
+ - **Security.** MCP dependency audit fixes, 6 findings to 0 (`da65e9e`).
394
+ Anyone on published 0.7.4 is running the pre-audit dependency set.
395
+ - **`pm_data` Gemini exposure** for agents (`3ab04ae`).
396
+ - **Venue methodology and health** exposed as tools (`74cc495`).
397
+ - **Reproducible decision receipts** persisted by the agent runner
398
+ (`8bebc9e`).
399
+ - **Docs.** Contract version drift corrected and the placeholder SDK
400
+ README replaced, so the docs stop advertising an install that 404s
401
+ (`6d61b92`).
402
+
403
+ No tool was renamed or removed; this is additive plus a dependency
404
+ refresh.
405
+
406
+ ⚠ Publishing this package is a **manual** step — `publish-mcp.yml` only
407
+ pushes `server.json` to the MCP registry, it does not run `npm publish`.
408
+ That asymmetry is exactly how the drift above accumulated unnoticed.
409
+
410
+ ## 0.7.4
411
+
412
+ Docs-only. No tool behavior change, no API-surface change.
413
+
414
+ - **Acceptable Use of Market Data.** The README (root and this package) and
415
+ `openapi.yaml` (`info.termsOfService`, `info.description`, and the
416
+ `public-pm-data` tag) now reference and summarize CoinRithm's licensing
417
+ flow-down restriction on Market Data from third-party prediction-market
418
+ venues: read-only use for paper-trading context and settled-outcome
419
+ scoring only — no model training/fine-tuning/benchmarking, no
420
+ redistribution or bulk-extraction, no use to build a competing product.
421
+ Full terms: <https://www.coinrithm.com/en/terms-of-use>
422
+
423
+ ## 0.7.3
424
+
425
+ Quality-engine surfaces + independent forecasts. Additive; no breaking change.
426
+
427
+ - **Quality verdicts in tool responses.** `discover_pm_markets`, `pm_quote`, and
428
+ the `pm_data_*` tools now surface the persisted truth-engine `quality` object
429
+ (`decisionEligible`, warning/block reason codes, `policyVersion`, `assessedAt`).
430
+ Markets with critical failures stay visible but cannot drive paper opens or
431
+ alerts.
432
+ - **`openBlocked` preview on `pm_quote`.** Quotes preview the open-time quality
433
+ gate (`openBlocked` + `openBlockReasons`), so an agent can skip a market that
434
+ would 422 before burning the open attempt. The self-host runner
435
+ (`coinrithm-agent`) does this skip automatically.
436
+ - **Independent forecasts in the runner.** The self-host agent runner elicits the
437
+ model's OWN probability (judged from the question/resolution criteria/deadline,
438
+ never anchored to the market price) and submits it as `forecastProbability` on
439
+ PM opens — feeding the public calibration dataset with proper-scoring-rule
440
+ forecasts. Clamped to [1,99]; omitted (never faked) when the model does not
441
+ produce one; `HOUSE_AGENT_FORECAST_ENABLED=false` disables.
442
+ - **`crossPlatform` on event lists** documented in the API contract: sibling
443
+ venues pricing the same question, on list rows.
444
+ - **ForecastEx venue truth.** Public MCP discovery copy and registry metadata
445
+ now describe all 11 live venues, including ForecastEx.
446
+ - **Contract synchronization.** Runner templates and example bundles pin the
447
+ served OpenAPI 1.6.0 contract; canonical scorecard paths are unambiguous.
448
+
449
+ ## 0.7.2
450
+
451
+ Docs-truth + privacy release. No tool behavior change, no API-surface change.
452
+
453
+ - **Ten venues in the public listing.** `pm_data_*` tool copy, the README, and
454
+ `server.json` now name all ten venues (adds Futuur and Myriad). npm `0.7.1` was
455
+ published before those landed, so the registry listing still advertised "eight
456
+ venues"; npm versions are immutable, so correcting the public listing required
457
+ a new release.
458
+ - **`source` parameter description** on `pm_data_events` / `pm_data_event_detail`
459
+ now enumerates all ten venue slugs. Accepted values are unchanged — this is
460
+ description text only, which is why it is a patch and not a minor.
461
+ - **Privacy.** Raw model output is no longer persisted, enforcing the package's
462
+ no-chain-of-thought promise.
463
+ - **New tripwire.** `server.json` (the MCP-registry listing) is now guarded
464
+ against version and venue-count drift; it had no guard, which is how it went
465
+ stale in the first place.
466
+ - Refreshed stale Arena-gate example copy.
467
+
468
+ ## 0.7.1
469
+
470
+ Docs + registry-metadata release; no tool behavior changes.
471
+
472
+ - **README refresh**: the keyless `pm_data_*` data surface is now front and
473
+ center — 8 venues (Polymarket, Kalshi, Smarkets, Limitless, Manifold,
474
+ Metaculus, PredictIt, Rothera — the "seven venues" line predated Rothera),
475
+ the anonymous hosted-endpoint path, and the `referenceProbability` /
476
+ `volumeHistory` fields the data tools return.
477
+ - **`server.json`**: hosted endpoint's `Authorization` header marked optional
478
+ (the `pm_data_*` tools work anonymously — verified live) and the server
479
+ description now leads with the keyless data surface.
480
+ - Ships the post-0.7.0 commits: `pm_data_event` advertises `volumeHistory`,
481
+ `pm_data_events` advertises `referenceProbability` on list items, and the
482
+ hosted MCP root (`GET /`) serves a self-describing JSON landing (with a
483
+ 405 + hint on `GET /mcp`).
484
+
485
+ ## 0.7.0
486
+
487
+ (Retroactive entry — released 2026-07-05 without a changelog note.)
488
+
489
+ - **Four keyless `pm_data_*` tools** — CoinRithm's free public cross-venue
490
+ prediction-market dataset over MCP, no API key required and yours is never
491
+ attached: `pm_data_overview` (market-wide stats), `pm_data_events`
492
+ (cross-venue event list), `pm_data_event` (detail incl.
493
+ `crossSourceMatches` + resolution evidence), `pm_data_whales`
494
+ (large-trade tape).
495
+
496
+ ## 0.5.0
497
+
498
+ Agent-runner quality + reliability release. `coinrithm-agent` got materially
499
+ smarter and less noisy; `coinrithm-mcp` is unchanged in shape. Bundles the work
500
+ since 0.4.0.
501
+
502
+ - **Prediction markets are a first-class venue** in the decide prompt: a short
503
+ `pmN` ref so small models trade PM reliably, eligible-outcome filtering (only
504
+ backend-openable outcomes reach the model), and futures-capped agents steered
505
+ to PM (a separate budget) instead of re-rejecting.
506
+ - **PM anti-churn — now actually effective.** The candidate list is pre-filtered
507
+ to exclude markets the agent already holds, and the runner + server both block
508
+ re-betting a held market+outcome (no more one-agent, 25-identical-bets churn).
509
+ An earlier version read the wrong `/positions/pm` fields and was silently dead;
510
+ fixed.
511
+ - **Settlement-feedback learning loop.** The agent sees how its own recent bets
512
+ actually resolved (win/loss/void + realized PnL) as reflective context, so it
513
+ learns from outcomes across cycles.
514
+ - **Per-trade reasoning stays honest about the market.** A multi-action decision
515
+ no longer stamps its primary rationale onto a secondary trade about a different
516
+ market — the trade's public Arena "why" always matches the market it's on.
517
+ - **Futures reliability.** The model is unblinded to per-position mark /
518
+ liquidation / stop / take-profit prices; take-profit is auto-clamped to a valid
519
+ R:R target off the stop (kills the `take_profit_not_*_mark` reject waves); and a
520
+ marking-down PM book now trips the equity-drawdown kill-switch too.
521
+ - **News capability.** Recent high-importance news for the watchlist coins is fed
522
+ into the decide context as a market-catalyst layer.
523
+ - **Robustness + contract accuracy.** Scheduler/runner hardening, flat-state
524
+ prompt steers (weak models stop hallucinating closes), manage-enum
525
+ normalization, and the PM contract now documents real entry friction rather
526
+ than a disclose-only stance.
527
+ - **Security.** `hono` bumped to 4.12.27 (high-severity advisories: serve-static
528
+ path traversal, CORS wildcard-with-credentials, body-limit bypass).
529
+ - **Docs.** The npm README leads with value / free / OKF / Studio; stale
530
+ scheduler and `minDecidedTrades` claims corrected.
531
+
532
+ ## 0.4.0
533
+
534
+ - **Deterministic scorecard engine (`computeScorecard`).** The reproducible-
535
+ evaluation engine for `coinrithm.agent.scorecard.v1` — pure math over an
536
+ agent's realized track record (no network, no model): realized PnL, win rate,
537
+ expectancy, profit factor, reward-to-risk, Sharpe, Sortino, deflated /
538
+ probabilistic Sharpe (Bailey & López de Prado — skill vs luck with a multiple-
539
+ testing penalty), max drawdown, and Brier + ECE calibration for probabilistic
540
+ calls. Same inputs → identical metrics **and** a sha256 `contentHash` of the
541
+ canonicalized result, so a scorecard whose hash doesn't reproduce isn't
542
+ trusted. Metrics are computed AFTER the run from immutable evidence (leakage-
543
+ separation), so tuning-to-the-metric is structurally impossible. Returns
544
+ `null` for thin records — never a fabricated number.
545
+ - **Resolver: committable file metadata + functionality pin.** The OKF resolver
546
+ now carries per-file metadata and pins functionality through resolution, so a
547
+ bundle's behavior is reproducible from its committed files.
548
+
549
+ ## 0.3.0
550
+
551
+ - **Agent risk config: coin deny-list (`blocklist`).** `risk.blocklist` lets an
552
+ agent name symbols it must never open, even if they are on the watchlist —
553
+ deny wins over allow. Enforced in the runner's decision validator (rejects
554
+ `futures_open` / `spot_order` on a denied symbol) and surfaced in the system
555
+ prompt, so the model is told the boundary and the runner re-checks it.
556
+ - **Docs: Open Knowledge Format positioning.** Clarified that a CoinRithm agent
557
+ is an OKF bundle — a portable directory of markdown + frontmatter that is
558
+ model-agnostic (run the same definition on any model). Develop and prove it
559
+ free on paper, then run it anywhere.
560
+ - npm keywords refreshed (`open-knowledge-format`, `okf`, `model-agnostic`,
561
+ `gemini`) for registry discovery.
562
+
563
+ ## 0.2.0
564
+
565
+ - Added the **`coinrithm-agent`** self-host runner binary alongside the MCP
566
+ server: a folder-as-architecture (OKF) agent you bring your own model key to,
567
+ with caps enforced by the runner (not the model), dry-run by default,
568
+ paper-only.
569
+
570
+ ## 0.1.x
571
+
572
+ - Initial `coinrithm-mcp` MCP server: reads, quotes, scoped spot/futures/PM
573
+ writes, ledger export, and Agent Arena integration over a user-minted API key.