@battlegrid/mcp-server 31.2.5 → 31.2.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +24 -16
- package/SKILL.md +48 -60
- package/dist/index.d.ts +1 -1
- package/dist/index.js +1 -1
- package/package.json +1 -1
- package/skills/EXPORT.json +14 -0
- package/skills/battlegrid-agent-management/SKILL.md +173 -0
- package/skills/battlegrid-arena-play/SKILL.md +132 -0
- package/skills/battlegrid-market-analysis/SKILL.md +76 -0
- package/skills/battlegrid-radar-deployment/SKILL.md +180 -0
- package/skills/battlegrid-strategy-authoring/SKILL.md +283 -0
- package/skills/battlegrid-strategy-doctor/SKILL.md +142 -0
- package/skills/battlegrid-strategy-examples/SKILL.md +401 -0
- package/skills/battlegrid-trade-analysis/SKILL.md +74 -0
- package/skills/battlegrid-strategy-studio/SKILL.md +0 -254
- package/skills/battlegrid-strategy-studio/references/playbooks.md +0 -370
- package/skills/battlegrid-strategy-studio/references/recipes.md +0 -173
- package/skills/battlegrid-strategy-studio/references/tradingview-ports.md +0 -250
|
@@ -0,0 +1,180 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: battlegrid-radar-deployment
|
|
3
|
+
description: Deploy the player's agents to standing duty — per-coin Radar policies that fire real trades on confirmed regime flips, and per-preset Arena deployment policies that enter sessions automatically. Reads what is deployed now, previews what a draft would actually resolve to, writes it only after the player has seen that preview, and un-deploys with the blast radius stated. Activate whenever the player wants an agent put on duty, wants to change or pause a deployment, asks what would fire right now or why nothing is firing, or wants to stop a coin or a preset being traded automatically.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Radar & Deployment
|
|
7
|
+
|
|
8
|
+
You are handing an agent **standing authority**. A Radar policy fires real trades on confirmed
|
|
9
|
+
regime flips with no further human action; a deployment policy enters real sessions the same way.
|
|
10
|
+
Nobody is watching when it happens — which is why the player has to see what a policy resolves to
|
|
11
|
+
*before* it exists, not after.
|
|
12
|
+
|
|
13
|
+
Radar (per coin) and Arena deployment (per preset) are **separate bounded contexts**. They share
|
|
14
|
+
no slot shape, no condition union and no resolution type. Never carry a fact from one to the other,
|
|
15
|
+
and never describe them as one thing with two modes.
|
|
16
|
+
|
|
17
|
+
## The five failures this flow exists to prevent
|
|
18
|
+
|
|
19
|
+
1. **A write with no fresh preview.** *Cue: any `upsert_radar_deployment` or
|
|
20
|
+
`upsert_deployment_policy`.* The preview is the only place the player sees which agent actually
|
|
21
|
+
goes on duty and why. → steps 2–3.
|
|
22
|
+
2. **A blind CAS retry.** *Cue: a CONFLICT on a write.* The revision moved because the stored
|
|
23
|
+
policy changed. Bumping the number and retrying overwrites an edit you never read. → step 3.
|
|
24
|
+
3. **Previewing one thing and writing another.** *Cue: the player adjusts a slot, a bar, a window,
|
|
25
|
+
a regime, or the enabled flag after you previewed.* A preview vouches only for the parameters it
|
|
26
|
+
was given. → step 2.
|
|
27
|
+
4. **A replacement that silently drops slots.** *Cue: any upsert on a coin or preset that already
|
|
28
|
+
has a policy.* `slots` is the COMPLETE set and replaces what is stored — a slot you did not
|
|
29
|
+
resend is deleted. → step 3.
|
|
30
|
+
5. **Answering "why isn't it firing?" by paging the journal.** *Cue: any question about why a
|
|
31
|
+
deployed agent has been quiet — "is it working?", "it hasn't traded all day", "what's blocking
|
|
32
|
+
it?".* `get_radar_activity_summary` now answers the whole question in one small call — which
|
|
33
|
+
cause recurs, how far the score sits from firing, and what just happened — so paging the journal
|
|
34
|
+
for it tallies an aggregate the server already computed and re-derives a proximity reading it
|
|
35
|
+
already serves. The journal read is for ONE occurrence, or for more rows than the summary's ten.
|
|
36
|
+
→ step 5.
|
|
37
|
+
|
|
38
|
+
## Sequence
|
|
39
|
+
|
|
40
|
+
### 1. Resolve the target, then read current state
|
|
41
|
+
|
|
42
|
+
**Radar is keyed on `coinId` — a text `coins.id`, not a ticker and not a UUID.** Resolve it
|
|
43
|
+
through `get_coin_metadata`. Never derive a `coinId` from the player's ticker text, however
|
|
44
|
+
obvious it looks: the identifier the tools take is the coin table's own, and a guessed one either
|
|
45
|
+
misses or hits the wrong coin.
|
|
46
|
+
|
|
47
|
+
Arena deployment is keyed on `presetId` (a UUID) — from `list_game_presets`.
|
|
48
|
+
|
|
49
|
+
Then read what exists:
|
|
50
|
+
|
|
51
|
+
- `get_radar_deployment` (one coin) or `list_radar_deployments` (the fleet, plus the platform
|
|
52
|
+
`radarPaused` kill-switch — if that is set, say so up front: nothing will fire whatever you
|
|
53
|
+
write).
|
|
54
|
+
- `get_deployment_policy` / `list_deployment_policies` for Arena.
|
|
55
|
+
|
|
56
|
+
The read is what supplies `expectedRevision` for the write. **`null` only for a first deploy** —
|
|
57
|
+
a coin or preset with no policy at all. Anything else carries the revision you just read.
|
|
58
|
+
|
|
59
|
+
`get_regime_snapshot` / `get_regime_history` when the policy turns on regime conditions.
|
|
60
|
+
|
|
61
|
+
If the question is *why did my radar agent not fire*, do not start from the journal — **go to step
|
|
62
|
+
5**, which reads state first and pattern second.
|
|
63
|
+
|
|
64
|
+
### 2. Preview the exact parameters under discussion
|
|
65
|
+
|
|
66
|
+
- **Radar** → `preview_radar_resolution` with the draft slots.
|
|
67
|
+
- **Arena** → `preview_deployment_resolution` with the draft slots.
|
|
68
|
+
|
|
69
|
+
Both run the **same resolver the live sweep runs**, so the preview is the real outcome, not an
|
|
70
|
+
estimate. Neither writes anything, neither needs a revision, and neither costs the player an LLM
|
|
71
|
+
call — so previewing repeatedly while iterating is free and correct.
|
|
72
|
+
|
|
73
|
+
Render the preview as a card and read it out: which agent goes on duty, which slot matched and at
|
|
74
|
+
what priority, the regime and conviction used, the qualification verdict, and any typed idle or
|
|
75
|
+
blocked reason. **Render `section` as the server sends it — never re-derive it.** A non-null reason
|
|
76
|
+
does not mean idle (`ON_DUTY_BUT_POSITION_BLOCKED` carries a reason and is not idle), so deciding
|
|
77
|
+
the headline yourself gets it wrong.
|
|
78
|
+
|
|
79
|
+
**If any parameter changes after a preview, re-preview.** A preview never vouches for parameters
|
|
80
|
+
it did not see. That includes a changed conviction bar, an added or removed slot, an edited time
|
|
81
|
+
window or regime set, and flipping `enabled`.
|
|
82
|
+
|
|
83
|
+
### 3. Confirm against the preview, then write
|
|
84
|
+
|
|
85
|
+
One confirmation whose question **references what the preview showed** — the agent it put on duty and
|
|
86
|
+
the reason — and **names what the replacement changes about the stored policy**, read from step 1:
|
|
87
|
+
|
|
88
|
+
- every slot the replacement **removes** (by agent and priority), because `slots` is the whole set;
|
|
89
|
+
- a bar, window or regime that **changes** on a slot that survives;
|
|
90
|
+
- `enabled` going false (**paused, slots kept — nothing fires**) or true (**resumed — it can fire
|
|
91
|
+
again from the next confirmed flip**).
|
|
92
|
+
|
|
93
|
+
Then `upsert_radar_deployment` / `upsert_deployment_policy` with the revision from step 1.
|
|
94
|
+
|
|
95
|
+
**On a typed CONFLICT: re-read, re-preview, re-confirm.** In that order, and all three. The policy
|
|
96
|
+
changed under you, so the state your preview resolved and the radius you stated are both stale.
|
|
97
|
+
Never retry with a bumped revision.
|
|
98
|
+
|
|
99
|
+
### 4. Deletes: read, state what stops, confirm
|
|
100
|
+
|
|
101
|
+
`delete_radar_deployment` / `delete_deployment_policy` remove the **entire** policy — every slot and
|
|
102
|
+
condition — and revoke the standing authority. There is no preview here, and correctly so: there is
|
|
103
|
+
no resolution to preview once the policy is gone. The evidence is **the deployment's own read**.
|
|
104
|
+
|
|
105
|
+
State what stops, from that read: which agents were on duty or eligible, what the policy was
|
|
106
|
+
firing on, and that the slots are not recoverable. Both tools carry a schema-level `confirm: true`.
|
|
107
|
+
|
|
108
|
+
**If the player wants to stop trading without losing the slots, that is not a delete** — it is an
|
|
109
|
+
upsert with `enabled: false`. Offer that whenever the ask sounds like "pause", "stop for now", or
|
|
110
|
+
"take it off duty for a while".
|
|
111
|
+
|
|
112
|
+
### 5. "Why isn't it firing?" — state, then pattern, then rows
|
|
113
|
+
|
|
114
|
+
Three reads, cheapest first. Stop as soon as the player's question is answered.
|
|
115
|
+
|
|
116
|
+
1. **`preview_radar_resolution` → `resolvesNow`** for what is true **right now**: the section, the
|
|
117
|
+
blocked reason and since when, the qualification block, the cooldown, the last fire, and which
|
|
118
|
+
agent is on duty. Most "is it working?" questions end here.
|
|
119
|
+
2. **`get_radar_activity_summary`** for everything else about "why is it quiet", in ONE call. It
|
|
120
|
+
carries three parts and they answer three different questions, each on its own scope:
|
|
121
|
+
- `groups` — which cause recurs and how often, over the window the response names. Quote the
|
|
122
|
+
counts with that window; never sum them across calls.
|
|
123
|
+
- `curveDigest` — how FAR the score sits from firing, over the on-duty agent's ring. This is what
|
|
124
|
+
separates "lower the minimum two points and it fires" from "this strategy does not fit this
|
|
125
|
+
coin": `bestUnqualifiedScorePercent` against `latestThresholdPercent`. Its `ringStartAt` /
|
|
126
|
+
`ringEndAt` describe the ring, NOT the cause window above.
|
|
127
|
+
- `recentEvents` — the last ten key events, lean. Deliberately NOT bounded by the cause window, so
|
|
128
|
+
a pair quiet for longer returns no groups beside populated older rows. That is correct, not a
|
|
129
|
+
contradiction; each row carries its own `occurredAt`.
|
|
130
|
+
|
|
131
|
+
For about one sweep after an agent rotation the digest can name the incoming agent while the rows
|
|
132
|
+
still show the outgoing one — the two halves are read from different stores. Say "just rotated"
|
|
133
|
+
rather than reporting a contradiction.
|
|
134
|
+
3. **`get_radar_activity`** only for what step 2 cannot do: more rows than its ten, a FIRES-only
|
|
135
|
+
view, paging back through history, or ONE occurrence's full margins. Rows are lean by default —
|
|
136
|
+
pass `detail: 'FULL'` for the margin surface, and `includeCurve: true` only if something will
|
|
137
|
+
actually plot the points.
|
|
138
|
+
|
|
139
|
+
**Neither of the first two substitutes for the other, and the reason is structural.** The journal is
|
|
140
|
+
a TRANSITION log, not a state log: a gate that has blocked continuously without crossing again
|
|
141
|
+
inside the window produces no group in the rollup at all. It shows up in `resolvesNow` and nowhere
|
|
142
|
+
else. So an empty or quiet rollup NEVER means "nothing is blocking it" — check step 1 before saying
|
|
143
|
+
anything of the sort.
|
|
144
|
+
|
|
145
|
+
Two negatives, both checkable:
|
|
146
|
+
|
|
147
|
+
- **Never page journal rows in order to count causes yourself.** The counts are served. Re-deriving
|
|
148
|
+
them spends the player's op budget on work the server already did and floods the transcript.
|
|
149
|
+
- **Never sum counts across pages, and never report a windowed count as a lifetime one.** The
|
|
150
|
+
rollup's counts span every matching row inside `windowStartAt`–`windowEndAt` — quote them with
|
|
151
|
+
that window ("41 times in the last 7 days"), never as a total.
|
|
152
|
+
|
|
153
|
+
## Preview-before-commit is a rule of this flow, and only that
|
|
154
|
+
|
|
155
|
+
There is no server-side preview receipt: the upsert takes no token proving a preview happened, and
|
|
156
|
+
the CAS revision is the only cross-call state. So the sequence above is what holds the invariant —
|
|
157
|
+
not a mechanism that could refuse you. Treat it as binding anyway. A write that reaches the player's
|
|
158
|
+
confirm with no preview in the conversation is a failure of this skill even if the server accepts it.
|
|
159
|
+
|
|
160
|
+
## When a write's outcome is unknown
|
|
161
|
+
|
|
162
|
+
An interrupted or timed-out upsert or delete may have landed. **Read the deployment first**
|
|
163
|
+
(`get_radar_deployment` / `get_deployment_policy`) and report what is actually stored. A committed
|
|
164
|
+
change is a success to report, not a call to repeat. Never blind-retry a write whose outcome you
|
|
165
|
+
did not see.
|
|
166
|
+
|
|
167
|
+
## `test_generate_deployment_grid`
|
|
168
|
+
|
|
169
|
+
This one runs a **billed LLM generation** against the player's intelligence credits and writes
|
|
170
|
+
thought and activity records. It is a composition aid for tuning a draft deployment — say that it
|
|
171
|
+
is billed **before** invoking it, and only invoke it when the player is actually iterating on slots
|
|
172
|
+
and wants to see what the resolved agent would produce. It is never a diagnostic read.
|
|
173
|
+
|
|
174
|
+
## Reporting discipline
|
|
175
|
+
|
|
176
|
+
- Report the preview's fields exactly as served — the section, the reason, the verdicts. Never
|
|
177
|
+
recompute or soften them.
|
|
178
|
+
- Radar facts and Arena facts stay separate. Never present one policy's resolution as the other's.
|
|
179
|
+
- A tool that fails is reported as failed, naming what could not be checked.
|
|
180
|
+
- Be concise. Everything written here trades without asking again.
|
|
@@ -0,0 +1,283 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: battlegrid-strategy-authoring
|
|
3
|
+
description: Build a trading strategy with the player from a plain-English idea — gather evidence, lock the spec with them, compile it against the platform grammar, show them exactly what will run, and apply it only once they confirm. Also forks, tunes, restores, archives and previews existing strategies. Activate whenever the player wants to create, change, copy, retire or inspect a strategy, or describes a trading idea they want built.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Strategy Authoring
|
|
7
|
+
|
|
8
|
+
You are turning an idea into a strategy that will trade real money on the player's live agents.
|
|
9
|
+
The whole point of this flow is that **nothing exists until they have seen what will actually
|
|
10
|
+
run and said yes to it.**
|
|
11
|
+
|
|
12
|
+
## The three failures this flow exists to prevent
|
|
13
|
+
|
|
14
|
+
All three come from real authoring sessions, and all three were discovered only after the strategy
|
|
15
|
+
was already built. None of them is allowed to happen here.
|
|
16
|
+
|
|
17
|
+
1. **Built ≠ picked.** The player selected a "≥ +10%" trigger and a +5% trigger was implemented.
|
|
18
|
+
Nobody noticed until the results looked wrong. You prevent this by showing the *compiled*
|
|
19
|
+
rules next to their *locked* picks, before applying, and naming any contradiction yourself.
|
|
20
|
+
2. **Silent zero-trigger.** A strategy was built that could never fire, and produced nothing for
|
|
21
|
+
days before anyone diagnosed it. You prevent this by reading the compile's own preview and the
|
|
22
|
+
per-signal preview before applying, and flagging a plan that shows no passing conditions.
|
|
23
|
+
|
|
24
|
+
3. **Silent substitution.** A player asked for a strategy that triggers on the daily chart and
|
|
25
|
+
executes on the 4-hour. The grammar carries no such semantics. Instead of saying so, the flow
|
|
26
|
+
asked *what they meant by it* and offered four readings — one of which was the nearest
|
|
27
|
+
expressible thing, presented as an equal alternative rather than as a substitute. They picked
|
|
28
|
+
it, and got a confluence strategy labelled as the thing they asked for. **They had no way to
|
|
29
|
+
learn otherwise.** You prevent this by testing expressibility BEFORE the spec lock, and by
|
|
30
|
+
naming the gap in the question itself — see step 2.
|
|
31
|
+
|
|
32
|
+
A strategy that reaches apply without all three checks having been made and reported is a failure
|
|
33
|
+
of this skill, even if the player is happy with it.
|
|
34
|
+
|
|
35
|
+
## Sequence
|
|
36
|
+
|
|
37
|
+
### 1. Evidence first — never propose from the idea alone
|
|
38
|
+
|
|
39
|
+
Before you propose any shape, measure: `get_regime_snapshot` (and `get_regime_history` when the
|
|
40
|
+
idea depends on how we got here), then `get_coin_candles` and `get_coin_performance_history` for
|
|
41
|
+
the coins in scope.
|
|
42
|
+
|
|
43
|
+
If what you fetched contradicts the player's stated direction, do **not** proceed on their
|
|
44
|
+
premise and do **not** silently substitute your own. Put the evidence in front of them and ask
|
|
45
|
+
them to choose the direction again, with your recommendation stated in the question itself. They
|
|
46
|
+
may know something you cannot measure — the point is that they choose with the contradiction
|
|
47
|
+
visible.
|
|
48
|
+
|
|
49
|
+
**Read each thing once.** A strategy, agent or signal log you have already fetched in this
|
|
50
|
+
conversation is still in front of you — do not fetch it again unless something in this
|
|
51
|
+
conversation has changed it. A re-read returns the same bytes, and both copies then ride every
|
|
52
|
+
later step, so the player pays for the same payload twice and keeps paying for it. If you need a
|
|
53
|
+
detail you did not keep, scroll back rather than re-fetching.
|
|
54
|
+
|
|
55
|
+
### 2. Lock the spec before you build anything
|
|
56
|
+
|
|
57
|
+
**First, check the ask is expressible at all — and note that you cannot know until you have looked.**
|
|
58
|
+
Expressibility is a fact about the vocabulary, which step 3 discovers. So when the ask names
|
|
59
|
+
anything the grammar might not carry, **do step 3 before this one** and lock the spec against what
|
|
60
|
+
you found. The cues, none of them subtle:
|
|
61
|
+
|
|
62
|
+
- **two timeframes in one rule** — "trigger on the daily, execute on the 4-hour"
|
|
63
|
+
- **a relation between assets** — "when BTC leads and ETH lags"
|
|
64
|
+
- **an ordering between events** — "after X fires, then wait for Y"
|
|
65
|
+
- **anything phrased as a sequence, a delay, or a dependency**
|
|
66
|
+
|
|
67
|
+
Locking first and discovering second is how the third failure happens: the form goes out while the
|
|
68
|
+
gap is still invisible, so the ask arrives as a product question — *what did you mean?* — when the
|
|
69
|
+
honest answer was *the grammar cannot do that*. The arc reads 1 → 2 → 3 for an ordinary ask; for
|
|
70
|
+
one of these it reads 1 → 3 → 2.
|
|
71
|
+
|
|
72
|
+
If the gap is real, **the question text itself names it**, and the nearest expressible option is
|
|
73
|
+
labelled as the substitute it is — never as one reading among several. A player choosing between
|
|
74
|
+
four equal-looking options cannot tell that none of them is what they asked for.
|
|
75
|
+
|
|
76
|
+
One round of questions to the player, at most five, covering: **direction, trigger definition,
|
|
77
|
+
exit policy, universe, sizing.** Put the one-line reason for each option beside it so they are
|
|
78
|
+
choosing between real alternatives, not guessing.
|
|
79
|
+
|
|
80
|
+
When the picks come back, restate the locked spec in one line before any build call — literally
|
|
81
|
+
"Locked in. Building: …". That line is what the review is later checked against.
|
|
82
|
+
|
|
83
|
+
The answered form is the record of what they picked. Do not write a second copy of the spec
|
|
84
|
+
anywhere; if the two ever disagree, you have created the exact ambiguity this step removes.
|
|
85
|
+
|
|
86
|
+
### 3. Discover the grammar — never guess it
|
|
87
|
+
|
|
88
|
+
`list_strategy_categories` → `list_strategy_vocabulary` → `get_strategy_column_contract` and
|
|
89
|
+
`get_strategy_section_template` → `list_strategy_signals`, plus
|
|
90
|
+
`get_strategy_signal_definition` for each signal you intend to use.
|
|
91
|
+
|
|
92
|
+
Compose sections, columns, conditions and rules **only** from vocabulary returned in this
|
|
93
|
+
conversation. A field you remember from another strategy is not discovery.
|
|
94
|
+
|
|
95
|
+
**Read the answer, not just the call.** Each of these returns one field that decides a composition
|
|
96
|
+
question you would otherwise guess — and a refusal you would otherwise earn:
|
|
97
|
+
|
|
98
|
+
- **A CREATE omits `sectionKey`.** It is derived from the section itself, so the same submitted
|
|
99
|
+
section at the same position yields the same key on every compile. Minting a `custom:<uuid>`
|
|
100
|
+
yourself on a CREATE is refused with a hint telling you to omit `sectionKey` on CREATE; on an
|
|
101
|
+
UPDATE, send back the key the compile returned. This is the one composition field whose right
|
|
102
|
+
answer is "leave it out".
|
|
103
|
+
- **The `entry` axis is required on every CREATE**, all seven keys, no defaults — `trigger`,
|
|
104
|
+
`confirmTf`, `closes`, `bandAtrMultiple`, `levelSource`, `levelOffsetAtrMultiple`,
|
|
105
|
+
`validForBars`. It decides when (and for the level triggers, where) an entry is taken. The
|
|
106
|
+
`strategy-examples` skill carries the vocabulary and the one-directional legality matrix; a
|
|
107
|
+
CREATE without it is refused outright.
|
|
108
|
+
- `get_strategy_column_contract` → `outputs[].conditionOperators`. An empty array means that
|
|
109
|
+
rendered header has no comparison semantics and cannot appear in a condition clause at all.
|
|
110
|
+
Legality is per rendered header, not per column: a trajectory's slot header and its `_trend`
|
|
111
|
+
header answer differently.
|
|
112
|
+
- `get_metric_construction_hints` → `rankOrderings`. Present only when rank is composable on that
|
|
113
|
+
metric, already range-gated server-side — read the offered set rather than deriving one from
|
|
114
|
+
the metric's native output.
|
|
115
|
+
|
|
116
|
+
For full-surface composition patterns — custom and benchmark sections, condition trees with
|
|
117
|
+
verdicts and enforcement gates, weight pyramids and gate math, trade-level and
|
|
118
|
+
position-management presets, worked desk-grade playbooks — activate `strategy-examples`. It
|
|
119
|
+
teaches what to compose; this skill stays the authority on the flow.
|
|
120
|
+
|
|
121
|
+
`derive_strategy_rule_view` belongs here, at composition time, and only here: it reports report
|
|
122
|
+
membership and registry-default allocations for a draft. It does **not** return a compiled plan's
|
|
123
|
+
values, so it can never stand in for the review in step 5.
|
|
124
|
+
|
|
125
|
+
`simulate_aggregate_score` does **not** belong here. It is a review tool (step 5) and running it
|
|
126
|
+
now answers a question about your draft, not about the plan the player will be asked to approve.
|
|
127
|
+
|
|
128
|
+
### 4. Compile, and iterate on the diagnostics
|
|
129
|
+
|
|
130
|
+
**An UPDATE carries only the axes that change.** The server preserves every axis you omit and
|
|
131
|
+
every signal you do not name, and it re-derives the complete post-state, the dense scorecard and
|
|
132
|
+
the diff itself. Restating the whole post-state is never required, changes nothing about the
|
|
133
|
+
result, and the player pays for every byte of it on this call and on every later step of the
|
|
134
|
+
conversation. Send the axes you are changing; send no other axis.
|
|
135
|
+
|
|
136
|
+
**Three fields are not axes, and every compile requires all three:** `intentSummary`,
|
|
137
|
+
`assumptions` and `coinSelection`. They describe *this call*, not the strategy, so "send no other
|
|
138
|
+
axis" never reaches them — omitting one is a typed error, not a saving.
|
|
139
|
+
|
|
140
|
+
`coinSelection` is the cohort the report preview renders over, so the player can read the plan
|
|
141
|
+
against live values. It is non-authoritative, never persisted, and **not recoverable from
|
|
142
|
+
`get_strategy`**: it is not strategy state, and no strategy has one. Do not go looking for it —
|
|
143
|
+
choose it. For a single-gate edit, a short explicit list of the tickers the change is about is
|
|
144
|
+
right; for a broad change, a `ranked` cohort is. Any reasonable cohort is correct, and searching
|
|
145
|
+
for "the strategy's" cohort is a search that cannot end.
|
|
146
|
+
|
|
147
|
+
`compile_strategy_plan` changes no strategy, agent or revision, and a failed compile has cost the
|
|
148
|
+
player nothing. It is not read-only, though: a successful compile parks the plan it approved in
|
|
149
|
+
server-side custody so its own apply can read it back. Compile once per reviewed payload — do not
|
|
150
|
+
retry a compile that succeeded, and never fire two in parallel for the same edit.
|
|
151
|
+
|
|
152
|
+
On a typed authoring error, the response names the offending path, the value it received and the
|
|
153
|
+
domain it allows. Fix the payload and recompile. On advisory `mismatches[]` or a `viability` that
|
|
154
|
+
is not viable, do the same.
|
|
155
|
+
|
|
156
|
+
**Stop after three consecutive autonomous recompiles.** A fourth means you are guessing — show
|
|
157
|
+
the player the diagnostics and ask. Never present a plan for confirmation while it still carries
|
|
158
|
+
diagnostics you have not explained to them.
|
|
159
|
+
|
|
160
|
+
### 5. Review — show the compiled truth, not your summary of it
|
|
161
|
+
|
|
162
|
+
Everything here comes from the compile response you just received. Render:
|
|
163
|
+
|
|
164
|
+
- **What will actually run.** `approvedPlan.postState.signalRules`, `approvedPlan.diff` and the
|
|
165
|
+
review columns — the rules as the server will execute them — beside the picks they locked in
|
|
166
|
+
step 2. **If any compiled rule or diff entry contradicts a locked pick, say so in words before
|
|
167
|
+
you ask for anything.** Do not make them spot it.
|
|
168
|
+
- **Whether it can fire.** `reviewContext.reportPreview` is the server's own preview, rendered
|
|
169
|
+
over this compiled draft against live market. Support it with `get_coin_signal_preview` on the
|
|
170
|
+
locked universe's main coin(s).
|
|
171
|
+
- **A routing what-if**, optionally, via `simulate_aggregate_score` — **after the compile, never
|
|
172
|
+
before it.** This is a calculator, not a verdict: it computes over whatever inputs you hand it,
|
|
173
|
+
so a draft-fed simulation reports on something the player is not being asked to approve. Feed it
|
|
174
|
+
the compiled values — the gate from `approvedPlan.postState.minAggregateScore`, the allocations
|
|
175
|
+
from `approvedPlan.postState.signalRules`, the per-signal scores from the preview you just read
|
|
176
|
+
— and show those inputs beside its output so a copy slip is visible on the card itself.
|
|
177
|
+
|
|
178
|
+
**Flag before confirming** if the preview shows no passing conditions, the coin preview shows
|
|
179
|
+
zero triggered signals, or the simulation reports `wouldRoute: false`. Any one of those means the
|
|
180
|
+
plan probably never fires — ask whether to revise rather than presenting it as healthy.
|
|
181
|
+
|
|
182
|
+
State the blast radius too: the bound-agent count and the open-position observation the compile
|
|
183
|
+
returned, as numbers, not buried in prose.
|
|
184
|
+
|
|
185
|
+
There is no backtest here and no expected-frequency figure. Do not imply one. What you have is a
|
|
186
|
+
point-in-time reading, and you say so.
|
|
187
|
+
|
|
188
|
+
### 6. Confirm, then apply
|
|
189
|
+
|
|
190
|
+
One confirmation carrying the plan's own `confirmationSummary`, offering Apply / Revise / Cancel.
|
|
191
|
+
|
|
192
|
+
On **Apply**, call `apply_strategy_plan` with two values and nothing else:
|
|
193
|
+
|
|
194
|
+
- the compile's `planToken`
|
|
195
|
+
- `confirm: true`
|
|
196
|
+
|
|
197
|
+
There is no `plan` member, and one is rejected as an unrecognized key. The server keeps the plan your
|
|
198
|
+
compile approved and reads it back, so **nothing is copied from the compile response** — the whole
|
|
199
|
+
class of mis-transcription is gone rather than guarded against.
|
|
200
|
+
|
|
201
|
+
`planToken` is opaque: forward it byte-for-byte exactly as received. Never retype, abbreviate or
|
|
202
|
+
reconstruct it. A mangled token addresses no approved plan and is refused with the real one intact.
|
|
203
|
+
|
|
204
|
+
On **Revise**, return to step 4. On **Cancel**, stop and let the token lapse.
|
|
205
|
+
|
|
206
|
+
If the player types free text while the confirm form is open, that is **not** consent and not a
|
|
207
|
+
cancellation. Answer what they said, then present the same plan's confirmation again, unchanged.
|
|
208
|
+
Prose never triggers an apply.
|
|
209
|
+
|
|
210
|
+
Do not pre-check expiry, digests, ownership, viability or quota before calling. The server is the
|
|
211
|
+
only authority on all of them; your job is to react to what it returns.
|
|
212
|
+
|
|
213
|
+
### 7. Lifecycle
|
|
214
|
+
|
|
215
|
+
`fork_strategy`, `update_strategy_signal_rule`, `restore_strategy`, `archive_strategy` and
|
|
216
|
+
`preview_strategy_report` are part of this flow.
|
|
217
|
+
|
|
218
|
+
Before any destructive one, state the blast radius from the server's own fields and confirm it
|
|
219
|
+
with the player:
|
|
220
|
+
|
|
221
|
+
- **Archiving** — how many agents are bound, and that their configuration stays byte-identical
|
|
222
|
+
and open positions are unaffected.
|
|
223
|
+
- **Tuning a single rule** — how many agents are bound, and that the change reaches every one of
|
|
224
|
+
them immediately.
|
|
225
|
+
|
|
226
|
+
## When apply is refused
|
|
227
|
+
|
|
228
|
+
Each of these is a specific typed code. Read it and take the cheapest correct step — never retry
|
|
229
|
+
the same call blindly.
|
|
230
|
+
|
|
231
|
+
- **`TOKEN_EXPIRED`** — a plan token lives five minutes and a human-paced review often outlives
|
|
232
|
+
it. Recompile, present the review again, and **ask for confirmation again.** Never apply a
|
|
233
|
+
recompiled plan on your own judgement that it matches the one they already approved; they
|
|
234
|
+
approve the plan that will actually be applied.
|
|
235
|
+
- **`PLAN_APPROVAL_NOT_FOUND`** — no approved plan answers to this token. It was already applied, it
|
|
236
|
+
lapsed, or it was never issued; these are deliberately one code, because the recovery is the same
|
|
237
|
+
in each case. **Read the committed state first** with `get_strategy` — if the change is already
|
|
238
|
+
there, the apply succeeded and this is a duplicate confirmation, so report success rather than
|
|
239
|
+
rebuilding. Otherwise recompile, re-present, re-confirm.
|
|
240
|
+
- **A refusal that is not about the token at all** — quota, a name collision, a bound agent that
|
|
241
|
+
moved, a shifted catalog. The approved plan **survives** these: clear the cause and confirm again
|
|
242
|
+
with the same token while it still lives. Recompiling works too, but costs the player a second
|
|
243
|
+
review they did not need.
|
|
244
|
+
- **`TOKEN_BINDING_MISMATCH` while the token is still fresh** — there is nothing left to mis-copy,
|
|
245
|
+
so this is real drift: a bound agent moved, the catalog changed, or the signing key rotated. Do
|
|
246
|
+
not resubmit. Recompile, re-present, re-confirm.
|
|
247
|
+
- **An apply whose outcome you never saw** (interrupted or timed out) — read the committed state
|
|
248
|
+
first with `get_strategy` or `list_strategies`. If it committed, report success. **Never
|
|
249
|
+
re-apply and never recompile over an unverified outcome** — a blind retry duplicates the
|
|
250
|
+
strategy or dies on the name check.
|
|
251
|
+
- **A run that ended on a budget guard mid-build** — when they nudge you, recompile and re-enter
|
|
252
|
+
at the review. Never apply a plan compiled in a run that was cut short; treat it as stale.
|
|
253
|
+
|
|
254
|
+
## When the grammar cannot express the ask
|
|
255
|
+
|
|
256
|
+
The gate is step 2, and it is a step rather than advice for a reason: this rule lived here alone
|
|
257
|
+
once, as a section after the sequence, and a real arc walked 1 → 2 → 3 straight past it.
|
|
258
|
+
|
|
259
|
+
Say so, and name the exact capability that is missing — for example multi-timeframe trigger
|
|
260
|
+
semantics the discovered vocabulary does not carry. Then offer the nearest thing it *can* express,
|
|
261
|
+
**labelled as the substitute**, never as one option among equals.
|
|
262
|
+
|
|
263
|
+
Never approximate an inexpressible ask and present it as the thing they asked for. Three things
|
|
264
|
+
make that concrete, and all three have to hold:
|
|
265
|
+
|
|
266
|
+
- the gap is named in the **question you put to the player**, not only in prose around it
|
|
267
|
+
- the nearest expressible option says what it gives up
|
|
268
|
+
- nothing is compiled for the original ask, because there is nothing to compile
|
|
269
|
+
|
|
270
|
+
## Keep the strategy small enough to apply
|
|
271
|
+
|
|
272
|
+
Apply carries only a token, so a large authored surface no longer blocks it. The compiled plan is
|
|
273
|
+
still capped at 256,000 UTF-8 bytes, which compile enforces and reports; splitting a build into a
|
|
274
|
+
lean strategy plus follow-up edits is now a choice about reviewability, not a workaround for a limit
|
|
275
|
+
apply used to impose.
|
|
276
|
+
|
|
277
|
+
## Reporting discipline
|
|
278
|
+
|
|
279
|
+
- Report numbers exactly as the tools return them. Never recompute or re-derive.
|
|
280
|
+
- Compile changes no strategy, agent or revision — it does park the plan its own apply will read.
|
|
281
|
+
Apply is the only write to the strategy itself. Say which one you are about to do.
|
|
282
|
+
- A tool that fails is reported as failed. Never fill a gap with a plausible value.
|
|
283
|
+
- Be concise. They are deciding whether to point real money at this.
|
|
@@ -0,0 +1,142 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: battlegrid-strategy-doctor
|
|
3
|
+
description: Diagnose an agent that is not doing what the player expected — why it has not traded, why it stopped, whether it is actually healthy — from the typed fields that carry health, then rank what would fix it with the exact lever each item needs. Read-only: it explains and recommends, and any change it proposes is applied by the flow that owns it. Activate whenever the player asks why an agent has not traded, why it stopped or is blocked, whether an agent is OK, what is wrong with a strategy's live behaviour, or asks for a check-up or a review of how an agent could be improved.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Strategy Doctor
|
|
7
|
+
|
|
8
|
+
The player configured an agent, pointed real money at it, and something did not happen. Your job is
|
|
9
|
+
to find out what, from the fields that actually say so — and then to say what would change it, in
|
|
10
|
+
terms of levers that exist.
|
|
11
|
+
|
|
12
|
+
Everything here is a **read**. You have no write path, and you do not acquire one by finding a
|
|
13
|
+
problem: the fix runs through the flow that owns it, with that flow's own confirms.
|
|
14
|
+
|
|
15
|
+
## The five failures this flow exists to prevent
|
|
16
|
+
|
|
17
|
+
1. **Diagnosing from prose.** *Cue: reaching for a journal entry's wording, or the agent's
|
|
18
|
+
overlay text, to explain a block.* The typed reason codes exist; use them. → step 2.
|
|
19
|
+
2. **Pathologizing healthy.** *Cue: the reads come back clean and the answer feels too short.* An
|
|
20
|
+
agent that evaluated and correctly declined to trade is working. → step 3.
|
|
21
|
+
3. **Mixing reason vocabularies.** *Cue: two different-looking codes that seem to mean the same
|
|
22
|
+
thing.* Six overlapping vocabularies describe why an evaluation ended. You read exactly one. →
|
|
23
|
+
step 2.
|
|
24
|
+
4. **A private write path.** *Cue: "shall I just fix it?" once you know the answer.* → step 5.
|
|
25
|
+
5. **Running the billed deployment test.** *Cue: wanting to see what the agent "would have
|
|
26
|
+
picked".* → the negatives below.
|
|
27
|
+
|
|
28
|
+
## Sequence
|
|
29
|
+
|
|
30
|
+
### 1. Triage from the fields that carry health
|
|
31
|
+
|
|
32
|
+
Three reads, and they are not interchangeable:
|
|
33
|
+
|
|
34
|
+
- **`get_agent_budget`** — the health read. `haltReason` (why it stopped, if it did),
|
|
35
|
+
`blockedReason` + `blockedSince` (the typed pipeline block currently standing, and since when),
|
|
36
|
+
and `gauges` — four guardrail meters (dailyTrades, exposure, drawdown, dailyLoss), each with a
|
|
37
|
+
server-computed `breached` flag. **Read `breached`; never compare fill against limit yourself.**
|
|
38
|
+
- **`list_gate_blocks`** — its `summary[]` groups every rejection by (stage, reason) with a `count`
|
|
39
|
+
and a `latestAt`. **That `count` spans all of the agent's rows, not the requested page** — never
|
|
40
|
+
sum counts across pages.
|
|
41
|
+
- **`get_agent_automation_status`** — **deployment coverage only.** Its payload is assignments and
|
|
42
|
+
assignable presets; it carries **no health verdict**. Use it to answer "is this agent deployed
|
|
43
|
+
anywhere at all" — an agent with no assignment cannot trade in Arena no matter how healthy the
|
|
44
|
+
rest reads — and never as evidence that automation is or is not "fine".
|
|
45
|
+
|
|
46
|
+
### 2. Diagnose from typed sources
|
|
47
|
+
|
|
48
|
+
- **The one vocabulary is `TradeEvaluationAttemptReasonCode`** — 25 members, two of them
|
|
49
|
+
`@deprecated` and historical-only (`TRADING_MODE_OFF`, `TRADING_MODE_INELIGIBLE`): if either
|
|
50
|
+
turns up, it is an old row, not a live cause, and you say so. Do not translate a code into a
|
|
51
|
+
different enum's wording, and do not try to unify the platform's overlapping reason vocabularies
|
|
52
|
+
— they describe different stages and merging them invents a cause.
|
|
53
|
+
- Render every code through its display meta. A raw `SETUP_GATES` or `OPEN_POSITION_CONFLICT` on
|
|
54
|
+
screen is a system identifier leaking into an explanation.
|
|
55
|
+
- **`get_agent_decision_context` is keyed by COIN**, not by agent. Use it for "why did nothing
|
|
56
|
+
happen on SOL", once you know which coin the blocks are about.
|
|
57
|
+
- **`get_agent_coin_qualification` answers the forward-looking half.** The reason codes above say
|
|
58
|
+
why an agent did NOT trade in the past; this says whether a coin would route for it RIGHT NOW,
|
|
59
|
+
and which gate stops it — candidate levels, aggregate score, required-signal count, the ATR%
|
|
60
|
+
volatility floor and the agent's own required conditions, without spending an LLM call. Read the
|
|
61
|
+
four-member verdict as written: `NOT_ENFORCED` means the agent's own threshold switches the gate
|
|
62
|
+
off, `UNMEASURABLE` means the input was missing and the gate fail-opened. Neither is a pass, and
|
|
63
|
+
reporting either as "cleared" is the conflation the verdict vocabulary exists to prevent.
|
|
64
|
+
- **Gate blocks link to their thought log** through `sourceThoughtLogId` — follow it with
|
|
65
|
+
`get_agent_thought_log` when the block's reason needs the evaluation behind it.
|
|
66
|
+
- `get_signal_performance` / `list_signal_logs` when the question is whether the signals fired, as
|
|
67
|
+
distinct from whether the trades made money. `list_trade_outcomes` and `get_agent_journal` for
|
|
68
|
+
what did happen.
|
|
69
|
+
- **`get_agent_conviction_calibration` honours `readiness`.** An `INSUFFICIENT_DATA` calibration
|
|
70
|
+
carries no win rate — report that there is not yet enough history, never a rate derived from a
|
|
71
|
+
handful of trades.
|
|
72
|
+
|
|
73
|
+
**Every stated cause cites the typed reason or journal entry it came from.** "It is blocked" is not
|
|
74
|
+
a diagnosis; "OPEN_POSITION_CONFLICT, 14 times, most recently 2h ago" is.
|
|
75
|
+
|
|
76
|
+
### 3. Halted agents, and healthy ones
|
|
77
|
+
|
|
78
|
+
**If `haltReason` is set, name the branch that actually clears it:**
|
|
79
|
+
|
|
80
|
+
- **MANUAL** — resume lifts it directly.
|
|
81
|
+
- **DRAWDOWN_BREACH** — clears by raising `maxCumulativeDrawdownUsd` **or** by the drawdown
|
|
82
|
+
baseline reset, then resuming.
|
|
83
|
+
- **DAILY_LOSS** — clears by raising `maxDailyLossUsd` **or** by the UTC-day rollover. **The
|
|
84
|
+
baseline reset cannot clear it. Never offer it here.**
|
|
85
|
+
|
|
86
|
+
In every case, say that a resume attempted while the stop is still breached is **refused by the
|
|
87
|
+
server**, with the current figure against the limit.
|
|
88
|
+
|
|
89
|
+
**If the reads are clean, say so.** An agent whose gauges are unbreached, whose blocks are ordinary
|
|
90
|
+
no-trade verdicts, and whose deployment covers what the player expected, is working — report the
|
|
91
|
+
healthy status and the ordinary reasons and stop. Do not manufacture a finding to have something to
|
|
92
|
+
recommend.
|
|
93
|
+
|
|
94
|
+
### 4. Improvements: ranked, and every one names its lever
|
|
95
|
+
|
|
96
|
+
Present a ranked list. **Each item names the concrete thing that would apply it:**
|
|
97
|
+
|
|
98
|
+
- a **strategy-rule** change → the `strategy-authoring` skill (its compile → review → confirm →
|
|
99
|
+
apply arc);
|
|
100
|
+
- a **deployment or radar policy** change → the `radar-deployment` skill's tools, named;
|
|
101
|
+
- a **risk-limit or budget** change → `agent-management`'s update verb, and note that it is a
|
|
102
|
+
**whole-object** write: the current limits are read and the complete object written back;
|
|
103
|
+
- a **halt recovery** → the specific lever from step 3.
|
|
104
|
+
|
|
105
|
+
An improvement with **no lever on this platform** is labelled as such, explicitly, and never
|
|
106
|
+
presented as actionable. Rank by what would change the observed behaviour most, not by how easy it
|
|
107
|
+
is to say.
|
|
108
|
+
|
|
109
|
+
### 5. Close with the three-option ask
|
|
110
|
+
|
|
111
|
+
One question to the player, offering exactly:
|
|
112
|
+
|
|
113
|
+
1. **Explain only** — the diagnosis stands as the answer.
|
|
114
|
+
2. **Apply via the named tools** — you activate the owning skill and its arc takes over, with its
|
|
115
|
+
own reads, blast-radius statements and confirms.
|
|
116
|
+
3. **Draft the change for review** — you write out what would change, and nothing runs.
|
|
117
|
+
|
|
118
|
+
**Applying routes through the owning flow.** You never write directly, never skip the owning
|
|
119
|
+
flow's confirm, and never treat "apply" as consent for a change that flow would have asked about
|
|
120
|
+
separately.
|
|
121
|
+
|
|
122
|
+
## UNDETERMINED, never "no issue"
|
|
123
|
+
|
|
124
|
+
If a backing tool call fails, returns nothing, or does not cover what is being asked about, report
|
|
125
|
+
that item as **UNDETERMINED** and name the surface you could not check. Never convert an absence of
|
|
126
|
+
data into a clean bill of health, and never let the overall verdict claim completeness when part of
|
|
127
|
+
it is undetermined.
|
|
128
|
+
|
|
129
|
+
## Negatives
|
|
130
|
+
|
|
131
|
+
- **Never call `test_generate_deployment_grid`.** It runs a billed LLM generation against the
|
|
132
|
+
player's intelligence credits and writes thought and activity records. It is the deployment
|
|
133
|
+
flow's composition aid for tuning a draft — it is not a diagnostic read, and running it as one
|
|
134
|
+
charges the player to answer a question the journals already answer.
|
|
135
|
+
- Never diagnose from the agent's own prose or overlay text where a typed field exists.
|
|
136
|
+
- Never present a gate-block count as a rate, a trend, or a percentage — report it as served.
|
|
137
|
+
|
|
138
|
+
## Reporting discipline
|
|
139
|
+
|
|
140
|
+
- Report numbers and codes exactly as served, through their display metas.
|
|
141
|
+
- Worst news first: a halt or a standing block outranks a tuning suggestion.
|
|
142
|
+
- Be concise. The player wants to know what is wrong and what to do about it.
|