@cyanheads/mcp-ts-core 0.12.8 → 0.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +12 -11
- package/CLAUDE.md +12 -11
- package/README.md +2 -2
- package/biome.json +1 -1
- package/changelog/0.12.x/0.12.9.md +36 -0
- package/changelog/0.13.x/0.13.0.md +48 -0
- package/changelog/template.md +7 -24
- package/dist/cli/init.js +2 -2
- package/dist/cli/init.js.map +1 -1
- package/dist/config/envValue.d.ts +18 -0
- package/dist/config/envValue.d.ts.map +1 -0
- package/dist/config/envValue.js +35 -0
- package/dist/config/envValue.js.map +1 -0
- package/dist/config/index.d.ts.map +1 -1
- package/dist/config/index.js +5 -7
- package/dist/config/index.js.map +1 -1
- package/dist/config/parseEnvConfig.d.ts +7 -0
- package/dist/config/parseEnvConfig.d.ts.map +1 -1
- package/dist/config/parseEnvConfig.js +9 -1
- package/dist/config/parseEnvConfig.js.map +1 -1
- package/dist/linter/validate.js +2 -2
- package/dist/linter/validate.js.map +1 -1
- package/dist/mcp-server/transports/http/httpTransport.d.ts.map +1 -1
- package/dist/mcp-server/transports/http/httpTransport.js +70 -2
- package/dist/mcp-server/transports/http/httpTransport.js.map +1 -1
- package/dist/mcp-server/transports/http/landing-page/sections/connect.d.ts.map +1 -1
- package/dist/mcp-server/transports/http/landing-page/sections/connect.js +9 -2
- package/dist/mcp-server/transports/http/landing-page/sections/connect.js.map +1 -1
- package/dist/mcp-server/transports/http/sessionStore.d.ts +10 -2
- package/dist/mcp-server/transports/http/sessionStore.d.ts.map +1 -1
- package/dist/mcp-server/transports/http/sessionStore.js.map +1 -1
- package/dist/services/mirror/sqlite/handle.d.ts.map +1 -1
- package/dist/services/mirror/sqlite/handle.js +14 -12
- package/dist/services/mirror/sqlite/handle.js.map +1 -1
- package/dist/services/mirror/sqlite/sqliteMirrorStore.js +8 -9
- package/dist/services/mirror/sqlite/sqliteMirrorStore.js.map +1 -1
- package/dist/services/mirror/types.d.ts +5 -1
- package/dist/services/mirror/types.d.ts.map +1 -1
- package/dist/utils/internal/performance.d.ts +1 -1
- package/dist/utils/internal/performance.js +2 -2
- package/dist/utils/network/fetchWithTimeout.js +1 -1
- package/dist/utils/network/retry.js +1 -1
- package/dist/utils/security/idGenerator.d.ts +3 -1
- package/dist/utils/security/idGenerator.d.ts.map +1 -1
- package/dist/utils/security/idGenerator.js +12 -1
- package/dist/utils/security/idGenerator.js.map +1 -1
- package/framework-skills/README.md +40 -0
- package/{skills → framework-skills}/add-app-tool/SKILL.md +2 -2
- package/{skills → framework-skills}/add-resource/SKILL.md +2 -2
- package/{skills → framework-skills}/add-service/SKILL.md +2 -2
- package/{skills → framework-skills}/add-test/SKILL.md +2 -2
- package/{skills → framework-skills}/add-tool/SKILL.md +7 -7
- package/{skills → framework-skills}/api-config/SKILL.md +3 -1
- package/{skills → framework-skills}/api-context/SKILL.md +3 -3
- package/{skills → framework-skills}/api-errors/SKILL.md +2 -1
- package/{skills → framework-skills}/api-linter/SKILL.md +4 -4
- package/{skills → framework-skills}/api-mirror/SKILL.md +3 -1
- package/{skills → framework-skills}/design-mcp-server/SKILL.md +59 -101
- package/{skills → framework-skills}/field-test/SKILL.md +10 -5
- package/{skills → framework-skills}/git-wrapup/SKILL.md +5 -3
- package/{skills → framework-skills}/maintenance/SKILL.md +30 -21
- package/{skills → framework-skills}/orchestrations/SKILL.md +2 -2
- package/{skills → framework-skills}/orchestrations/workflows/field-test-fix.md +8 -8
- package/{skills → framework-skills}/orchestrations/workflows/fix-wrapup-release.md +5 -5
- package/{skills → framework-skills}/orchestrations/workflows/greenfield-build.md +11 -11
- package/{skills → framework-skills}/orchestrations/workflows/maintenance-release.md +12 -12
- package/{skills → framework-skills}/polish-docs-meta/SKILL.md +18 -10
- package/{skills → framework-skills}/polish-docs-meta/references/agent-protocol.md +1 -1
- package/{skills → framework-skills}/polish-docs-meta/references/package-meta.md +1 -1
- package/{skills → framework-skills}/polish-docs-meta/references/readme.md +88 -72
- package/{skills → framework-skills}/release-and-publish/SKILL.md +4 -1
- package/{skills → framework-skills}/release-pr-review/SKILL.md +2 -2
- package/{skills → framework-skills}/report-issue-framework/SKILL.md +25 -25
- package/{skills → framework-skills}/report-issue-local/SKILL.md +22 -24
- package/{skills → framework-skills}/security-pass/SKILL.md +2 -2
- package/{skills → framework-skills}/setup/SKILL.md +10 -8
- package/package.json +13 -13
- package/scripts/check-framework-antipatterns.ts +1 -1
- package/scripts/check-skill-versions.ts +16 -9
- package/scripts/check-skills-sync.ts +64 -13
- package/scripts/clean-mcpb.ts +3 -3
- package/scripts/devcheck.ts +37 -27
- package/scripts/lint-packaging.ts +158 -24
- package/scripts/list-skills.ts +2 -2
- package/templates/.claude-plugin/plugin.json +5 -1
- package/templates/.env.example +1 -1
- package/templates/.github/CONTRIBUTING.md +4 -5
- package/templates/.github/ISSUE_TEMPLATE/bug_report.yml +5 -4
- package/templates/.github/ISSUE_TEMPLATE/config.yml +6 -1
- package/templates/.github/ISSUE_TEMPLATE/feature_request.yml +1 -2
- package/templates/AGENTS.md +16 -15
- package/templates/CLAUDE.md +16 -15
- package/templates/_.mcpbignore +1 -1
- package/templates/changelog/template.md +7 -24
- package/templates/package.json +4 -3
- package/templates/src/mcp-server/resources/definitions/echo-app-ui.app-resource.ts +1 -1
- package/skills/README.md +0 -38
- /package/{skills → framework-skills}/add-export/SKILL.md +0 -0
- /package/{skills → framework-skills}/add-prompt/SKILL.md +0 -0
- /package/{skills → framework-skills}/add-provider/SKILL.md +0 -0
- /package/{skills → framework-skills}/api-auth/SKILL.md +0 -0
- /package/{skills → framework-skills}/api-canvas/SKILL.md +0 -0
- /package/{skills → framework-skills}/api-services/SKILL.md +0 -0
- /package/{skills → framework-skills}/api-services/references/graph.md +0 -0
- /package/{skills → framework-skills}/api-services/references/llm.md +0 -0
- /package/{skills → framework-skills}/api-services/references/speech.md +0 -0
- /package/{skills → framework-skills}/api-telemetry/SKILL.md +0 -0
- /package/{skills → framework-skills}/api-testing/SKILL.md +0 -0
- /package/{skills → framework-skills}/api-utils/SKILL.md +0 -0
- /package/{skills → framework-skills}/api-utils/references/formatting.md +0 -0
- /package/{skills → framework-skills}/api-utils/references/parsing.md +0 -0
- /package/{skills → framework-skills}/api-utils/references/security.md +0 -0
- /package/{skills → framework-skills}/api-workers/SKILL.md +0 -0
- /package/{skills → framework-skills}/code-simplifier/SKILL.md +0 -0
- /package/{skills → framework-skills}/polish-docs-meta/references/server-json.md +0 -0
- /package/{skills → framework-skills}/techniques/SKILL.md +0 -0
- /package/{skills → framework-skills}/techniques/references/outline-on-overflow.md +0 -0
- /package/{skills → framework-skills}/tool-defs-analysis/SKILL.md +0 -0
|
@@ -4,7 +4,7 @@ description: >
|
|
|
4
4
|
Design the tool surface, resources, and service layer for a new MCP server. Use when starting a new server, planning a major feature expansion, or when the user describes a domain/API they want to expose via MCP. Produces a design doc at docs/design.md that drives implementation.
|
|
5
5
|
metadata:
|
|
6
6
|
author: cyanheads
|
|
7
|
-
version: "2.
|
|
7
|
+
version: "2.25"
|
|
8
8
|
audience: external
|
|
9
9
|
type: workflow
|
|
10
10
|
---
|
|
@@ -14,9 +14,9 @@ metadata:
|
|
|
14
14
|
- User says "I want to build a ___ MCP server"
|
|
15
15
|
- User has an API, database, or system they want to expose to LLMs
|
|
16
16
|
- User wants to plan tools before scaffolding
|
|
17
|
-
- Existing server needs a new capability area (design the addition, not just a single tool)
|
|
17
|
+
- Existing server needs a new capability area — two or more tools sharing a new noun or service, or any new upstream source (design the addition, not just a single tool)
|
|
18
18
|
|
|
19
|
-
Do NOT use for single
|
|
19
|
+
Do NOT use for a single tool on an existing noun — use `add-tool` directly.
|
|
20
20
|
|
|
21
21
|
## Inputs
|
|
22
22
|
|
|
@@ -42,19 +42,13 @@ The unit of a server is a *user workflow*, not an API. A single rich API can ear
|
|
|
42
42
|
| Single API with rich surface, large audience | Standalone server named for the platform (`pubmed-mcp-server`, `secedgar-mcp-server`) |
|
|
43
43
|
| Multiple APIs serving the same workflow | One server named for the workflow (`threat-intel-mcp-server`), APIs are internal sources |
|
|
44
44
|
| Domain with distinct sub-audiences | Consider splitting — a pentester and a SOC analyst have different workflows even in the same domain |
|
|
45
|
-
| Pure computation, no external deps | Standalone server named for the capability (`calculator-mcp-server`, `
|
|
45
|
+
| Pure computation, no external deps | Standalone server named for the capability (`calculator-mcp-server`, `pentest-mcp-server`) |
|
|
46
46
|
|
|
47
47
|
When multiple APIs collapse into one server, the tool surface is organized around what the user is doing, not which API gets called. The agent says "investigate this domain" and the server routes to the best available source internally. Individual APIs become service-layer implementation details, not tool-surface identities.
|
|
48
48
|
|
|
49
49
|
## Server Naming
|
|
50
50
|
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
- **Use the canonical platform/brand name, not abbreviations.** `libofcongress-mcp-server` not `loc-mcp-server` ("loc" reads as lines-of-code or location). `federal-reserve-mcp-server` not `fred-mcp-server` ("fred" reads as a person's name).
|
|
54
|
-
- **Add a descriptive suffix when the base name is a non-obvious acronym.** Pattern: `{acronym}-{domain}-mcp-server` — e.g., `eia-energy-mcp-server`, `bls-labor-mcp-server`, `nhtsa-vehicle-safety-mcp-server`. Skip when the name is already self-descriptive (`earthquake-mcp-server`, `wikidata-mcp-server`).
|
|
55
|
-
- **Don't overclaim scope.** A name asserts breadth on two independent axes, and must be honest on both. *Source breadth:* is this a first-party wrapper around one provider's API, or genuine aggregation across several independent sources behind one normalized surface? *Domain or jurisdiction breadth:* one, or many? A generic name claims breadth on whichever axis it leaves unqualified. `threat-intel-mcp-server` earns its generic name by aggregating independent sources for one workflow; that same name over a single vendor's API would be a defect — name that for the vendor. When the real scope is one jurisdiction, put it in the name (`uk-legislation-mcp-server`, `statistics-canada-mcp-server`) rather than letting a generic noun imply worldwide coverage a user will only discover is absent after installing. Note that aggregating several bodies *within* one jurisdiction earns a cross-source name, not a cross-jurisdiction one.
|
|
56
|
-
- **Aggregation is a claim about entities, not endpoints.** One provider publishing five APIs is still first-party. Wrapping several endpoints of the same upstream does not make a server an aggregator and does not earn a generic name.
|
|
57
|
-
- **The tool prefix derives from the name but is a separate identifier.** Every tool is `{prefix}_{verb}_{noun}`, so the prefix shows up in every tool call an agent sees, and a descriptive one gives agents domain context without reading the server's instructions. The prefix names the *source*, dropping the descriptive qualifier the repo name carries for human browsing: `eia-energy-mcp-server` → `eia_`, `nhtsa-vehicle-safety-mcp-server` → `nhtsa_`. Because it names the source rather than the repo, the two can diverge — a later repo rename does not have to move the prefix, and usually shouldn't: renaming an advertised tool surface is a breaking change for every existing client, while renaming a package is not.
|
|
51
|
+
Usually settled before this skill runs; confirm it passes one test before designing against it. A name is earned when someone reading only the name — an npm result, a marketplace grid — forms no wrong expectation about scope. Name a wrapper for its source (`pubmed-mcp-server`, `secedgar-mcp-server`); a generic domain name is earned only by aggregating independent sources (`threat-intel-mcp-server`), and a single-jurisdiction scope goes in the name (`uk-legislation-mcp-server`). Spell out an acronym that reads as something else (`libofcongress`, not `loc`), or pair it with its domain (`eia-energy`, `bls-labor`). The tool prefix is a separate, stable identifier (see the Design table) — it need not move when the package is renamed, and renaming it is a breaking change for every client.
|
|
58
52
|
|
|
59
53
|
## Steps
|
|
60
54
|
|
|
@@ -109,7 +103,7 @@ The user-goal list shapes the tool surface; the operation list fills in the gaps
|
|
|
109
103
|
|
|
110
104
|
### 3. Classify into MCP Primitives
|
|
111
105
|
|
|
112
|
-
**Tools are the primary interface.** Not all MCP clients expose resources
|
|
106
|
+
**Tools are the primary interface.** Not all MCP clients expose resources, and those that do rarely surface one to the model without a human selecting it. Design the tool surface to be self-sufficient: an agent with only tool access should be able to do everything the server is built for. Resources add convenience for clients that support them (injectable context, stable URIs), but are not a reliable access path.
|
|
113
107
|
|
|
114
108
|
| Primitive | Use when | Examples |
|
|
115
109
|
|:----------|:---------|:--------|
|
|
@@ -117,9 +111,10 @@ The user-goal list shapes the tool surface; the operation list fills in the gaps
|
|
|
117
111
|
| **App Tool** | **Rare — default to a standard tool.** Only when a human will actively interact with the result in real time *and* the target client supports MCP Apps. Most clients are tool-only and most agent workflows are read-by-LLM, not viewed-by-human. App tools add an iframe + CSP, `app.ontoolresult`/`callServerTool` plumbing, host-context wiring, and a `format()` text twin that still has to be content-complete (since most clients only see that). Two surfaces to keep in sync, two failure modes per change. | Dense tabular state a human scrubs through; form-based human approval in an MCP Apps-capable client |
|
|
118
112
|
| **Resource** | *Additionally* expose as a resource when the data is addressable by stable URI, read-only, and useful as injectable context. | Config, schemas, status, entity-by-ID lookups |
|
|
119
113
|
| **Prompt** | Reusable message template that structures how the LLM approaches a task | Analysis framework, report template, review checklist |
|
|
114
|
+
| **Client round-trip** | Not a registered primitive — a handler *returns* `ctx.requestInput(...)` to ask the client for what only it has, and is re-entered with the answer: a confirmation or form (`inputRequired.elicit`), an authorization or hosted-form URL (`inputRequired.elicitUrl`), the client model's judgment (`inputRequired.createMessage` — borrow the caller's model rather than bundling one), filesystem roots (`inputRequired.listRoots`). Design it into the tool that needs it; see Workflow tool safety and `api-context`. | Destructive-arm confirmation, OAuth consent, summarize-with-the-client's-model |
|
|
120
115
|
| **Neither** | Internal detail, admin-only, not useful to an LLM | Token refresh, webhook setup, migrations |
|
|
121
116
|
|
|
122
|
-
What the tool surface needs to cover depends on the server: a read-only research server has different economics than a CRUD project management server. Consider the domain, the expected agent workflows, whether it wraps one API or many, and what data relationships exist.
|
|
117
|
+
What the tool surface needs to cover depends on the server: a read-only research server has different economics than a CRUD project management server. Consider the domain, the expected agent workflows, whether it wraps one API or many, and what data relationships exist.
|
|
123
118
|
|
|
124
119
|
**Common traps:**
|
|
125
120
|
|
|
@@ -141,11 +136,11 @@ Most tools follow the `{server}_{verb}_{noun}` default — one focused responsib
|
|
|
141
136
|
|
|
142
137
|
| Shape | Purpose | Typical form | Examples |
|
|
143
138
|
|:------|:--------|:-------------|:---------|
|
|
144
|
-
| **Workflow** | Multi-step orchestration that replaces a common agent chain | N upstream calls (often parallelized); may request confirmation; may need mid-flow cleanup | `
|
|
139
|
+
| **Workflow** | Multi-step orchestration that replaces a common agent chain | N upstream calls (often parallelized); may request confirmation; may need mid-flow cleanup | `clinicaltrials_find_eligible` (search → filter → rank) |
|
|
145
140
|
| **Instruction** | State-aware procedural guidance — advice, not action | Static markdown + a few live-state fetches, `readOnlyHint: true`, outputs `nextToolSuggestions` pre-filling the recommended follow-up. No writes. | `git_wrapup_instructions` |
|
|
146
141
|
| **Reference** | Decode opaque domain vocabulary — codes, enums, identifier formats, coverage windows — so agents can build valid inputs for the rest of the surface | Static tables or one cached fetch (often zero upstream calls); consolidate N lists under one `topic` enum; `readOnlyHint: true`, `openWorldHint: false` when offline | `medcode_list_systems`, `osv_list_ecosystems` |
|
|
147
142
|
|
|
148
|
-
These aren't boxes every tool must fit into — some blend shapes — but the design pressures differ enough that naming them helps avoid re-discovering the patterns per server.
|
|
143
|
+
These aren't boxes every tool must fit into — some blend shapes — but the design pressures differ enough that naming them helps avoid re-discovering the patterns per server.
|
|
149
144
|
|
|
150
145
|
#### Think in workflows, not endpoints
|
|
151
146
|
|
|
@@ -155,38 +150,23 @@ A single tool can call multiple APIs internally, apply local filtering, reshape
|
|
|
155
150
|
|
|
156
151
|
```ts
|
|
157
152
|
// Workflow tool — search + local filter pipeline, not a raw API proxy
|
|
158
|
-
const
|
|
159
|
-
description: '
|
|
153
|
+
const findEligible = tool('clinicaltrials_find_eligible', {
|
|
154
|
+
description: 'Match a patient profile to eligible clinical trials, filtering by age, sex, conditions, location, and healthy volunteer status. Results are ranked and carry a per-study eligibility explanation.',
|
|
160
155
|
// handler: listStudies() → filter by eligibility → rank by location proximity → slice
|
|
161
156
|
});
|
|
162
157
|
```
|
|
163
158
|
|
|
164
|
-
> **Tip — mode consolidation.** When a tool has several related operations on the same noun, you can consolidate them under one tool with a `mode`/`operation` enum. This affects both naming (noun-led, e.g., `github_pull_request`) and handler design (dispatch by mode). Use when it tightens the surface; skip when ops diverge enough to warrant separate tools.
|
|
159
|
+
> **Tip — mode consolidation.** When a tool has several related operations on the same noun, you can consolidate them under one tool with a `mode`/`operation` enum. This affects both naming (noun-led, e.g., `github_pull_request`) and handler design (dispatch by mode). Use when it tightens the surface; skip when ops diverge enough to warrant separate tools. When the arms need *different* required fields (look up by ID vs. search by name), declare the input as `z.discriminatedUnion('mode', [...])` rather than making every field optional and checking the combination by hand — each arm advertises its own `required`, the handler narrows on the discriminator, and mixed arguments are rejected. Two constraints: `output` stays a flat `z.object`, and a union root rules out `headerParam`. See `add-tool` § *Multi-mode tools*.
|
|
165
160
|
|
|
166
161
|
#### Multi-source tools and fallback chains
|
|
167
162
|
|
|
168
163
|
**Applies when:** a server aggregates multiple data sources for the same workflow, and the "best" source varies by input type, availability, or coverage. Skip for single-API servers.
|
|
169
164
|
|
|
170
|
-
When a tool's goal can be served by multiple sources, design it as a **multi-source tool** — the agent calls one tool, the handler routes to the best source (or fans out to several) internally. This is the difference between a "PubMed wrapper" and a "literature research server": `
|
|
165
|
+
When a tool's goal can be served by multiple sources, design it as a **multi-source tool** — the agent calls one tool, the handler routes to the best source (or fans out to several) internally. This is the difference between a "PubMed wrapper" and a "literature research server": a hypothetical `literature_search_articles` tries PubMed first, falls back to EuropePMC for broader coverage, then Unpaywall for open access. The agent doesn't choose which API to hit — the server makes that decision based on what works.
|
|
171
166
|
|
|
172
167
|
Two patterns:
|
|
173
168
|
|
|
174
|
-
**Source fallback chains** — try sources in priority order, fall through on failure or empty results. Best when sources cover the same
|
|
175
|
-
|
|
176
|
-
```ts
|
|
177
|
-
// Handler pseudocode — not a real implementation
|
|
178
|
-
async handler(input, ctx) {
|
|
179
|
-
// Primary: PubMed E-utilities (authoritative, best metadata)
|
|
180
|
-
const result = await pubmedService.search(input.query);
|
|
181
|
-
if (result.items.length > 0) return { ...result, source: 'pubmed' };
|
|
182
|
-
|
|
183
|
-
// Fallback: EuropePMC (broader coverage, includes preprints)
|
|
184
|
-
const epmcResult = await epmcService.search(input.query);
|
|
185
|
-
if (epmcResult.items.length > 0) return { ...epmcResult, source: 'europepmc' };
|
|
186
|
-
|
|
187
|
-
return { items: [], source: 'none', message: 'No results from any source.' };
|
|
188
|
-
}
|
|
189
|
-
```
|
|
169
|
+
**Source fallback chains** — try sources in priority order, fall through on failure or empty results. Best when sources cover the *same corpus* with different depth or availability. The output should indicate which source provided the data so the agent (and human) can assess provenance. When the fallback changes what is being searched — a different corpus, different identifiers, different licensing — don't chain: expose the second source as a sibling tool so the agent chooses the corpus knowingly (the shipped `pubmed-mcp-server` keeps `pubmed_europepmc_search` separate for exactly this reason).
|
|
190
170
|
|
|
191
171
|
**Multi-source fan-out** — query multiple sources in parallel, merge results. Best when sources provide complementary data about the same entity. Use `Promise.allSettled` so one failing source doesn't tank the whole call.
|
|
192
172
|
|
|
@@ -213,11 +193,9 @@ async handler(input, ctx) {
|
|
|
213
193
|
|
|
214
194
|
In both patterns, the tool surface is organized around what the user is doing. Sources are service-layer details — the agent sees `threat_enrich_indicator`, not `virustotal_lookup` + `abuseipdb_check` + `greynoise_query`. Mode-based dispatch by input type (e.g., `indicator_type: 'ip' | 'domain' | 'hash'`) naturally routes to different source chains per mode, since different sources cover different indicator types.
|
|
215
195
|
|
|
216
|
-
There is no fixed ceiling on tool count — tools need to earn their keep, but don't artificially limit the surface. If the domain genuinely has 20 distinct workflows, expose 20 tools.
|
|
217
|
-
|
|
218
196
|
#### Cut the surface
|
|
219
197
|
|
|
220
|
-
After mapping tools, review the full list critically. A tool that covers a niche use case, serves a tiny fraction of agents, or duplicates what another tool already handles is a candidate for deferral. Drop it from the design and note it as a future addition if demand warrants. Every tool in the surface is cognitive load for tool selection — a tight surface outperforms a comprehensive one.
|
|
198
|
+
There is no fixed ceiling on tool count and no target either. The ceiling is workflow coverage — if the domain genuinely has 20 distinct workflows, expose 20 tools; the cut is per-tool overlap and reach. After mapping tools, review the full list critically. A tool that covers a niche use case, serves a tiny fraction of agents, or duplicates what another tool already handles is a candidate for deferral. Drop it from the design and note it as a future addition if demand warrants. Every tool in the surface is cognitive load for tool selection — a tight surface outperforms a comprehensive one.
|
|
221
199
|
|
|
222
200
|
#### Instruction tools
|
|
223
201
|
|
|
@@ -229,13 +207,13 @@ Characteristics:
|
|
|
229
207
|
|
|
230
208
|
- **Output is markdown guidance**, not structured data (though the output schema still has fields — typically `guidance`, `diagnostics`, and `nextToolSuggestions`)
|
|
231
209
|
- **Merges static procedural content with live state** — the value is the tailoring. "You have 12 staged files spanning 4 unrelated changes — split them into separate commits before pushing" beats a generic best-practices article. The same shape works in other domains: "Your slowest query is 2.3s on `orders.customer_id` — add the index before tuning the planner" (database advisor), "Error rate spiked 4× at 14:32 UTC, 4 minutes after the `web@a3f9c2` deploy — roll back before chasing the upstream provider" (incident triage).
|
|
232
|
-
- **`readOnlyHint: true
|
|
210
|
+
- **`readOnlyHint: true`; `openWorldHint` follows where the diagnostics come from** — `false` when the live state is local (a repo on disk), `true` when it is fetched from an external API. No writes either way.
|
|
233
211
|
- **Outputs `nextToolSuggestions`** — an array of recommended follow-up tool calls with arguments **pre-filled** from the diagnostics, not just tool names. The agent consumes the playbook, then executes steps with other tools.
|
|
234
212
|
- **Consolidate by `topic` enum** — what could be N separate per-topic tools collapses into one
|
|
235
213
|
|
|
236
214
|
```ts
|
|
237
215
|
const wrapupInstructions = tool('git_wrapup_instructions', {
|
|
238
|
-
description: '
|
|
216
|
+
description: 'Get procedural guidance tailored to the current repo state: best-practice markdown merged with live diagnostics (staged/unstaged files, branch info, recent commits) and pre-filled follow-up tool calls. Read-only; execute the steps with other tools.',
|
|
239
217
|
annotations: { readOnlyHint: true, openWorldHint: false },
|
|
240
218
|
input: z.object({
|
|
241
219
|
topic: z.enum(['review-changes', 'stage-and-commit', 'push-to-remote'])
|
|
@@ -273,42 +251,17 @@ Tools that perform multi-step mutations (the Workflow shape) have two safety con
|
|
|
273
251
|
|
|
274
252
|
**Confirmation-gated destructive modes, with an annotation fallback.** When a workflow's `mode` parameter switches between safe and destructive arms (`draft` vs `send`, `plan` vs `apply`), gate the destructive arm on a confirmation the handler asks for via `ctx.requestInput(...)`, so a human approves before the irreversible step fires. The handler is re-entered with the answer on `ctx.inputs`; it does not `await` mid-call.
|
|
275
253
|
|
|
276
|
-
The gate is always *reachable* — `ctx.requestInput` is present on every transport and both protocol
|
|
277
|
-
|
|
278
|
-
```ts
|
|
279
|
-
annotations: { destructiveHint: true }, // client-side approval flows still see the risk
|
|
280
|
-
// ...
|
|
281
|
-
const Confirm = z.object({ confirmed: z.literal(true).describe('Type true to apply.') });
|
|
282
|
-
|
|
283
|
-
handler(input, ctx) {
|
|
284
|
-
if (input.mode === 'apply') {
|
|
285
|
-
// A decline is terminal — re-asking would loop until the round budget runs out.
|
|
286
|
-
const view = ctx.inputs.view('confirm');
|
|
287
|
-
if (view.kind === 'elicit' && view.action !== 'accept') {
|
|
288
|
-
throw validationError('Migration cancelled by user.');
|
|
289
|
-
}
|
|
290
|
-
if (!ctx.inputs.accepted('confirm', Confirm)) {
|
|
291
|
-
return ctx.requestInput({
|
|
292
|
-
inputRequests: {
|
|
293
|
-
confirm: inputRequired.elicit({
|
|
294
|
-
message: `Apply migration affecting ${affectedRowCount} rows in production? Cannot be rolled back automatically.`,
|
|
295
|
-
requestedSchema: Confirm,
|
|
296
|
-
}),
|
|
297
|
-
},
|
|
298
|
-
});
|
|
299
|
-
}
|
|
300
|
-
}
|
|
301
|
-
// destructive step proceeds
|
|
302
|
-
}
|
|
303
|
-
```
|
|
254
|
+
The gate is always *reachable* — `ctx.requestInput` is present on every transport and both protocol revisions (2025-11-25 legacy, 2026-07-28 current) — but it is not always *answerable*: a client that never fulfils the `input_required` result simply doesn't retry, and the destructive step never runs. The same holds for a 2025-11-25 HTTP client when the server runs `MCP_SESSION_MODE=stateless`, which disables the legacy round-trip shim — the gate refuses and the destructive step never fires. That is the safe outcome, but it makes the tool unusable for those clients, so weigh it before defaulting such a server to `stateless` (`api-context` § `ctx.requestInput`). Keep `destructiveHint: true` in annotations so those clients' own approval flows still surface the risk. A decline is terminal — the handler fails the call rather than re-asking, which would loop until the round budget runs out. The handler shape is in `api-context` § *The shape of a multi-round-trip handler*.
|
|
304
255
|
|
|
305
256
|
**Safe defaults on parameters that determine blast radius.** When a workflow accepts a parameter that controls how far-reaching a mutation is, default to the safer value. A bulk file-update tool defaulting `mode: 'preview'` (no writes) means a sloppy agent call shows a diff rather than blasting changes; an apply-plan tool defaulting `dryRun: true` means a misread plan previews rather than executes; an object-delete tool requiring an explicit `confirmCount` matching the result-set size means an unscoped query can't silently nuke a million rows. Agents that genuinely want the destructive behavior have to name it explicitly, which surfaces intent in the tool call and in logs.
|
|
306
257
|
|
|
258
|
+
**Make retried writes safe.** Agents re-issue a call that timed out. A write that can double-apply (create, send, charge, enqueue) takes a caller-supplied idempotency key or resolves to upsert semantics — and only then earns `idempotentHint: true`.
|
|
259
|
+
|
|
307
260
|
#### Tool descriptions
|
|
308
261
|
|
|
309
262
|
The description is the LLM's primary signal for tool selection. It must answer: *what does this do, and when should I use it?*
|
|
310
263
|
|
|
311
|
-
- **
|
|
264
|
+
- **Imperative present tense, capability first.** "Search for clinical trial studies using queries and filters" beats "Interact with studies", and beats "Searches for…", "This tool…", "Allows you to…". The `tool-defs-analysis` skill audits this across a finished surface.
|
|
312
265
|
- **Include operational guidance when it matters.** If the tool has prerequisites, constraints, or gotchas the LLM needs to know, say so in the description. Don't add boilerplate workflow hints when the tool is self-explanatory.
|
|
313
266
|
- **Prefer a single cohesive paragraph.** Pack operational guidance into prose sentences (separated by periods or em-dashes) rather than bullet lists or blank-line-separated sections. Descriptions render inline in most clients, and bullet structure reads as visual noise rather than signal. Operation-by-operation bullets also duplicate info that already lives in the `operation` enum's `.describe()`.
|
|
314
267
|
- **Don't leak.** Descriptions are for the consumer, not the author. Three categories to audit against:
|
|
@@ -324,7 +277,7 @@ description: 'Set the session working directory for all git operations. This all
|
|
|
324
277
|
description: 'Show the working tree status including staged, unstaged, and untracked files.'
|
|
325
278
|
|
|
326
279
|
// Good — warns about constraints
|
|
327
|
-
description: '
|
|
280
|
+
description: 'Fetch trial results data for completed studies. Only available for studies where hasResults is true.'
|
|
328
281
|
```
|
|
329
282
|
|
|
330
283
|
Descriptions should be as long as needed — concise but complete. Don't artificially truncate, and don't pad with filler.
|
|
@@ -334,6 +287,7 @@ Descriptions should be as long as needed — concise but complete. Don't artific
|
|
|
334
287
|
Every `.describe()` is prompt text the LLM reads. Parameters should convey: what the value is, what it affects, and (where non-obvious) how to use it well.
|
|
335
288
|
|
|
336
289
|
- **Constrain the type.** Enums and literals over free strings. Regex validation for formatted IDs. Ranges for numeric bounds.
|
|
290
|
+
- **The input root is already strict.** `tool()` applies `.strict()` at the root and advertises `additionalProperties: false`, so an unknown top-level key is rejected by name instead of silently stripped; nested objects still strip unless made strict themselves. Declare `.passthrough()` only on a tool that deliberately proxies arbitrary upstream parameters (a raw-query tool), and say so in its description.
|
|
337
291
|
- **Use JSON-Schema-serializable types only.** The MCP SDK serializes schemas to JSON Schema for `tools/list`. Types like `z.custom()`, `z.date()`, `z.transform()`, `z.bigint()`, `z.symbol()`, `z.void()`, `z.map()`, `z.set()` throw at runtime. Use structural equivalents (e.g., `z.string().describe('ISO 8601 date')` instead of `z.date()`).
|
|
338
292
|
- **Explain costs and tradeoffs** when a parameter choice has meaningful consequences.
|
|
339
293
|
- **Name alternative approaches** when a simpler path exists.
|
|
@@ -362,27 +316,19 @@ The output schema and `format` function control what the LLM reads back. Design
|
|
|
362
316
|
- **Server reports what only the server can know; agent decides what only the agent can know.** Schema, scopes, rate limits, and raw observable state belong to the server. Semantic correctness, intent-vs-effect matching, and recovery choice belong to the agent. For mutators, this means surfacing pre/post observable state rather than throwing on synthetic deltas the server can't authoritatively classify — `file shrunk` could be deliberate truncation or a bug; only the agent knows. See `add-tool` skill's **Mutator response design**.
|
|
363
317
|
- **Include IDs and references for chaining.** If the agent might act on a result, return the identifiers it needs for follow-up tool calls.
|
|
364
318
|
- **Curate vs. pass-through depends on domain.** Medical/scientific data — don't trim fields that could alter correctness. CRUD responses — return what the agent needs, not the full API payload. Match fidelity to consequence.
|
|
319
|
+
- **Absent upstream data stays absent.** Sparse APIs omit fields; declare those output fields `.optional()` and render "Not available" in `format()` rather than coercing to `false`, `0`, or `""` — a fabricated value is worse than a gap.
|
|
320
|
+
- **Image and audio bytes ride `ctx.content`, never `output`.** `ctx.content.image(data, mimeType)` / `.audio(...)` emit a `content[]` block once; `output` keeps the metadata the agent reasons over (dimensions, duration, a reference). Base64 in a typed output field ships the bytes twice.
|
|
365
321
|
- **Surface what was done, not just results.** After a write operation, include the post-state so the LLM can chain without an extra round trip.
|
|
322
|
+
- **When the effect lands after the call, wait for it by default.** Some actions are fire-and-forget at the wire — send a wake packet, trigger a job, dispatch a notification, provision a resource — and their immediate result ("sent", "queued", "accepted") answers nothing the agent asked; the agent wants to know whether the machine is up, the job ran, the resource exists. Design the tool to confirm: pre-probe the observable state (cheap; if it is already in the target state, skip the action and say so), act, then poll with early return until the state is observed or a bounded window elapses. Express the window as **one numeric parameter with a default** (`wait_for_s: 30`), where `0` means act and return — not a boolean plus a timeout, which is two parameters for one decision and an awkward default. Every terminal outcome is a *result*, never a throw: `already_<state>`, `<state>` (with elapsed time), `not_<state>` within the window (with `guidance` naming the read-only check tool to re-poll), and `unverified` (window `0`, or nothing to probe). Size the default to cover the common cases while staying inside client tool timeouts, and pair the action with a read-only sibling that probes the same state so an agent can re-check without re-acting.
|
|
366
323
|
- **Seed orientation context alongside the primary result.** When a tool's call position makes the agent's next moves predictable, attaching a compact snapshot of relevant state — recent activity, tracked state, a couple of reference items — both saves round-trips *and* **primes the LLM on the project's patterns**. Surfacing recent commits teaches the commit-message style the agent should match when it later writes one; recent tags teach the versioning convention; reference records teach the naming format. Common fits: tools that open or close a session (set working dir, wrap-up), state-changing verbs where the caller wants post-action confirmation (commit, push, merge), entry points that drop the agent into a new scope (clone, checkout). Gather sub-operations in parallel with `Promise.allSettled` so a single failure degrades to a warning rather than tanking the outer call.
|
|
367
|
-
- **Communicate filtering.** If the tool silently excluded content, tell the LLM what was excluded and how to get it back. The agent can't act on what it doesn't know about.
|
|
368
|
-
|
|
369
|
-
```ts
|
|
370
|
-
// git_diff — when lock files are filtered, the output tells the LLM
|
|
371
|
-
output: z.object({
|
|
372
|
-
diff: z.string().describe('Unified diff output.'),
|
|
373
|
-
excludedFiles: z.array(z.string()).optional()
|
|
374
|
-
.describe('Files automatically excluded from the diff (e.g., lock files). Call again with autoExclude=false to include them.'),
|
|
375
|
-
}),
|
|
376
|
-
```
|
|
377
|
-
|
|
324
|
+
- **Communicate filtering.** If the tool silently excluded content, tell the LLM what was excluded and how to get it back — an `excludedFiles` list whose description says "call again with `autoExclude=false`". The agent can't act on what it doesn't know about.
|
|
378
325
|
- **Empty results are a designed surface, not a fallthrough.** For every search/list tool, spec the zero-hit behavior at design time: zero hits are success + an `enrichment` notice, never an error, and the notice is composed from condition → fragment pairs (too-narrow filter, a defaulted date window, a syntax trap), each routing to a concrete next call — relax a named filter, switch to the named sibling tool, or consult the reference tool. Relatedly, when the server applies a default that changes result *semantics* (an as-of date, an implicit status filter), echo the applied value in the output — an agent can't reason about a filter it can't see.
|
|
379
|
-
- **Capped lists disclose truncation.** When a tool accepts a cap-like input (`limit`, `per_page`, `page_size`, `max_results`, `max_items`) and returns an array, the handler must disclose when the cap was hit
|
|
380
|
-
- **
|
|
381
|
-
- **Spill big *analytical* results to a queryable surface.** When a tool's row set is something an agent would run SQL over
|
|
382
|
-
- **Outline one large *document* into sections.** When a single tool call returns one document-shaped record (not many rows) that can exceed context — a ~130KB FDA drug label, a big API entity dominated by a few fat fields — return a section *outline* (top-level keys + per-section byte size) instead of truncating, and let the agent re-call with `sections: [...]` to pull only what it needs.
|
|
383
|
-
- **Mirror a bulk upstream instead of paginating it live.** When the server wraps a large or slow API whose corpus is queried far more than it changes, sync it once into a persistent local index and query that as the primary data path — not the live API per request. Match the backend to corpus size:
|
|
384
|
-
-
|
|
385
|
-
- **Agent-facing context must reach both client surfaces — put it in `enrichment`.** `structuredContent` (from `output`) and `content[]` (from `format()`) are read by different clients. Empty-result notices, the query/filter as the server parsed it, and pagination totals — the context the agent *reasons with*, distinct from the domain payload — reach only `content[]` if hand-authored into `format()` text alone, leaving `structuredContent`-only clients (Claude Code) blind. (The reverse can't happen: `format-parity` drags every `output` field into `format()`, so `output`-authored context already reaches both.) An `enrichment` block — the success-path counterpart to `errors[]`, populated via `ctx.enrich(...)` — reaches both automatically: merged into `structuredContent`, advertised as `output.extend(enrichment)`, mirrored into a `content[]` trailer, no `format()` entry needed. How each field renders in that trailer is a per-tool call — a kind-tag (`notice`/`total`/`echo`/`delta`) when a canonical form fits, a domain key like `totalFound` otherwise, and an `enrichmentTrailer.render` for any structured (object/array) field so it doesn't ship as a JSON blob. See `add-tool`'s **Tool Response Design**.
|
|
326
|
+
- **Capped lists disclose truncation.** When a tool accepts a cap-like input (`limit`, `per_page`, `page_size`, `max_results`, `max_items`) and returns an array, the handler must disclose when the cap was hit — `truncated: true`, `shown`, `cap` in the `enrichment` block via `ctx.enrich.truncated({ shown, cap })`, and `ctx.enrich.total(n)` for the full count. The same applies to a display cut in `format()`: show the top N and say "...and X more". Silent caps leave the agent treating a partial set as complete; the `capped-list-no-truncation` lint rule enforces the input-cap case.
|
|
327
|
+
- **Continuation is a designed field.** Truncation says the cap was hit; continuation says how to get the rest. Return an opaque `cursor` plus `has_more` (via `extractCursor`/`paginateArray` for local sets), and never invent page numbers over a cursor-based upstream — a page the agent can't ask for is a page it will never see.
|
|
328
|
+
- **Spill big *analytical* results to a queryable surface.** When a tool's row set is something an agent would run SQL over *and* can exceed any reasonable context budget — paginated APIs, streamed exports, big query results — pair an inline preview with a `DataCanvas` table holding the full set (`spillover()` in `api-canvas`), and compute distributions or refinement hints across the full result, not the preview, so aggregate signal stays honest. The gates on when a canvas earns its keep are in Step 7.
|
|
329
|
+
- **Outline one large *document* into sections.** When a single tool call returns one document-shaped record (not many rows) that can exceed context — a ~130KB FDA drug label, a big API entity dominated by a few fat fields — return a section *outline* (top-level keys + per-section byte size) instead of truncating, and let the agent re-call with `sections: [...]` to pull only what it needs. `outlineOnOverflow()` (`@cyanheads/mcp-ts-core/utils`) returns a `full | outline` result; pure measure + key-slice, so Cloudflare Workers-portable, unlike canvas-bound `spillover()`. Distinct from spillover on *shape*: spillover splits a row collection, this outlines one fat record. Schema shape and `format()` parity are in the `techniques` skill's `outline-on-overflow` reference.
|
|
330
|
+
- **Mirror a bulk upstream instead of paginating it live.** When the server wraps a large or slow API whose corpus is queried far more than it changes, sync it once into a persistent local index and query that as the primary data path — not the live API per request. Match the backend to corpus size: below ~10⁴ rows → an in-memory index (server-level, no primitive); ~10⁴–10⁷ → the `MirrorService` (embedded SQLite + FTS5; declare a schema + a `sync` ingester via `defineMirror`/`sqliteMirrorStore`, then `runSync`/`query`, see `api-mirror`); above ~10⁷ → an external store. Distinct lifecycle from DataCanvas: a mirror is long-lived and cross-session, refreshed on a schedule; canvas is ephemeral and per-session.
|
|
331
|
+
- **Two client surfaces, both content-complete.** Different MCP clients forward different surfaces to the model: some (e.g., Claude Code) read `structuredContent` from `output`, others (e.g., Claude Desktop) read `content[]` from `format()`. `format()` is the markdown twin of `structuredContent`, not a summary — a thin `format()` that returns only a count or title leaves `content[]`-only clients blind (the `format-parity` lint catches this). Agent-facing context that is *not* domain payload — empty-result notices, the query as the server parsed it, echoed defaults, totals — goes in the `enrichment` block via `ctx.enrich(...)`, which reaches both surfaces automatically; hand-authored into `format()` text alone it reaches only one. Field-by-field rendering of that block is in the Design table's Enrichment row.
|
|
386
332
|
|
|
387
333
|
#### Batch input design
|
|
388
334
|
|
|
@@ -412,8 +358,6 @@ output: z.object({
|
|
|
412
358
|
|
|
413
359
|
Single-item tools don't need this — they either succeed or throw. The partial success question only arises when the tool can partially complete.
|
|
414
360
|
|
|
415
|
-
**Telemetry:** The framework automatically detects partial success — when a handler returns a result with a non-empty `failed` array, the span gets `mcp.tool.partial_success`, `mcp.tool.batch.succeeded_count`, and `mcp.tool.batch.failed_count` attributes. No manual instrumentation needed.
|
|
416
|
-
|
|
417
361
|
#### Convenience shortcuts for complex inputs
|
|
418
362
|
|
|
419
363
|
**Applies when:** a tool wraps a structured query language or filter system where the 80% case is a simple string. Skip when the primary input is already simple.
|
|
@@ -443,7 +387,7 @@ Two params, two behaviors — keep them named distinctly:
|
|
|
443
387
|
|
|
444
388
|
**Correctness: filter the *complete* bounded set, not the current page.** Fetch up to the cap (or page through) before filtering — filtering one page returns a misleading partial slice.
|
|
445
389
|
|
|
446
|
-
**Matching: strict token match is the default.** Normalize (lowercase, strip punctuation/diacritics) and require every query token to appear, so word order and missing interior words still match. That strict core is the ~90% case
|
|
390
|
+
**Matching: strict token match is the default.** Normalize (lowercase, strip punctuation/diacritics) and require every query token to appear, so word order and missing interior words still match. That strict core is the ~90% case and needs no fuzzy library. Add a fuzzy fallback **only when a caller genuinely needs typo tolerance** (an LLM caller rarely does): fire it only when the strict match is empty, score against the best-matching *token* in each name (not the whole string) and **cap** the results — or one short query clears the threshold against dozens of long multi-word names — and label its hits `approximate`. Often a bare "no match — call the unfiltered list to browse" beats an `approximate` guess: it lets the model self-correct instead of committing to the wrong record. See `add-tool` for the param + handler implementation.
|
|
447
391
|
|
|
448
392
|
#### Error design
|
|
449
393
|
|
|
@@ -456,8 +400,10 @@ Errors are part of the tool's interface — design them during the design phase,
|
|
|
456
400
|
| Origin | Examples | Error code | Agent can recover? |
|
|
457
401
|
|:-------|:---------|:-----------|:-------------------|
|
|
458
402
|
| **Client input** | Bad ID format, invalid params, missing required field, out-of-range value | `ValidationError` | Yes — fix the input and retry |
|
|
459
|
-
| **Upstream API** | 5xx,
|
|
403
|
+
| **Upstream API** | 5xx, timeout, network error | `ServiceUnavailable` | Maybe — retry later, or the upstream is down |
|
|
404
|
+
| **Rate limit** | 429, quota exhausted, queue full | `RateLimited` (`retryable: true`; `withRetry` honors `Retry-After`) | Yes — wait, then retry or reduce frequency |
|
|
460
405
|
| **Not found** | Valid ID format but entity doesn't exist | `NotFound` (or `ValidationError` if ambiguous) | Yes — check the ID, try a search |
|
|
406
|
+
| **Conflict** | Duplicate key, version mismatch, concurrent modification on a write | `Conflict` | Yes — re-read current state, then retry with it |
|
|
461
407
|
| **Auth/permissions** | Insufficient scopes, expired token | `Forbidden` / `Unauthorized` | Maybe — escalate or re-auth |
|
|
462
408
|
| **Server internal** | Parse failure, missing config, unexpected state | `InternalError` | No — server-side issue |
|
|
463
409
|
|
|
@@ -465,7 +411,7 @@ Errors are part of the tool's interface — design them during the design phase,
|
|
|
465
411
|
|
|
466
412
|
The framework auto-classifies many of these at runtime (HTTP status codes, JS error types, common patterns), but explicit classification in the handler gives better error messages. For declared contract failures, throw via `ctx.fail('reason', …)`. For ad-hoc throws outside the contract, use error factories (`notFound()`, `validationError()`, etc.) when the code matters; plain `throw new Error()` when the framework's auto-classification is good enough.
|
|
467
413
|
|
|
468
|
-
**Expected misses are results, not errors.** When a tool's whole job is resolving one identifier — a citation, a code, a name → ID — a no-match is an expected outcome the agent must reason about, not a failure. Return `{ found: false, guidance }` instead of throwing, and treat `guidance` as a first-class recovery surface: per miss outcome, say what didn't parse or resolve and route to the named tool that can recover (the broader search tool, the reference tool). Agents self-correct better from a structured miss than from a throw — a throw reads as "something broke," a miss result reads as "adjust and retry."
|
|
414
|
+
**Expected misses are results, not errors.** When a tool's whole job is resolving one identifier — a citation, a code, a name → ID — a no-match is an expected outcome the agent must reason about, not a failure. Return `{ found: false, guidance }` instead of throwing, and treat `guidance` as a first-class recovery surface: per miss outcome, say what didn't parse or resolve and route to the named tool that can recover (the broader search tool, the reference tool). Agents self-correct better from a structured miss than from a throw — a throw reads as "something broke," a miss result reads as "adjust and retry." The split with search tools: a search's empty result is a valid empty collection, so its notice rides `enrichment`; a resolver's miss *is* the primary result, so `found` and its `guidance` live in `output`.
|
|
469
415
|
|
|
470
416
|
**Write error messages as recovery instructions.** The message is the agent's only signal for what to do next — and the strongest recovery instruction ends in a named tool call, never a bare "check your input."
|
|
471
417
|
|
|
@@ -503,14 +449,15 @@ Summarize each tool:
|
|
|
503
449
|
|
|
504
450
|
| Aspect | Decision |
|
|
505
451
|
|:-------|:---------|
|
|
506
|
-
| **Name** | Lowercase snake_case with a canonical server prefix. **3 segments is the strong default** (`{server}_{verb}_{noun}` — e.g., `pubmed_search_articles`, `
|
|
452
|
+
| **Name** | Lowercase snake_case with a canonical server prefix. **3 segments is the strong default** (`{server}_{verb}_{noun}` — e.g., `pubmed_search_articles`, `clinicaltrials_find_eligible`). **2 is fine when the operation name is canonical** and no noun adds signal (`git_pull`, `git_status` — "pull" already implies the remote). Don't invent a word to pad to 3. **4 is fine when the noun is inherently two words** (`openfda_search_device_clearances`) or the prefix is multi-part. The prefix is judged on clarity, not length: the brand name or the plain well-known word for the domain both pass (`pubmed_`, `patents_`, `earthquake_`); an abbreviation fails only when it reads as something else out of context (`loc_` → lines of code, `ct_` → CT scan). The verb+noun pair should be unambiguous within the server — if two tools could plausibly share a name, the noun isn't specific enough (`read_fulltext` not `read_text` when structured metadata is a separate concept). **Treat name length as a scope smell only when** the extra segment is the *verb* overreaching (e.g., `foo_create_and_send_notification` → split or use modes). |
|
|
507
453
|
| **Granularity** | Scope each tool to one coherent agent action. The implementation can be a single API call (`pubmed_search_articles`), a multi-step workflow, or internal-only — match the unit to the work, don't constrain by call count. |
|
|
508
454
|
| **Description** | Concrete capability statement. Add operational guidance (prerequisites, constraints, gotchas) when non-obvious. |
|
|
509
455
|
| **Input schema** | `.describe()` on every field. Constrained types (enums, literals, regex). Explain costs/tradeoffs of parameter choices. |
|
|
510
456
|
| **Output schema** | Designed for the LLM's next action. Include chaining IDs. Communicate filtering. Post-write state where useful. |
|
|
511
457
|
| **Errors** | Declare domain failure modes as a typed contract (`errors: [{ reason, code, when, recovery, retryable? }]`) so `ctx.fail` is type-checked and capable clients can preview failures via `tools/list`. Every `recovery` string follows the no-dead-ends rule — it names the next tool call. |
|
|
512
|
-
| **
|
|
513
|
-
| **
|
|
458
|
+
| **Enrichment** | The success-path counterpart to `errors`: declare the agent-facing context fields the handler populates via `ctx.enrich(...)` — zero-hit notice, echoed defaults, totals, truncation — with a kind-tag (`notice`/`total`/`echo`/`delta`) where one fits and an `enrichmentTrailer.render` for any structured field. Keys stay disjoint from `output`. |
|
|
459
|
+
| **Annotations** | `readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint`. Helps clients auto-approve safely. `destructiveHint` defaults to **true** on any tool that isn't read-only, so a benign write must set `destructiveHint: false` explicitly; a read-only tool omits it entirely (`annotation-coherence` lint). |
|
|
460
|
+
| **Auth scopes** | `tool:<snake_tool_name>:<verb>` or `resource:<kebab-resource-name>:<verb>` (e.g., `tool:inventory_search:read`, `resource:echo-app-ui:read`). Domain-led `<domain>:<verb>` (e.g., `inventory:read`) is an acceptable alternative — pick one convention per server and stay consistent. Skip when the server runs `MCP_AUTH_MODE=none` (stdio-only, local). |
|
|
514
461
|
|
|
515
462
|
### 5. Design Resources
|
|
516
463
|
|
|
@@ -524,6 +471,8 @@ For each resource:
|
|
|
524
471
|
| **Params** | Minimal — typically just an identifier. Complex queries belong in tools. |
|
|
525
472
|
| **Pagination** | Needed if lists exceed ~50 items. Opaque cursors via `extractCursor`/`paginateArray`. |
|
|
526
473
|
| **list()** | Provide if discoverable. Top-level categories or recent items, not exhaustive dumps. |
|
|
474
|
+
| **Cache hint** | `cacheHint: { ttlMs, cacheScope }` on 2026-07-28 connections. Static reference data is `public` with a long TTL; per-tenant or per-session data is `private` or uncached. |
|
|
475
|
+
| **Completion** | Template params with a bounded vocabulary (a species list, a dataset code) get argument completion so clients can offer valid values; the same `completable()` wrapper applies to prompt args. |
|
|
527
476
|
| **Tool coverage** | Verify the data is reachable via tools — either a dedicated tool, included in another tool's output, or not needed for tool-only agents. |
|
|
528
477
|
|
|
529
478
|
### 6. Design Prompts (if needed)
|
|
@@ -532,7 +481,7 @@ Optional. Use when the server has recurring interaction patterns worth structuri
|
|
|
532
481
|
|
|
533
482
|
- Analysis frameworks, report templates, multi-step workflows
|
|
534
483
|
|
|
535
|
-
Skip for purely data/action-oriented servers.
|
|
484
|
+
Wrap enum-like prompt args in `completable()` so clients can offer valid values. Skip prompts entirely for purely data/action-oriented servers.
|
|
536
485
|
|
|
537
486
|
### 7. Plan Services and Config
|
|
538
487
|
|
|
@@ -540,7 +489,7 @@ Skip for purely data/action-oriented servers.
|
|
|
540
489
|
|
|
541
490
|
**Server-as-service.** When the server IS the source of truth (knowledge graph, in-memory task tracker, local scratchpad, embedded inference wrapper), the resilience table below doesn't apply — there's no upstream to retry. The design questions shift to state management: what's tenant-scoped vs. global, what TTLs apply, what survives a restart, what the storage backend is. Plan persistence via `ctx.state` for tenant-scoped KV (auto-namespaced by `tenantId`), or use a `StorageService` provider directly when data must cross tenants. Service init still happens in `setup()`, accessed via `getMyService()` at request time. Calls within the server are local and synchronous-ish — the API-efficiency table below also doesn't apply.
|
|
542
491
|
|
|
543
|
-
**Analytical API servers: DataCanvas is one option.** For servers that fetch **analytical** data — result sets an agent runs SQL over (aggregate, group, join, time-series) — and want to expose a SQL workspace, the framework's optional `DataCanvas` primitive (Tier 3, opt-in via `CANVAS_PROVIDER_TYPE=duckdb`) handles lifecycle, ID generation, eviction, and export wiring so you don't design your own. **It earns its keep on shape, not size:** a discovery/search surface returning categorical metadata (titles, IDs, types) — where the workflow is find-the-record-then-drill-in — does *not* qualify even when the result is large; resolve names over a bounded set with [MCP-side list filtering](#mcp-side-list-filtering) instead. **If you opt in, the consumer tools are mandatory:** a tool that emits a `canvas_id` MUST be paired with a `dataframe_query` (and `dataframe_describe`) tool in the same surface — a `canvas_id` with no query tool is dead output the agent can't reach. Surface `canvas_id` as an optional input on register/query/export tools; the framework mints on omit and resolves on match.
|
|
492
|
+
**Analytical API servers: DataCanvas is one option.** For servers that fetch **analytical** data — result sets an agent runs SQL over (aggregate, group, join, time-series) — and want to expose a SQL workspace, the framework's optional `DataCanvas` primitive (Tier 3, opt-in via `CANVAS_PROVIDER_TYPE=duckdb`) handles lifecycle, ID generation, eviction, and export wiring so you don't design your own. **It earns its keep on shape, not size:** a discovery/search surface returning categorical metadata (titles, IDs, types) — where the workflow is find-the-record-then-drill-in — does *not* qualify even when the result is large; resolve names over a bounded set with [MCP-side list filtering](#mcp-side-list-filtering) instead. **If you opt in, the consumer tools are mandatory:** a tool that emits a `canvas_id` MUST be paired with a `dataframe_query` (and `dataframe_describe`) tool in the same surface — a `canvas_id` with no query tool is dead output the agent can't reach. Surface `canvas_id` as an optional input on register/query/export tools; the framework mints on omit and resolves on match. The accessor is wired once in `setup()` via `setCanvas(core.canvas)` (undefined when disabled or running on Cloudflare Workers — DuckDB has no V8-isolate build). See `api-canvas` for the full reference.
|
|
544
493
|
|
|
545
494
|
For services wrapping external APIs, plan the resilience layer.
|
|
546
495
|
|
|
@@ -551,6 +500,8 @@ For services wrapping external APIs, plan the resilience layer.
|
|
|
551
500
|
| **HTTP status check** | `fetchWithTimeout` already handles this — non-OK → `ServiceUnavailable`. |
|
|
552
501
|
| **Parse failure classification** | Response handler detects HTML error pages and throws transient errors, not `SerializationError`. |
|
|
553
502
|
| **Exhausted retry messaging** | `withRetry` enriches the final error with attempt count automatically. |
|
|
503
|
+
| **Pacing** | No framework primitive paces requests. When the upstream mandates a rate (one request per second, one per five seconds, N concurrent), decide per service how the tool surface honors it — a queue in the service, a concurrency cap on fan-out, or a documented ceiling in the server instructions — and say which. |
|
|
504
|
+
| **Caller-supplied URLs or hosts** | Route through `fetchWithTimeout`, which carries the SSRF guard (private ranges, DNS rebinding). Never a bare `fetch` on a caller-controlled destination; `security-pass` audits this sink. |
|
|
554
505
|
|
|
555
506
|
For API efficiency, design the service methods to minimize upstream calls:
|
|
556
507
|
|
|
@@ -604,6 +555,12 @@ list, the error contract table (reason | code | when | recovery — verbatim str
|
|
|
604
555
|
zero-hit notice fragments for search tools. This is the section implementation reads
|
|
605
556
|
tool-by-tool.
|
|
606
557
|
|
|
558
|
+
## Resources — detail / ## Prompts — detail
|
|
559
|
+
|
|
560
|
+
Same treatment when the server has them: one subsection per resource (URI template, params,
|
|
561
|
+
cache hint, the tool that also covers the data) and per prompt (args, completions, the
|
|
562
|
+
message shape).
|
|
563
|
+
|
|
607
564
|
## Services
|
|
608
565
|
| Service | Wraps | Used By |
|
|
609
566
|
|:--------|:------|:--------|
|
|
@@ -687,7 +644,7 @@ Items without an `If …:` prefix apply to every design. Conditional items only
|
|
|
687
644
|
- [ ] Tool surface audited — niche, overlapping, or low-value tools cut or deferred
|
|
688
645
|
- [ ] Tool surface is self-sufficient — a tool-only agent can accomplish everything the server is for
|
|
689
646
|
- [ ] Workflow and Instruction variants considered where they add value (single-action tools are the default)
|
|
690
|
-
- [ ] Tool descriptions are concrete and include operational guidance where non-obvious
|
|
647
|
+
- [ ] Tool descriptions are imperative present tense, concrete, and include operational guidance where non-obvious
|
|
691
648
|
- [ ] Parameter `.describe()` text explains what the value is, what it affects, and tradeoffs
|
|
692
649
|
- [ ] Input schemas use constrained types (enums, literals, regex) over free strings
|
|
693
650
|
- [ ] Output schemas designed for LLM's next action — chaining IDs, post-write state, filtering communicated
|
|
@@ -697,17 +654,18 @@ Items without an `If …:` prefix apply to every design. Conditional items only
|
|
|
697
654
|
- [ ] **If the server has search/list tools:** zero-hit notices specced (condition → fragment, each routing to a named next call); server-applied defaults that change result semantics echoed in output
|
|
698
655
|
- [ ] **If a tool resolves a single identifier:** no-match returns `{ found: false, guidance }` — a result, not a throw — with guidance routing per miss outcome
|
|
699
656
|
- [ ] **If the domain has opaque vocabulary (codes, identifier formats, coverage windows):** reference tool designed (`topic` enum), implemented first, and used as the routing target in recovery strings and notices
|
|
700
|
-
- [ ] Annotations set correctly (`readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint`)
|
|
657
|
+
- [ ] Annotations set correctly (`readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint`) — benign writes set `destructiveHint: false` explicitly, read-only tools omit it
|
|
701
658
|
- [ ] Server-level `instructions` string drafted — workflow chain, identifier semantics, rate-limit posture (ships via `createApp()` on every initialize)
|
|
702
659
|
- [ ] Design doc written to `docs/design.md`
|
|
703
660
|
- [ ] Design confirmed with user (or user pre-authorized implementation)
|
|
704
|
-
- [ ] **If ops share a noun:** related operations consolidated under one tool with `mode`/`operation` enum
|
|
661
|
+
- [ ] **If ops share a noun:** related operations consolidated under one tool with a `mode`/`operation` enum — as a `z.discriminatedUnion` input when the arms need different required fields
|
|
705
662
|
- [ ] **If an upstream API has no native search but the relevant set is bounded:** MCP-side list filtering considered — a distinct local filter param (`filter`/`nameContains`, not `query`), filtering the full set, strict token match (fuzzy only when a caller needs typo tolerance)
|
|
706
663
|
- [ ] **If the server has workflow tools:** call-flow documented (upstream sequence + mode arms) in design doc's Workflow Analysis
|
|
707
664
|
- [ ] **If state-aware procedural guidance adds value:** instruction tool considered with `nextToolSuggestions` pre-filled from diagnostics
|
|
708
665
|
- [ ] **If any tool is config-gated:** nothing routes to it while the gate is off — recovery strings, notices, and `guidance` name a callable target or state the capability is unavailable, and structured follow-ups naming it are emitted only under the config that registers it
|
|
709
666
|
- [ ] **If workflow tools have destructive modes:** destructive arm gated on a `ctx.requestInput` confirmation read back from `ctx.inputs`, with `destructiveHint` annotation so clients that never fulfil the round still surface the risk
|
|
710
667
|
- [ ] **If a parameter determines blast radius:** safe default set (e.g., `mode: 'preview'`, `dryRun: true`, `confirmCount` required)
|
|
668
|
+
- [ ] **If an action's effect is observable only after the call (wake, trigger, dispatch, provision):** confirmation on by default through one numeric window param (`0` = act and return), pre-probe then poll with early return, every outcome a result rather than a throw, and a read-only sibling tool that probes the same state
|
|
711
669
|
- [ ] **App tools default to no.** If one was proposed, verified there's a real human-in-the-loop in an MCP Apps-capable client justifying the iframe/CSP/`format()`-twin maintenance cost — otherwise dropped in favor of a standard tool
|
|
712
670
|
- [ ] **If the server exposes resources:** URIs use `{param}` templates, pagination planned for large lists
|
|
713
671
|
- [ ] **If the server is itself the source of truth (no external API):** state lifecycle planned — tenant-scoped vs. global, TTLs, what survives restart, storage backend chosen
|
|
@@ -4,7 +4,7 @@ description: >
|
|
|
4
4
|
Exercise tools, resources, and prompts against a live HTTP server via MCP JSON-RPC over curl. Starts the server, surfaces the catalog, runs real and adversarial inputs, measures every call (bytes, token estimate, wall-clock) and weighs the catalog, and produces a tight report with concrete findings and numbered follow-up options. Use after adding or modifying definitions, or when the user asks to test, try out, or verify their MCP surface.
|
|
5
5
|
metadata:
|
|
6
6
|
author: cyanheads
|
|
7
|
-
version: "2.
|
|
7
|
+
version: "2.14"
|
|
8
8
|
audience: external
|
|
9
9
|
type: debug
|
|
10
10
|
---
|
|
@@ -19,7 +19,7 @@ Unit tests (`add-test` skill) verify handler logic with mocked context. Field te
|
|
|
19
19
|
|
|
20
20
|
This skill drives an HTTP server because curl + JSON-RPC is the most reliable harness for shell-based agents. The same handlers run on both transports — only the framing differs — so HTTP exercises the full functional surface. Both HTTP session modes are covered: a durable `Mcp-Session-Id` session, and the sessionless initialization a `MCP_SESSION_MODE=stateless` server performs.
|
|
21
21
|
|
|
22
|
-
**Stdio coverage is a boot check only — run this before Step 1.** Run `bun run rebuild && bun run start:stdio`, confirm the startup logs look clean (banner, expected tool/resource counts, no errors/warnings, no missing-config gripes), then
|
|
22
|
+
**Stdio coverage is a boot check only — run this before Step 1.** Run `bun run rebuild && bun run start:stdio < /dev/null`, and confirm the startup logs look clean (banner, expected tool/resource counts, no errors/warnings, no missing-config gripes). Redirecting stdin is what ends the run: the server treats EOF as a shutdown signal, boots fully, then exits on its own, so the log also shows the graceful-shutdown path. Do not background it and reach for `pkill` — a pattern like `pkill -f dist/index.js` matches every other stdio MCP server on the machine, including the ones the calling agent's own session is connected to. Pino logs go to stderr in stdio mode (stdout is reserved for JSON-RPC), so they print straight to the terminal when you run interactively. No need to call tools over stdio — the HTTP pass already covered handler behavior.
|
|
23
23
|
|
|
24
24
|
---
|
|
25
25
|
|
|
@@ -140,7 +140,10 @@ mcp_init() {
|
|
|
140
140
|
# Unwrap SSE framing when present; a plain JSON body is used as-is.
|
|
141
141
|
local payload; payload=$(sed -n 's/^data: //p' "$body_file")
|
|
142
142
|
[ -z "$payload" ] && payload=$(cat "$body_file")
|
|
143
|
-
|
|
143
|
+
# Pick the reply frame by structure, not by substring: a server that logs to the
|
|
144
|
+
# client emits `notifications/message` frames first, and a `"level":"error"` or a
|
|
145
|
+
# log string containing `result` matches a text grep and gets read as the reply.
|
|
146
|
+
local reply; reply=$(printf '%s\n' "$payload" | jq -c 'select(type=="object" and (has("result") or has("error")))' 2>/dev/null | tail -1)
|
|
144
147
|
[ -z "$reply" ] && reply="$payload"
|
|
145
148
|
if printf '%s' "$reply" | grep -q '"error"'; then
|
|
146
149
|
_mcp_init_fail "server returned a JSON-RPC error" "$body_file" "$hdr"
|
|
@@ -240,7 +243,9 @@ mcp_call() {
|
|
|
240
243
|
local reply
|
|
241
244
|
local sse; sse=$(sed -n 's/^data: //p' "$resp_file")
|
|
242
245
|
if [ -n "$sse" ]; then
|
|
243
|
-
|
|
246
|
+
# Structural pick, same reason as in mcp_init: log-notification frames precede
|
|
247
|
+
# the reply and can carry the literal tokens a text grep keys on.
|
|
248
|
+
reply=$(printf '%s\n' "$sse" | jq -c 'select(type=="object" and (has("result") or has("error")))' 2>/dev/null | tail -1)
|
|
244
249
|
reply="${reply:-$sse}"
|
|
245
250
|
else
|
|
246
251
|
reply=$(cat "$resp_file")
|
|
@@ -494,7 +499,7 @@ End with:
|
|
|
494
499
|
|
|
495
500
|
## Checklist
|
|
496
501
|
|
|
497
|
-
- [ ] Stdio boot check completed — `bun run rebuild && bun run start:stdio` shows clean startup (banner, expected counts, no errors)
|
|
502
|
+
- [ ] Stdio boot check completed — `bun run rebuild && bun run start:stdio < /dev/null` shows clean startup (banner, expected counts, no errors) and a graceful shutdown on EOF
|
|
498
503
|
- [ ] HTTP server built and started; real port parsed from log
|
|
499
504
|
- [ ] Session initialized (a stateless server returns an empty `sid` — still a pass); `notifications/initialized` sent; negotiated protocol version matches the requested one (a downgrade is a finding)
|
|
500
505
|
- [ ] Catalog surfaced and presented; descriptions audited for leaks (implementation details, meta-coaching, consumer-aware phrasing)
|
|
@@ -4,7 +4,7 @@ description: >
|
|
|
4
4
|
Land working-tree changes as logical commits — the work grouped by concern, topped by a release commit (version bump, changelog, regenerated artifacts). Verify, commit. Stops at "committed locally on main" — or, when the project releases through a release PR, at "release branch pushed, PR open". No tag, no push to main, no publish: the release-and-publish skill merges, tags, and ships from here. Distilled from the git_wrapup_instructions protocol.
|
|
5
5
|
metadata:
|
|
6
6
|
author: cyanheads
|
|
7
|
-
version: "1.
|
|
7
|
+
version: "1.16"
|
|
8
8
|
audience: external
|
|
9
9
|
type: workflow
|
|
10
10
|
---
|
|
@@ -116,6 +116,8 @@ security: false # true ONLY for a security fix in this server's own source
|
|
|
116
116
|
|
|
117
117
|
**Tone:** Terse, fact-dense. Bullet = **symbol** + what changed + at most one consumer-facing caveat; one sentence by default, two max — a bullet past ~40 words or three sentences is wrong. The linked issue carries the why and the commit diff the how; the changelog names what changed and what a consumer does about it. Cut: history/justification narration, design-rationale defense, "X unchanged" clauses (short parenthetical only where a misread is likely), edge-case inventories. **Verified ≠ included** — the diff-is-source-of-truth rule bounds the truth of what you write, never the amount. Model length on `changelog/template.md`'s authoring guide, never on the previous entry (entries modeled on entries compound). `agent-notes` carries adoption steps only, never a second rendering of the body; a consequence shared by many bullets is stated once, not per bullet. Full conventions: the authoring guide in `changelog/template.md`.
|
|
118
118
|
|
|
119
|
+
**Re-read the entry file after writing it, then sweep for harness markup:** `grep -rlF -e '</invoke>' -e '</content>' changelog/` must print nothing. A stray closing tag at EOF is the authoring tool's own syntax bleeding into the file; `changelog/` is in `package.json` `files`, so it ships inside the npm tarball, and `changelog:check` cannot catch it — the rollup drops the trailing line, so a clean `CHANGELOG.md` proves nothing about the entry.
|
|
120
|
+
|
|
119
121
|
### 5. Regenerate derived artifacts
|
|
120
122
|
|
|
121
123
|
```bash
|
|
@@ -161,7 +163,7 @@ git commit -m "<subject>"
|
|
|
161
163
|
# repeat per concern; version + changelog + tree are the final commit
|
|
162
164
|
```
|
|
163
165
|
|
|
164
|
-
**The file is the atomic boundary:** NEVER split a single file's changes across commits. When one file serves two concerns, it ships whole in the commit of its dominant concern.
|
|
166
|
+
**The file is the atomic boundary:** NEVER split a single file's working-tree changes across commits, regardless of mechanism — not `git add -p`, not an index-only patch (`git apply --cached`), not editing the file between commits to remove-then-re-add a hunk. When one file serves two concerns, it ships whole in the commit of its dominant concern; a later commit may touch the file again only for changes made AFTER the first commit (a version badge bumped after the fix landed).
|
|
165
167
|
|
|
166
168
|
**Subject format:** Conventional Commits.
|
|
167
169
|
- Work commits (no version): `feat: hosted server endpoint`, `fix: handle empty SPARQL result sets`, `feat(linter): enrichment contract rules`, `docs: document the enrichment block`
|
|
@@ -267,7 +269,7 @@ If the working tree isn't clean or the release commit isn't at HEAD, something w
|
|
|
267
269
|
## Checklist
|
|
268
270
|
|
|
269
271
|
- [ ] Diff reviewed end-to-end before version bump
|
|
270
|
-
- [ ] Version bumped in every declaring file (`package.json`, `server.json`, `manifest.json`, `.claude-plugin/plugin.json`, `.codex-plugin/plugin.json`, README badge, `CLAUDE.md`/`AGENTS.md` if they pin a version)
|
|
272
|
+
- [ ] Version bumped in every declaring file (`package.json`, `server.json`, `manifest.json`, `.claude-plugin/plugin.json`, `.codex-plugin/plugin.json`, README badge, `CLAUDE.md`/`AGENTS.md` if they pin a version) — verify by command, not by eye: `v=$(jq -r .version package.json); grep -rl "$v" package.json server.json manifest.json .claude-plugin/plugin.json .codex-plugin/plugin.json README.md | wc -l` must equal the count of files that exist, and `grep -c "Version-$v-" README.md` must print `1`. The README badge is the one no lint reads, so it is the one that ships stale
|
|
271
273
|
- [ ] GH issues addressed by this work commented with what landed (if working from GH issues)
|
|
272
274
|
- [ ] Docs updated for any new or changed features
|
|
273
275
|
- [ ] Changelog authored at `changelog/<major.minor>.x/<version>.md`
|