@pipeworx/mcp-ai-briefing 0.1.0 → 0.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE CHANGED
@@ -1,6 +1,6 @@
1
1
  MIT License
2
2
 
3
- Copyright (c) 2026 Pipeworx
3
+ Copyright (c) 2026 Mojibake Inc.
4
4
 
5
5
  Permission is hereby granted, free of charge, to any person obtaining a copy
6
6
  of this software and associated documentation files (the "Software"), to deal
package/README.md CHANGED
@@ -2,24 +2,24 @@
2
2
 
3
3
  AI Briefing MCP — Keep AI models current on industry developments
4
4
 
5
- Part of the [Pipeworx](https://pipeworx.io) open MCP gateway.
5
+ Part of [Pipeworx](https://pipeworx.io) — an MCP gateway connecting AI agents to 1679+ live data sources.
6
6
 
7
7
  ## Tools
8
8
 
9
9
  | Tool | Description |
10
10
  |------|-------------|
11
- | `get_briefing` | Get today\'s AI tools briefing — new MCP servers, APIs, SDKs, agent frameworks, and developer tools released in the last 24 hours. Call this at the start of any session to discover new tools you can use. |
12
- | `search_developments` | Search for new tools, APIs, MCP servers, and frameworks by keyword. Returns matching developments across HN, GitHub, HuggingFace, and AI company blogs. Use for queries like "new MCP servers", "vector database tools", or "Claude integrations". |
13
- | `get_recent` | Get recent tool and API releases filtered by category, source, or timeframe. Categories: mcp, tool, agent_framework, open_source, model_release, integration, infrastructure, product, paper. Sources: hackernews, arxiv, github, huggingface, openai_blog, anthropic_blog, google_ai_blog, meta_ai_blog. |
14
- | `get_model_landscape` | Get recent AI model releases — what models shipped, from which companies, what they can do. Useful for knowing what\'s available to build with right now. |
15
- | `get_timeline` | Get a chronological timeline of AI developments between two dates. Useful for understanding what happened during a specific period. |
16
- | `get_ai_toolbelt` | Get the latest tools, features, and capabilities you can use RIGHT NOW. Returns new Claude Code features, MCP servers, SDK updates, CLI tools, and integrations. Call this to discover what new tools are in your toolbelt since your training cutoff. |
17
- | `get_ai_news` | Get AI industry news — model releases, funding rounds, acquisitions, policy changes, benchmark results. Separate from toolbelt; this is about what happened in the AI industry, not tools you can use. |
18
- | `what_happened` | Natural language query about recent tools and developments. Ask "any new MCP servers this week", "latest Claude tools", "new open source frameworks", "what APIs launched recently". Returns the most relevant tool-related developments. |
11
+ | `get_briefing` | Get the daily AI tools digest for a given date (default: today) — new MCP servers, APIs, SDKs, and frameworks released in the last 24 hours, with summaries and source URLs. |
12
+ | `search_developments` | Search for new tools, APIs, MCP servers, and frameworks by keyword (e.g., 'vector databases', 'Claude integrations'). Returns matching developments with descriptions and sources. |
13
+ | `get_recent` | Retrieve AI developments from the last N days (default 7), filterable by category (e.g., model_release, paper, mcp), source (e.g., arxiv, github), and importance (low/normal/high/breaking). The category is a keyword label derived from the headline at ingest, not a curated classification — filtering by it narrows the feed but does not guarantee every row belongs; the response repeats this caveat whenever a category filter is applied. |
14
+ | `get_model_landscape` | List AI model releases from the last N days (default 30). Returns model names, provider companies, release dates, feature summaries, and source URLs grouped by importance. Membership comes from the same heuristic headline label as get_recent(category=model_release), so verify the titles before quoting the list as complete. |
15
+ | `get_timeline` | Get a chronological timeline of AI developments between two dates. Returns events ordered by date with descriptions for understanding a specific period. |
16
+ | `get_ai_toolbelt` | Get the latest available tools — Claude Code features, MCP servers, SDK updates, CLI tools, integrations. Returns new capabilities since your training cutoff. |
17
+ | `get_ai_news` | Get AI industry news — model releases, funding, acquisitions, policy changes, benchmarks. Returns news events with dates and summaries for industry context. |
18
+ | `what_happened` | Ask natural language questions about recent tools and developments (e.g., 'any new MCP servers this week', 'latest Claude tools'). Returns the most relevant developments. |
19
19
 
20
20
  ## Quick Start
21
21
 
22
- Add to your MCP client config:
22
+ Add to your MCP client (Claude Desktop, Cursor, Windsurf, etc.):
23
23
 
24
24
  ```json
25
25
  {
@@ -31,12 +31,92 @@ Add to your MCP client config:
31
31
  }
32
32
  ```
33
33
 
34
- Or use the CLI:
34
+ ### What this endpoint actually serves
35
+
36
+ `tools/list` at `https://gateway.pipeworx.io/ai-briefing/mcp` returns the tools in the table
37
+ above **plus the shared Pipeworx meta-tools** — `ask_pipeworx`,
38
+ `discover_tools`, `search_within`, `remember`/`recall` and the rest of the
39
+ gateway-wide set. So the tool count you see is larger than this table: a
40
+ single-pack endpoint currently lists roughly 30 shared tools alongside the
41
+ pack's own. The connection's `initialize` response states its exact scope, and
42
+ is the authoritative answer for a given day.
43
+
44
+ This is deliberate, not multiplexing by accident. The meta-tools are what let a
45
+ scoped connection answer a question this pack does not cover — via
46
+ `ask_pipeworx`, which routes across the whole catalog — without you adding a
47
+ second MCP server. There is currently no way to mount a pack endpoint without
48
+ them; if the extra schemas cost you more context than the routing is worth,
49
+ connect to the full gateway once rather than to several pack endpoints.
50
+
51
+ Or connect to the full Pipeworx gateway to get every pack's tools listed
52
+ directly, instead of just this one's:
53
+
54
+ ```json
55
+ {
56
+ "mcpServers": {
57
+ "pipeworx": {
58
+ "url": "https://gateway.pipeworx.io/mcp"
59
+ }
60
+ }
61
+ }
62
+ ```
63
+
64
+ Both URLs reach the same gateway and the same 1679+ data sources. The
65
+ only difference is which pack's tools are listed **directly**; `ask_pipeworx`
66
+ reaches all of them from either one.
67
+
68
+ ## No MCP client? Call it over HTTP
69
+
70
+ ```bash
71
+ curl -X POST https://gateway.pipeworx.io/v1/tools/get_briefing \
72
+ -H 'Content-Type: application/json' \
73
+ -d '{}'
74
+ ```
75
+
76
+ No account needed for the first calls. Inspect any tool: `GET https://gateway.pipeworx.io/v1/tools/get_briefing`. Find one: `POST https://gateway.pipeworx.io/v1/tools/search_packs` with `{"query":"..."}`.
77
+
78
+ ## Standalone (no gateway account)
79
+
80
+ This package also runs as a local stdio MCP server — no Pipeworx account, no
81
+ gateway round-trip:
82
+
83
+ ```json
84
+ {
85
+ "mcpServers": {
86
+ "ai-briefing": {
87
+ "command": "npx",
88
+ "args": ["-y", "@pipeworx/mcp-ai-briefing"]
89
+ }
90
+ }
91
+ }
92
+ ```
93
+
94
+ Or run it directly to confirm it starts:
35
95
 
36
96
  ```bash
37
- npx pipeworx use ai-briefing
97
+ npx -y @pipeworx/mcp-ai-briefing
38
98
  ```
39
99
 
100
+ It speaks MCP over stdin/stdout and answers `initialize`/`tools/list`/`tools/call`
101
+ for **only** this pack's tools — none of the shared meta-tools the gateway
102
+ connection above adds. Same source, same tools, no ask_pipeworx routing.
103
+
104
+ ## Using with ask_pipeworx
105
+
106
+ Instead of calling tools directly, you can ask questions in plain English —
107
+ this works on the pack endpoint above as well as on the full gateway:
108
+
109
+ ```
110
+ ask_pipeworx({ question: "your question about Ai Briefing data" })
111
+ ```
112
+
113
+ The gateway picks the right tool and fills the arguments automatically.
114
+
115
+ ## More
116
+
117
+ - [Docs and guides](https://pipeworx.io/docs)
118
+ - [pipeworx.io](https://pipeworx.io)
119
+
40
120
  ## License
41
121
 
42
122
  MIT
package/bin/cli.js ADDED
@@ -0,0 +1,17 @@
1
+ #!/usr/bin/env node
2
+ //
3
+ // Entry point for `npx @pipeworx/mcp-<slug>`.
4
+ //
5
+ // Packs ship as raw TypeScript (no build step — see publish-pack.sh for why:
6
+ // tsx sidesteps every extensionless-import / bare-JSON-import edge case a
7
+ // per-pack tsc build would have to solve one pack at a time). This file
8
+ // registers tsx's ESM loader programmatically, then hands off to src/server.ts,
9
+ // which wraps the pack's {tools, callTool} export in a stdio MCP server.
10
+ //
11
+ // Copied verbatim into every published pack repo by scripts/publish-pack.sh —
12
+ // edit this file, not a per-pack copy.
13
+ import { register } from 'tsx/esm/api';
14
+
15
+ register();
16
+
17
+ await import('../src/server.ts');
package/package.json CHANGED
@@ -1,20 +1,32 @@
1
1
  {
2
2
  "name": "@pipeworx/mcp-ai-briefing",
3
- "version": "0.1.0",
3
+ "version": "0.1.1",
4
4
  "description": "AI Briefing MCP — Keep AI models current on industry developments",
5
5
  "type": "module",
6
6
  "main": "src/index.ts",
7
7
  "types": "src/index.ts",
8
+ "bin": {
9
+ "mcp-ai-briefing": "bin/cli.js"
10
+ },
8
11
  "keywords": ["mcp", "mcp-server", "model-context-protocol", "pipeworx", "ai-briefing"],
9
12
  "license": "MIT",
10
13
  "repository": {
11
14
  "type": "git",
12
- "url": "https://github.com/pipeworx-io/mcp-ai-briefing"
15
+ "url": "git+https://github.com/pipeworx-io/mcp-ai-briefing.git"
13
16
  },
14
17
  "scripts": {
15
18
  "typecheck": "tsc --noEmit"
16
19
  },
20
+ "dependencies": {
21
+ "@modelcontextprotocol/sdk": "^1.30.0",
22
+ "tsx": "^4.19.0"
23
+ },
17
24
  "devDependencies": {
18
- "typescript": "^5.7.0"
25
+ "typescript": "^5.9.3",
26
+ "@cloudflare/workers-types": "^4.20260405.1"
27
+ },
28
+ "pipeworx": {
29
+ "sourceHash": "v1-c0263f801c7031df0a0e8af3598f3a990afeaa992c2b7b0ab7130e4b520c8608",
30
+ "sourceCommit": "60f04b68c347e970c6be94a53439d1364061fc06"
19
31
  }
20
32
  }
package/server.json CHANGED
@@ -1,9 +1,9 @@
1
1
  {
2
2
  "$schema": "https://static.modelcontextprotocol.io/schemas/2025-12-11/server.schema.json",
3
3
  "name": "io.github.pipeworx-io/ai-briefing",
4
- "title": "ai uriefing",
4
+ "title": "Ai Briefing",
5
5
  "description": "AI Briefing MCP — Keep AI models current on industry developments",
6
- "version": "0.1.0",
6
+ "version": "0.1.1",
7
7
  "websiteUrl": "https://pipeworx.io/packs/ai-briefing",
8
8
  "repository": {
9
9
  "url": "https://github.com/pipeworx-io/mcp-ai-briefing",
package/src/index.ts CHANGED
@@ -1,18 +1,640 @@
1
1
  interface McpToolDefinition {
2
2
  name: string;
3
3
  description: string;
4
+ /** Human-facing one-liner (fleet #1967). Optional; consumers fall back to
5
+ * description. Kept in step with shared/src/types.ts — scripts/lib/
6
+ * check-inlined-types.mjs reports drift at publish time. */
7
+ summary?: string;
4
8
  inputSchema: {
5
9
  type: 'object';
6
10
  properties: Record<string, unknown>;
7
11
  required?: string[];
12
+ anyOf?: Array<{ required: string[] }>;
13
+ oneOf?: Array<{ required: string[] }>;
14
+ allOf?: Array<{ required: string[] }>;
8
15
  };
16
+ outputSchema?: Record<string, unknown>;
9
17
  }
10
18
 
11
19
  interface McpToolExport {
12
20
  tools: McpToolDefinition[];
13
21
  callTool: (name: string, args: Record<string, unknown>) => Promise<unknown>;
22
+ meter?: { credits: number };
23
+ cost?: Record<string, unknown>;
24
+ provider?: string;
14
25
  }
15
26
 
27
+ /**
28
+ * Was this failure OUR OWN web service? — the other half of `internal-db-class.ts`.
29
+ *
30
+ * fleet #1089 pulled failures from our own Postgres out of `upstream_down` by
31
+ * keying on the SQLSTATE inside PostgREST's four-key error envelope. That
32
+ * covered the majority and structurally could not cover the rest: the rest
33
+ * never reach Postgres, so they carry no SQLSTATE. What was left, measured over
34
+ * the 24h to 2026-09-02T15:00Z (fleet #1096):
35
+ *
36
+ * 5 pipeworx-catalog get_pack_tools Pipeworx catalog error: 522 — error code: 522
37
+ * 3 fleet fleet_list_open … upstream_down: Fleet task queue did not respond within 25s
38
+ *
39
+ * 521/522/523/526 are Cloudflare saying its edge could not reach an ORIGIN, and
40
+ * in both of those rows the origin is ours — `gateway.pipeworx.io` for the
41
+ * catalog pack (it self-fetches when the gateway hasn't injected a manifest),
42
+ * our own Supabase for fleet. There is no third party anywhere in either call.
43
+ * Same defect as #1089: our own outage filed under `upstream_down`, the one
44
+ * class that means "the source is unreachable and there is nothing for us to
45
+ * fix", which is why the problem-tools triage skips it.
46
+ *
47
+ * WHY NOT A WORDING RULE. The obvious fix is to match `fleet db error:` and
48
+ * `Pipeworx catalog error:` in classifyToolError. Each is emitted from exactly
49
+ * one site today, so it would work today. It would also rot the first time
50
+ * somebody rewords a label — silently, and in the direction of hiding our own
51
+ * outage, which is worse than the bug being fixed. Every prose rule in
52
+ * error-class.ts has needed widening as packs invented new wording (#409/#450/
53
+ * #584); that history is most of that file's comment budget.
54
+ *
55
+ * WHAT THIS KEYS ON INSTEAD: **the host the call actually reached.** A URL's
56
+ * hostname is a fact about the call, not a guess about its prose. Two
57
+ * consequences that a pack-level flag could not give us, and the reason the
58
+ * flag was rejected:
59
+ *
60
+ * - It describes the CALL, not the pack. `govcon-intel` fans out to our own
61
+ * Supabase AND to genuine third parties; `court-listener` holds our cache
62
+ * in Supabase and fetches courtlistener.com. An `internallyHosted: true` on
63
+ * either pack would relabel a real third-party outage as ours — inventing
64
+ * work, which is the same class of error in the opposite direction.
65
+ * - It covers every future internal pack for free, instead of one declared
66
+ * slug at a time.
67
+ *
68
+ * WHY IT SURVIVES A REWORD. The marker below is not matched as a literal by two
69
+ * separate files. `markInternalOrigin()` writes it and `internalHostMetricsClass()`
70
+ * reads it, both from the single exported `INTERNAL_ORIGIN_MARKER` constant in
71
+ * this module — so changing the wording changes both sides in the same edit and
72
+ * cannot desynchronise them. The pack's own label (`fleet db error:`,
73
+ * `Pipeworx catalog error:`) is not read at all: reword it freely, the class is
74
+ * unaffected. That is the property `stripClassPrefix` lacked when it drifted
75
+ * from its own classifier three times and needed a CI gate to hold them
76
+ * together.
77
+ *
78
+ * WHERE THE 5xx TEST LIVES. `markInternalOrigin` is called from the places that
79
+ * hold the real `Response` — `httpError`/`httpErrorMessage` and the timeout
80
+ * branch of `fetchWithTimeout` in `shared/src/http.ts` — so "is this an
81
+ * availability failure" is decided from the actual status code, never re-derived
82
+ * by scraping a number out of a sentence. A 404 from our own registry for a slug
83
+ * that does not exist is a caller's bad argument and is deliberately NOT marked.
84
+ */
85
+
86
+ /**
87
+ * OUR OWN web service was unreachable — not an upstream, and never `upstream_down`.
88
+ *
89
+ * ONE value, not three, unlike `internal_db_*`. That split existed because a
90
+ * slow query, an exhausted pool and an unknown SQLSTATE have different owners
91
+ * and different fixes. Here there is only one story to tell — an origin we run
92
+ * did not answer the edge — and one owner. A bucket with no distinct owner per
93
+ * value is decoration; #724 is what happens when a class holds several
94
+ * situations, and inventing sub-values ahead of a reason to act on them
95
+ * differently is the same mistake with the sign flipped.
96
+ *
97
+ * METRICS ONLY, exactly like PLATFORM_KEY_ERROR_CLASS and the internal_db
98
+ * values. `classifyToolError` still answers `upstream_down` for the retry and
99
+ * hint paths, which only care whether retrying or a sibling tool might work —
100
+ * and it might. Nothing a caller sees or is charged changes here.
101
+ *
102
+ * READ SIDE: this value is in BROKEN_TOOL_CLASSES, FAULT_CLASSES and
103
+ * ALL_ERROR_CLASSES in `workers/registry-api/src/index.ts`. All three, or it
104
+ * lands on no dashboard — fleet #721 is the warning, where the #719 split
105
+ * worked on the write side and was invisible for weeks.
106
+ */
107
+ const INTERNAL_SERVICE_UNREACHABLE_CLASS = 'internal_service_unreachable';
108
+
109
+ /**
110
+ * The token that carries "this origin is ours" from the call site to the
111
+ * classifier.
112
+ *
113
+ * Appended to the error message rather than attached to the Error object,
114
+ * because the object does not survive the trip: 275 packs return `{ error:
115
+ * string }` instead of throwing, the gateway reads `observedError` as a string,
116
+ * and the fleet pack rebuilds its error from a captured status + body across a
117
+ * retry loop. A property on an Error would be dropped by every one of those
118
+ * paths and the class would work in tests and vanish in production.
119
+ *
120
+ * WORDING IS LOAD-BEARING, same rule as labelAge's note in authority.ts. This
121
+ * string is appended to a pack's thrown Error message (shared/src/http.ts),
122
+ * and a thrown Error's message is exactly what the gateway hands back to the
123
+ * caller as `content[0].text` when nothing rewrites it (workers/gateway/src
124
+ * catches the throw and sets `rawResult.message = stripClassPrefix(error)`,
125
+ * which does not touch this suffix) — so the original wording,
126
+ * " [pipeworx-hosted origin — our own service, not a third party]", was not a
127
+ * theoretical leak: it shipped live on pipeworx-catalog's 522s, 7 times in 6
128
+ * hours on 2026-09-02 (see tests/golden-internal-service.test.ts), verbatim
129
+ * naming Pipeworx as the host. check:hosting-claims never caught it because it
130
+ * did not scan shared/ at all (task #2009). Reworded to describe the
131
+ * OBSERVATION (the origin did not answer) without a claim about who runs it —
132
+ * the identical fix labelAge got: drop the possessive, keep the fact.
133
+ */
134
+ const INTERNAL_ORIGIN_MARKER = ' [origin did not respond — retry before concluding the named source is down]';
135
+
136
+ /**
137
+ * Supabase's data plane for a project is `<ref>.supabase.co`, where the ref is
138
+ * exactly twenty lowercase letters (ours is `pqauisounztsgdgfkhke`).
139
+ *
140
+ * Matching the shape rather than listing the ref keeps this correct when we add
141
+ * a project — `supabaseEnv` on a pack entry already points some packs at a
142
+ * second one — while still excluding `status.supabase.co`, which is Supabase's
143
+ * own status page and emphatically not our database. Verified 2026-09-02 by
144
+ * `grep -rhoE '[a-z0-9-]+\.supabase\.(co|in)' mcps shared workers scripts`: the
145
+ * only real project ref anywhere in the tree is ours, the rest are doc
146
+ * placeholders (`abc`, `xyz`, `example`) which this pattern also excludes. Same
147
+ * finding internal-db-class.ts relies on for the PostgREST envelope being ours
148
+ * by construction.
149
+ */
150
+ const SUPABASE_PROJECT_HOST = /^[a-z]{20}\.supabase\.(co|in)$/;
151
+
152
+ /**
153
+ * Is this a host WE run?
154
+ *
155
+ * Deliberately NOT including `*.workers.dev`: plenty of third-party APIs are
156
+ * hosted on workers.dev, so the suffix says where something runs and not who
157
+ * owns it. Every internal call we actually make goes to a `pipeworx.io`
158
+ * hostname or to our Supabase project, both of which are ownership facts.
159
+ *
160
+ * `workers/gateway/src/provenance.ts`'s `OUR_HOSTS` answers the same
161
+ * question and DOES include `workers.dev` — a documented divergence
162
+ * (task #2051), not a bug to converge. That list decides what a response may
163
+ * cite as a data SOURCE, where a false negative (citing our own worker as an
164
+ * external source) is the hosting-disclosure leak this whole file exists to
165
+ * prevent, so it errs broad. This one decides who gets BLAMED for a 5xx in
166
+ * outage metrics read by on-call, where a false positive (crediting our own
167
+ * infra with a third party's outage) hides the real failure, so it errs
168
+ * narrow. Same suffix, opposite direction, because they are never called for
169
+ * the same reason.
170
+ *
171
+ * Returns false on anything unparseable rather than throwing — this runs inside
172
+ * an error path, and an error path that can itself throw turns a diagnosable
173
+ * failure into a mystery.
174
+ */
175
+ function isPipeworxOrigin(url: string | URL | undefined | null): boolean {
176
+ if (!url) return false;
177
+ let host: string;
178
+ try {
179
+ host = new URL(url instanceof URL ? url.href : url).hostname.toLowerCase();
180
+ } catch {
181
+ return false;
182
+ }
183
+ if (host === 'pipeworx.io' || host.endsWith('.pipeworx.io')) return true;
184
+ return SUPABASE_PROJECT_HOST.test(host);
185
+ }
186
+
187
+ /**
188
+ * Append the marker when this failure was OUR origin failing to answer.
189
+ *
190
+ * `status` is the HTTP status when there is one, and omitted for a timeout —
191
+ * where there is no response at all, and "the origin did not answer" is the
192
+ * whole observation. Statuses below 500 are left alone: a 404 from our own
193
+ * registry for a slug that does not exist is the caller's argument, not our
194
+ * outage, and marking it would put ordinary 404s on the incident dashboard.
195
+ *
196
+ * Idempotent, so a message that is wrapped and re-marked on the way up (the
197
+ * fleet pack's retry loop re-throws through two layers) carries the marker once.
198
+ */
199
+ function markInternalOrigin(
200
+ message: string,
201
+ url: string | URL | undefined | null,
202
+ status?: number,
203
+ ): string {
204
+ if (status !== undefined && status < 500) return message;
205
+ if (!isPipeworxOrigin(url)) return message;
206
+ if (message.includes(INTERNAL_ORIGIN_MARKER)) return message;
207
+ return message + INTERNAL_ORIGIN_MARKER;
208
+ }
209
+
210
+ /**
211
+ * Which blob4 value a failure from our own web services books as, or undefined
212
+ * if this is not one.
213
+ *
214
+ * Ordered AFTER `internalDbMetricsClass` at the call site: a PostgREST envelope
215
+ * from our own Supabase is a strictly more specific statement about the same
216
+ * row (which of our services, and why), and the two cannot disagree about
217
+ * whether the failure is ours.
218
+ */
219
+ function internalHostMetricsClass(error: string): string | undefined {
220
+ return error.includes(INTERNAL_ORIGIN_MARKER) ? INTERNAL_SERVICE_UNREACHABLE_CLASS : undefined;
221
+ }
222
+
223
+
224
+ /**
225
+ * One place to turn a failed `fetch` into an error a caller can act on.
226
+ *
227
+ * Nearly every pack was written the same way:
228
+ *
229
+ * if (!res.ok) throw new Error(`Unsplash: ${res.status}`);
230
+ *
231
+ * which discards the response body — and the body is usually where the upstream
232
+ * says what was actually wrong ("**symbol** not found: GBP", "parameter `year`
233
+ * out of range", "unknown taxonomy id"). The caller gets a number, cannot
234
+ * self-correct, and retries the same broken call. A 2026-07-31 sweep found this
235
+ * shape in 481 of 1,400 packs, 47 of them PLATFORM-keyed.
236
+ *
237
+ * It also hides bugs one level down. Two of the first three packs audited had a
238
+ * second defect that only existed because of this line: unsplash's rate-limit
239
+ * branch sat BELOW a catch-all and was unreachable, and bea-gov parsed
240
+ * `BEAAPI.Error.APIErrorDescription` below a `!res.ok` throw that made the
241
+ * parsing dead code for every non-200.
242
+ *
243
+ * DELIBERATELY NOT A CLASSIFIER. It does not add `user_error:` /
244
+ * `upstream_down:` prefixes. Those decide which tier a failure lands in, and the
245
+ * `error` tier is what the daily problem-tools list is built from — it means
246
+ * "Pipeworx has a defect". A 400 is genuinely ambiguous: often a caller's bad
247
+ * argument, but sometimes a query WE built wrong (ted-eu comma-joined its CPV
248
+ * values into something TED rejected, and that bug was found only because it sat
249
+ * in `error`). Blanket-classifying 400s as caller mistakes would have hidden it.
250
+ * A pack that KNOWS which it is should keep saying so explicitly; this helper is
251
+ * for the 481 that say nothing at all.
252
+ */
253
+
254
+ /** Longest upstream explanation we'll pass through. Enough for a real message,
255
+ * short enough that an HTML page or a stack trace can't swamp the error. */
256
+
257
+ const MAX_DETAIL = 300;
258
+
259
+ /**
260
+ * Default bound for `fetchWithTimeout` when a pack doesn't state its own.
261
+ *
262
+ * 25s mirrors the number `epo-ops` landed on after measuring the real failure:
263
+ * a degraded upstream that doesn't error, it just never answers, and a Worker
264
+ * sits in `await fetch()` until ITS OWN execution budget kills the request —
265
+ * which can take minutes, not seconds (epo_ops_search_patents measured 4-8
266
+ * MINUTE hangs before this existed). 25s is short enough that a caller gets a
267
+ * fast, actionable error instead of holding the connection, and long enough
268
+ * that it doesn't false-trip on a merely-slow-but-alive upstream.
269
+ */
270
+ const DEFAULT_FETCH_TIMEOUT_MS = 25_000;
271
+
272
+ /**
273
+ * Read the body of a failed response and fold it into a throwable Error.
274
+ *
275
+ * Usage — note the `await`, which is the one thing that makes this a mechanical
276
+ * change rather than a drop-in:
277
+ *
278
+ * if (!res.ok) throw await httpError(res, 'Unsplash');
279
+ *
280
+ * Safe to call on any non-ok response: a body that is missing, empty, unreadable
281
+ * or HTML degrades to exactly the old `Name: 404` string rather than throwing
282
+ * something new from inside the error path.
283
+ */
284
+ async function httpError(res: Response, name: string): Promise<Error> {
285
+ return new Error(await httpErrorMessage(res, name));
286
+ }
287
+
288
+ /** The message text without constructing an Error — for packs that need to wrap
289
+ * it in their own envelope or add an explicit classification prefix. */
290
+ async function httpErrorMessage(res: Response, name: string): Promise<string> {
291
+ // The one place a 5xx from a host WE run gets stamped as ours. `res.url` is
292
+ // the URL the fetch actually resolved to (after redirects), so this is a fact
293
+ // about the call rather than a guess from the `name` the pack passed in —
294
+ // reword that label freely, the class does not move. See
295
+ // internal-host-class.ts; no-op for every third-party upstream, which is why
296
+ // this touches 481 packs' error text and changes none of it.
297
+ return markInternalOrigin(
298
+ `${name}: ${res.status}${detailSuffix(await readDetail(res))}`,
299
+ res.url,
300
+ res.status,
301
+ );
302
+ }
303
+
304
+ /**
305
+ * Just the upstream's own explanation — no name, no status.
306
+ *
307
+ * For a pack that has already said both in its own sentence. epo-ops reads
308
+ * `EPO rejected this search as too large (HTTP 413) — ${httpErrorMessage(…)}`,
309
+ * which rendered as `… (HTTP 413) — EPO: 413.` once the XML detail was being
310
+ * dropped: the upstream named twice, the status twice, and the one thing EPO
311
+ * actually said ("Not enough characters before truncation character") nowhere
312
+ * (fleet #712). Returns '' when the body carries nothing readable, so a caller
313
+ * can fall back to its own wording.
314
+ */
315
+ async function upstreamDetail(res: Response): Promise<string> {
316
+ return readDetail(res);
317
+ }
318
+
319
+ /**
320
+ * Read a SUCCESSFUL response as JSON, failing loudly when it isn't JSON.
321
+ *
322
+ * `httpError` above only ever runs on `!res.ok`, which leaves the nastier half
323
+ * of the problem unhandled: an upstream that answers **HTTP 200 with an HTML
324
+ * page**. A bot wall, a login redirect, a maintenance interstitial and a CDN
325
+ * error page are all 200s, so `res.ok` is true, and `res.json()` then throws
326
+ * `Unexpected token '<', "<!DOCTYPE "... is not valid JSON`.
327
+ *
328
+ * That string is the problem. It names no upstream, carries no status, and
329
+ * reads like a parser bug in Pipeworx — so it lands in the `error` tier, which
330
+ * means "we have a defect", and the caller is told nothing they can act on.
331
+ * data.govt.nz sat dead behind an Imperva challenge this way and every
332
+ * status-code health check we own reported it green (7889a845). A zero-length
333
+ * body has the same shape: `Unexpected end of JSON input`, seen this week on
334
+ * uk-gazette (83% of external calls) and census.
335
+ *
336
+ * UNLIKE `httpError`, this one DOES classify, and the asymmetry is deliberate.
337
+ * A 400 is genuinely ambiguous — often the caller's bad argument, sometimes a
338
+ * query we built wrong — so blanket-classifying it would hide our own bugs.
339
+ * There is no such ambiguity here: **no argument a caller can pass makes a JSON
340
+ * API return an HTML page.** It is always the upstream, so `upstream_down:` is
341
+ * a statement of fact rather than a guess, and it keeps these out of the
342
+ * problem-tools list where they crowd out real defects.
343
+ *
344
+ * const data = await parseJson<Feed>(res, 'UK Gazette');
345
+ *
346
+ * Call it only after the `!res.ok` check — on a failed response you want
347
+ * `httpError`, which mines the body for the upstream's own explanation.
348
+ */
349
+ async function parseJson<T>(res: Response, name: string): Promise<T> {
350
+ let raw: string;
351
+ try {
352
+ raw = await res.text();
353
+ } catch {
354
+ throw new Error(
355
+ `upstream_down: ${name} returned a body that could not be read (HTTP ${res.status}). ` +
356
+ 'The connection most likely dropped mid-response; retrying is reasonable.',
357
+ );
358
+ }
359
+
360
+ const type = res.headers.get('content-type') ?? 'no content-type';
361
+
362
+ if (!raw.trim()) {
363
+ throw new Error(
364
+ `upstream_down: ${name} answered HTTP ${res.status} with an EMPTY body where JSON was expected (${type}). ` +
365
+ 'Nothing about the request can cause this — it is an upstream fault, and the same call may well work on retry.',
366
+ );
367
+ }
368
+
369
+ // Checked before parsing rather than in the catch, because knowing it is
370
+ // markup is what turns "we failed to parse something" into "they served a
371
+ // web page" — the second is diagnosable, the first is not.
372
+ const head = raw.slice(0, 200).trimStart().toLowerCase();
373
+ if (head.startsWith('<!doctype') || head.startsWith('<html') || head.startsWith('<?xml')) {
374
+ const kind = head.startsWith('<?xml') ? 'an XML document' : 'an HTML page';
375
+ // The summary, not the source. Pasting the first 120 characters of a web
376
+ // page handed the agent `<!DOCTYPE html><html lang="en"…` — the same leak
377
+ // this branch exists to describe (fleet #712).
378
+ throw new Error(
379
+ `upstream_down: ${name} answered HTTP ${res.status} with ${kind} instead of JSON (${type}). ` +
380
+ 'That is typically a bot wall, a login redirect or a maintenance page — it is returned as a SUCCESS, ' +
381
+ `so status-code health checks read it as fine. No argument change will get past it. ` +
382
+ `The page says: ${summarizeErrorBody(raw) || 'nothing readable'}`,
383
+ );
384
+ }
385
+
386
+ try {
387
+ return JSON.parse(raw) as T;
388
+ } catch {
389
+ throw new Error(
390
+ `upstream_down: ${name} answered HTTP ${res.status} with a body that is not valid JSON (${type}). ` +
391
+ `It begins: ${stripMarkup(raw).slice(0, 120) || '(unreadable)'}`,
392
+ );
393
+ }
394
+ }
395
+
396
+ /**
397
+ * `fetch`, but bounded — the fix for a systemic gap found 2026-08-30: a grep
398
+ * audit of every pack's `mcps/*\/src/index.ts` found 1,339 of ~1,500 call
399
+ * `fetch()` with NO timeout guard anywhere in the file. Two of those
400
+ * (epo-ops, statcan) were confirmed live-hanging for 4-8 minutes before this
401
+ * existed — every unguarded call carries the same risk, just unconfirmed.
402
+ *
403
+ * Mirrors the `epoFetch` wrapper `mcps/epo-ops/src/index.ts` shipped first:
404
+ * bound the request with `AbortSignal.timeout`, and on a timeout/abort throw
405
+ * an `upstream_down:` error that names the upstream and the bound rather than
406
+ * letting the raw `TimeoutError`/`AbortError` (which names neither) propagate.
407
+ * `upstream_down:` is deliberate, same reasoning as `parseJson` above — no
408
+ * argument a caller passes can make an upstream hang, so it is always the
409
+ * upstream's fault, and marking it that way keeps a slow API off the
410
+ * problem-tools list where it would crowd out our own defects.
411
+ *
412
+ * Usage — a mechanical swap for a bare `fetch(url, init)`:
413
+ *
414
+ * const res = await fetchWithTimeout(url, init, 'Some API');
415
+ *
416
+ * Pass `timeoutMs` as a fourth argument to override the default for a pack
417
+ * with a known-slower upstream; the label should be the same short name you'd
418
+ * pass to `httpError`/`httpErrorMessage` for that call.
419
+ */
420
+ async function fetchWithTimeout(
421
+ url: string | URL,
422
+ init: RequestInit = {},
423
+ name: string,
424
+ timeoutMs: number = DEFAULT_FETCH_TIMEOUT_MS,
425
+ ): Promise<Response> {
426
+ try {
427
+ return await fetch(url, { ...init, signal: AbortSignal.timeout(timeoutMs) });
428
+ } catch (err) {
429
+ if (err instanceof Error && (err.name === 'TimeoutError' || err.name === 'AbortError')) {
430
+ // States the OBSERVATION (no response in N seconds), not a diagnosis.
431
+ // "appears to be degraded" is an inference about the vendor that we have
432
+ // not checked, and it is wrong in a way that misdirects whoever reads it:
433
+ // a timeout from a Worker can equally mean OUR egress is blocked.
434
+ //
435
+ // Measured today (2026-09-01, fleet #1047): every call to
436
+ // mainnet.base.org failed from the x402 facilitator while the identical
437
+ // request from a laptop returned 200. Base was entirely healthy; the
438
+ // public RPC refuses Cloudflare Worker egress. Had this message fired
439
+ // there it would have blamed Base by name, and the next person would have
440
+ // waited for a vendor outage to clear that did not exist.
441
+ // A timeout has no status to test — there is no response at all — so
442
+ // `markInternalOrigin` is called without one: an origin we run that never
443
+ // answered is an availability failure by definition. This is the half of
444
+ // fleet #1096 with neither a SQLSTATE nor a status code to key on.
445
+ throw new Error(
446
+ markInternalOrigin(
447
+ `upstream_down: ${name} did not respond within ${timeoutMs / 1000}s. ` +
448
+ `That can be ${name} being slow or down, or this environment being unable to reach it ` +
449
+ `(some hosts refuse datacenter/Worker egress) — retry shortly, and check reachability ` +
450
+ `from elsewhere before concluding ${name} is down.`,
451
+ url,
452
+ ),
453
+ );
454
+ }
455
+ throw err;
456
+ }
457
+ }
458
+
459
+ function detailSuffix(detail: string): string {
460
+ return detail ? ` — ${detail}` : '';
461
+ }
462
+
463
+ async function readDetail(res: Response): Promise<string> {
464
+ let raw: string;
465
+ try {
466
+ raw = await res.text();
467
+ } catch {
468
+ // Body already consumed, or the connection died mid-read. The status alone
469
+ // is still worth throwing — never let the error path throw its own error.
470
+ return '';
471
+ }
472
+ return summarizeErrorBody(raw);
473
+ }
474
+
475
+ /**
476
+ * Turn ANY error body — JSON, HTML, XML or plain text — into one short phrase
477
+ * that never contains markup.
478
+ *
479
+ * This used to just drop an HTML or XML body on the floor, on the reasoning
480
+ * that markup crowds out the status. That was half right. Dropping it loses the
481
+ * one sentence a caller could have acted on: an `Access Denied` title, an SDMX
482
+ * `<message:Error>` text, an OPS fault string. A 2026-08-30 support sweep
483
+ * measured 13 of 291 caller-facing error rows carrying a raw page or document
484
+ * verbatim, across 11 packs, and in every one of them the useful content —
485
+ * "Access Denied", "Invalid country code", "SCRAPE_TIMEOUT" — was in there,
486
+ * buried in markup the agent had to parse out of a string (fleet #712).
487
+ *
488
+ * So: extract the meaning, discard the markup. The output is passed through
489
+ * `stripMarkup` unconditionally, which is what lets `check:error-body-leak`
490
+ * assert mechanically that no caller-facing message can contain `<?xml`,
491
+ * `<!DOCTYPE` or `<html`.
492
+ */
493
+ function summarizeErrorBody(raw: string): string {
494
+ if (!raw || !raw.trim()) return '';
495
+
496
+ const head = raw.slice(0, 400).trimStart().toLowerCase();
497
+
498
+ // An HTML error page (Cloudflare interstitial, nginx default, a login
499
+ // redirect) says what it is in its <title>, and almost nowhere else.
500
+ if (head.startsWith('<!doctype') || head.startsWith('<html')) {
501
+ const title = htmlTitle(raw);
502
+ return title
503
+ ? `${title} (upstream returned an HTML error page, not an API response)`
504
+ : 'upstream returned an HTML error page, not an API response';
505
+ }
506
+
507
+ // XML fault documents — EPO OPS, SDMX (`<message:Error>`), SOAP faults. The
508
+ // human sentence sits in a child element whose tag name says what it is.
509
+ if (head.startsWith('<?xml') || head.startsWith('<')) {
510
+ const fault = xmlFaultText(raw);
511
+ return fault
512
+ ? `${stripMarkup(fault).slice(0, MAX_DETAIL)} (from the upstream's XML error document)`
513
+ : 'upstream returned an XML error document with no readable message';
514
+ }
515
+
516
+ // Most JSON error bodies bury one human sentence among ids and echoed request
517
+ // params. Prefer that sentence; fall back to the whole body when the shape is
518
+ // unfamiliar, since an unfamiliar shape is exactly when we can least afford to
519
+ // guess wrong and show nothing.
520
+ const fromJson = messageFromJson(raw);
521
+ return stripMarkup(fromJson ?? raw).slice(0, MAX_DETAIL);
522
+ }
523
+
524
+ /** The `<title>` of an HTML error page, or its first `<h1>` — the two places a
525
+ * bot wall, a 502 and an "Access Denied" all state what happened. */
526
+ function htmlTitle(raw: string): string | null {
527
+ const head = raw.slice(0, 4000);
528
+ for (const re of [/<title[^>]*>([\s\S]*?)<\/title>/i, /<h1[^>]*>([\s\S]*?)<\/h1>/i]) {
529
+ const m = re.exec(head);
530
+ const text = m ? stripMarkup(m[1]) : '';
531
+ if (text) return text.slice(0, 160);
532
+ }
533
+ return null;
534
+ }
535
+
536
+ /** Tag names that carry the explanation in an XML fault document, namespace
537
+ * prefix optional (`<message:Error>`, `<com:Text>`, `<faultstring>`). */
538
+ const XML_FAULT_TAG_RE =
539
+ /<(?:[A-Za-z0-9_.-]+:)?(?:text|message|description|faultstring|reason|detail|title|errormessage|error)\b[^>]*>([^<]{2,400})</i;
540
+
541
+ function xmlFaultText(raw: string): string | null {
542
+ const head = raw.slice(0, 8000);
543
+ const tagged = XML_FAULT_TAG_RE.exec(head);
544
+ if (tagged && tagged[1].trim()) return tagged[1];
545
+
546
+ // Nothing conventionally named — take the longest text node instead. A fault
547
+ // document with one sentence in an oddly named element is still readable;
548
+ // returning nothing at all is not.
549
+ let best = '';
550
+ for (const m of head.matchAll(/>([^<>]{8,400})</g)) {
551
+ const text = m[1].trim();
552
+ if (text.length > best.length) best = text;
553
+ }
554
+ return best || null;
555
+ }
556
+
557
+ /**
558
+ * Remove every tag and stray angle bracket, then collapse whitespace.
559
+ *
560
+ * Applied to everything on the way out, including the JSON and plain-text
561
+ * paths, because an upstream is free to embed markup in a JSON string field —
562
+ * and a leak is a leak regardless of which branch produced it.
563
+ */
564
+ function stripMarkup(s: string): string {
565
+ return collapse(decodeEntities(s.replace(/<[^>]*>/g, ' ')).replace(/[<>]/g, ' '));
566
+ }
567
+
568
+ /** The handful of entities that show up in error-page titles. Decoded AFTER
569
+ * tags are stripped and BEFORE the angle-bracket sweep, so `&lt;script&gt;`
570
+ * in a title cannot decode into markup that survives — EMBL-EBI's ChEMBL 500
571
+ * page renders as `500 Internal Server Error &lt; EMBL-EBI` otherwise. */
572
+ function decodeEntities(s: string): string {
573
+ return s
574
+ .replace(/&(?:amp|#0*38);/gi, '&')
575
+ .replace(/&(?:lt|#0*60);/gi, '<')
576
+ .replace(/&(?:gt|#0*62);/gi, '>')
577
+ .replace(/&(?:quot|#0*34);/gi, '"')
578
+ .replace(/&(?:#0*39|apos|#x0*27);/gi, "'")
579
+ .replace(/&nbsp;/gi, ' ');
580
+ }
581
+
582
+ /** The conventional "what went wrong" field, under any of the names upstreams
583
+ * actually use. Checked in order; first non-empty string wins. */
584
+ const MESSAGE_KEYS = [
585
+ 'message', 'error_message', 'errorMessage', 'detail', 'details',
586
+ 'description', 'error_description', 'reason', 'title', 'fault',
587
+ ];
588
+
589
+ function messageFromJson(raw: string): string | null {
590
+ let parsed: unknown;
591
+ try {
592
+ parsed = JSON.parse(raw);
593
+ } catch {
594
+ return null;
595
+ }
596
+ return pickMessage(parsed, 0);
597
+ }
598
+
599
+ function pickMessage(node: unknown, depth: number): string | null {
600
+ // Two levels covers `{error: {message}}` and `{errors: [{detail}]}`, the two
601
+ // shapes that account for nearly all of them, without walking a large payload.
602
+ if (depth > 2 || node == null) return null;
603
+
604
+ if (typeof node === 'string') return node.trim() || null;
605
+
606
+ if (Array.isArray(node)) {
607
+ for (const item of node) {
608
+ const found = pickMessage(item, depth + 1);
609
+ if (found) return found;
610
+ }
611
+ return null;
612
+ }
613
+
614
+ if (typeof node !== 'object') return null;
615
+ const obj = node as Record<string, unknown>;
616
+
617
+ for (const key of MESSAGE_KEYS) {
618
+ const v = obj[key];
619
+ if (typeof v === 'string' && v.trim()) return v.trim();
620
+ }
621
+ // `{error: …}` where error is itself an object or a string — the single most
622
+ // common wrapper, so it is worth descending into by name rather than scanning
623
+ // every key and risking picking up an echoed request parameter.
624
+ for (const key of ['error', 'errors', 'fault', 'Error', 'data']) {
625
+ if (key in obj) {
626
+ const found = pickMessage(obj[key], depth + 1);
627
+ if (found) return found;
628
+ }
629
+ }
630
+ return null;
631
+ }
632
+
633
+ /** Errors are read in a single line of log output; newlines and runs of
634
+ * whitespace make a multi-line body unreadable there. */
635
+ function collapse(s: string): string {
636
+ return s.replace(/\s+/g, ' ').trim();
637
+ }
16
638
  /**
17
639
  * AI Briefing MCP — Keep AI models current on industry developments
18
640
  *
@@ -30,12 +652,20 @@ interface McpToolExport {
30
652
  */
31
653
 
32
654
 
655
+ // Bound every fetch() in this pack to a fixed timeout — an upstream that
656
+ // degrades without erroring would otherwise hold the Worker in `await fetch()`
657
+ // until its own execution budget kills the request (minutes, not seconds).
658
+ // Mirrors the epoFetch / usaspending retryFetch pattern (fleet #685).
659
+ async function pwFetch(url: string | URL, init?: RequestInit): Promise<Response> {
660
+ return fetchWithTimeout(url, init ?? {}, 'AI Briefing');
661
+ }
662
+
33
663
  const SUPABASE_URL = 'https://pqauisounztsgdgfkhke.supabase.co';
34
664
  const SUPABASE_KEY = 'eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJzdXBhYmFzZSIsInJlZiI6InBxYXVpc291bnp0c2dkZ2ZraGtlIiwicm9sZSI6InNlcnZpY2Vfcm9sZSIsImlhdCI6MTc3NDM3MDMxMSwiZXhwIjoyMDg5OTQ2MzExfQ.gciWwNdmss8ko0ThAeUjQiFgwlMWxEP6LSucyyjTfcA';
35
665
 
36
666
  async function querySupabase(path: string): Promise<unknown> {
37
667
  const url = `${SUPABASE_URL}/rest/v1/${path}`;
38
- const res = await fetch(url, {
668
+ const res = await pwFetch(url, {
39
669
  headers: {
40
670
  apikey: SUPABASE_KEY,
41
671
  Authorization: `Bearer ${SUPABASE_KEY}`,
@@ -43,7 +673,7 @@ async function querySupabase(path: string): Promise<unknown> {
43
673
  });
44
674
  if (!res.ok) {
45
675
  const body = await res.text().catch(() => '');
46
- throw new Error(`Supabase ${res.status}: ${body.slice(0, 200)}`);
676
+ throw new Error(`data query ${res.status}: ${body.slice(0, 200)}`);
47
677
  }
48
678
  return res.json();
49
679
  }
@@ -51,7 +681,7 @@ async function querySupabase(path: string): Promise<unknown> {
51
681
  const tools: McpToolExport['tools'] = [
52
682
  {
53
683
  name: 'get_briefing',
54
- description: 'Get today\'s AI tools briefing — new MCP servers, APIs, SDKs, agent frameworks, and developer tools released in the last 24 hours. Call this at the start of any session to discover new tools you can use.',
684
+ description: 'Get the daily AI tools digest for a given date (default: today) — new MCP servers, APIs, SDKs, and frameworks released in the last 24 hours, with summaries and source URLs.',
55
685
  inputSchema: {
56
686
  type: 'object' as const,
57
687
  properties: {
@@ -62,7 +692,7 @@ const tools: McpToolExport['tools'] = [
62
692
  },
63
693
  {
64
694
  name: 'search_developments',
65
- description: 'Search for new tools, APIs, MCP servers, and frameworks by keyword. Returns matching developments across HN, GitHub, HuggingFace, and AI company blogs. Use for queries like "new MCP servers", "vector database tools", or "Claude integrations".',
695
+ description: 'Search for new tools, APIs, MCP servers, and frameworks by keyword (e.g., \'vector databases\', \'Claude integrations\'). Returns matching developments with descriptions and sources.',
66
696
  inputSchema: {
67
697
  type: 'object' as const,
68
698
  properties: {
@@ -74,11 +704,11 @@ const tools: McpToolExport['tools'] = [
74
704
  },
75
705
  {
76
706
  name: 'get_recent',
77
- description: 'Get recent tool and API releases filtered by category, source, or timeframe. Categories: mcp, tool, agent_framework, open_source, model_release, integration, infrastructure, product, paper. Sources: hackernews, arxiv, github, huggingface, openai_blog, anthropic_blog, google_ai_blog, meta_ai_blog.',
707
+ description: 'Retrieve AI developments from the last N days (default 7), filterable by category (e.g., model_release, paper, mcp), source (e.g., arxiv, github), and importance (low/normal/high/breaking). The category is a keyword label derived from the headline at ingest, not a curated classification — filtering by it narrows the feed but does not guarantee every row belongs; the response repeats this caveat whenever a category filter is applied.',
78
708
  inputSchema: {
79
709
  type: 'object' as const,
80
710
  properties: {
81
- category: { type: 'string', description: 'Filter by category (e.g., model_release, paper, funding)' },
711
+ category: { type: 'string', description: 'Filter by category (e.g., model_release, paper, funding). Heuristic headline label, not a curated taxonomy — check the titles.' },
82
712
  source: { type: 'string', description: 'Filter by source (e.g., arxiv, hackernews, github)' },
83
713
  days: { type: 'number', description: 'Look back N days (default 7)' },
84
714
  importance: { type: 'string', description: 'Filter by importance: low, normal, high, breaking' },
@@ -89,7 +719,7 @@ const tools: McpToolExport['tools'] = [
89
719
  },
90
720
  {
91
721
  name: 'get_model_landscape',
92
- description: 'Get recent AI model releases — what models shipped, from which companies, what they can do. Useful for knowing what\'s available to build with right now.',
722
+ description: 'List AI model releases from the last N days (default 30). Returns model names, provider companies, release dates, feature summaries, and source URLs grouped by importance. Membership comes from the same heuristic headline label as get_recent(category=model_release), so verify the titles before quoting the list as complete.',
93
723
  inputSchema: {
94
724
  type: 'object' as const,
95
725
  properties: {
@@ -100,7 +730,7 @@ const tools: McpToolExport['tools'] = [
100
730
  },
101
731
  {
102
732
  name: 'get_timeline',
103
- description: 'Get a chronological timeline of AI developments between two dates. Useful for understanding what happened during a specific period.',
733
+ description: 'Get a chronological timeline of AI developments between two dates. Returns events ordered by date with descriptions for understanding a specific period.',
104
734
  inputSchema: {
105
735
  type: 'object' as const,
106
736
  properties: {
@@ -114,7 +744,7 @@ const tools: McpToolExport['tools'] = [
114
744
  },
115
745
  {
116
746
  name: 'get_ai_toolbelt',
117
- description: 'Get the latest tools, features, and capabilities you can use RIGHT NOW. Returns new Claude Code features, MCP servers, SDK updates, CLI tools, and integrations. Call this to discover what new tools are in your toolbelt since your training cutoff.',
747
+ description: 'Get the latest available tools — Claude Code features, MCP servers, SDK updates, CLI tools, integrations. Returns new capabilities since your training cutoff.',
118
748
  inputSchema: {
119
749
  type: 'object' as const,
120
750
  properties: {
@@ -126,7 +756,7 @@ const tools: McpToolExport['tools'] = [
126
756
  },
127
757
  {
128
758
  name: 'get_ai_news',
129
- description: 'Get AI industry news — model releases, funding rounds, acquisitions, policy changes, benchmark results. Separate from toolbelt; this is about what happened in the AI industry, not tools you can use.',
759
+ description: 'Get AI industry news — model releases, funding, acquisitions, policy changes, benchmarks. Returns news events with dates and summaries for industry context.',
130
760
  inputSchema: {
131
761
  type: 'object' as const,
132
762
  properties: {
@@ -138,7 +768,7 @@ const tools: McpToolExport['tools'] = [
138
768
  },
139
769
  {
140
770
  name: 'what_happened',
141
- description: 'Natural language query about recent tools and developments. Ask "any new MCP servers this week", "latest Claude tools", "new open source frameworks", "what APIs launched recently". Returns the most relevant tool-related developments.',
771
+ description: 'Ask natural language questions about recent tools and developments (e.g., \'any new MCP servers this week\', \'latest Claude tools\'). Returns the most relevant developments.',
142
772
  inputSchema: {
143
773
  type: 'object' as const,
144
774
  properties: {
@@ -192,12 +822,39 @@ async function getBriefing(date?: string): Promise<unknown> {
192
822
  };
193
823
  }
194
824
 
195
- // No digest yet — build one from raw developments
196
- const since = new Date(new Date(targetDate).getTime() - 24 * 60 * 60 * 1000).toISOString();
825
+ // No digest yet — build one from raw developments published AROUND the
826
+ // requested date. Bound BOTH ends of the window: with only a lower bound, a
827
+ // past date (e.g. 2025-01-15) returned the LATEST 30 developments regardless
828
+ // of the date asked for — an ungrounded answer ("asked Jan 2025, got 2026").
829
+ const center = new Date(targetDate).getTime();
830
+ if (Number.isNaN(center)) {
831
+ throw new Error(`Invalid date "${targetDate}". Use YYYY-MM-DD.`);
832
+ }
833
+ const windowStart = new Date(center - 24 * 60 * 60 * 1000).toISOString();
834
+ const windowEnd = new Date(center + 24 * 60 * 60 * 1000).toISOString();
197
835
  const devs = await querySupabase(
198
- `ai_developments?published_at=gte.${since}&order=published_at.desc&limit=30`
836
+ `ai_developments?published_at=gte.${windowStart}&published_at=lt.${windowEnd}&order=published_at.desc&limit=30`
199
837
  ) as Array<Record<string, unknown>>;
200
838
 
839
+ if (devs.length === 0) {
840
+ // Nothing for that date — say so honestly (with the actual coverage range)
841
+ // rather than silently returning the latest briefing.
842
+ const earliest = await querySupabase(
843
+ `ai_developments?select=published_at&order=published_at.asc&limit=1`
844
+ ) as Array<Record<string, unknown>>;
845
+ const latest = await querySupabase(
846
+ `ai_developments?select=published_at&order=published_at.desc&limit=1`
847
+ ) as Array<Record<string, unknown>>;
848
+ const lo = (earliest[0]?.published_at as string | undefined)?.slice(0, 10) ?? '?';
849
+ const hi = (latest[0]?.published_at as string | undefined)?.slice(0, 10) ?? '?';
850
+ return {
851
+ date: targetDate,
852
+ found: false,
853
+ note: `No AI briefing or developments recorded around ${targetDate}. Coverage runs ${lo} to ${hi}. For the latest, call get_briefing with no date; for a range use get_recent({days}) or get_timeline({start_date,end_date}).`,
854
+ developments: [],
855
+ };
856
+ }
857
+
201
858
  return {
202
859
  date: targetDate,
203
860
  note: 'Live query (no cached digest available)',
@@ -235,6 +892,23 @@ async function searchDevelopments(query: string, limit: number): Promise<unknown
235
892
  };
236
893
  }
237
894
 
895
+ /**
896
+ * `ai_developments.category` is assigned at ingest by keyword rules over the
897
+ * HEADLINE ALONE — there is no body text and no model in the loop. It is a
898
+ * hint, not a taxonomy, and every payload that filters on it has to say so.
899
+ *
900
+ * fleet #2301: filtering to `model_release` returned an umbrella-insurance
901
+ * Launch HN, a resolved status-page incident and a Linux alpha, while the
902
+ * actual releases in the same window ("Claude Opus 5.5", "Grok 4.7") sat under
903
+ * `news`. The rules were tightened in workers/ai-briefing-scraper, but rows
904
+ * already stored keep the label they were written with — the ingest path uses
905
+ * `resolution=ignore-duplicates`, so nothing is ever recategorised in place.
906
+ * The caveat is therefore permanent, not a note about one bad week.
907
+ */
908
+ const CATEGORY_CAVEAT =
909
+ 'category is a keyword label applied to the headline at ingest, not a curated classification — ' +
910
+ 'rows can be mislabelled in both directions, so read the titles rather than trusting the filter.';
911
+
238
912
  async function getRecent(args: Record<string, unknown>): Promise<unknown> {
239
913
  const days = (args.days as number) ?? 7;
240
914
  const limit = (args.limit as number) ?? 20;
@@ -249,9 +923,41 @@ async function getRecent(args: Record<string, unknown>): Promise<unknown> {
249
923
  `ai_developments?${filters.join('&')}&order=published_at.desc&limit=${limit}`
250
924
  ) as Array<Record<string, unknown>>;
251
925
 
926
+ // Say what this feed is ABOUT (fleet #2238). get_recent takes no entity, so
927
+ // a caller who asked about one specific model gets the general AI feed —
928
+ // and the payload never said so, which is why 39/79 answers scored
929
+ // `unverifiable`. `subject` names it; the prose restates it with the scope.
930
+ const scope = [
931
+ args.category ? `category ${String(args.category)}` : null,
932
+ args.source ? `source ${String(args.source)}` : null,
933
+ args.importance ? `importance ${String(args.importance)}` : null,
934
+ ].filter(Boolean);
935
+ const newest = devs[0];
936
+ const subject = `AI industry developments${scope.length ? ` (${scope.join(', ')})` : ''}, last ${days} days`;
937
+ const statement =
938
+ `Recent artificial intelligence (AI) industry developments from the ai-briefing feed over the last ${days} days` +
939
+ (scope.length ? `, filtered to ${scope.join(', ')}` : ', all categories and sources') +
940
+ `: ${devs.length} item(s)` +
941
+ (newest ? `, newest "${String(newest.title)}" (${String(newest.source)}, ${String(newest.published_at).slice(0, 10)})` : '') +
942
+ '.' +
943
+ // Do not let the prose assert the filter worked. It is true about the
944
+ // query we ran and false about what the caller gets back (fleet #2301).
945
+ (args.category ? ` Note: ${CATEGORY_CAVEAT}` : '');
946
+
252
947
  return {
948
+ subject,
949
+ statement,
253
950
  period: `last ${days} days`,
254
951
  total: devs.length,
952
+ ...(args.category
953
+ ? {
954
+ category_filter: {
955
+ applied: String(args.category),
956
+ label_quality: 'heuristic',
957
+ note: CATEGORY_CAVEAT,
958
+ },
959
+ }
960
+ : {}),
255
961
  developments: devs.map((d) => ({
256
962
  title: d.title,
257
963
  summary: d.summary,
@@ -273,6 +979,13 @@ async function getModelLandscape(days: number): Promise<unknown> {
273
979
  return {
274
980
  period: `last ${days} days`,
275
981
  model_releases: devs.length,
982
+ // Same column, same caveat (fleet #2301) — this tool IS the model_release
983
+ // filter, so it inherits the label's accuracy wholesale.
984
+ category_filter: {
985
+ applied: 'model_release',
986
+ label_quality: 'heuristic',
987
+ note: CATEGORY_CAVEAT,
988
+ },
276
989
  models: devs.map((d) => ({
277
990
  title: d.title,
278
991
  summary: d.summary,
@@ -301,6 +1014,15 @@ async function getTimeline(args: Record<string, unknown>): Promise<unknown> {
301
1014
  start_date: startDate,
302
1015
  end_date: endDate,
303
1016
  total: devs.length,
1017
+ ...(args.category
1018
+ ? {
1019
+ category_filter: {
1020
+ applied: String(args.category),
1021
+ label_quality: 'heuristic',
1022
+ note: CATEGORY_CAVEAT,
1023
+ },
1024
+ }
1025
+ : {}),
304
1026
  timeline: devs.map((d) => ({
305
1027
  date: (d.published_at as string)?.slice(0, 10),
306
1028
  title: d.title,
@@ -367,9 +1089,14 @@ async function getAiNews(days: number, limit: number): Promise<unknown> {
367
1089
 
368
1090
  async function whatHappened(question: string, days: number): Promise<unknown> {
369
1091
  // Extract keywords from the question for search
1092
+ // Punctuation SPLITS words rather than being deleted: "haiku-4-5" used to
1093
+ // collapse to "haiku45", which matches nothing in the feed, so the single
1094
+ // most-asked what_happened question over the 7 days to 2026-09-18 (37
1095
+ // calls) returned 0 items every time (fleet #2238). Same for
1096
+ // "claude-code", "gpt-5" and every other hyphenated product name.
370
1097
  const keywords = question
371
1098
  .toLowerCase()
372
- .replace(/[^a-z0-9\s]/g, '')
1099
+ .replace(/[^a-z0-9\s]/g, ' ')
373
1100
  .split(/\s+/)
374
1101
  .filter((w) => w.length > 2 && !['what', 'the', 'any', 'new', 'latest', 'recent', 'happened', 'with', 'about', 'has', 'have', 'been', 'there'].includes(w));
375
1102
 
@@ -407,12 +1134,24 @@ async function whatHappened(question: string, days: number): Promise<unknown> {
407
1134
  return new Date(b.published_at as string).getTime() - new Date(a.published_at as string).getTime();
408
1135
  });
409
1136
 
1137
+ const shown = allDevs.slice(0, 15);
1138
+ const newest = shown[0];
1139
+ // What the answer is about, in words (fleet #2238). Only the keywords the
1140
+ // database actually matched are named as the subject: an item is in this
1141
+ // list because its title or summary contains one of them, so the claim is
1142
+ // checkable against the rows below rather than an echo of the question.
1143
+ const statement = shown.length
1144
+ ? `AI developments from the last ${days} days whose title or summary matched ${keywords.slice(0, 3).join(', ')}: ${allDevs.length} item(s), showing ${shown.length}` +
1145
+ (newest ? `, top "${String(newest.title)}" (${String(newest.source)}, ${String(newest.published_at).slice(0, 10)})` : '') +
1146
+ '.'
1147
+ : `No AI developments in the last ${days} days matched ${keywords.slice(0, 3).join(', ')}.`;
410
1148
  return {
411
1149
  question,
412
1150
  keywords_used: keywords,
1151
+ statement,
413
1152
  period: `last ${days} days`,
414
1153
  total: allDevs.length,
415
- developments: allDevs.slice(0, 15).map((d) => ({
1154
+ developments: shown.map((d) => ({
416
1155
  title: d.title,
417
1156
  summary: d.summary,
418
1157
  url: d.url,
package/src/server.ts ADDED
@@ -0,0 +1,45 @@
1
+ /**
2
+ * Stdio MCP server entry point for @pipeworx/mcp-ai-briefing.
3
+ * Generated by scripts/publish-pack.sh — do not hand-edit in the pack repo;
4
+ * edit scripts/publish-pack.sh (the server.ts heredoc) and republish instead.
5
+ */
6
+ import { Server } from '@modelcontextprotocol/sdk/server/index.js';
7
+ import { StdioServerTransport } from '@modelcontextprotocol/sdk/server/stdio.js';
8
+ import { CallToolRequestSchema, ListToolsRequestSchema } from '@modelcontextprotocol/sdk/types.js';
9
+ import pack from './index.js';
10
+
11
+ const server = new Server(
12
+ { name: '@pipeworx/mcp-ai-briefing', version: '0.1.1' },
13
+ { capabilities: { tools: {} } },
14
+ );
15
+
16
+ server.setRequestHandler(ListToolsRequestSchema, async () => ({
17
+ tools: pack.tools.map((t) => ({
18
+ name: t.name,
19
+ description: t.description,
20
+ inputSchema: t.inputSchema,
21
+ })),
22
+ }));
23
+
24
+ server.setRequestHandler(CallToolRequestSchema, async (request) => {
25
+ const { name, arguments: args } = request.params;
26
+ try {
27
+ const result = await pack.callTool(name, (args ?? {}) as Record<string, unknown>);
28
+ return { content: [{ type: 'text', text: JSON.stringify(result, null, 2) }] };
29
+ } catch (err) {
30
+ return {
31
+ content: [{ type: 'text', text: err instanceof Error ? err.message : String(err) }],
32
+ isError: true,
33
+ };
34
+ }
35
+ });
36
+
37
+ async function main() {
38
+ const transport = new StdioServerTransport();
39
+ await server.connect(transport);
40
+ }
41
+
42
+ main().catch((err) => {
43
+ console.error('Fatal error running server:', err);
44
+ process.exit(1);
45
+ });
package/tsconfig.json CHANGED
@@ -3,12 +3,16 @@
3
3
  "target": "ES2022",
4
4
  "module": "ESNext",
5
5
  "moduleResolution": "bundler",
6
+ "lib": ["ES2022"],
7
+ "types": ["@cloudflare/workers-types"],
6
8
  "strict": true,
7
9
  "esModuleInterop": true,
8
10
  "skipLibCheck": true,
11
+ "resolveJsonModule": true,
9
12
  "outDir": "dist",
10
13
  "rootDir": "src",
11
14
  "declaration": true
12
15
  },
13
- "include": ["src"]
16
+ "include": ["src"],
17
+ "exclude": ["src/server.ts"]
14
18
  }