mcp-scraper 0.38.2 → 0.40.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (111) hide show
  1. package/README.md +5 -2
  2. package/package.json +5 -6
  3. package/dist/bin/api-server.cjs +0 -58752
  4. package/dist/bin/api-server.cjs.map +0 -1
  5. package/dist/bin/api-server.d.cts +0 -1
  6. package/dist/bin/api-server.d.ts +0 -1
  7. package/dist/bin/api-server.js +0 -38
  8. package/dist/bin/api-server.js.map +0 -1
  9. package/dist/bin/mcp-scraper-cli.cjs +0 -2671
  10. package/dist/bin/mcp-scraper-cli.cjs.map +0 -1
  11. package/dist/bin/mcp-scraper-cli.d.cts +0 -1
  12. package/dist/bin/mcp-scraper-cli.d.ts +0 -1
  13. package/dist/bin/mcp-scraper-cli.js +0 -742
  14. package/dist/bin/mcp-scraper-cli.js.map +0 -1
  15. package/dist/bin/mcp-scraper-install.cjs +0 -129
  16. package/dist/bin/mcp-scraper-install.cjs.map +0 -1
  17. package/dist/bin/mcp-scraper-install.d.cts +0 -1
  18. package/dist/bin/mcp-scraper-install.d.ts +0 -1
  19. package/dist/bin/mcp-scraper-install.js +0 -27
  20. package/dist/bin/mcp-scraper-install.js.map +0 -1
  21. package/dist/bin/mcp-stdio-server.cjs +0 -12264
  22. package/dist/bin/mcp-stdio-server.cjs.map +0 -1
  23. package/dist/bin/mcp-stdio-server.d.cts +0 -1
  24. package/dist/bin/mcp-stdio-server.d.ts +0 -1
  25. package/dist/bin/mcp-stdio-server.js +0 -135
  26. package/dist/bin/mcp-stdio-server.js.map +0 -1
  27. package/dist/bin/paa-harvest.cjs +0 -3808
  28. package/dist/bin/paa-harvest.cjs.map +0 -1
  29. package/dist/bin/paa-harvest.d.cts +0 -1
  30. package/dist/bin/paa-harvest.d.ts +0 -1
  31. package/dist/bin/paa-harvest.js +0 -44
  32. package/dist/bin/paa-harvest.js.map +0 -1
  33. package/dist/chunk-345BQXZH.js +0 -712
  34. package/dist/chunk-345BQXZH.js.map +0 -1
  35. package/dist/chunk-44HZLHDV.js +0 -52
  36. package/dist/chunk-44HZLHDV.js.map +0 -1
  37. package/dist/chunk-AZRPG43B.js +0 -617
  38. package/dist/chunk-AZRPG43B.js.map +0 -1
  39. package/dist/chunk-CB5C3BPB.js +0 -135
  40. package/dist/chunk-CB5C3BPB.js.map +0 -1
  41. package/dist/chunk-EGJKUB4Q.js +0 -276
  42. package/dist/chunk-EGJKUB4Q.js.map +0 -1
  43. package/dist/chunk-FQI5PFE7.js +0 -1866
  44. package/dist/chunk-FQI5PFE7.js.map +0 -1
  45. package/dist/chunk-FRYT3ID4.js +0 -684
  46. package/dist/chunk-FRYT3ID4.js.map +0 -1
  47. package/dist/chunk-G3P3ZDB4.js +0 -69
  48. package/dist/chunk-G3P3ZDB4.js.map +0 -1
  49. package/dist/chunk-K443GQY5.js +0 -24
  50. package/dist/chunk-K443GQY5.js.map +0 -1
  51. package/dist/chunk-N7KUTTCC.js +0 -3007
  52. package/dist/chunk-N7KUTTCC.js.map +0 -1
  53. package/dist/chunk-NGM237OO.js +0 -3410
  54. package/dist/chunk-NGM237OO.js.map +0 -1
  55. package/dist/chunk-NKCCGADE.js +0 -11285
  56. package/dist/chunk-NKCCGADE.js.map +0 -1
  57. package/dist/chunk-NNW3O6ZD.js +0 -108
  58. package/dist/chunk-NNW3O6ZD.js.map +0 -1
  59. package/dist/chunk-QZXKQB7Y.js +0 -414
  60. package/dist/chunk-QZXKQB7Y.js.map +0 -1
  61. package/dist/chunk-SFRMFGQ6.js +0 -158
  62. package/dist/chunk-SFRMFGQ6.js.map +0 -1
  63. package/dist/chunk-YCI2PNCS.js +0 -499
  64. package/dist/chunk-YCI2PNCS.js.map +0 -1
  65. package/dist/chunk-YODBNTTN.js +0 -7
  66. package/dist/chunk-YODBNTTN.js.map +0 -1
  67. package/dist/db-C5KVCOYT.js +0 -239
  68. package/dist/db-C5KVCOYT.js.map +0 -1
  69. package/dist/extract-bundle-KUBX6N6Z.js +0 -568
  70. package/dist/extract-bundle-KUBX6N6Z.js.map +0 -1
  71. package/dist/index.cjs +0 -4160
  72. package/dist/index.cjs.map +0 -1
  73. package/dist/index.d.cts +0 -413
  74. package/dist/index.d.ts +0 -413
  75. package/dist/index.js +0 -338
  76. package/dist/index.js.map +0 -1
  77. package/dist/location-data-repository-TTWF3OTM.js +0 -35
  78. package/dist/location-data-repository-TTWF3OTM.js.map +0 -1
  79. package/dist/server-5EX6XBIA.js +0 -33596
  80. package/dist/server-5EX6XBIA.js.map +0 -1
  81. package/dist/site-extract-repository-XSPJTCIL.js +0 -62
  82. package/dist/site-extract-repository-XSPJTCIL.js.map +0 -1
  83. package/dist/worker-XCPU4YSN.js +0 -142
  84. package/dist/worker-XCPU4YSN.js.map +0 -1
  85. package/docs/adr/0001-in-page-graphql-interception-for-anti-bot-scraping.md +0 -58
  86. package/docs/adr/0002-hybrid-smart-rag-vault-retrieval.md +0 -62
  87. package/docs/adr/0003-waive-unrecoverable-scheduled-model-cost.md +0 -22
  88. package/docs/adr/README.md +0 -13
  89. package/docs/final-tooling-spec.md +0 -206
  90. package/docs/hosted-location-data.md +0 -108
  91. package/docs/kernel-proxy-future-enhancements.md +0 -80
  92. package/docs/mcp-tool-craft-lint.generated.md +0 -183
  93. package/docs/mcp-tool-design-guide.md +0 -225
  94. package/docs/mcp-tool-manifest.generated.json +0 -22871
  95. package/docs/mcp-tool-quality-spec.md +0 -240
  96. package/docs/oauth-legal-review.md +0 -38
  97. package/docs/seo-crawl-report-spec.md +0 -287
  98. package/docs/specs/api-forge-spec.md +0 -234
  99. package/docs/specs/connected-services-control-plane-decoupling-spec.md +0 -1044
  100. package/docs/specs/deferred-work-spec.md +0 -86
  101. package/docs/specs/google-drive-bulk-access-and-mcp-schema-passthrough-spec.md +0 -1689
  102. package/docs/specs/kernel-stealth-captcha-test-matrix.md +0 -278
  103. package/docs/specs/main-mcp-integration-ownership-spec.md +0 -1164
  104. package/docs/specs/mcp-tool-definition-quality-audit-spec.md +0 -1602
  105. package/docs/specs/meta-ad-creative-media-resolution-spec.md +0 -31
  106. package/docs/specs/multimodal-image-memory-architecture-spec.md +0 -1022
  107. package/docs/specs/oauth-mcp-spec.md +0 -213
  108. package/docs/specs/query-fanout-transport-contract-fix.md +0 -45
  109. package/docs/specs/relationship-workspace-ai-behavior-plan.md +0 -26
  110. package/docs/specs/unified-credit-and-scheduled-execution-billing-spec.md +0 -995
  111. package/docs/tool-catalog-spec.md +0 -388
@@ -1,108 +0,0 @@
1
- # Hosted location data
2
-
3
- Production directory workflows and `location_markets` read versioned location snapshots from Turso. They do not read a server-local CSV and do not call Census during a customer request.
4
-
5
- Two datasets are active independently:
6
-
7
- - `us_census_places:<STATE>` stores every Census place and its 2020–2025 population estimates for one state.
8
- - `us_zip_groups` stores normalized city ZIP and county groups for all imported states.
9
-
10
- Each import receives an immutable dataset ID. An active pointer is switched only after every normalized row is durable, so readers see either the old complete version or the new complete version. Responses include both active dataset IDs, source URLs, and refresh timestamps.
11
-
12
- ## Required environment
13
-
14
- ```dotenv
15
- TURSO_DATABASE_URL=...
16
- TURSO_AUTH_TOKEN=...
17
- MCP_SCRAPER_LOCATION_ADMIN_KEY=...
18
- ```
19
-
20
- `MCP_SCRAPER_USZIPS_CSV_URL` may hold the default public HTTPS ZIP source. The import endpoint also accepts a source URL explicitly. `MCP_SCRAPER_USZIPS_CSV_PATH` remains available only for local tests and local CLI runs.
21
-
22
- ZIP imports are fail-closed: a request URL must exactly match `MCP_SCRAPER_USZIPS_CSV_URL` or one entry in `MCP_SCRAPER_LOCATION_IMPORT_SOURCE_ALLOWLIST`. Redirects are rejected, private-network targets are rejected, and response bytes are streamed through a 64 MiB in-flight limit. Census sync does not accept a URL; it constructs the trusted `www2.census.gov` state URL server-side.
23
-
24
- ## Publish the nationwide GeoNames source
25
-
26
- The production ZIP/county snapshot is derived from the official GeoNames US postal-code archive at `https://download.geonames.org/export/zip/US.zip`. GeoNames postal-code data is licensed under [Creative Commons Attribution 4.0](https://creativecommons.org/licenses/by/4.0/); retain attribution to [GeoNames](https://www.geonames.org/) anywhere this derived dataset is redistributed.
27
-
28
- Publish a normalized, immutable public CSV to Vercel Blob:
29
-
30
- ```bash
31
- export LOCATION_DATA_READ_WRITE_TOKEN='...'
32
- npm run locations:publish-geonames
33
- ```
34
-
35
- `BLOB_READ_WRITE_TOKEN` is also accepted for projects that use their default
36
- public Blob store. Production uses a dedicated public store connected with the
37
- `LOCATION_DATA_` environment-variable prefix so this dataset cannot collide
38
- with private connected-data artifacts.
39
-
40
- The publisher downloads the direct HTTPS ZIP without following redirects, applies status, timeout, compressed-size, and extracted-size limits, extracts only `US.txt`, and converts the documented tab-separated fields to `state_abbr,zipcode,county,city`. It then runs the same production CSV parser and nationwide gate used by the import route: all 50 states plus DC and at least 25,000 distinct five-digit ZIPs must be present.
41
-
42
- The Blob pathname includes the normalized CSV's SHA-256 digest. Random suffixes and overwrites are disabled, so a published URL is an immutable snapshot. The command prints only non-secret counts, hashes, paths, and public URLs; copy its `url` value into `MCP_SCRAPER_USZIPS_CSV_URL`. A repeat publication of byte-identical data intentionally fails instead of replacing the existing content-addressed object.
43
-
44
- ## Seed and refresh
45
-
46
- Bootstrap the nationwide ZIP snapshot plus all 50 states and DC in one command:
47
-
48
- ```bash
49
- export MCP_SCRAPER_LOCATION_ADMIN_KEY='...'
50
- export MCP_SCRAPER_USZIPS_CSV_URL='https://...blob.vercel-storage.com/location-data/geonames/us/geonames-us-<sha256>.csv'
51
- npm run locations:sync
52
- ```
53
-
54
- The bootstrap uses three concurrent state requests by default, retries retryable HTTP/network failures up to three times, prints a per-state result and final failed-state list, and exits nonzero if any dataset failed. It never prints the admin key or ZIP source URL. Useful bounded refreshes include:
55
-
56
- ```bash
57
- # Refresh only selected Census snapshots and leave the active ZIP version unchanged.
58
- npm run locations:sync -- --skip-zips --states=TN,CO,CA
59
-
60
- # Point the operator at a preview API and lower concurrency.
61
- npm run locations:sync -- --base-url=https://preview.example.com --concurrency=2
62
- ```
63
-
64
- The individual admin calls below remain available for targeted recovery.
65
-
66
- Sync one state's trusted Census snapshot:
67
-
68
- ```bash
69
- curl -sS https://mcpscraper.dev/locations/admin/sync-census \
70
- -H "x-location-admin-key: $MCP_SCRAPER_LOCATION_ADMIN_KEY" \
71
- -H 'content-type: application/json' \
72
- --data '{"state":"TN"}'
73
- ```
74
-
75
- Import the nationwide ZIP/county CSV from a public HTTPS object:
76
-
77
- ```bash
78
- curl -sS https://mcpscraper.dev/locations/admin/import \
79
- -H "x-location-admin-key: $MCP_SCRAPER_LOCATION_ADMIN_KEY" \
80
- -H 'content-type: application/json' \
81
- --data '{"sourceUrl":"https://data.example.com/uszips.csv"}'
82
- ```
83
-
84
- The ZIP import accepts the common header aliases `state_abbr|state|state_id`, `zipcode|zip|zip_code`, `city|primary_city`, and `county|county_name`. A replacement must cover all 50 states plus DC and contain at least 25,000 distinct five-digit ZIPs. Unsupported state codes and malformed rows are counted as rejected rows in the dataset record. A partial or undersized snapshot is rejected without changing the active pointer. Re-importing byte-identical active data is idempotent.
85
-
86
- Census imports verify that every place row's `STATE` FIPS matches the requested state before activation. `/locations/status` reports missing population states, required state count, ZIP coverage, rejected rows, and readiness; `available` becomes true only when all 51 population snapshots and a nationwide ZIP snapshot are active.
87
-
88
- ## Read and verify
89
-
90
- Dataset readiness requires an API key:
91
-
92
- ```bash
93
- curl -sS https://mcpscraper.dev/locations/status \
94
- -H "x-api-key: $MCP_SCRAPER_API_KEY"
95
- ```
96
-
97
- Query hosted markets directly:
98
-
99
- ```bash
100
- curl -sS 'https://mcpscraper.dev/locations/markets?state=TN&populationYear=2025&minPopulation=100000&maxResults=25&includeZipGroups=true' \
101
- -H "x-api-key: $MCP_SCRAPER_API_KEY"
102
- ```
103
-
104
- The same read is exposed to agents as `location_markets`. Filters are applied against the hosted population and ZIP rows before the population-sorted result limit. If a required state snapshot or ZIP snapshot has not been activated, the API returns a typed `503` instead of silently falling back to a local path or a live Census fetch.
105
-
106
- ## Operational follow-up
107
-
108
- The active-pointer design prevents partial data from becoming visible and preserves old dataset rows for rollback. A future maintenance pass should add an explicit cross-process import lease, an admin rollback/history endpoint, and retention cleanup for superseded snapshot rows. These do not weaken current read isolation or coverage gates, but they will make concurrent operator refreshes and long-term storage management more ergonomic.
@@ -1,80 +0,0 @@
1
- # Kernel / proxy retry — future enhancement notes
2
-
3
- **Status:** ideas only — NOT implemented, NO code from here is in the tree.
4
- **Origin:** salvaged from an abandoned automated branch `codex/kernel-proxy-retry-cleanup`
5
- (a stale `~/Projects/mcp-scraper` clone, since deleted). That branch was evaluated against the
6
- current Desktop code and found **superseded** — Desktop independently built the same retry +
7
- error-classification machinery and went further (fresh disposable proxy per attempt, state
8
- escalation, rollout to maps/Facebook/site-crawl). **Do not `git merge` that branch** — it
9
- diverged onto an older base and a merge would regress the current proxy layer.
10
-
11
- The two items below are the *only* things Desktop genuinely lacked. They are documented here as
12
- "if we want to build X, here's how" — to be rebuilt from scratch onto current code, never pasted.
13
-
14
- ---
15
-
16
- ## A. SERP-timeout salvage (best candidate)
17
-
18
- **Problem it solves:** when a Google page load hits a navigation timeout (e.g.
19
- `domcontentloaded` slow) but the page is *actually* a usable `/search` results page with no
20
- CAPTCHA, the current code throws and triggers a wasted retry. We can instead detect "good page,
21
- just slow" and continue.
22
-
23
- **Where:** `src/driver/BrowserDriver.ts`, in the `goto` flow.
24
-
25
- **How to build:**
26
- 1. Add a predicate `currentGoogleSerpIsUsable(page)` — returns true when the current URL is a
27
- Google `/search` page, results containers are present, and no CAPTCHA/consent interstitial is
28
- detected.
29
- 2. In `goto`, catch the navigation-timeout case (an `isNavigationTimeout(err)` guard). Before
30
- rethrowing, if `currentGoogleSerpIsUsable(page)` is true, log a "salvaged slow SERP" telemetry
31
- line and proceed instead of throwing.
32
- 3. Keep the throw for every other timeout (non-SERP pages, CAPTCHA present, empty page).
33
-
34
- **Tools it improves:** `search_serp`, `harvest_paa`, `capture_serp_*` — fewer discarded captures
35
- and fewer retry credits burned on pages that were fine.
36
-
37
- **Watch out for:** only salvage when results are genuinely present — a timed-out blank/consent
38
- page must still fail so the normal retry/rotation kicks in.
39
-
40
- ---
41
-
42
- ## B. Opt-in proxy connectivity-check bypass (situational)
43
-
44
- **Problem it solves:** the browser-service residential proxies run an internal connectivity ping
45
- (an S3-style reachability check) before use. When that ping falsely fails, proxy creation is
46
- rejected even though the proxy would work — surfacing as `proxy_unavailable`.
47
-
48
- **Where:** `src/kernel-proxy-resolver.ts`, in the residential-proxy creation path.
49
-
50
- **How to build:**
51
- 1. Gate behind an env flag `KERNEL_PROXY_BYPASS_CONNECTIVITY_CHECK` (default off).
52
- 2. When set, pass a `bypass_hosts` / skip-connectivity-check option into the proxy-create call so
53
- the internal ping is skipped.
54
- 3. Add a `shouldRetryWithoutBypass` fallback: if a bypassed proxy then fails for real, retry once
55
- with the check re-enabled so we don't mask a genuinely dead proxy.
56
-
57
- **Build ONLY if** we actually observe proxy-create failures from the connectivity ping in
58
- production. Otherwise leave it — it's complexity with no payoff until that failure mode is real.
59
-
60
- ---
61
-
62
- ## C. Minor niceties (low priority, build only if convenient)
63
-
64
- - **Per-attempt hard timeout:** a fresh `AbortController` per retry attempt with its own
65
- `perAttemptTimeoutMs` (~120s), chained to the parent signal. Desktop bounds attempt *count* but
66
- has no per-attempt hard cap.
67
- - **Typed `KernelProxyUnavailableError`:** Desktop detects proxy-unavailable via regex on the
68
- message; a dedicated error class would be cleaner to branch on.
69
- - **Request-level retry knobs:** expose `maxAttempts`, `retryOnCaptcha`, `retryOnProxyFailure`,
70
- `allowLocationFallback` as request schema fields instead of hardcoded constants, for per-call
71
- tuning / debugging.
72
-
73
- ---
74
-
75
- ## What NOT to take
76
-
77
- Everything else from the codex branch — the retry loop, the `proxy_tunnel_failed` /
78
- `proxy_unavailable` / `location_mismatch` error classification, `LocationMismatchError`, the
79
- 503/retryable mappings — is **already in Desktop** (`src/harvest.ts`, `src/errors.ts`,
80
- `src/api/harvest-problems.ts`) and in a more advanced form. Do not reintroduce it.
@@ -1,183 +0,0 @@
1
- # MCP Tool Craft Lint
2
-
3
- Generated: 2026-07-28T19:01:51.307Z
4
- Manifest: `/private/tmp/mcp-scraper-archive-release/docs/mcp-tool-manifest.generated.json`
5
-
6
- Tools: 172
7
- Checks per tool: 7
8
- Failing tools: 0
9
-
10
- | Tool | Score | Missing |
11
- |---|---:|---|
12
- | `access-accept-share` | 7/7 | none |
13
- | `access-approve-sender` | 7/7 | none |
14
- | `access-decline-share` | 7/7 | none |
15
- | `access-inbox-settings` | 7/7 | none |
16
- | `access-invite-account` | 7/7 | none |
17
- | `access-issue-key` | 7/7 | none |
18
- | `access-list-approved-senders` | 7/7 | none |
19
- | `access-list-keys` | 7/7 | none |
20
- | `access-note-inbox` | 7/7 | none |
21
- | `access-remove-approved-sender` | 7/7 | none |
22
- | `access-revoke-key` | 7/7 | none |
23
- | `access-revoke-share` | 7/7 | none |
24
- | `access-set-scope` | 7/7 | none |
25
- | `access-share-note` | 7/7 | none |
26
- | `access-share-vault` | 7/7 | none |
27
- | `access-swap-vault` | 7/7 | none |
28
- | `access-switch-account` | 7/7 | none |
29
- | `access-unlink-share` | 7/7 | none |
30
- | `add-vault` | 7/7 | none |
31
- | `archive_read` | 7/7 | none |
32
- | `audit_site` | 7/7 | none |
33
- | `browser_click` | 7/7 | none |
34
- | `browser_close` | 7/7 | none |
35
- | `browser_extension_delete` | 7/7 | none |
36
- | `browser_extension_import` | 7/7 | none |
37
- | `browser_extension_list` | 7/7 | none |
38
- | `browser_goto` | 7/7 | none |
39
- | `browser_list_replays` | 7/7 | none |
40
- | `browser_list_sessions` | 7/7 | none |
41
- | `browser_locate` | 7/7 | none |
42
- | `browser_open` | 7/7 | none |
43
- | `browser_press` | 7/7 | none |
44
- | `browser_profile_connect` | 7/7 | none |
45
- | `browser_profile_list` | 7/7 | none |
46
- | `browser_read` | 7/7 | none |
47
- | `browser_replay_annotate` | 7/7 | none |
48
- | `browser_replay_download` | 7/7 | none |
49
- | `browser_replay_mark` | 7/7 | none |
50
- | `browser_replay_start` | 7/7 | none |
51
- | `browser_replay_stop` | 7/7 | none |
52
- | `browser_screenshot` | 7/7 | none |
53
- | `browser_scroll` | 7/7 | none |
54
- | `browser_type` | 7/7 | none |
55
- | `bulk-delete-notes` | 7/7 | none |
56
- | `call_service_connection_action` | 7/7 | none |
57
- | `capture_serp_page_snapshots` | 7/7 | none |
58
- | `capture_serp_snapshot` | 7/7 | none |
59
- | `check_site_export` | 7/7 | none |
60
- | `cost-usage` | 7/7 | none |
61
- | `create-channel` | 7/7 | none |
62
- | `create-scheduled-action` | 7/7 | none |
63
- | `create-secure-vault` | 7/7 | none |
64
- | `create-webhook` | 7/7 | none |
65
- | `credits_info` | 7/7 | none |
66
- | `delete-note` | 7/7 | none |
67
- | `delete-scheduled-action` | 7/7 | none |
68
- | `delete-vault` | 7/7 | none |
69
- | `describe_service_connection_tool` | 7/7 | none |
70
- | `diff_page` | 7/7 | none |
71
- | `directory_workflow` | 7/7 | none |
72
- | `directory_workflow_status` | 7/7 | none |
73
- | `export_connected_service_data` | 7/7 | none |
74
- | `export_search_console_table_data` | 7/7 | none |
75
- | `extract_site` | 7/7 | none |
76
- | `extract_url` | 7/7 | none |
77
- | `facebook_ad_search` | 7/7 | none |
78
- | `facebook_ad_transcribe` | 7/7 | none |
79
- | `facebook_page_intel` | 7/7 | none |
80
- | `facebook_video_transcribe` | 7/7 | none |
81
- | `fact-history` | 7/7 | none |
82
- | `g2_reviews` | 7/7 | none |
83
- | `get-chat-link` | 7/7 | none |
84
- | `get-message-note` | 7/7 | none |
85
- | `get-schedule-link` | 7/7 | none |
86
- | `get-schedule-status` | 7/7 | none |
87
- | `get-vault-app-link` | 7/7 | none |
88
- | `get-vault-contract` | 7/7 | none |
89
- | `gmail_search_contacts` | 7/7 | none |
90
- | `gmail_send_message` | 7/7 | none |
91
- | `google_ads_page_intel` | 7/7 | none |
92
- | `google_ads_search` | 7/7 | none |
93
- | `google_ads_transcribe` | 7/7 | none |
94
- | `google_calendar_create_event` | 7/7 | none |
95
- | `harvest_paa` | 7/7 | none |
96
- | `import_service_connection_to_memory` | 7/7 | none |
97
- | `instagram_media_download` | 7/7 | none |
98
- | `instagram_profile_content` | 7/7 | none |
99
- | `library-ingest` | 7/7 | none |
100
- | `list_service_connections` | 7/7 | none |
101
- | `list-channel-members` | 7/7 | none |
102
- | `list-channel-messages` | 7/7 | none |
103
- | `list-memory-tags` | 7/7 | none |
104
- | `list-scheduled-actions` | 7/7 | none |
105
- | `list-shared-with-me` | 7/7 | none |
106
- | `list-vaults` | 7/7 | none |
107
- | `list-webhooks` | 7/7 | none |
108
- | `location_markets` | 7/7 | none |
109
- | `map_site_urls` | 7/7 | none |
110
- | `map_wayback_snapshots` | 7/7 | none |
111
- | `maps_place_intel` | 7/7 | none |
112
- | `maps_search` | 7/7 | none |
113
- | `memory-backlinks` | 7/7 | none |
114
- | `memory-capture` | 7/7 | none |
115
- | `memory-export` | 7/7 | none |
116
- | `memory-get` | 7/7 | none |
117
- | `memory-graph-path` | 7/7 | none |
118
- | `memory-graph-universe` | 7/7 | none |
119
- | `memory-list` | 7/7 | none |
120
- | `memory-put` | 7/7 | none |
121
- | `memory-questions` | 7/7 | none |
122
- | `memory-search` | 7/7 | none |
123
- | `memory-suggest` | 7/7 | none |
124
- | `memory-upload` | 7/7 | none |
125
- | `merge-memory-tags` | 7/7 | none |
126
- | `meta_ad_creative_media` | 7/7 | none |
127
- | `my-mentions` | 7/7 | none |
128
- | `pause-scheduled-action` | 7/7 | none |
129
- | `poll-channel` | 7/7 | none |
130
- | `post-message` | 7/7 | none |
131
- | `prepare-memory-write` | 7/7 | none |
132
- | `propose-scheduled-action` | 7/7 | none |
133
- | `provision-defaults` | 7/7 | none |
134
- | `query_fanout_workflow` | 7/7 | none |
135
- | `rank_tracker_workflow` | 7/7 | none |
136
- | `react-message` | 7/7 | none |
137
- | `read_service_connection` | 7/7 | none |
138
- | `record-fact` | 7/7 | none |
139
- | `reddit_thread` | 7/7 | none |
140
- | `reddit_trending` | 7/7 | none |
141
- | `remove-channel-member` | 7/7 | none |
142
- | `renew_connected_data_download` | 7/7 | none |
143
- | `reply-message` | 7/7 | none |
144
- | `report_artifact_read` | 7/7 | none |
145
- | `resolve-memory-tags` | 7/7 | none |
146
- | `resume-scheduled-action` | 7/7 | none |
147
- | `revoke-chat-link` | 7/7 | none |
148
- | `revoke-schedule-link` | 7/7 | none |
149
- | `revoke-vault-app-link` | 7/7 | none |
150
- | `revoke-webhook` | 7/7 | none |
151
- | `route-memory` | 7/7 | none |
152
- | `search_serp` | 7/7 | none |
153
- | `set_scheduled_action_connections` | 7/7 | none |
154
- | `set-agent-identity` | 7/7 | none |
155
- | `set-schedule-defaults` | 7/7 | none |
156
- | `set-schedule-entitlement` | 7/7 | none |
157
- | `slack_send_message` | 7/7 | none |
158
- | `storage-usage` | 7/7 | none |
159
- | `table-create` | 7/7 | none |
160
- | `table-delete-rows` | 7/7 | none |
161
- | `table-describe` | 7/7 | none |
162
- | `table-drop` | 7/7 | none |
163
- | `table-insert-rows` | 7/7 | none |
164
- | `table-list` | 7/7 | none |
165
- | `table-query` | 7/7 | none |
166
- | `temporal-recall` | 7/7 | none |
167
- | `test_service_connection` | 7/7 | none |
168
- | `trustpilot_reviews` | 7/7 | none |
169
- | `upsert-memory-tag` | 7/7 | none |
170
- | `validate-memory-write` | 7/7 | none |
171
- | `video_frame_analysis` | 7/7 | none |
172
- | `video_frame_analysis_status` | 7/7 | none |
173
- | `video-analyze-start` | 7/7 | none |
174
- | `video-analyze-status` | 7/7 | none |
175
- | `workflow_artifact_read` | 7/7 | none |
176
- | `workflow_list` | 7/7 | none |
177
- | `workflow_run` | 7/7 | none |
178
- | `workflow_status` | 7/7 | none |
179
- | `workflow_step` | 7/7 | none |
180
- | `workflow_suggest` | 7/7 | none |
181
- | `youtube_harvest` | 7/7 | none |
182
- | `youtube_transcribe` | 7/7 | none |
183
- | `zoom_create_meeting` | 7/7 | none |
@@ -1,225 +0,0 @@
1
- # Designing MCP Tools an AI Can Actually Find and Use
2
-
3
- A practical guide to building MCP servers whose tools an LLM (Claude, but the mechanics generalize)
4
- will **discover, select correctly, and invoke well**. It is written from how the model *actually*
5
- perceives tools at runtime — not how a human reads your docs page.
6
-
7
- The single thesis: **the model routes on what it can see, and it can see almost nothing until it asks.**
8
- Everything below follows from that.
9
-
10
- ---
11
-
12
- ## 1. How the model actually sees your tools (the mechanic that governs everything)
13
-
14
- Modern MCP clients **defer** tool schemas to save context. The lifecycle in a session:
15
-
16
- 1. **At rest — names only.** When the server connects, the model gets a flat list of tool
17
- **identifiers** (`mcp__<server>__<tool>`) — *names, nothing else*. No parameters, no descriptions,
18
- no input/output shapes. For ~50 tools that's a cheap, scannable block.
19
- 2. **Server-instructions block — the one thing loaded upfront.** An MCP server may ship a top-level
20
- `instructions` string. This **is** in context from the start, alongside the names. It's the only
21
- place you can influence routing *before* the model fetches anything.
22
- 3. **Discovery — the model fetches a schema.** To use a tool, the model runs a tool-search step
23
- (exact `select:<name>` or a keyword query) that loads that tool's **full definition** — description
24
- prose + parameter JSONSchema — into context. *This is the first moment the model sees how the tool works.*
25
- 4. **Call → result.** The model invokes with arguments; your formatted result returns into context.
26
-
27
- Two consequences that drive the whole guide:
28
-
29
- - **Names do the *pre-fetch* routing.** The decision "which tool do I even look at?" is made from names
30
- (+ the server-instructions block). A name that doesn't signal intent never gets fetched.
31
- - **Descriptions do the *post-fetch* disambiguation.** They guide correct parameter choice and break
32
- ties between near-neighbors — but only *after* the model already chose to fetch. They can't rescue a
33
- tool the model never looked at.
34
-
35
- > Mental model: the **name** is the search keyword; the **description** is the manual you read after you
36
- > already pulled the book off the shelf; the **server-instructions block** is the sign on the shelf.
37
-
38
- ---
39
-
40
- ## 2. Principle 1 — The name is the routing surface
41
-
42
- Because names carry pre-fetch routing, treat them as the most important UX you ship.
43
-
44
- - **Front-load the discriminator.** Put the distinguishing token first: `youtube_transcribe`,
45
- `facebook_video_transcribe`. The platform/object is what the model scans for.
46
- - **`platform_object_action` (or `object_action`).** A consistent scheme lets the model *predict*
47
- names it hasn't seen and disambiguate neighbors.
48
- - **Never ship a bare generic verb.** `search` is a trap — search *what*? `search_serp`,
49
- `maps_search`, `facebook_ad_search`. The qualifier prevents cross-fetch and sharpens keyword ranking.
50
- - **Disambiguate near-neighbors on their real axis.** `extract_url` vs `extract_site` vs
51
- `map_site_urls` work because the names encode *single page / whole site / URLs-only*. If those blur,
52
- the model picks the wrong granularity.
53
- - **Match the user's vocabulary, or bridge it.** Users say "scrape a page"; if the tool is
54
- `extract_url`, that gap is only bridged *after* fetch (via description/`symptomTriggers`). Where a
55
- phrasing is common and high-value, make the name or its triggers reflect it.
56
- - **Mind the namespace stutter.** The client prepends `mcp__`. Naming the server `mcp-scraper` yields
57
- `mcp__mcp-scraper__…`. Name it `scraper` → `mcp__scraper__…`. Cosmetic, but it shows in permission
58
- prompts/logs, and renaming later is breaking — do it early.
59
-
60
- ### The decisive test: named tool vs. generic engine
61
- A user says *"collect ICP info from Reddit."*
62
- - A **`reddit_research`** tool: the word "reddit" hits the name directly → fetched immediately, one hop.
63
- - The generic **`browser_*`** agent: matches *none* of the user's words. To land there the model must
64
- *infer* "no reddit tool → reddit is a website → drive a browser → navigate → search → read." Several
65
- hops, each a chance to mis-route (e.g. grab `search_serp` and Google it instead).
66
-
67
- **Lesson:** if there's a phrase users will say that names a platform or intent, you want a **named tool
68
- catching it** — even if the implementation is a thin wrapper over a generic engine. *The name triggers;
69
- the engine executes.* A category that merely "points at the browser" serves a human clicking a button;
70
- it does **not** serve the model that has to *find* the capability from words.
71
-
72
- ---
73
-
74
- ## 3. Principle 2 — Split by capability, parameterize the variations
75
-
76
- It's tempting to think "names are cheap, so ship hundreds of hyper-specific tools." Don't. Three reasons:
77
-
78
- 1. **The name list isn't free at scale.** All names load at rest. 50 is fine; 300 is a wall the model
79
- must scan and a real token cost. There's a ceiling.
80
- 2. **Search ranking *degrades* with volume.** More near-neighbors → more ambiguous keyword ranking →
81
- *higher* wrong-fetch rate. Over-splitting makes the exact failure you're trying to avoid worse.
82
- 3. **It's the combinatorial trap.** `extract_url_screenshot`, `extract_url_branding`,
83
- `extract_url_mobile`… is N×M×K variants of one capability. Parameters exist to prevent this.
84
-
85
- **The rule:** one tool per *distinct thing a user wants to do*; **parameters** for *variations* of that thing.
86
-
87
- The split test:
88
- > Would a user think of these as **different things to do**, or the **same thing with options**?
89
-
90
- - Different platform/object → different tool (`youtube_transcribe` vs `facebook_video_transcribe`).
91
- - Different granularity / mental model → different tool (`extract_url` vs `extract_site`).
92
- - A knob on the same action → **parameter** (`screenshot`, `formats`, `maxPages`, `rotateProxies`).
93
-
94
- Post-fetch param-learning is cheap and reliable for the model, so you lose almost nothing by using
95
- params for variations — and you keep the name list scannable and the ranking sharp.
96
-
97
- ---
98
-
99
- ## 4. Principle 3 — Descriptions are the post-fetch safety net
100
-
101
- Loaded on fetch, the description's job is **correct usage + tie-breaking**, not discovery. Make it earn that:
102
-
103
- - **Lead with when-to-use and when-NOT-to-use.** Negative space ("use `extract_url` for one page; use
104
- `extract_site` for a whole site") is what separates near-neighbors once both are fetched.
105
- - **State side effects and cost plainly.** Especially for tools that spend money, open sessions, or
106
- post — the model should be able to refuse a wrong pick *before* executing.
107
- - **Encode parameter intent, bounds, and defaults** that mirror your actual schema, so the model fills
108
- the form correctly the first time.
109
- - **Carry symptom triggers / sample prompts.** Phrases that should route here. These help both the
110
- keyword search ranking and the post-fetch confirmation.
111
-
112
- Where the net fails (design against this): two tools with **similar names AND overlapping
113
- descriptions** — the wrong one reads "plausible enough" and the model proceeds. Distinct names +
114
- explicit negative space prevent it.
115
-
116
- ---
117
-
118
- ## 5. Principle 4 — Use the server-instructions block as the routing map
119
-
120
- It's the only content loaded *before* fetch besides names, so it's your lever on pre-fetch routing.
121
- Keep it short and put **cross-tool routing**, not per-tool detail, there:
122
-
123
- - A decision map: "single page → `extract_url`; whole site → `extract_site`; just URLs → `map_site_urls`."
124
- - Setup/auth prerequisites that span tools.
125
- - How to batch related fetches (clients often let the model load several schemas in one search call —
126
- tell it which tools group together).
127
-
128
- This is where you catch granularity and family mistakes *before* the model ever fetches the wrong one.
129
-
130
- ---
131
-
132
- ## 6. Principle 5 — Two surfaces: human categories vs. AI tools
133
-
134
- Your frontend and your MCP tool list are **different products with different optimal shapes.**
135
-
136
- - **Humans** benefit from *consolidation*: one "Search" card with a source dropdown
137
- (Web / Maps / Ads / YouTube). Fewer, richer entries; group by platform or verb as fits the UI.
138
- - **The model** benefits from *distinction*: separate, well-named tools (`search_serp`, `maps_search`,
139
- `facebook_ad_search`). A single `search(source=…)` tool forces the model to fetch it just to learn the
140
- source enum, and the name `search` is ambiguous at discovery.
141
-
142
- **Do not let the human IA pressure you into merging MCP tool names.** Map *one frontend card → many
143
- distinct MCP tools*. Two surfaces, one backend.
144
-
145
- Corollary — **a category can exist without a tool, but the AI path can't.** A "Reddit" card can be a
146
- guided browser flow with no `reddit_*` tool; a human clicks it fine. But the model gets a clean path
147
- *only* if a named tool exists (see §2). Decide per capability: human-only/occasional → category over the
148
- generic engine; want strong AI routing → ship the named wrapper.
149
-
150
- Platform-first vs verb-first grouping: for a multi-platform tool, **platform-first**
151
- (YouTube / Facebook / Google Maps / Websites / Browser / Workflows) often wins — it matches how users
152
- think *and* mirrors `platform_verb` naming. The category is the umbrella (covers search *and*
153
- transcribe *and* …); the tools inside are the named verbs.
154
-
155
- ---
156
-
157
- ## 7. Principle 6 — Tool *results* are context too; keep them lean
158
-
159
- Discovery isn't the only context cost — **what you return floods the window if you let it.**
160
-
161
- - **Summarize inline, persist the bulk, hand back a reference.** For large outputs, return a concise
162
- summary + a saved file / artifact / URL the model can read on demand, not the whole payload.
163
- - **Truncate with a signpost.** "showing first N of M; full data at `<path>`" beats silently dumping or
164
- silently dropping.
165
- - **Put the full-fidelity data somewhere the model can fetch** (disk file, object storage URL), so the
166
- model gets the gist now and the detail only if needed.
167
- - **Structured + human-readable.** Return both a structured object (for programmatic chaining) and a
168
- short readable block (for the model to reason over) — don't make it parse a wall of text.
169
-
170
- A tool that returns 200 KB of HTML per call will blow out a multi-step task; one that returns a 2 KB
171
- summary + a file path scales.
172
-
173
- ---
174
-
175
- ## 8. Failure modes → the fix
176
-
177
- | Failure | Cause | Fix |
178
- |---|---|---|
179
- | **Tool never found** | Name doesn't match user vocabulary | Front-loaded, intent-matching name; `symptomTriggers` |
180
- | **Wrong tool fetched & believed** | Similar names + overlapping descriptions | Distinct names; explicit negative space; bounds in description |
181
- | **Routed to generic engine the hard way** | Capability has no named tool | Thin named wrapper over the engine (§2 Reddit) |
182
- | **Wrong granularity** | `*_url` / `*_site` / `*_map` blur | Encode the axis in the name; routing map in server instructions |
183
- | **Picked god-tool, wrong sub-mode** | One tool, mega-param surface | Split distinct capabilities into separate tools |
184
- | **Context blown out** | Tool returns full payload | Summarize + persist + reference (§7) |
185
- | **Expensive wrong call** | Side-effecting tool, ambiguous name/desc | Clear name; state cost/side-effects up top; confirm before execute |
186
- | **Namespace stutter / churn** | Server renamed late | Pick the clean server name on day one |
187
-
188
- ---
189
-
190
- ## 9. Naming checklist (per tool)
191
-
192
- - [ ] Discriminator (platform/object) is **first** in the name.
193
- - [ ] Follows a consistent `platform_object_action` scheme used across the server.
194
- - [ ] No bare generic verb; qualifier present.
195
- - [ ] Distinct from every neighbor on a real axis (object / granularity / platform).
196
- - [ ] Matches at least one phrase a real user would say (or `symptomTriggers` covers it).
197
- - [ ] It's a *distinct capability*, not a variation that should be a parameter.
198
- - [ ] Side effects / cost surfaced in the first lines of the description.
199
- - [ ] Description leads with when-to-use **and** when-not-to-use (negative space).
200
- - [ ] Params mirror the real schema (names, bounds, enums, defaults).
201
- - [ ] Result is lean: summary + reference, not a raw dump.
202
-
203
- ## 10. Server-level checklist
204
-
205
- - [ ] Clean server name (no `mcp__mcp-…` stutter), chosen before launch.
206
- - [ ] `instructions` block carries a short **cross-tool routing map**, not per-tool detail.
207
- - [ ] Tool count is at the *capability* grain — not exploded into variants.
208
- - [ ] Frontend categories map to *multiple distinct* MCP tools (two surfaces).
209
- - [ ] High-value, user-named capabilities have a **named tool**, even if engine-backed.
210
- - [ ] Related tools are documented as a batch the model can fetch together.
211
-
212
- ---
213
-
214
- ## 11. The one-paragraph version
215
-
216
- The model sees only tool **names** until it fetches a schema, so **names carry routing** — front-load
217
- the discriminator, never ship a bare verb, and give every user-named capability its own named tool
218
- (even a thin wrapper over a generic engine), because a name that matches the user's words is a one-hop
219
- hit while a generic engine is a multi-hop inference. Split tools by **distinct capability** and use
220
- **parameters** for variations — more tools past a point *degrades* selection, it doesn't help.
221
- **Descriptions** are the post-fetch safety net (when-to-use + when-not, cost, bounds); the
222
- **server-instructions block** is your only pre-fetch routing lever (put a cross-tool decision map
223
- there). Keep the **human frontend** (consolidated cards) and the **AI tool list** (distinct names)
224
- as separate surfaces over one backend. And remember tool **results are context** — return a summary
225
- plus a reference, never a raw dump.