mcp-scraper 0.38.2 → 0.40.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (111) hide show
  1. package/README.md +5 -2
  2. package/package.json +5 -6
  3. package/dist/bin/api-server.cjs +0 -58752
  4. package/dist/bin/api-server.cjs.map +0 -1
  5. package/dist/bin/api-server.d.cts +0 -1
  6. package/dist/bin/api-server.d.ts +0 -1
  7. package/dist/bin/api-server.js +0 -38
  8. package/dist/bin/api-server.js.map +0 -1
  9. package/dist/bin/mcp-scraper-cli.cjs +0 -2671
  10. package/dist/bin/mcp-scraper-cli.cjs.map +0 -1
  11. package/dist/bin/mcp-scraper-cli.d.cts +0 -1
  12. package/dist/bin/mcp-scraper-cli.d.ts +0 -1
  13. package/dist/bin/mcp-scraper-cli.js +0 -742
  14. package/dist/bin/mcp-scraper-cli.js.map +0 -1
  15. package/dist/bin/mcp-scraper-install.cjs +0 -129
  16. package/dist/bin/mcp-scraper-install.cjs.map +0 -1
  17. package/dist/bin/mcp-scraper-install.d.cts +0 -1
  18. package/dist/bin/mcp-scraper-install.d.ts +0 -1
  19. package/dist/bin/mcp-scraper-install.js +0 -27
  20. package/dist/bin/mcp-scraper-install.js.map +0 -1
  21. package/dist/bin/mcp-stdio-server.cjs +0 -12264
  22. package/dist/bin/mcp-stdio-server.cjs.map +0 -1
  23. package/dist/bin/mcp-stdio-server.d.cts +0 -1
  24. package/dist/bin/mcp-stdio-server.d.ts +0 -1
  25. package/dist/bin/mcp-stdio-server.js +0 -135
  26. package/dist/bin/mcp-stdio-server.js.map +0 -1
  27. package/dist/bin/paa-harvest.cjs +0 -3808
  28. package/dist/bin/paa-harvest.cjs.map +0 -1
  29. package/dist/bin/paa-harvest.d.cts +0 -1
  30. package/dist/bin/paa-harvest.d.ts +0 -1
  31. package/dist/bin/paa-harvest.js +0 -44
  32. package/dist/bin/paa-harvest.js.map +0 -1
  33. package/dist/chunk-345BQXZH.js +0 -712
  34. package/dist/chunk-345BQXZH.js.map +0 -1
  35. package/dist/chunk-44HZLHDV.js +0 -52
  36. package/dist/chunk-44HZLHDV.js.map +0 -1
  37. package/dist/chunk-AZRPG43B.js +0 -617
  38. package/dist/chunk-AZRPG43B.js.map +0 -1
  39. package/dist/chunk-CB5C3BPB.js +0 -135
  40. package/dist/chunk-CB5C3BPB.js.map +0 -1
  41. package/dist/chunk-EGJKUB4Q.js +0 -276
  42. package/dist/chunk-EGJKUB4Q.js.map +0 -1
  43. package/dist/chunk-FQI5PFE7.js +0 -1866
  44. package/dist/chunk-FQI5PFE7.js.map +0 -1
  45. package/dist/chunk-FRYT3ID4.js +0 -684
  46. package/dist/chunk-FRYT3ID4.js.map +0 -1
  47. package/dist/chunk-G3P3ZDB4.js +0 -69
  48. package/dist/chunk-G3P3ZDB4.js.map +0 -1
  49. package/dist/chunk-K443GQY5.js +0 -24
  50. package/dist/chunk-K443GQY5.js.map +0 -1
  51. package/dist/chunk-N7KUTTCC.js +0 -3007
  52. package/dist/chunk-N7KUTTCC.js.map +0 -1
  53. package/dist/chunk-NGM237OO.js +0 -3410
  54. package/dist/chunk-NGM237OO.js.map +0 -1
  55. package/dist/chunk-NKCCGADE.js +0 -11285
  56. package/dist/chunk-NKCCGADE.js.map +0 -1
  57. package/dist/chunk-NNW3O6ZD.js +0 -108
  58. package/dist/chunk-NNW3O6ZD.js.map +0 -1
  59. package/dist/chunk-QZXKQB7Y.js +0 -414
  60. package/dist/chunk-QZXKQB7Y.js.map +0 -1
  61. package/dist/chunk-SFRMFGQ6.js +0 -158
  62. package/dist/chunk-SFRMFGQ6.js.map +0 -1
  63. package/dist/chunk-YCI2PNCS.js +0 -499
  64. package/dist/chunk-YCI2PNCS.js.map +0 -1
  65. package/dist/chunk-YODBNTTN.js +0 -7
  66. package/dist/chunk-YODBNTTN.js.map +0 -1
  67. package/dist/db-C5KVCOYT.js +0 -239
  68. package/dist/db-C5KVCOYT.js.map +0 -1
  69. package/dist/extract-bundle-KUBX6N6Z.js +0 -568
  70. package/dist/extract-bundle-KUBX6N6Z.js.map +0 -1
  71. package/dist/index.cjs +0 -4160
  72. package/dist/index.cjs.map +0 -1
  73. package/dist/index.d.cts +0 -413
  74. package/dist/index.d.ts +0 -413
  75. package/dist/index.js +0 -338
  76. package/dist/index.js.map +0 -1
  77. package/dist/location-data-repository-TTWF3OTM.js +0 -35
  78. package/dist/location-data-repository-TTWF3OTM.js.map +0 -1
  79. package/dist/server-5EX6XBIA.js +0 -33596
  80. package/dist/server-5EX6XBIA.js.map +0 -1
  81. package/dist/site-extract-repository-XSPJTCIL.js +0 -62
  82. package/dist/site-extract-repository-XSPJTCIL.js.map +0 -1
  83. package/dist/worker-XCPU4YSN.js +0 -142
  84. package/dist/worker-XCPU4YSN.js.map +0 -1
  85. package/docs/adr/0001-in-page-graphql-interception-for-anti-bot-scraping.md +0 -58
  86. package/docs/adr/0002-hybrid-smart-rag-vault-retrieval.md +0 -62
  87. package/docs/adr/0003-waive-unrecoverable-scheduled-model-cost.md +0 -22
  88. package/docs/adr/README.md +0 -13
  89. package/docs/final-tooling-spec.md +0 -206
  90. package/docs/hosted-location-data.md +0 -108
  91. package/docs/kernel-proxy-future-enhancements.md +0 -80
  92. package/docs/mcp-tool-craft-lint.generated.md +0 -183
  93. package/docs/mcp-tool-design-guide.md +0 -225
  94. package/docs/mcp-tool-manifest.generated.json +0 -22871
  95. package/docs/mcp-tool-quality-spec.md +0 -240
  96. package/docs/oauth-legal-review.md +0 -38
  97. package/docs/seo-crawl-report-spec.md +0 -287
  98. package/docs/specs/api-forge-spec.md +0 -234
  99. package/docs/specs/connected-services-control-plane-decoupling-spec.md +0 -1044
  100. package/docs/specs/deferred-work-spec.md +0 -86
  101. package/docs/specs/google-drive-bulk-access-and-mcp-schema-passthrough-spec.md +0 -1689
  102. package/docs/specs/kernel-stealth-captcha-test-matrix.md +0 -278
  103. package/docs/specs/main-mcp-integration-ownership-spec.md +0 -1164
  104. package/docs/specs/mcp-tool-definition-quality-audit-spec.md +0 -1602
  105. package/docs/specs/meta-ad-creative-media-resolution-spec.md +0 -31
  106. package/docs/specs/multimodal-image-memory-architecture-spec.md +0 -1022
  107. package/docs/specs/oauth-mcp-spec.md +0 -213
  108. package/docs/specs/query-fanout-transport-contract-fix.md +0 -45
  109. package/docs/specs/relationship-workspace-ai-behavior-plan.md +0 -26
  110. package/docs/specs/unified-credit-and-scheduled-execution-billing-spec.md +0 -995
  111. package/docs/tool-catalog-spec.md +0 -388
@@ -1,62 +0,0 @@
1
- import {
2
- abandonExtractSettlement,
3
- claimFailedExtractJobForRefinalize,
4
- completeExtractJob,
5
- countSuccessfulPages,
6
- createExtractJob,
7
- createOrGetExtractJob,
8
- extractJobLimitInfo,
9
- failExtractJob,
10
- failStaleRunningExtractJob,
11
- failUnfundedExtractJob,
12
- finishExtractJob,
13
- getExtractJob,
14
- getExtractJobByIdempotencyKey,
15
- getExtractedImageLinks,
16
- getExtractedPages,
17
- getExtractedUrls,
18
- listExtractJobs,
19
- listFundedPendingExtractJobs,
20
- listStaleRunningExtractJobs,
21
- listUnsettledExtractJobs,
22
- markExtractJobDispatchAttempt,
23
- recordExtractJobDispatchFailure,
24
- recordExtractSettlementFailure,
25
- saveExtractPages,
26
- setExtractJobTotal,
27
- settleExtractJob,
28
- terminalExtractJobStatus
29
- } from "./chunk-YCI2PNCS.js";
30
- import "./chunk-AZRPG43B.js";
31
- import "./chunk-44HZLHDV.js";
32
- import "./chunk-N7KUTTCC.js";
33
- export {
34
- abandonExtractSettlement,
35
- claimFailedExtractJobForRefinalize,
36
- completeExtractJob,
37
- countSuccessfulPages,
38
- createExtractJob,
39
- createOrGetExtractJob,
40
- extractJobLimitInfo,
41
- failExtractJob,
42
- failStaleRunningExtractJob,
43
- failUnfundedExtractJob,
44
- finishExtractJob,
45
- getExtractJob,
46
- getExtractJobByIdempotencyKey,
47
- getExtractedImageLinks,
48
- getExtractedPages,
49
- getExtractedUrls,
50
- listExtractJobs,
51
- listFundedPendingExtractJobs,
52
- listStaleRunningExtractJobs,
53
- listUnsettledExtractJobs,
54
- markExtractJobDispatchAttempt,
55
- recordExtractJobDispatchFailure,
56
- recordExtractSettlementFailure,
57
- saveExtractPages,
58
- setExtractJobTotal,
59
- settleExtractJob,
60
- terminalExtractJobStatus
61
- };
62
- //# sourceMappingURL=site-extract-repository-XSPJTCIL.js.map
@@ -1 +0,0 @@
1
- {"version":3,"sources":[],"sourcesContent":[],"mappings":"","names":[]}
@@ -1,142 +0,0 @@
1
- import {
2
- classifyHarvestProblem,
3
- createHarvestAttemptRecorder,
4
- harvestProblemResponse,
5
- serializeHarvestProblem
6
- } from "./chunk-SFRMFGQ6.js";
7
- import {
8
- harvest
9
- } from "./chunk-NGM237OO.js";
10
- import {
11
- browserServiceApiKey,
12
- runWithCostContext
13
- } from "./chunk-EGJKUB4Q.js";
14
- import "./chunk-K443GQY5.js";
15
- import "./chunk-CB5C3BPB.js";
16
- import {
17
- MC_COSTS,
18
- serpActualCostMc
19
- } from "./chunk-AZRPG43B.js";
20
- import "./chunk-44HZLHDV.js";
21
- import {
22
- claimPendingJob,
23
- completeJob,
24
- creditMc,
25
- debitMc,
26
- failJob,
27
- listHarvestAttempts
28
- } from "./chunk-N7KUTTCC.js";
29
-
30
- // src/api/webhook.ts
31
- async function deliverWebhook(url, payload, retries = 3) {
32
- for (let attempt = 1; attempt <= retries; attempt++) {
33
- try {
34
- const res = await fetch(url, {
35
- method: "POST",
36
- headers: { "content-type": "application/json" },
37
- body: JSON.stringify(payload),
38
- signal: AbortSignal.timeout(1e4)
39
- });
40
- if (res.ok) return;
41
- console.warn(`[webhook] attempt ${attempt} \u2192 ${res.status} from ${url}`);
42
- } catch (err) {
43
- console.warn(`[webhook] attempt ${attempt} failed:`, err instanceof Error ? err.message : err);
44
- }
45
- if (attempt < retries) await new Promise((r) => setTimeout(r, 1e3 * attempt * 2));
46
- }
47
- console.error(`[webhook] gave up after ${retries} attempts for ${url}`);
48
- }
49
-
50
- // src/api/worker.ts
51
- var MAX_CONCURRENT = 2;
52
- var running = 0;
53
- function countPaaQuestions(result) {
54
- if (!result || typeof result !== "object") return 0;
55
- const value = result;
56
- if (typeof value.totalQuestions === "number") return value.totalQuestions;
57
- return Array.isArray(value.flat) ? value.flat.length : 0;
58
- }
59
- function paaCostForQuestionCount(questionCount) {
60
- return MC_COSTS.paa_base + Math.max(1, questionCount) * MC_COSTS.paa;
61
- }
62
- async function processJob(job) {
63
- running++;
64
- try {
65
- const opts = typeof job.options === "string" ? JSON.parse(job.options) : job.options;
66
- const headlessSentOut = { value: null };
67
- const result = await runWithCostContext(
68
- { op: opts.serpOnly ? "serp" : "paa", userId: Number(job.user_id), headlessSentOut },
69
- () => harvest({
70
- ...opts,
71
- kernelApiKey: browserServiceApiKey(),
72
- headless: true,
73
- format: "json",
74
- outputDir: "/tmp/paa-output-api",
75
- onAttemptEvent: createHarvestAttemptRecorder(job.id, job.user_id)
76
- })
77
- );
78
- await completeJob(job.id, result);
79
- const attempts = await listHarvestAttempts(job.id, job.user_id);
80
- if (!opts.serpOnly && typeof opts.billingHoldMc === "number") {
81
- const actualCost = paaCostForQuestionCount(countPaaQuestions(result));
82
- const diff = opts.billingHoldMc - actualCost;
83
- if (diff > 0) await creditMc(job.user_id, diff, "paa_refund", "overestimate refund");
84
- else if (diff < 0) await debitMc(job.user_id, -diff, "paa", opts.query ?? job.query);
85
- } else if (opts.serpOnly && typeof opts.billingHoldMc === "number") {
86
- const actualCost = serpActualCostMc(headlessSentOut.value);
87
- const diff = opts.billingHoldMc - actualCost;
88
- if (diff > 0) await creditMc(job.user_id, diff, "serp_refund", "headless-mode pricing settle");
89
- else if (diff < 0) await debitMc(job.user_id, -diff, "serp", opts.query ?? job.query);
90
- }
91
- if (job.callback_url) {
92
- await deliverWebhook(job.callback_url, { job_id: job.id, status: "done", result, attempts });
93
- }
94
- } catch (err) {
95
- const problem = classifyHarvestProblem(err);
96
- await failJob(job.id, serializeHarvestProblem(problem));
97
- const attempts = await listHarvestAttempts(job.id, job.user_id);
98
- try {
99
- const opts = typeof job.options === "string" ? JSON.parse(job.options) : job.options;
100
- if (typeof opts.billingHoldMc === "number" && opts.billingHoldMc > 0) {
101
- await creditMc(job.user_id, opts.billingHoldMc, "refund", "failed call");
102
- }
103
- } catch {
104
- }
105
- if (job.callback_url) {
106
- await deliverWebhook(job.callback_url, { job_id: job.id, status: "failed", ...harvestProblemResponse(problem), attempts });
107
- }
108
- } finally {
109
- running--;
110
- }
111
- }
112
- async function tickOnce() {
113
- const job = await claimPendingJob();
114
- if (!job) return { claimed: false };
115
- const startedAt = Date.now();
116
- await processJob(job);
117
- return { claimed: true, jobId: job.id, completed: true, durationMs: Date.now() - startedAt };
118
- }
119
- async function drainQueue(budget) {
120
- const results = [];
121
- for (let i = 0; i < budget.maxJobs; i++) {
122
- if (Date.now() >= budget.deadlineMs) break;
123
- const r = await tickOnce();
124
- results.push(r);
125
- if (!r.claimed) break;
126
- }
127
- return results;
128
- }
129
- function startWorker() {
130
- setInterval(async () => {
131
- if (running >= MAX_CONCURRENT) return;
132
- const job = await claimPendingJob();
133
- if (job) void processJob(job);
134
- }, 2e3);
135
- console.log(`[worker] started \u2014 polling every 2s, max ${MAX_CONCURRENT} concurrent`);
136
- }
137
- export {
138
- drainQueue,
139
- startWorker,
140
- tickOnce
141
- };
142
- //# sourceMappingURL=worker-XCPU4YSN.js.map
@@ -1 +0,0 @@
1
- {"version":3,"sources":["../src/api/webhook.ts","../src/api/worker.ts"],"sourcesContent":["export async function deliverWebhook(url: string, payload: object, retries = 3): Promise<void> {\n for (let attempt = 1; attempt <= retries; attempt++) {\n try {\n const res = await fetch(url, {\n method: 'POST',\n headers: { 'content-type': 'application/json' },\n body: JSON.stringify(payload),\n signal: AbortSignal.timeout(10_000),\n })\n if (res.ok) return\n console.warn(`[webhook] attempt ${attempt} → ${res.status} from ${url}`)\n } catch (err) {\n console.warn(`[webhook] attempt ${attempt} failed:`, err instanceof Error ? err.message : err)\n }\n if (attempt < retries) await new Promise((r) => setTimeout(r, 1000 * attempt * 2))\n }\n console.error(`[webhook] gave up after ${retries} attempts for ${url}`)\n}\n","import { claimPendingJob, completeJob, failJob, creditMc, debitMc, listHarvestAttempts } from './db.js'\nimport { browserServiceApiKey } from '../lib/browser-service-env.js'\nimport { harvest } from '../harvest.js'\nimport { deliverWebhook } from './webhook.js'\nimport type { HarvestOptions } from '../types.js'\nimport { MC_COSTS, serpActualCostMc } from './rates.js'\nimport { classifyHarvestProblem, harvestProblemResponse, serializeHarvestProblem } from './harvest-problems.js'\nimport { createHarvestAttemptRecorder } from './harvest-attempt-events.js'\nimport { runWithCostContext } from './cost-context.js'\n\nexport type TickResult = {\n claimed: boolean\n jobId?: string\n completed?: boolean\n durationMs?: number\n}\nexport type DrainBudget = {\n maxJobs: number\n deadlineMs: number\n}\n\nconst MAX_CONCURRENT = 2\nlet running = 0\n\nfunction countPaaQuestions(result: unknown): number {\n if (!result || typeof result !== 'object') return 0\n const value = result as { totalQuestions?: unknown; flat?: unknown }\n if (typeof value.totalQuestions === 'number') return value.totalQuestions\n return Array.isArray(value.flat) ? value.flat.length : 0\n}\n\nfunction paaCostForQuestionCount(questionCount: number): number {\n return MC_COSTS.paa_base + Math.max(1, questionCount) * MC_COSTS.paa\n}\n\nasync function processJob(job: Awaited<ReturnType<typeof claimPendingJob>> & object) {\n running++\n try {\n const opts = typeof job.options === 'string' ? JSON.parse(job.options) as Partial<HarvestOptions> & { billingHoldMc?: number } : job.options as Partial<HarvestOptions> & { billingHoldMc?: number }\n const headlessSentOut: { value: boolean | null } = { value: null }\n const result = await runWithCostContext(\n { op: opts.serpOnly ? 'serp' : 'paa', userId: Number(job.user_id), headlessSentOut },\n () => harvest({\n ...opts,\n kernelApiKey: browserServiceApiKey(),\n headless: true,\n format: 'json',\n outputDir: '/tmp/paa-output-api',\n onAttemptEvent: createHarvestAttemptRecorder(job.id, job.user_id),\n }),\n )\n await completeJob(job.id, result)\n const attempts = await listHarvestAttempts(job.id, job.user_id)\n if (!opts.serpOnly && typeof opts.billingHoldMc === 'number') {\n const actualCost = paaCostForQuestionCount(countPaaQuestions(result))\n const diff = opts.billingHoldMc - actualCost\n if (diff > 0) await creditMc(job.user_id, diff, 'paa_refund', 'overestimate refund')\n else if (diff < 0) await debitMc(job.user_id, -diff, 'paa', opts.query ?? job.query)\n } else if (opts.serpOnly && typeof opts.billingHoldMc === 'number') {\n const actualCost = serpActualCostMc(headlessSentOut.value)\n const diff = opts.billingHoldMc - actualCost\n if (diff > 0) await creditMc(job.user_id, diff, 'serp_refund', 'headless-mode pricing settle')\n else if (diff < 0) await debitMc(job.user_id, -diff, 'serp', opts.query ?? job.query)\n }\n if (job.callback_url) {\n await deliverWebhook(job.callback_url, { job_id: job.id, status: 'done', result, attempts })\n }\n } catch (err) {\n const problem = classifyHarvestProblem(err)\n await failJob(job.id, serializeHarvestProblem(problem))\n const attempts = await listHarvestAttempts(job.id, job.user_id)\n try {\n const opts = typeof job.options === 'string' ? JSON.parse(job.options) as { billingHoldMc?: number } : job.options as { billingHoldMc?: number }\n if (typeof opts.billingHoldMc === 'number' && opts.billingHoldMc > 0) {\n await creditMc(job.user_id, opts.billingHoldMc, 'refund', 'failed call')\n }\n } catch {}\n if (job.callback_url) {\n await deliverWebhook(job.callback_url, { job_id: job.id, status: 'failed', ...harvestProblemResponse(problem), attempts })\n }\n } finally {\n running--\n }\n}\n\nexport async function tickOnce(): Promise<TickResult> {\n const job = await claimPendingJob()\n if (!job) return { claimed: false }\n const startedAt = Date.now()\n await processJob(job as NonNullable<typeof job>)\n return { claimed: true, jobId: (job as { id: string }).id, completed: true, durationMs: Date.now() - startedAt }\n}\n\nexport async function drainQueue(budget: DrainBudget): Promise<TickResult[]> {\n const results: TickResult[] = []\n for (let i = 0; i < budget.maxJobs; i++) {\n if (Date.now() >= budget.deadlineMs) break\n const r = await tickOnce()\n results.push(r)\n if (!r.claimed) break\n }\n return results\n}\n\nexport function startWorker(): void {\n setInterval(async () => {\n if (running >= MAX_CONCURRENT) return\n const job = await claimPendingJob()\n if (job) void processJob(job as NonNullable<typeof job>)\n }, 2000)\n console.log(`[worker] started — polling every 2s, max ${MAX_CONCURRENT} concurrent`)\n}\n"],"mappings":";;;;;;;;;;;;;;;;;;;;;;;;;;;;;;AAAA,eAAsB,eAAe,KAAa,SAAiB,UAAU,GAAkB;AAC7F,WAAS,UAAU,GAAG,WAAW,SAAS,WAAW;AACnD,QAAI;AACF,YAAM,MAAM,MAAM,MAAM,KAAK;AAAA,QAC3B,QAAQ;AAAA,QACR,SAAS,EAAE,gBAAgB,mBAAmB;AAAA,QAC9C,MAAM,KAAK,UAAU,OAAO;AAAA,QAC5B,QAAQ,YAAY,QAAQ,GAAM;AAAA,MACpC,CAAC;AACD,UAAI,IAAI,GAAI;AACZ,cAAQ,KAAK,qBAAqB,OAAO,WAAM,IAAI,MAAM,SAAS,GAAG,EAAE;AAAA,IACzE,SAAS,KAAK;AACZ,cAAQ,KAAK,qBAAqB,OAAO,YAAY,eAAe,QAAQ,IAAI,UAAU,GAAG;AAAA,IAC/F;AACA,QAAI,UAAU,QAAS,OAAM,IAAI,QAAQ,CAAC,MAAM,WAAW,GAAG,MAAO,UAAU,CAAC,CAAC;AAAA,EACnF;AACA,UAAQ,MAAM,2BAA2B,OAAO,iBAAiB,GAAG,EAAE;AACxE;;;ACIA,IAAM,iBAAiB;AACvB,IAAI,UAAU;AAEd,SAAS,kBAAkB,QAAyB;AAClD,MAAI,CAAC,UAAU,OAAO,WAAW,SAAU,QAAO;AAClD,QAAM,QAAQ;AACd,MAAI,OAAO,MAAM,mBAAmB,SAAU,QAAO,MAAM;AAC3D,SAAO,MAAM,QAAQ,MAAM,IAAI,IAAI,MAAM,KAAK,SAAS;AACzD;AAEA,SAAS,wBAAwB,eAA+B;AAC9D,SAAO,SAAS,WAAW,KAAK,IAAI,GAAG,aAAa,IAAI,SAAS;AACnE;AAEA,eAAe,WAAW,KAA2D;AACnF;AACA,MAAI;AACF,UAAM,OAAO,OAAO,IAAI,YAAY,WAAW,KAAK,MAAM,IAAI,OAAO,IAA4D,IAAI;AACrI,UAAM,kBAA6C,EAAE,OAAO,KAAK;AACjE,UAAM,SAAS,MAAM;AAAA,MACnB,EAAE,IAAI,KAAK,WAAW,SAAS,OAAO,QAAQ,OAAO,IAAI,OAAO,GAAG,gBAAgB;AAAA,MACnF,MAAM,QAAQ;AAAA,QACZ,GAAG;AAAA,QACH,cAAc,qBAAqB;AAAA,QACnC,UAAU;AAAA,QACV,QAAQ;AAAA,QACR,WAAW;AAAA,QACX,gBAAgB,6BAA6B,IAAI,IAAI,IAAI,OAAO;AAAA,MAClE,CAAC;AAAA,IACH;AACA,UAAM,YAAY,IAAI,IAAI,MAAM;AAChC,UAAM,WAAW,MAAM,oBAAoB,IAAI,IAAI,IAAI,OAAO;AAC9D,QAAI,CAAC,KAAK,YAAY,OAAO,KAAK,kBAAkB,UAAU;AAC5D,YAAM,aAAa,wBAAwB,kBAAkB,MAAM,CAAC;AACpE,YAAM,OAAO,KAAK,gBAAgB;AAClC,UAAI,OAAO,EAAG,OAAM,SAAS,IAAI,SAAS,MAAM,cAAc,qBAAqB;AAAA,eAC1E,OAAO,EAAG,OAAM,QAAQ,IAAI,SAAS,CAAC,MAAM,OAAO,KAAK,SAAS,IAAI,KAAK;AAAA,IACrF,WAAW,KAAK,YAAY,OAAO,KAAK,kBAAkB,UAAU;AAClE,YAAM,aAAa,iBAAiB,gBAAgB,KAAK;AACzD,YAAM,OAAO,KAAK,gBAAgB;AAClC,UAAI,OAAO,EAAG,OAAM,SAAS,IAAI,SAAS,MAAM,eAAe,8BAA8B;AAAA,eACpF,OAAO,EAAG,OAAM,QAAQ,IAAI,SAAS,CAAC,MAAM,QAAQ,KAAK,SAAS,IAAI,KAAK;AAAA,IACtF;AACA,QAAI,IAAI,cAAc;AACpB,YAAM,eAAe,IAAI,cAAc,EAAE,QAAQ,IAAI,IAAI,QAAQ,QAAQ,QAAQ,SAAS,CAAC;AAAA,IAC7F;AAAA,EACF,SAAS,KAAK;AACZ,UAAM,UAAU,uBAAuB,GAAG;AAC1C,UAAM,QAAQ,IAAI,IAAI,wBAAwB,OAAO,CAAC;AACtD,UAAM,WAAW,MAAM,oBAAoB,IAAI,IAAI,IAAI,OAAO;AAC9D,QAAI;AACF,YAAM,OAAO,OAAO,IAAI,YAAY,WAAW,KAAK,MAAM,IAAI,OAAO,IAAkC,IAAI;AAC3G,UAAI,OAAO,KAAK,kBAAkB,YAAY,KAAK,gBAAgB,GAAG;AACpE,cAAM,SAAS,IAAI,SAAS,KAAK,eAAe,UAAU,aAAa;AAAA,MACzE;AAAA,IACF,QAAQ;AAAA,IAAC;AACT,QAAI,IAAI,cAAc;AACpB,YAAM,eAAe,IAAI,cAAc,EAAE,QAAQ,IAAI,IAAI,QAAQ,UAAU,GAAG,uBAAuB,OAAO,GAAG,SAAS,CAAC;AAAA,IAC3H;AAAA,EACF,UAAE;AACA;AAAA,EACF;AACF;AAEA,eAAsB,WAAgC;AACpD,QAAM,MAAM,MAAM,gBAAgB;AAClC,MAAI,CAAC,IAAK,QAAO,EAAE,SAAS,MAAM;AAClC,QAAM,YAAY,KAAK,IAAI;AAC3B,QAAM,WAAW,GAA8B;AAC/C,SAAO,EAAE,SAAS,MAAM,OAAQ,IAAuB,IAAI,WAAW,MAAM,YAAY,KAAK,IAAI,IAAI,UAAU;AACjH;AAEA,eAAsB,WAAW,QAA4C;AAC3E,QAAM,UAAwB,CAAC;AAC/B,WAAS,IAAI,GAAG,IAAI,OAAO,SAAS,KAAK;AACvC,QAAI,KAAK,IAAI,KAAK,OAAO,WAAY;AACrC,UAAM,IAAI,MAAM,SAAS;AACzB,YAAQ,KAAK,CAAC;AACd,QAAI,CAAC,EAAE,QAAS;AAAA,EAClB;AACA,SAAO;AACT;AAEO,SAAS,cAAoB;AAClC,cAAY,YAAY;AACtB,QAAI,WAAW,eAAgB;AAC/B,UAAM,MAAM,MAAM,gBAAgB;AAClC,QAAI,IAAK,MAAK,WAAW,GAA8B;AAAA,EACzD,GAAG,GAAI;AACP,UAAQ,IAAI,iDAA4C,cAAc,aAAa;AACrF;","names":[]}
@@ -1,58 +0,0 @@
1
- # ADR 0001: In-page API interception for anti-bot scraping (Facebook Ad Library)
2
-
3
- - **Status:** Accepted
4
- - **Date:** 2026-06-03
5
- - **Deciders:** Andrew (operator), engineering
6
- - **Applies to:** `facebook_ad_search`, `facebook_page_intel`, and future scrapers of JS-heavy, anti-bot-protected sites
7
-
8
- ## Context
9
-
10
- `facebook_ad_search` was unreliable — it usually returned `soft-block: no results (refunded)`, and when it did return, advertiser names came back `undefined` with `—` for library IDs. Two independent root causes:
11
-
12
- 1. **DOM scraping a hostile SPA.** The Facebook Ad Library is a React app. The ads are *not* in the initial HTML — the page fires its own background GraphQL request and renders the JSON into DOM cards asynchronously. Our scraper waited for those cards to render, then re-parsed the rendered text with regexes (`See ad details … Sponsored`). That is fragile (breaks on any UI change) and slow, and it conflated "page still loading" with "no results."
13
-
14
- 2. **A single static proxy.** Every Facebook scrape launched with the static `KERNEL_PROXY_ID`. Facebook flags/rate-limits a repeatedly-used IP and serves a near-empty / login-gated page to it. A real residential browser loaded the same search fine — confirming the block was IP-reputation, not the URL or a login requirement. `facebook_page_intel` worked intermittently for the same reason (same proxy, sometimes flagged).
15
-
16
- The "undefined names / — library IDs" was a third, smaller bug: a field-name mismatch between the API response (`pageName`, `sampleLibraryId`) and the MCP formatter (`name`, `libraryId`).
17
-
18
- ## Decision
19
-
20
- **Stop scraping the rendered page. Intercept the JSON the page already fetches, on a browser-service proxy, with the DOM scrape kept only as a fallback.**
21
-
22
- Concretely, for the Ad Library:
23
-
24
- 1. **Intercept the in-page GraphQL response.** The SPA issues `AdLibrarySearchPaginationQuery` (a `doc_id`-based POST to `/api/graphql/`) and Facebook returns the ads as JSON. We attach a Playwright `page.on('response')` listener before navigation and parse the JSON directly (`data.ad_library_main.search_results_connection.edges[].node.collated_results[]`). No request is forged — the page makes the request itself, so it carries Facebook's own session/cookies and looks like a legitimate first-party call. Implemented in `src/extractor/FacebookAdGraphql.ts`.
25
- 2. **Configured proxy by default.** Both FB routes now launch via `kernelLaunchOptsResidential()`, which uses the configured browser-service proxy without city/ZIP targeting by default. Location-targeted residential proxying remains an explicit escalation mode rather than the default.
26
- 3. **DOM scrape as fallback.** If no GraphQL response is captured (e.g., the query shape drifts), the existing regex DOM parse still runs on the already-loaded page.
27
- 4. **Field aliasing for client compatibility.** The API now returns both `pageName`/`name` and `sampleLibraryId`/`libraryId`, so even already-installed MCP clients render correctly without a package update; the formatter was also fixed to read the canonical fields.
28
-
29
- This is the **same pattern already proven by the YouTube transcription InnerTube tier** (`CaptionFetcher.ts` does `page.evaluate(fetch('/youtubei/v1/player'…))` from inside the loaded page). This ADR names it as the preferred default for this class of target.
30
-
31
- ### The reusable pattern
32
-
33
- > For a JS-heavy, anti-bot-protected site whose data arrives via an internal API call: load the page on a clean (residential) IP, capture or replay that internal request **from within the page's own session**, parse the JSON, and keep DOM scraping only as a fallback.
34
-
35
- Prefer **intercepting the response** (`page.on('response')`) over **replaying the request** when the page fires the call itself — it needs zero token/`doc_id`/`fb_dtsg` reconstruction and is the most robust form.
36
-
37
- ## Consequences
38
-
39
- **Positive**
40
- - Parses structured JSON, not DOM → immune to CSS/markup changes (the usual scraper breakage).
41
- - Faster — no waiting for the grid to render/scroll before reading.
42
- - Unambiguous — an explicit ad list or an error, killing the "0 results = blocked?" guesswork.
43
- - Lower block rate — an in-session first-party request plus a clean residential IP looks legitimate.
44
- - Verified live: `Nike` returned 10 distinct, correctly-named advertisers with library IDs (Nike, Jordan, Nordstrom Rack, eBay, Whatnot…), end-to-end through the MCP, no soft-block.
45
-
46
- **Negative / risks**
47
- - The GraphQL `doc_id` and response shape are Facebook-internal and **drift every few months**. Mitigation: the DOM fallback, plus the capture probe technique below to re-derive the shape in minutes.
48
- - Still not 100% — the Ad Library is adversarial; residential IPs can occasionally be challenged.
49
- - Captured `doc_id` (`24922295957467452` as of 2026-06-03) is intentionally *not* hardcoded — we match on the `fb_api_req_friendly_name` (`AdLibrarySearchPaginationQuery`) so a `doc_id` rotation alone doesn't break us.
50
-
51
- **Maintenance — re-capturing the request shape when it drifts**
52
- Run a Kernel session on a residential proxy, attach `page.on('request'|'response')`, navigate to the Ad Library search URL, and log GraphQL `fb_api_req_friendly_name` / `doc_id` / `variables` / response keys. (A throwaway `fb-capture.ts` probe was used during this work; recreate it from this description.)
53
-
54
- ## Alternatives considered
55
-
56
- - **Official Graph Ad Library API (`/ads_archive`).** Rejected: for the US it only returns political/issue ads; commercial ads (the use case) aren't exposed outside the EU.
57
- - **Replay the GraphQL request manually** (rebuild `lsd`/`jazoest`/`__dyn`/`doc_id`/variables). Rejected as the default: more brittle than intercepting the page's own response, though viable for pagination beyond what scrolling triggers.
58
- - **Just rotate proxies, keep DOM scraping.** Rejected: fixes the block but leaves the fragile, slow DOM parse and the "0 results" ambiguity.
@@ -1,62 +0,0 @@
1
- # ADR 0002: Hybrid Smart RAG retrieval for memory vaults
2
-
3
- - **Status:** Accepted
4
- - **Date:** 2026-07-16
5
- - **Deciders:** Andrew (operator), engineering
6
- - **Applies to:** MCP Memory retrieval, memory writes, vault search, and the default Skills vault
7
-
8
- ## Context
9
-
10
- Semantic similarity alone can miss exact tags, dates, metadata, and the note graph that gives an Obsidian vault its useful local context. Expanding every link from every semantic result has the opposite problem: it creates a large, noisy candidate set and can make structural proximity look like evidence of relevance. The connected AI also needs a strong habit of reusing the live tag vocabulary without turning one advisory tool-call order into a persisted write precondition.
11
-
12
- Skills add a second vault-specific constraint. A stored skill is not just one note: its internal links may point to reference material, templates, and executable support files, so the vault needs a stable folder contract.
13
-
14
- ## Decision
15
-
16
- **Use bounded hybrid Smart RAG for memory retrieval, and treat tag and graph context as evidence-bearing candidates rather than automatic conclusions.**
17
-
18
- The default retrieval flow is:
19
-
20
- 1. Form **2–4 focused query variants**, with **3** as the default.
21
- 2. Combine exact **tag, metadata, vault, and date** retrieval with semantic retrieval.
22
- 3. Fuse and deduplicate those channels into a total candidate pool of **50**.
23
- 4. Select the top **8** preliminary seeds.
24
- 5. Expand each seed to **depth 1**, considering at most **5** outgoing links or backlinks per seed.
25
- 6. Apply **Jina reranking** to the combined candidates and retain the top **30** for reasoning.
26
-
27
- Graph-expanded items remain candidates. They are not automatically relevant and must never become automatic note links; the AI adds an internal link only when the note contents support it.
28
-
29
- Before memory search, creation, or update, the MCP instructions strongly direct the AI to inspect the complete accessible tag inventory, reuse existing tags, and add a tag only when the inventory does not contain an adequate equivalent. This is an automatic behavioral directive, **not** a persisted call-order precondition that makes writes fail solely because a separate tag-list call was not recorded.
30
-
31
- The default **Skills** vault stores reusable skill packages. Every skill requires a `scripts/` folder and at least one of `references/` or `templates/`; a skill may have both. Skill notes use normal internal links to those supporting files.
32
-
33
- ## Consequences
34
-
35
- **Positive**
36
- - Exact filters preserve high-precision matches while multi-query semantic retrieval recovers differently worded concepts.
37
- - Bounded one-hop expansion recovers directly connected context without recursively walking the vault.
38
- - Jina reranking makes semantic, exact, and graph-derived candidates comparable before the AI reasons over them.
39
- - The 50-candidate and 30-result bounds provide predictable retrieval and context costs.
40
- - Tag reuse becomes the default AI behavior without coupling persisted data validity to orchestration traces.
41
- - Skills have a portable, inspectable support-file layout.
42
-
43
- **Negative / risks**
44
- - Multi-query retrieval, graph expansion, and reranking add latency and Jina dependency cost.
45
- - A fixed pool of 50 may under-recall unusually broad vaults; 30 reranked notes may still exceed the useful reading budget for narrow questions.
46
- - Link quality depends on note contents and graph maintenance; a bad link can introduce a weak candidate even though reranking limits its impact.
47
- - The tag-inventory-first behavior remains advisory at the AI layer and can be skipped by a non-compliant client.
48
- - Skills that do not need executable logic must still carry a `scripts/` folder to satisfy the vault contract.
49
-
50
- ## Alternatives considered
51
-
52
- - **One semantic query only.** Rejected because it misses exact structured constraints and alternate phrasings.
53
- - **Expand every link from every retrieved note.** Rejected because candidate growth and graph noise become unbounded; expansion is limited to the strongest eight seeds.
54
- - **Treat graph neighbors as automatically relevant or automatically link them.** Rejected because structural proximity is not evidence that a note answers the current request.
55
- - **Persist tag-list call order as a write precondition.** Rejected because it couples durable writes to client orchestration history; forceful instructions provide the intended behavior without rejecting otherwise valid writes.
56
-
57
- ## Revisit when
58
-
59
- - Retrieval evaluation shows that the 50/8/5/30 bounds materially harm recall or precision.
60
- - Jina reranking becomes unavailable, too costly, or is outperformed by a tested replacement.
61
- - Vault graphs become dense enough that one-hop expansion regularly adds mostly irrelevant candidates.
62
- - The Skills packaging contract adopts an external standard that makes `scripts/`, `references/`, or `templates/` incompatible.
@@ -1,22 +0,0 @@
1
- # ADR-0003: Waive unrecoverable scheduled model cost during repair
2
-
3
- **Status:** Accepted
4
-
5
- **Date:** 2026-07-22
6
-
7
- ## Context
8
-
9
- An agent-mode scheduled run completed provider/model work on 2026-07-14, but OpenRouter did not return a model-cost value. The 75-Credit base charge was already captured. Its settlement remained `cost_pending`, which correctly prevented an unverified model delta from being guessed but also blocked every later scheduled run for the account. The original authorization token and exact vendor cost are not recoverable.
10
-
11
- ## Decision
12
-
13
- For this historical repair, preserve the captured base charge, waive the unknown model-cost delta, and close the settlement and authorization as settled with an explicit `model_cost_unavailable_waived` audit reason. Do not invent or estimate a vendor cost.
14
-
15
- Future runs continue to require exact reported OpenRouter cost under the existing billing contract. A new missing-cost occurrence must retain enough durable provider evidence for reconciliation or be explicitly waived by an audited operator repair; it must never be silently estimated.
16
-
17
- ## Consequences
18
-
19
- - The account is no longer permanently blocked by an unrecoverable historical delta.
20
- - MCP Scraper absorbs the unknown model expense for this run.
21
- - The repair remains visible in billing-event metadata and preserves the original base-charge receipt.
22
- - A productized recovery path is still needed if missing provider cost recurs.
@@ -1,13 +0,0 @@
1
- # Architecture Decision Records
2
-
3
- Short, durable records of *why* a non-obvious technical decision was made — the context, the choice, and the consequences — so future readers (and future us) don't re-litigate or accidentally undo it.
4
-
5
- Write one when a decision is hard to reverse, surprising, or encodes a constraint that isn't visible in the code (an anti-bot workaround, a vendor limitation, a deliberate fallback). Don't write one for routine changes.
6
-
7
- ## Format
8
- `NNNN-kebab-title.md`, Nygard-style: **Status · Context · Decision · Consequences**. Status is `Proposed` → `Accepted` → (later) `Superseded by ADR-XXXX`. Number sequentially; never renumber.
9
-
10
- ## Index
11
- - [0001 — In-page API interception for anti-bot scraping (Facebook Ad Library)](./0001-in-page-graphql-interception-for-anti-bot-scraping.md)
12
- - [0002 — Hybrid Smart RAG retrieval for memory vaults](./0002-hybrid-smart-rag-vault-retrieval.md)
13
- - [0003 — Waive unrecoverable scheduled model cost during repair](./0003-waive-unrecoverable-scheduled-model-cost.md)
@@ -1,206 +0,0 @@
1
- # Final Tooling Spec — rename, server instructions, description optimization, catalog
2
-
3
- Execution blueprint consolidating the decisions reached: (A) server namespace rename
4
- `mcp-scraper → scraper`, (B) `*_workflow` rename of workflow-class tools, (C) a front-loaded
5
- **server-instructions block** carrying the cross-tool routing map, (D) a per-tool description
6
- optimization pass that moves routing *out* of descriptions, and (E) catalog integration.
7
-
8
- Grounded in: the SDK's `instructions?: string` option
9
- (`node_modules/@modelcontextprotocol/sdk/dist/esm/server/index.d.ts:15`), verified tool registrations
10
- in `src/mcp/paa-mcp-server.ts` and `src/mcp/browser-agent-mcp-server.ts`, and the executor map in
11
- `src/mcp/http-mcp-tool-executor.ts`.
12
-
13
- Governing principle (from `docs/mcp-tool-design-guide.md`): **the model routes on names + the
14
- server-instructions block before it fetches any schema.** Therefore the cross-tool routing map must
15
- live in the instructions block, never solely in per-tool descriptions that may never be fetched.
16
-
17
- ---
18
-
19
- ## Part A — Server namespace rename: `mcp-scraper` → `scraper`
20
-
21
- The client prepends `mcp__`, so today every tool is `mcp__mcp-scraper__…` (stutter). Renaming the
22
- **server name** to `scraper` yields `mcp__scraper__…`.
23
-
24
- ### A.1 Exact edits — CORRECTED after verification
25
-
26
- > **The namespace `mcp__<key>__` comes from the CLIENT REGISTRATION KEY, not the McpServer `name`.**
27
- > Verified: `src/cli/agent-config.ts` registers the server with the key `mcp-scraper` via
28
- > `claude mcp add mcp-scraper …` and the Desktop config block `'mcp-scraper': {…}`. THAT key is what
29
- > Claude turns into `mcp__mcp-scraper__*`. Changing only `new McpServer({name})` does NOT change the
30
- > prefix. The full set of edits to actually rename the namespace:
31
-
32
- - `src/cli/agent-config.ts:54` — `['mcp', 'remove', 'mcp-scraper', …]` → `'scraper'`.
33
- - `src/cli/agent-config.ts:58` — `['mcp', 'add', 'mcp-scraper', …]` → `'scraper'`.
34
- - `src/cli/agent-config.ts:75` — Desktop config key `'mcp-scraper': {…}` → `'scraper': {…}`.
35
- - `scripts/build-mcpb.mjs:61` — `.mcpb` manifest server `name: 'mcp-scraper'` → `'scraper'` (Desktop
36
- extension registration key). Line 50 (`mcp-scraper-desktop-extension`) is the extension *package* id; leave it.
37
- - `src/mcp/paa-mcp-server.ts:152` and `bin/mcp-scraper-combined-stdio-server.ts:61` — the McpServer
38
- `name` is cosmetic self-description; update to `'scraper'` for consistency but it is NOT what drives the prefix.
39
- - **Tests:** `tests/unit/cli-agent-config.test.ts:32` asserts `'mcp-scraper'` — update.
40
- - **Migration is harder than a code change:** existing users already have the server registered under
41
- `mcp-scraper` in *their* client config. Renaming the key here only affects NEW installs; existing
42
- installs keep `mcp__mcp-scraper__*` until they `claude mcp remove mcp-scraper && claude mcp add scraper`
43
- (or reinstall the `.mcpb`). Document this explicitly; consider shipping a migration note/command.
44
-
45
- ### A.2 Leave alone (URL / artifact stability)
46
- - Bundle filenames/paths in `build-mcpb.mjs` (`mcp-scraper.mcpb`, `mcp-scraper-${version}.mcpb`,
47
- `build/mcpb/mcp-scraper`, `dist/bin/mcp-scraper-combined-stdio-server.js`) — **keep**. These are
48
- download URLs and build paths; changing them breaks existing download links for no namespace benefit.
49
- - Repo name, GitHub URLs — keep.
50
-
51
- ### A.3 Breaking-change handling (REQUIRED — do not skip)
52
- Renaming the server changes **every tool identifier** clients see → invalidates:
53
- - User permission allowlists referencing `mcp__mcp-scraper__*`.
54
- - Any saved agent configs / scripted tool calls by full name.
55
-
56
- Mitigations:
57
- - **Do it in a single release, called out in the changelog** as a breaking namespace change with a
58
- one-line migration ("re-allow tools; the namespace is now `mcp__scraper__…`").
59
- - Bump a **minor/major** version, not a patch.
60
- - Ship in the same release as Parts B–D so users re-onboard once, not repeatedly.
61
-
62
- ---
63
-
64
- ## Part B — Workflow-class tool renames (`*_workflow` suffix)
65
-
66
- Rule: a tool whose mental model is "run a multi-step workflow" carries `_workflow` so the keyword is in
67
- the name (matches a user saying "run the … workflow / research / directory"). Engine plumbing
68
- (`workflow_run/step/status/list/suggest/artifact_read`) is **not** renamed — it's infrastructure, not a
69
- user-named capability.
70
-
71
- ### B.1 Rename map
72
- | Current name | New name | File(s) |
73
- |---|---|---|
74
- | `rank_tracker_blueprint` | `rank_tracker_workflow` | `paa-mcp-server.ts:346`, schema, manifest, catalog |
75
- | `browser_capture_fanout` | `query_fanout_workflow` | `browser-agent-mcp-server.ts:791`, `browser-agent-tool-schemas.ts:189`, manifest, catalog |
76
- | `directory_workflow` | *(unchanged — already correct)* | — |
77
- | *(new, optional)* | `deep_research_workflow` | thin wrapper → `workflow_run` for the `deep-research` workflow; register in `paa-mcp-server.ts` |
78
-
79
- ### B.2 Per-tool change checklist (apply to each renamed tool)
80
- 1. **Registration** — `server.registerTool('<new_name>', { … })` (the string literal).
81
- 2. **Input schema** — rename the exported `…InputSchema` only if its name encodes the old tool name;
82
- the schema *shape* is unchanged.
83
- 3. **Manifest** — `src/mcp/mcp-tool-manifest.ts` `entry({ name: '<new_name>', … })`.
84
- 4. **Executor** — `src/mcp/http-mcp-tool-executor.ts` method name only if it mirrors the tool name; the
85
- **REST route string is unchanged** (rename is MCP-surface only, not the backend).
86
- 5. **Formatter** — `format<Tool>` function name optional; wiring in `registerTool` handler updated.
87
- 6. **Catalog** — update `mcp_name` in `docs/tool-catalog-spec.md` §3.2/§9 and the seed.
88
- 7. **Lint/manifest generators** — re-run `npm run mcp:manifest` and `npm run mcp:lint-tools` to
89
- regenerate `docs/mcp-tool-manifest.generated.*`.
90
-
91
- ### B.3 Optional backward-compat (recommended for 1–2 releases)
92
- Register the **old name as a hidden alias** pointing at the same handler, so in-flight callers don't
93
- hard-break: `server.registerTool('<old_name>', { ...sameConfig, annotations: { deprecated: true } }, handler)`.
94
- Remove after one release. If you'd rather hard-cut, fold it into the Part A breaking release.
95
-
96
- ---
97
-
98
- ## Part C — Server-instructions block (the pre-fetch routing map)
99
-
100
- This is the highest-leverage change: a short string loaded **at session start** (alongside tool names),
101
- so the model routes correctly *before* fetching any schema.
102
-
103
- ### C.1 New file — `src/mcp/server-instructions.ts`
104
- ```ts
105
- export const SERVER_INSTRUCTIONS = `
106
- This server scrapes and analyzes web, social, search, maps, and site data. Pick a tool by NAME using
107
- this routing map; load its schema before calling.
108
-
109
- ROUTING
110
- - One web page → extract_url. Whole site (crawl + SEO) → extract_site. Just the URL list → map_site_urls.
111
- - Web search (SERP) → search_serp. Full SERP + People-Also-Ask detail → harvest_paa.
112
- - Google Maps: find places → maps_search; one place deep-dive + reviews → maps_place_intel.
113
- - YouTube: find/list videos → youtube_harvest; transcribe one → youtube_transcribe.
114
- - Facebook: find ads → facebook_ad_search; transcribe an ad → facebook_ad_transcribe; transcribe a
115
- video → facebook_video_transcribe; page profile → facebook_page_intel.
116
- - Instagram: profile inventory → instagram_profile_content; one post/reel → instagram_media_download.
117
- - Workflows (multi-step): local directory → directory_workflow; rank blueprint → rank_tracker_workflow;
118
- AI-answer citation fan-out (AEO) → query_fanout_workflow; deep research → deep_research_workflow.
119
- - Anything without a dedicated tool (e.g. Reddit, arbitrary logged-in sites) → the browser_* agent
120
- (browser_open then navigate/read), and capture profiles via browser_profile_*.
121
-
122
- NOTES
123
- - Bulk/full-site crawls: call extract_site with rotateProxies:true for blocked/rate-limited sites; it
124
- discovers URLs and returns a saved folder/artifact, not the full content inline.
125
- - Tools that open a browser session or transcribe media cost credits and take longer — prefer the
126
- cheapest tool that answers the question (plain search/extract before browser agents).
127
- - Results may save large output to disk/artifact and return a summary + path; read the path for detail.
128
- `.trim()
129
- ```
130
- > Keep it **cross-tool only** — no per-tool parameter detail (that's the fetched schema's job). Use
131
- > neutral capability language; never name the underlying browser/proxy vendor (concealment policy).
132
-
133
- ### C.2 Wiring (both server construction sites)
134
- - `src/mcp/paa-mcp-server.ts:152`:
135
- `new McpServer({ name: 'scraper', version: PACKAGE_VERSION }, { instructions: SERVER_INSTRUCTIONS })`
136
- - `bin/mcp-scraper-combined-stdio-server.ts:61`: same second argument.
137
- - Import `SERVER_INSTRUCTIONS` from `../src/mcp/server-instructions.js` (and `./server-instructions.js`).
138
-
139
- ### C.3 The insight, enforced
140
- Per-tool descriptions must **NOT** be the only home for "which tool for which job." That guidance moves
141
- to `SERVER_INSTRUCTIONS` (always loaded). Per-tool descriptions keep only a **single** negative-space
142
- line for post-fetch tie-breaking (see Part D). This prevents the failure where the routing hint lives in
143
- a description the model never fetched.
144
-
145
- ---
146
-
147
- ## Part D — Tool-description optimization pass
148
-
149
- For every tool in `mcp-tool-schemas.ts` / `paa-mcp-server.ts` `registerTool` descriptions, rewrite to
150
- this shape (apply the `docs/mcp-tool-design-guide.md` checklist):
151
-
152
- 1. **Line 1 = when-to-use**, in the user's words. ("Transcribe one YouTube video to text…")
153
- 2. **Line 2 = side-effects/cost** if any. ("Opens a browser session; costs credits per minute.")
154
- 3. **Params**: each param's `.describe()` states intent + bounds + default, mirroring the Zod schema.
155
- 4. **One** negative-space line max. ("For a whole site use extract_site.") — the *broad* map lives in
156
- §C, so do not enumerate the full tool family here.
157
- 5. **Delete** redundant cross-tool prose and marketing. Shorter descriptions load faster and rank cleaner.
158
-
159
- Concrete trims to make now:
160
- - `extract_url` / `extract_site` / `map_site_urls`: each currently restates the others; cut to one
161
- negative-space line each (the trio map is in §C).
162
- - `search_serp` vs `harvest_paa`: both hit `/harvest/sync`; clarify in one line each that
163
- search_serp = organic results only, harvest_paa = + PAA/AI-Overview/SERP features.
164
- - Renamed workflow tools: rewrite line 1 to lead with the keyword now in the name.
165
-
166
- Re-run `npm run mcp:lint-tools` (the craft linter) after edits; fix any flagged tools.
167
-
168
- ---
169
-
170
- ## Part E — Catalog integration
171
-
172
- Reconcile `docs/tool-catalog-spec.md` with the renames, and ship the **static v1** (no DB) first:
173
-
174
- - Update catalog `mcp_name` values: `rank_tracker_blueprint→rank_tracker_workflow`,
175
- `browser_capture_fanout→query_fanout_workflow`; add `deep_research_workflow`.
176
- - Categories stay **platform-first** (YouTube / Facebook / Instagram / Google Search / Google Maps /
177
- Websites / Browser / Workflows), with **Reddit** served by the Browser category (no dedicated tool) —
178
- or add a `reddit_*` named wrapper later if AI routing for Reddit is wanted (per the design guide).
179
- - **v1 implementation = static**: `src/api/catalog-seed.ts` exports `CATALOG`; `GET /catalog` in
180
- `server.ts` returns it. Defer the 6-table DB schema (§2 of the catalog spec) to v2 (runtime editing).
181
- - Catalog params must mirror the live Zod schemas (field names/bounds) so the frontend posts valid bodies.
182
-
183
- ---
184
-
185
- ## Part F — Build order
186
- 1. **Part C** (server instructions) — highest leverage, lowest risk, non-breaking. Ship first.
187
- 2. **Part D** (description optimization) — non-breaking; pairs with C (move routing out → into C).
188
- 3. **Part B** (workflow renames) — with hidden aliases (B.3) so it's non-breaking.
189
- 4. **Part A** (server rename) — the one breaking change; batch into a single versioned release with a
190
- changelog migration note. Run after B/C/D so users re-onboard once.
191
- 5. **Part E** static catalog — independent; any time.
192
- 6. Regenerate `docs/mcp-tool-manifest.generated.*`; run `npm run mcp:lint-tools`; full test + live tests;
193
- then the release flow (3 version files, Vercel prod first, npm publish, rebuild `.mcpb`).
194
-
195
- ## Part G — Verification checklist
196
- - [ ] `grep -rn "mcp-scraper'" src bin scripts` → only intentional (bundle filenames), server `name` = `scraper`.
197
- - [ ] No tool identifier appears in two namespaces; `npm run mcp:manifest` regenerated.
198
- - [ ] `SERVER_INSTRUCTIONS` loads on `initialize` (verify via an MCP inspector / the live protocol test).
199
- - [ ] Every renamed tool: registration + manifest + catalog `mcp_name` updated; old name aliased or cut.
200
- - [ ] No per-tool description contains the full cross-tool map (only one negative-space line).
201
- - [ ] No vendor name (browser/proxy provider) in instructions or any description.
202
- - [ ] `npm run mcp:lint-tools` clean; `vitest run tests/unit tests/contract` green.
203
-
204
- ## Non-goals (this spec)
205
- DB-backed catalog (v2), the `reddit_*` / `deep_research_workflow` *implementations* beyond thin
206
- registration, per-plan tool gating, and the frontend rendering itself.