smart-web-mcp 0.44.1 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (162) hide show
  1. package/CHANGELOG.md +350 -0
  2. package/README.md +33 -568
  3. package/dist/assessment.d.ts +10 -3
  4. package/dist/assessment.js +21 -6
  5. package/dist/assessment.js.map +1 -1
  6. package/dist/body-export.d.ts +10 -0
  7. package/dist/body-export.js +28 -0
  8. package/dist/body-export.js.map +1 -0
  9. package/dist/browser-session.d.ts +32 -1
  10. package/dist/browser-session.js +373 -3
  11. package/dist/browser-session.js.map +1 -1
  12. package/dist/cli-fetch-options.d.ts +7 -0
  13. package/dist/cli-fetch-options.js +41 -0
  14. package/dist/cli-fetch-options.js.map +1 -0
  15. package/dist/cli.js +28 -11
  16. package/dist/cli.js.map +1 -1
  17. package/dist/comment-export.d.ts +10 -0
  18. package/dist/comment-export.js +39 -0
  19. package/dist/comment-export.js.map +1 -0
  20. package/dist/composition.d.ts +1 -22
  21. package/dist/composition.js +0 -16
  22. package/dist/composition.js.map +1 -1
  23. package/dist/extraction/budget.d.ts +3 -0
  24. package/dist/extraction/budget.js +88 -7
  25. package/dist/extraction/budget.js.map +1 -1
  26. package/dist/extraction/document.js +134 -9
  27. package/dist/extraction/document.js.map +1 -1
  28. package/dist/korea/coupang-filters.d.ts +5 -0
  29. package/dist/korea/coupang-filters.js +87 -0
  30. package/dist/korea/coupang-filters.js.map +1 -0
  31. package/dist/korea/coupang.d.ts +27 -6
  32. package/dist/korea/coupang.js +202 -82
  33. package/dist/korea/coupang.js.map +1 -1
  34. package/dist/korea/proxy-client.d.ts +1 -0
  35. package/dist/korea/proxy-client.js +6 -2
  36. package/dist/korea/proxy-client.js.map +1 -1
  37. package/dist/korea/smartvertical.d.ts +1 -0
  38. package/dist/korea/smartvertical.js +39 -25
  39. package/dist/korea/smartvertical.js.map +1 -1
  40. package/dist/lib.d.ts +4 -2
  41. package/dist/lib.js +2 -1
  42. package/dist/lib.js.map +1 -1
  43. package/dist/mcp-server.js +30 -15
  44. package/dist/mcp-server.js.map +1 -1
  45. package/dist/operator-browser-fixture.d.ts +7 -0
  46. package/dist/operator-browser-fixture.js +19 -0
  47. package/dist/operator-browser-fixture.js.map +1 -0
  48. package/dist/operator-canary-cli.d.ts +6 -0
  49. package/dist/operator-canary-cli.js +190 -0
  50. package/dist/operator-canary-cli.js.map +1 -0
  51. package/dist/operator-canary.d.ts +74 -0
  52. package/dist/operator-canary.js +135 -0
  53. package/dist/operator-canary.js.map +1 -0
  54. package/dist/operator-cdp.d.ts +15 -0
  55. package/dist/operator-cdp.js +195 -0
  56. package/dist/operator-cdp.js.map +1 -0
  57. package/dist/operator-frame.d.ts +3 -0
  58. package/dist/operator-frame.js +55 -0
  59. package/dist/operator-frame.js.map +1 -0
  60. package/dist/settings.d.ts +3 -0
  61. package/dist/settings.js +1 -0
  62. package/dist/settings.js.map +1 -1
  63. package/dist/shared.d.ts +33 -8
  64. package/dist/shared.js +15 -9
  65. package/dist/shared.js.map +1 -1
  66. package/dist/smartbrowser.d.ts +169 -0
  67. package/dist/smartbrowser.js +288 -0
  68. package/dist/smartbrowser.js.map +1 -0
  69. package/dist/smartcrawl.js +16 -21
  70. package/dist/smartcrawl.js.map +1 -1
  71. package/dist/smartfetch/academic-fallback.js +3 -5
  72. package/dist/smartfetch/academic-fallback.js.map +1 -1
  73. package/dist/smartfetch/archive-fallback.js +2 -5
  74. package/dist/smartfetch/archive-fallback.js.map +1 -1
  75. package/dist/smartfetch/assets.d.ts +1 -1
  76. package/dist/smartfetch/assets.js +28 -1
  77. package/dist/smartfetch/assets.js.map +1 -1
  78. package/dist/smartfetch/jina-reader.d.ts +1 -0
  79. package/dist/smartfetch/jina-reader.js +17 -3
  80. package/dist/smartfetch/jina-reader.js.map +1 -1
  81. package/dist/smartfetch/pipeline.js +15 -25
  82. package/dist/smartfetch/pipeline.js.map +1 -1
  83. package/dist/smartfetch/provider-policy.d.ts +3 -0
  84. package/dist/smartfetch/provider-policy.js +45 -2
  85. package/dist/smartfetch/provider-policy.js.map +1 -1
  86. package/dist/smartfetch/provider-types.d.ts +8 -0
  87. package/dist/smartfetch/providers/article.js +76 -10
  88. package/dist/smartfetch/providers/article.js.map +1 -1
  89. package/dist/smartfetch/providers/blind-post.d.ts +2 -0
  90. package/dist/smartfetch/providers/blind-post.js +167 -0
  91. package/dist/smartfetch/providers/blind-post.js.map +1 -0
  92. package/dist/smartfetch/providers/clien.d.ts +2 -0
  93. package/dist/smartfetch/providers/clien.js +157 -0
  94. package/dist/smartfetch/providers/clien.js.map +1 -0
  95. package/dist/smartfetch/providers/commerce.d.ts +15 -0
  96. package/dist/smartfetch/providers/commerce.js +123 -7
  97. package/dist/smartfetch/providers/commerce.js.map +1 -1
  98. package/dist/smartfetch/providers/dcinside.js +313 -20
  99. package/dist/smartfetch/providers/dcinside.js.map +1 -1
  100. package/dist/smartfetch/providers/github-issue.d.ts +3 -0
  101. package/dist/smartfetch/providers/github-issue.js +101 -0
  102. package/dist/smartfetch/providers/github-issue.js.map +1 -0
  103. package/dist/smartfetch/providers/hackernews-html.d.ts +21 -0
  104. package/dist/smartfetch/providers/hackernews-html.js +148 -0
  105. package/dist/smartfetch/providers/hackernews-html.js.map +1 -0
  106. package/dist/smartfetch/providers/hackernews.js +76 -14
  107. package/dist/smartfetch/providers/hackernews.js.map +1 -1
  108. package/dist/smartfetch/providers/index.js +9 -1
  109. package/dist/smartfetch/providers/index.js.map +1 -1
  110. package/dist/smartfetch/providers/linkedin.js +2 -1
  111. package/dist/smartfetch/providers/linkedin.js.map +1 -1
  112. package/dist/smartfetch/providers/naver-blog.d.ts +2 -0
  113. package/dist/smartfetch/providers/naver-blog.js +87 -45
  114. package/dist/smartfetch/providers/naver-blog.js.map +1 -1
  115. package/dist/smartfetch/providers/naver-cafe-native.d.ts +7 -0
  116. package/dist/smartfetch/providers/naver-cafe-native.js +188 -0
  117. package/dist/smartfetch/providers/naver-cafe-native.js.map +1 -0
  118. package/dist/smartfetch/providers/naver-cafe.js +165 -4
  119. package/dist/smartfetch/providers/naver-cafe.js.map +1 -1
  120. package/dist/smartfetch/providers/naver-map-search.js +1 -1
  121. package/dist/smartfetch/providers/naver-map-search.js.map +1 -1
  122. package/dist/smartfetch/providers/ppomppu.d.ts +2 -0
  123. package/dist/smartfetch/providers/ppomppu.js +289 -0
  124. package/dist/smartfetch/providers/ppomppu.js.map +1 -0
  125. package/dist/smartfetch/providers/reddit.js +24 -3
  126. package/dist/smartfetch/providers/reddit.js.map +1 -1
  127. package/dist/smartfetch/providers/threads.js +1 -1
  128. package/dist/smartfetch/providers/threads.js.map +1 -1
  129. package/dist/smartfetch/providers/wikipedia.js +138 -0
  130. package/dist/smartfetch/providers/wikipedia.js.map +1 -1
  131. package/dist/smartfetch.d.ts +7 -1
  132. package/dist/smartfetch.js +353 -76
  133. package/dist/smartfetch.js.map +1 -1
  134. package/dist/smartsearch-brave.d.ts +2 -0
  135. package/dist/smartsearch-brave.js +43 -0
  136. package/dist/smartsearch-brave.js.map +1 -0
  137. package/dist/smartsearch-browser.d.ts +5 -0
  138. package/dist/smartsearch-browser.js +116 -0
  139. package/dist/smartsearch-browser.js.map +1 -0
  140. package/dist/smartsearch-ddg.d.ts +2 -0
  141. package/dist/smartsearch-ddg.js +46 -0
  142. package/dist/smartsearch-ddg.js.map +1 -0
  143. package/dist/smartsearch-node.d.ts +6 -0
  144. package/dist/smartsearch-node.js +21 -0
  145. package/dist/smartsearch-node.js.map +1 -0
  146. package/dist/smartsearch-postgresql.d.ts +7 -0
  147. package/dist/smartsearch-postgresql.js +44 -0
  148. package/dist/smartsearch-postgresql.js.map +1 -0
  149. package/dist/smartsearch-python.d.ts +8 -0
  150. package/dist/smartsearch-python.js +87 -0
  151. package/dist/smartsearch-python.js.map +1 -0
  152. package/dist/smartsearch-quality.d.ts +6 -0
  153. package/dist/smartsearch-quality.js +47 -0
  154. package/dist/smartsearch-quality.js.map +1 -0
  155. package/dist/smartsearch-youtube.d.ts +7 -0
  156. package/dist/smartsearch-youtube.js +76 -0
  157. package/dist/smartsearch-youtube.js.map +1 -0
  158. package/dist/smartsearch.d.ts +2 -0
  159. package/dist/smartsearch.js +506 -84
  160. package/dist/smartsearch.js.map +1 -1
  161. package/package.json +8 -5
  162. package/python/requirements-undetected-chromedriver.txt +2 -2
package/README.md CHANGED
@@ -2,598 +2,63 @@
2
2
 
3
3
  # smart-web
4
4
 
5
- Local MCP for agentic web retrieval.
5
+ One local MCP entry point for finding, reading, and using the web.
6
6
 
7
- `smart-web` bundles four public tools in one stdio server:
7
+ **Goal:** finish the user's web task—not merely return a successful request or a browser handoff. Use fast retrieval where it works; continue through the approved local browser-use session when it does not.
8
8
 
9
- - `smartfetch` for direct URLs
10
- - `smartsearch` for discovery
11
- - `smartcrawl` for same-site multi-page traversal
12
- - `smartvertical` for Korean/local structured vertical data such as places, weather, transit, finance, law, commerce, and safety lookups
9
+ ## Principles
13
10
 
14
- It is designed for **agents and MCP hosts**, not for manual browser automation. The job is to return structured retrieval output first, then signal when a browser-task runtime would be a better next step.
11
+ - Judge success by the requested content or action, including comments and exact product identity.
12
+ - Measure search relevance, freshness, query constraints, and usable sources—not result count.
13
+ - Prioritize failures from real workflows across sites, not a single showcase domain.
14
+ - Keep output compact, with source evidence, attempt history, and explicit omissions.
15
+ - Isolate fragile site logic; detect drift and repair it through bounded, verified agent runs.
15
16
 
16
- If a page is login-gated, keep the login UI, session persistence, and capture flow in a companion runtime instead of moving that stateful behavior into the `smart-web` core.
17
+ ## Status
17
18
 
18
- - npm: <https://www.npmjs.com/package/smart-web-mcp>
19
- - MCP Registry: `io.github.jojo-labs/smart-web`
20
- - Issues: <https://github.com/jojo-labs/smart-web/issues>
19
+ The package exposes four retrieval tools plus the bounded `smartbrowser` action tool. The **operator-first revamp is in progress**: broader interactive coverage and automatic repair remain acceptance goals, not a claim of universal support. The legacy web-task-api backend is retired; browser actions use the approved local session. Cookies persist across batches, but tab state only persists within a batch.
21
20
 
22
- If `smart-web` returns reproducible incorrect or misleading output, open a GitHub issue with the exact input, tool arguments, observed output, expected behavior, and version. Skip obvious transient network, auth, or rate-limit failures unless the smart-web classification itself looks wrong.
21
+ Track shipped evidence and remaining work in [the revamp issue](https://github.com/jojo-labs/smart-web/issues/496). Browser-readable content should not require the user to switch tools; genuine login, permission, or unavailable-content limits remain explicit.
23
22
 
24
- ## Highlights
23
+ ## Quick start
25
24
 
26
- - one local MCP instead of separate search, fetch, crawl, and Korean vertical-data servers
27
- - explicit routing: URL → `smartfetch`, query → `smartsearch`, known site → `smartcrawl`
28
- - up-front smartfetch acquisition lanes instead of broad direct → browser retry chains, with `pipeline.acquire.attempt_map` showing seed/direct/relay/archive/browser phases and skip/failure reasons
29
- - optional Jina Reader relay for weak generic/article fetches when you explicitly enable it, with JSON mode and alternate-link preservation
30
- - shared document extraction for article-like pages keeps browser primary content separate from raw shell HTML, preserving headings, links, blocks, and Markdown when requested
31
- - docs export manifests are versioned and include per-document SHA-256 hashes plus resource-handle manifests for stable local Markdown mirrors; resource-handle manifests include both exported documents and metadata handles for the summary/manifest files, and successful export `structuredContent` embeds the same handle manifest for host follow-up calls
32
- - budgeted `smartfetch` projection records explicit truncation metadata instead of scattering silent hardcoded caps; MCP calls default to the compact preset for token-sensitive agent hosts
33
- - weak generic pages can recover through JSON-LD or Next.js payload extraction and same-origin RSS/Atom feed discovery before falling back to a handoff
34
- - optional local adaptive scraper fallback can call a Scrapling-compatible `scrapling` command for AI-targeted extraction on weak JavaScript shells before spending a full Playwright browser pass
35
- - paywalled article pages prefer lawful archive recovery and search-indexed reference previews before telling the host to escalate elsewhere
36
- - LinkedIn authwall handling prefers compliant public fallbacks such as archives and search-indexed reference metadata before giving up
37
- - legal paper fallback for academic URLs: DOI-aware OpenAlex, Unpaywall, Semantic Scholar, CORE discovery, and Europe PMC enrichment plus bioRxiv/medRxiv API fallback when a direct paper page is thin or blocked
38
- - site-native public search fallbacks cover Reddit (Arctic Shift, live RSS, PullPush, then JSON), GitHub repository discovery/repo-docs queries, npm search/exact packages, Hacker News, Stack Exchange, Wikipedia, Velog, MDN, and crates.io when a matching `site:` query is available
39
- - YouTube watch, shorts, embed, live, and youtu.be video URLs get best-effort public caption enrichment by default while non-video YouTube pages stay on the generic fetch path
40
- - structured output with `assessment`, `pipeline`, and partial-result `research_handoff` fields for host-side decision making
41
- - local-first defaults with privacy-aware search/fetch behavior
42
- - useful normalization for common public surfaces instead of raw HTML dumps
43
- - social unavailable shells such as Threads generic join/login pages stay partial and chrome-free instead of masquerading as verified post content
44
- - blocked commerce URLs can accept up to five contextual `reference_queries`; public catalog metadata is returned only when the resolved product, item, and vendor identifiers all match, so a title hint cannot silently select a different listing
45
-
46
- ## Tool routing
47
-
48
- - `smartfetch` first when the user already gave a URL or shortlink
49
- - `smartsearch` when the user gave a topic, keywords, or a `site:` query
50
- - `smartcrawl` when you already know the site and need multiple relevant pages
51
-
52
- | Situation | Tool |
53
- | ------------------------------------------------- | --------------- |
54
- | User already gave a URL or shortlink | `smartfetch` |
55
- | User gave a topic, keywords, or `site:` query | `smartsearch` |
56
- | You already know the site and need multiple pages | `smartcrawl` |
57
- | User needs Korean/local structured data | `smartvertical` |
58
-
59
- `smart-web` is a retrieval layer. It is not a general-purpose click/type/form-fill browser agent.[^1]
60
-
61
- ## What the tools return
62
-
63
- All public tools are shaped for agent consumption. Common high-signal fields include:
64
-
65
- - `assessment`: confidence, block/auth hints, recommended handoff
66
- - `pipeline`: resolve/acquire/normalize/assess stages, plus policy metadata for relay-style usage and authwall likelihood; `pipeline.acquire.attempt_map` describes the default seed, direct/impit, Jina Reader, archive, and browser phases as succeeded, failed, or skipped
67
- - `research_handoff`: compact metadata on blocked or partial `smartfetch` results telling downstream research workers and editable report renderers which evidence fields to consume
68
- - `budget`: output projection metadata showing which fields were truncated and by how much, plus final MCP `response_chars` counts for compact-JSON `structuredContent` and rendered text; use `smartfetch`/`smartsearch`/`smartcrawl`/`smartvertical` `output_budget: "compact"` for token-sensitive MCP hosts, `preferred_budget_chars` for host-budget-aware routing, and `"full"` when larger structured payloads are explicitly needed
69
- - `evidence.selected_output_budget*`: the selected `smartfetch` budget preset, winning control, and resolver reason so hosts can see which v0.x compatibility hint actually took effect
70
- - `errors`: structured failure reasons
71
- - `post`, `thread`, `comments`: normalized content surfaces
72
- - `assets` and `download`: file-like outputs when present
73
-
74
- This lets hosts decide whether deterministic retrieval was good enough or whether they should escalate to an auth-aware companion flow or a broader browser-task runtime.[^2]
75
-
76
- For `smartfetch`, MCP `structuredContent` always carries the projected JSON result. MCP calls default to `output_budget: "compact"` so agent hosts do not ingest large page/thread/link payloads unless they ask for them. The projection caps known body/list fields plus arbitrary provider record width, nesting, and aggregate string content, so unknown `post`, thread, or comment fields cannot bypass the preset; `budget.fields` reports omitted nested paths and primary URL/title/author/status/body fields are prioritized. The human-readable `content[0].text` defaults to compact text instead of duplicating that JSON payload, with separate caps for long body text, result, comment, link, and asset sections. Omitted counts are reported in the text view while the projected fields remain available through `structuredContent`; set `format: "json"` only when a host needs JSON repeated in the text channel. Set `output_budget: "balanced"` or `"full"` only for follow-up calls that explicitly need more structured content. For host-side integration, prefer `preferred_budget_chars`; use `context_mode` only as a backward-compatible coarse alias when numeric hints are not supplied. If neither is present, `headroom_tokens` keeps compact output by default and uses `balanced` only when there is enough room. `format: "markdown"` keeps the compact response shape and requests Markdown extraction when available.
77
-
78
- When a `smartfetch` result is blocked, partial, or low-confidence, `structuredContent.research_handoff` stays small and metadata-only. It points downstream tools at the source URL and the projected evidence bundle: `pipeline`, `assessment`, `errors`, `assets`, `outbound_links`, and the full projected `structuredContent`. Heavy research agents or editable report/rendering tools can consume that bundle downstream, but `smart-web` itself remains the unified retrieval kernel and does not spawn those tools.
79
-
80
- For `smartsearch`, MCP calls now project both `structuredContent` and rendered text through the same compact/balanced/full budget controls. Compact search keeps enough ranked links for first-pass discovery while capping every normalized string field, including query/engine metadata, result titles and URLs, snippets, notes, assessment, and search-to-crawl handoff guidance; `contextMaxCharacters` remains a final text-channel cap for legacy hosts. Each shortened field is recorded under `budget.fields`. Use `output_budget: "full"` only for a known follow-up when the host really wants every returned snippet/result in `structuredContent`.
81
-
82
- `smartsearch` reports provider execution separately from search matches under `structuredContent.execution`: the overall `status` is `succeeded` when at least one attempted provider completed, and each entry in `attempts` records its provider, status, result count, and any execution error. If every configured provider fails to execute, MCP returns `isError: true` and the CLI exits non-zero after preserving the structured diagnostics. A successful provider response with zero matching results remains successful even when the compatibility fields stay `engine: "none"` and `results: []`.
83
-
84
- For `smartcrawl`, MCP summary calls now project both `structuredContent` and rendered text through compact/balanced/full budget controls, so broad docs/forum crawls do not flood agent context with every page, candidate, note, or error by default. Summary projection caps all normalized titles, URLs, discovery/page metadata, notes, errors, assessment, and handoff strings and records each omission under `budget.fields`. `output_dir` export mode remains handle-first: successful export responses embed `resource_handle_manifest` in `structuredContent`, and export summaries include file:// re-open guidance for the summary, manifest, resource-handle JSON, and first Markdown document so agent hosts can resume from handles without parsing the full JSON first. The resource-handle manifest also carries ordered `recommended_resources` for hosts that need a ready-to-open summary, manifest, and first document without inferring role priority. Blocked exports label those paths as planned/unwritten and include retry guidance instead of presenting false file handles. Set `format: "json"` only for hosts that cannot consume `structuredContent` and intentionally need JSON repeated in the text channel.
85
-
86
- For `smartvertical`, MCP calls use the same budget-control precedence as the web tools. Compact mode caps rendered text plus arbitrary structured `result` arrays and long nested strings, and `structuredContent.budget.fields` reports exactly which vertical fields were shortened. Use `output_budget: "balanced"` or a larger `preferred_budget_chars` value when a Korean/local lookup needs more rows or full legal/safety record text.
87
-
88
- Use budget presets this way:
89
-
90
- - `compact`: default first pass for agent hosts, link triage, social/thread previews, and "is this page enough?" checks.
91
- - `balanced`: follow-up when the first pass found the right page but omitted useful body, result, comment, or link detail.
92
- - `full`: explicit retrieval pass for a known relevant URL when the host can afford the larger `structuredContent` payload.
93
-
94
- When compact output omits needed content, make a follow-up call for the same URL/query/start site with a larger budget control instead of asking the model to infer from the compact text. Prefer `preferred_budget_chars` when the host knows its available response budget, or use `output_budget: "balanced"` / `"full"` for manual retries.
95
-
96
- ## Supported high-signal surfaces
97
-
98
- - communities: Reddit, DCInside, LinkedIn public posts with public fallback recovery, Algumon
99
- - search-native surfaces: Reddit, GitHub repository discovery and repo-docs path scopes, npm search and exact package path scopes, Hacker News, Stack Exchange, Wikipedia, Velog, MDN, crates.io
100
- - competitive-programming surfaces: solved.ac problem/search/profile routes, BOJ problem/workbook/user pages, Codeforces problem/profile pages, AtCoder task/contest pages, QOJ problem pages, Jungol problem pages
101
- - local map place pages: Naver Map, Kakao Map, and redirected `naver.me` place-share links
102
- - reference/article pages: NamuWiki, Wikipedia, Naver Blog, Tistory, Velog
103
- - media/social: YouTube video URLs with best-effort public transcripts, X, Threads, Instagram, Telegram
104
- - commerce pages: Amazon, Coupang, Danawa, Aladin, AliExpress
105
- - archives/reference wrappers: Wayback snapshots, archive.md lookup pages, arXiv abstract pages with direct paper links
106
- - academic paper surfaces: arXiv, PubMed, PMC, bioRxiv, medRxiv, plus legal DOI-linked OA copies surfaced from OpenAlex, Unpaywall, Semantic Scholar, CORE, and Europe PMC when a publisher page is thin or paywalled
107
-
108
- When a source stays blocked or under-specified, `smart-web` prefers partial but honest output over pretending to have verified content.
109
-
110
- ## Install
25
+ Requires Node.js 24 and npm. Start the stdio server:
111
26
 
112
27
  ```bash
113
28
  npx -y smart-web-mcp
114
29
  ```
115
30
 
116
- When `smartfetch` or `smartcrawl` needs Playwright-backed page loading, `smart-web` checks whether the bundled Chromium revision is available. If the local Playwright browser cache is missing or outdated, `smart-web` warns at startup and will try `npx playwright install chromium` automatically on first browser use before surfacing a structured error.
117
-
118
- ## Host setup
119
-
120
- ### Claude Code
121
-
122
- ```bash
123
- claude mcp add smart-web -- npx -y smart-web-mcp
124
- ```
125
-
126
- Suggested tool-use profile: default to `smartfetch` without `format` and without a budget control for compact first-pass retrieval; retry with `output_budget: "balanced"` or `"full"` only when the user asks for deeper extraction. If Claude Code exposes a host response budget, send it as `preferred_budget_chars`.
127
-
128
- ### Codex
129
-
130
- ```bash
131
- codex mcp add smart-web -- npx -y smart-web-mcp
132
- ```
133
-
134
- Or via `~/.codex/config.toml` (or project-scoped `.codex/config.toml`):
135
-
136
- ```toml
137
- [mcp_servers.smart-web]
138
- command = "npx"
139
- args = ["-y", "smart-web-mcp"]
140
- ```
141
-
142
- Suggested tool-use profile: keep default compact text for normal browsing. For targeted follow-up retrieval, pass `preferred_budget_chars` from the remaining response/tool budget when available; otherwise pass one explicit `output_budget` preset.
143
-
144
- ### OpenCode
145
-
146
- ```json
147
- {
148
- "$schema": "https://opencode.ai/config.json",
149
- "mcp": {
150
- "smart-web": {
151
- "type": "local",
152
- "command": ["npx", "-y", "smart-web-mcp"],
153
- "enabled": true
154
- }
155
- }
156
- }
157
- ```
158
-
159
- Suggested tool-use profile: call `smartfetch` with no `format` for normal retrieval, `format: "markdown"` only when Markdown extraction is useful, and `format: "json"` only for hosts that intentionally require the full JSON duplicated into `content[0].text`.
160
-
161
- ### Hermes/Raon
162
-
163
- For Hermes/Raon-style hosts, keep the MCP server command standard and set budget controls per tool call:
164
-
165
- ```json
166
- {
167
- "tool": "smartfetch",
168
- "arguments": {
169
- "url": "https://example.com/post",
170
- "preferred_budget_chars": 40000
171
- }
172
- }
173
- ```
174
-
175
- Use `preferred_budget_chars` as the primary host-profile control when Hermes/Raon knows the response budget. Use `headroom_tokens` only when the host has token headroom but no character budget, and use `context_mode` only for older integrations that cannot send numeric budget metadata.
176
-
177
- Budget-control precedence is deterministic:
178
-
179
- 1. `output_budget`
180
- 2. `preferred_budget_chars`
181
- 3. `headroom_tokens`
182
- 4. `context_mode`
183
- 5. compact default
184
-
185
- Send only one control when possible. If older host profiles send more than one, `evidence.selected_output_budget_source` reports which control won.
186
-
187
- ### Custom settings file
188
-
189
- To use a non-default settings path, append `--settings-file` to the command:
190
-
191
- ```bash
192
- npx -y smart-web-mcp --settings-file /absolute/path/to/smart-web.settings.json
193
- ```
194
-
195
- For Claude Code:
196
-
197
- ```bash
198
- claude mcp add smart-web -- npx -y smart-web-mcp --settings-file /absolute/path/to/smart-web.settings.json
199
- ```
200
-
201
- For OpenCode, add `"args"` after `"command"`:
202
-
203
- ```json
204
- {
205
- "mcp": {
206
- "smart-web": {
207
- "type": "local",
208
- "command": ["npx", "-y", "smart-web-mcp", "--settings-file", "/absolute/path/to/smart-web.settings.json"],
209
- "enabled": true
210
- }
211
- }
212
- }
213
- ```
214
-
215
- ## Configuration
216
-
217
- ### Settings path
218
-
219
- Default: `~/.config/smart-web/settings.json`
220
-
221
- Print a template:
222
-
223
- ```bash
224
- npx -y smart-web-mcp --print-settings-example
225
- ```
226
-
227
- Initialize the default file:
228
-
229
- ```bash
230
- npx -y smart-web-mcp --init-settings
231
- ```
232
-
233
- `smart-web` treats the settings file as the single runtime config surface.
234
-
235
- ### Full settings reference
236
-
237
- ```jsonc
238
- {
239
- // "balanced" (default) or "private"
240
- // "private" disables relay-style providers and public search helpers
241
- "profile": "balanced",
242
-
243
- "runtime": {
244
- // Optional override for staging/export temp files
245
- // Default: platform cache root (e.g. ~/.cache/smart-web/tmp)
246
- "tempDir": ""
247
- },
248
-
249
- "search": {
250
- // Self-hosted SearXNG instance URL
251
- "searxngBaseUrl": "",
252
- "enableSearxng": true
253
- },
254
-
255
- "fetch": {
256
- // Browser overrides — most users leave these empty
257
- "chromeChannel": "",
258
- "chromePath": "",
259
- // Auto-install Playwright Chromium when missing (default: true)
260
- "autoInstallPlaywright": true,
261
- // Jina Reader relay — opt-in third-party path for weak article pages
262
- "enableJinaReader": false,
263
- "jinaReaderBaseUrl": "https://r.jina.ai/",
264
- // Academic fallback — legal OA enrichment for paper URLs
265
- "enableAcademicFallback": true,
266
- "enableOpenAlex": true,
267
- "enableEuropePmc": true,
268
- "enableBiorxivApi": true,
269
- // Unpaywall — requires a contact email
270
- "enableUnpaywall": true,
271
- "unpaywallEmail": "",
272
- // Semantic Scholar — optional API key for higher-rate access
273
- "enableSemanticScholar": true,
274
- "semanticScholarApiKey": "",
275
- // CORE — optional API key for richer search-backed enrichment
276
- "enableCoreDiscovery": true,
277
- "coreApiKey": "",
278
- // FxTwitter — transparent x.com → fxtwitter redirect
279
- "enableFxTwitter": true,
280
- // Undetected-chromedriver — optional Python Selenium fallback for tough anti-bot pages
281
- "enableUndetectedChromedriver": true,
282
- "undetectedChromedriverPython": "python3",
283
- // Reddit personal retrieval chain (legacy key name retained for compatibility)
284
- "enableRedditJson": true,
285
- // YouTube video transcript enrichment for public captions (default: true)
286
- "enableYoutubeTranscript": true,
287
- // Archive fallback — Wayback and archive.md recovery
288
- "enableArchiveFallback": true,
289
- "enableWayback": true,
290
- "enableArchiveMd": true,
291
- // Optional local adaptive scraper fallback. This is disabled by default;
292
- // install Scrapling separately and opt in only on deployments that want an
293
- // extra no-secret local shell/bot-challenge extraction attempt before the
294
- // built-in Playwright fallback.
295
- "enableAdaptiveScraper": false,
296
- "adaptiveScraperCommand": "scrapling",
297
- "adaptiveScraperTimeoutMs": 20000,
298
- // Output projection budget for smartfetch responses. MCP, CLI, and library
299
- // calls honor this setting unless the caller passes a per-call budget control.
300
- // With no setting or per-call control, MCP calls default to compact.
301
- // preset: "compact" | "balanced" | "full"; numeric fields override the preset.
302
- "outputBudget": {
303
- "preset": "balanced",
304
- "maxPostTextChars": 8000,
305
- "maxPostMarkdownChars": 10000,
306
- "maxThreadItems": 10,
307
- "maxCommentItems": 10,
308
- "maxOutboundLinks": 60,
309
- "maxAssets": 15,
310
- "maxBlocks": 20,
311
- "maxBlockTextChars": 500,
312
- "maxItemTextChars": 500
313
- },
314
- // Optional default named browser profile for authenticated smartfetch calls.
315
- // Names are 1-64 characters: letters, numbers, dots, underscores, or hyphens;
316
- // they must start with a letter or number and cannot end with a dot.
317
- "browserProfile": ""
318
- },
319
-
320
- "network": {
321
- // Allow localhost/private/reserved fetch targets (default: false)
322
- // Applies to smartfetch, smartcrawl, and Korean public API helpers.
323
- "allowPrivateHosts": false
324
- }
325
- }
326
- ```
327
-
328
- ### Common setups
329
-
330
- #### Default — no config file needed
331
-
332
- Out of the box `smart-web` works with sensible defaults. Create a settings file only when you need to change something.
333
-
334
- #### Minimal: profile only
335
-
336
- ```json
337
- {
338
- "profile": "balanced"
339
- }
340
- ```
341
-
342
- #### Private / local-first
343
-
344
- ```json
345
- {
346
- "profile": "private",
347
- "search": {
348
- "searxngBaseUrl": "http://localhost:8080"
349
- }
350
- }
351
- ```
352
-
353
- In `private` mode, relay-style provider requests are blocked, public no-key search fallbacks are disabled, and SearXNG bases must resolve to localhost or private addresses. Hosted API-key search providers such as Exa, Tavily, and Brave Search are intentionally not part of the default smartsearch runtime; use local/site-native search, DuckDuckGo/Brave HTML fallbacks, or self-hosted SearXNG instead.
354
-
355
- #### Disable Playwright auto-install
356
-
357
- ```json
358
- {
359
- "fetch": {
360
- "autoInstallPlaywright": false
361
- }
362
- }
363
- ```
364
-
365
- The server will return an actionable `playwright_browsers_missing` or `playwright_browsers_outdated` error instead.
366
-
367
- #### Opt in to Jina Reader for weak generic article pages
368
-
369
- ```json
370
- {
371
- "fetch": {
372
- "enableJinaReader": true
373
- }
374
- }
375
- ```
376
-
377
- This stays opt-in because it is a third-party relay path. When enabled, it only applies to weak generic/article direct fetches and prefers Jina JSON mode when that improves the page. `profile: "private"` hard-disables relay-style providers.
378
-
379
- Relay eligibility is provider-policy based. Generic and document-like providers can opt in, while browser/auth/social/commerce providers stay off unless their provider declares support. `pipeline.policy.relay_style_allowed` and `pipeline.policy.relay_style_used` tell hosts whether a relay path was allowed and whether it was used.
380
-
381
- #### Opt in to a local adaptive scraper command
382
-
383
- ```json
384
- {
385
- "fetch": {
386
- "enableAdaptiveScraper": true,
387
- "adaptiveScraperCommand": "scrapling",
388
- "adaptiveScraperTimeoutMs": 20000
389
- }
390
- }
391
- ```
392
-
393
- This lane is for local, no-secret deployments that install Scrapling or a compatible CLI themselves. It runs only after direct retrieval still looks weak, blocked, or shell-like and before the built-in browser fallback. The attempt is recorded as `adaptive_scraper` in `pipeline.acquire.attempt_map`; if the command is missing, times out, returns empty text, or does not improve the active result, smartfetch keeps the existing result and continues through the normal fallback chain.
394
-
395
- #### Tune smartfetch output budgets
31
+ MCP host configuration:
396
32
 
397
33
  ```json
398
34
  {
399
- "fetch": {
400
- "outputBudget": {
401
- "maxPostTextChars": 16000,
402
- "maxThreadItems": 25,
403
- "maxOutboundLinks": 80
404
- }
35
+ "mcpServers": {
36
+ "smart-web": { "command": "npx", "args": ["-y", "smart-web-mcp"] }
405
37
  }
406
38
  }
407
39
  ```
408
40
 
409
- Extraction runs before budgeting so providers can work from complete page content. The budget applies to returned `smartfetch` projections and records truncation under `budget.fields`.
41
+ Use `--init-settings` to create the default settings file, or `--print-settings-example` to inspect it. The published npm package can lag source; verify the runtime you actually use.
410
42
 
411
- #### Force archive fallback off
412
-
413
- ```json
414
- {
415
- "fetch": {
416
- "enableArchiveFallback": false
417
- }
418
- }
419
- ```
420
-
421
- #### Legal OA DOI recovery for paywalled paper pages
422
-
423
- ```json
424
- {
425
- "fetch": {
426
- "enableUnpaywall": true,
427
- "unpaywallEmail": "research@example.com",
428
- "enableSemanticScholar": true,
429
- "enableCoreDiscovery": true
430
- }
431
- }
432
- ```
433
-
434
- #### Enable undetected-chromedriver from a dedicated virtualenv
435
-
436
- ```json
437
- {
438
- "fetch": {
439
- "enableUndetectedChromedriver": true,
440
- "undetectedChromedriverPython": "/absolute/path/to/venv/bin/python"
441
- }
442
- }
443
- ```
444
-
445
- This helper is optional. `smart-web` still works without it, but when installed and reachable it can act as a second browser engine for stubborn pages where Playwright alone is not enough.
446
-
447
- #### Custom temp directory
448
-
449
- ```json
450
- {
451
- "runtime": {
452
- "tempDir": "/absolute/path/to/smart-web-tmp"
453
- }
454
- }
455
- ```
456
-
457
- ### Advanced overrides
458
-
459
- These keys live inside `settings.json` when you need to force one provider on or off:
460
-
461
- **Search**
462
-
463
- - `search.enableSearxng`
464
- - `search.enableDuckDuckGo`
465
- - `search.enableBraveHtml`
466
-
467
- **Fetch**
468
-
469
- - `fetch.enableFxTwitter`
470
- - `fetch.enableXOembed`
471
- - `fetch.autoInstallPlaywright`
472
- - `fetch.enableJinaReader`
473
- - `fetch.jinaReaderBaseUrl`
474
- - `fetch.enableAcademicFallback`
475
- - `fetch.enableOpenAlex`
476
- - `fetch.enableEuropePmc`
477
- - `fetch.enableBiorxivApi`
478
- - `fetch.enableUnpaywall`
479
- - `fetch.unpaywallEmail`
480
- - `fetch.enableSemanticScholar`
481
- - `fetch.semanticScholarApiKey`
482
- - `fetch.enableCoreDiscovery`
483
- - `fetch.coreApiKey`
484
- - `fetch.enableUndetectedChromedriver`
485
- - `fetch.undetectedChromedriverPython`
486
- - `fetch.enableRedditJson`
487
- - `fetch.enableYoutubeTranscript`
488
- - `fetch.enableArchiveFallback`
489
- - `fetch.enableWayback`
490
- - `fetch.enableArchiveMd`
491
- - `fetch.outputBudget.maxPostTextChars`
492
- - `fetch.outputBudget.maxPostMarkdownChars`
493
- - `fetch.outputBudget.maxThreadItems`
494
- - `fetch.outputBudget.maxCommentItems`
495
- - `fetch.outputBudget.maxOutboundLinks`
496
- - `fetch.outputBudget.maxAssets`
497
- - `fetch.outputBudget.maxBlocks`
498
- - `fetch.outputBudget.maxBlockTextChars`
499
- - `fetch.outputBudget.maxItemTextChars`
500
- - `fetch.outputBudget.maxEnvelopeArrayItems`
501
- - `fetch.outputBudget.maxAssessmentSignals`
502
- - `fetch.outputBudget.maxTotalEnvelopeEntries`
503
- - `fetch.outputBudget.maxTotalEnvelopeStringChars`
504
-
505
- **Compatibility**
506
-
507
- - `network.localOnly`: advanced override that forces local-only behavior regardless of profile
508
-
509
- ## Provider benchmark
510
-
511
- Use the provider benchmark when deciding whether to build, wrap, or replace retrieval lanes:
512
-
513
- ```bash
514
- npm run benchmark:providers
515
- npm run benchmark:providers:measure
516
- npm run benchmark:providers:compare
517
- npm run --silent benchmark:providers:json
518
- npm run benchmark:providers:json-file
519
- npm run benchmark:providers:report
520
- npm run benchmark:providers:gaps
521
- npm run benchmark:providers:live
522
- npm run benchmark:providers:live:jina
523
- npm run benchmark:providers:live:cp
524
- npm run benchmark:providers:live:browser
525
- npm run benchmark:providers:live:youtube
526
- npm run benchmark:providers:live:blocked
527
- npm run benchmark:providers:live:naver-blog
528
- npm run benchmark:providers:live:naver-map
529
- npm run benchmark:providers:live:naver-cafe
530
- npm run benchmark:providers:live:x-twitter
531
- npm run benchmark:providers:live:no-secret
532
- npm run benchmark:providers:live:no-secret:gaps
533
- npm run watchdog:no-secret:live
534
- npm run benchmark:providers:trend
535
- npm run benchmark:smartsearch
536
- npm run --silent benchmark:smartsearch:json
537
- npm run watchdog:no-secret
538
- npm run watchdog:partial-handoff
539
- ```
540
-
541
- The default report is availability-only and safe for CI. `--run-measurements` (`npm run benchmark:providers:measure`) adds deterministic fixture records for available local/core and optional no-key lanes: success, useful character counts, projected/text character counts, useful-to-projected and useful-to-text yield ratios, latency, extractor/provider id, confidence, partial status, error class, lane label/family/category metadata, provider-family/category lane-status rollups, surface coverage, and recommendation sections grouped by lane, case, and surface.
542
-
543
- Use `--compare-external` (`npm run benchmark:providers:compare`) to add an `externalComparisons` section that compares current smart-web fixture evidence against optional no-secret/local candidates. The comparison includes the local Crawl4AI command lane, the public Jina Reader relay fixture, and available direct/Playwright baselines. If `crawl4ai` is not installed, the Crawl4AI row is `skipped` with `action: "install-local-prerequisite"` and `prerequisite.command: "crawl4ai"`; this is an expected setup signal, not a benchmark failure or required dependency.
544
-
545
- Reports carry the `provider-benchmark.v1` schema version plus non-fatal git and runtime metadata so archived JSON, Markdown, and NDJSON trend artifacts are parseable, traceable, and comparable as the harness evolves. Markdown, text, JSON, and NDJSON outputs include aggregate fixture yield averages, compact yield rollups by provider family/category/lane, and a best-current-evidence surface table that ranks measured/skipped candidates so review jobs can compare token efficiency and build-vs-wrap posture between core, browser, relay, and specialist paths without reprocessing the full measurement table.
546
-
547
- Measurement artifacts also include a live-readiness manifest listing each lane's env/command/settings prerequisites, a deterministic fixture smoke, a safe MCP live-smoke candidate, and surfaces that still lack live evidence. `--run-live-smokes` (`npm run benchmark:providers:live`) is an explicit gated mode that runs only no-secret, core-surface live checks for static articles, long docs/Wikipedia, Hacker News, and docs export; optional no-secret lanes remain gated until explicitly allowlisted: `--allow-jina-live-smoke` (`npm run benchmark:providers:live:jina`) exercises the public Jina Reader relay with a secret-stripped settings file, `--allow-competitive-programming-live-smoke` (`npm run benchmark:providers:live:cp`) adds a public BOJ/Codeforces representative smoke for the competitive-programming specialist surface, `--allow-browser-live-smoke` (`npm run benchmark:providers:live:browser`) runs a Playwright `--force-dynamic` smoke for the JS-rendered article surface only when the local browser lane is available, `--allow-youtube-live-smoke` (`npm run benchmark:providers:live:youtube`) adds the public YouTube seed-lane representative smoke, `--allow-blocked-authwall-live-smoke` (`npm run benchmark:providers:live:blocked`) adds a deterministic HTTP Basic Auth challenge smoke that treats honest blocked/error/handoff signals as positive calibration without private credentials, `--allow-naver-blog-live-smoke` (`npm run benchmark:providers:live:naver-blog`) adds a stable public Naver Blog representative smoke for Korean specialist-lane calibration, `--allow-naver-map-live-smoke` (`npm run benchmark:providers:live:naver-map`) adds a stable public Naver Map place smoke with browser-preflight diagnostics, `--allow-naver-cafe-live-smoke` (`npm run benchmark:providers:live:naver-cafe`) adds a stable public Naver Cafe shared-link smoke that treats useful no-credential partial output as honest calibration, and `--allow-x-twitter-live-smoke` (`npm run benchmark:providers:live:x-twitter`) adds a public X/Twitter post smoke that keeps no-secret social-provider drift visible without private cookies. `--allow-all-no-secret-live-smokes` (`npm run benchmark:providers:live:no-secret`) enables every no-secret gate in one run; hosted API-key search lanes are excluded from the benchmark surface by product policy.
548
-
549
- Live smoke artifacts record per-case `present`/`missing`/`skipped` evidence next to fixture evidence without printing secret values. When fixture rows exist but all are skipped by missing local commands, comparison and gap artifacts report fixture evidence as `skipped` with a `fixture-skipped-*` comparison instead of implying the fixture rows are absent. Browser-lane live smoke rows also include a secret-safe Playwright preflight diagnostic with executable path, launch timing, `domcontentloaded` navigation timing, HTTP status/final URL, and truncated failure reason so runtime/cache/network failures can be separated from smartfetch normalization failures.
550
-
551
- The same JSON, text, Markdown, and NDJSON outputs include a compact live-vs-fixture comparison rollup so replacement reviews can distinguish calibrated surfaces from fixture-only, skipped-live, live-disagreeing, and fixture-skipped-by-prerequisite surfaces. Secret-safe provider diagnostics list prerequisite names with satisfied/missing status and never print env values. Use `npm run benchmark:providers:json-file` or add `--json-file <path>` to write the full JSON artifact without relying on stdout capture. Use `npm run benchmark:providers:report` or add `--report-file <path>` with measurement mode to write a Markdown review artifact containing lane availability, diagnostics, surface/corpus coverage, measurement highlights, yield comparisons, per-surface evidence, live readiness, live-smoke diagnostics, and build-vs-wrap recommendations without changing stdout or JSON behavior. Use `npm run benchmark:providers:gaps` or add `--gap-file <path>` with measurement mode to write a compact `provider-benchmark-gaps.v1` JSON artifact containing only the ranked provider evidence gap queue, summary counts, git/runtime metadata, and safe smoke candidates. Use `npm run benchmark:providers:trend` or add `--trend-file <path>` with measurement mode to append compact NDJSON trend records with aggregate lane, prerequisite, measurement, recommendation, surface-evidence, and fixture-yield counts. Unknown `--flag` options fail fast so typoed benchmark jobs do not produce misleading artifacts. The corpus covers static articles, JS-rendered articles, blocked/authwall-like pages, long docs/Wikipedia pages, Naver Blog/Cafe/Map, Hacker News, YouTube, X/Twitter, BOJ/Codeforces, relay, and docs export surfaces. Optional local/no-secret lanes without commands remain explicit skips instead of silently failing; API-key provider candidates are excluded from the default corpus and reports by product policy.
552
-
553
- Provider benchmark reports also include `providerEvidenceGaps`: a ranked next-action queue built from the live-vs-fixture rollup. Each gap records the surface, priority score, selected and candidate lanes, follow-up action, unlock prerequisites, fixture/live useful chars, fixture skipped-case lanes/reasons, live-smoke failure error classes/reasons, and safe live smoke candidate so replacement reviews can pick the next provider experiment without scanning every measurement row or re-running a timed live calibration. Use `npm run benchmark:providers:live:no-secret:gaps` to refresh only the compact all-no-secret gap artifact at `reports/provider-benchmark-live-no-secret-gaps.json`. When fixture rows exist but were skipped by missing local commands, the live-vs-fixture and gap actions are `unlock-skipped-fixture-prerequisites` rather than fixture-creation actions.
554
-
555
- `npm run benchmark:smartsearch` is the deterministic quality gate for site-native smartsearch routing. It patches `fetch`, disables generic web-search fallbacks with a temporary settings file, and exercises representative `site:` queries for Reddit, GitHub repo discovery, GitHub repo docs, npm, Hacker News, Stack Exchange, Wikipedia, Velog, MDN, and crates.io without live network access. The human report is compact; `npm run --silent benchmark:smartsearch:json` emits `smartsearch-quality-benchmark.v1` JSON with expected engine, top URL, URL-substring checks, actual engine, top URL, notes, and pass/fail status per case.
556
-
557
- See `reports/README.md` for the curated report archive index. It identifies `provider-benchmark-live-no-secret.{json,md}` and `provider-benchmark-live-no-secret-gaps.json` as the current maintenance artifacts, and labels older cumulative live-smoke reports as historical calibration snapshots.
558
-
559
- `npm run watchdog:no-secret` is the deterministic current-surface gate for no-secret maintenance. It fails when `reports/provider-benchmark-live-no-secret-gaps.json` contains current gaps, when runtime/docs/examples/scripts reintroduce hosted API-key provider markers for Exa, Tavily, Firecrawl, Browserless, or Brave Search API, or when the partial/handoff watchdog no longer has a compact text ceiling. It intentionally ignores historical `CHANGELOG.md` entries and archived `reports/` output, and does not run live network smokes.
560
-
561
- `npm run watchdog:no-secret:live` is the live regression gate for no-secret maintenance. It first refreshes `reports/provider-benchmark-live-no-secret-gaps.json` with `npm run benchmark:providers:live:no-secret:gaps`, then runs `npm run watchdog:no-secret` so non-empty live provider gaps fail locally or in CI. The scheduled/manual GitHub Actions workflow `.github/workflows/no-secret-live-watchdog.yml` runs the same alias weekly with read-only permissions, no secrets, and a short-retention gap artifact.
562
-
563
- `npm run watchdog:partial-handoff` builds the package and runs a compact no-secret drift watchdog over representative partial/blocked/handoff `smartfetch` outputs (httpbin authwall, LinkedIn profile authwall, Naver Blog unavailable/deleted posts, Threads unavailable/join-login output, DCInside unavailable/deleted posts, Telegram missing/private posts, a non-existent Wikidocs article, Naver Cafe, Naver Map search, and BOJ unavailable/shutdown output). It exits non-zero when expected handoff cases are silently reclassified as non-partial, when partial/handoff `content[0].text` grows beyond the compact watchdog ceiling, or when compact text or structured links/assets leak chrome URLs, script globals, or style-shell tokens, and prints a `partial-handoff-watchdog.v1` JSON report without raw page dumps or credentials. For runner artifacts without JSON stdout, run `node scripts/partial-handoff-watchdog.mjs --report-file <path>` after `npm run build`; the command keeps stdout human-readable and writes the full JSON report to the requested path. Fixture cases can set `expectPartialHandoff: true` to make classifier drift fail even if the output would otherwise be skipped.
564
-
565
- ## Quick verification
566
-
567
- Sanity-check the server with a few real calls:
568
-
569
- - `smartfetch` on a normal article URL
570
- - `smartfetch` on a Medium or other member-only article URL — confirm it returns either archive-backed content or an honest `reference_only` preview
571
- - `smartfetch` on a Naver Map, Kakao Map, or `naver.me` place-share URL
572
- - `smartfetch` on an arXiv abstract URL
573
- - `smartfetch` on a PubMed, PMC, or bioRxiv/medRxiv paper URL
574
- - `smartfetch` on a known product URL
575
- - `smartsearch` on a `site:` query
576
- - `smartcrawl` on a docs site or board you actually use
577
-
578
- ## Reporting issues
579
-
580
- If `smart-web` returns reproducible incorrect or misleading output, open a GitHub issue with:
581
-
582
- - exact input and tool arguments
583
- - observed output
584
- - expected behavior
585
- - `smart-web-mcp` version
586
-
587
- Skip obvious transient network, auth, or rate-limit failures unless the classification itself looks wrong.
43
+ ## Tool routing
588
44
 
589
- ## License
45
+ - `smartfetch` first when the user already gave a URL or shortlink.
46
+ - `smartsearch` when the user gave a topic, keywords, or a `site:` query.
47
+ - `smartcrawl` when you already know the site and need multiple relevant pages.
48
+ - `smartvertical` for Korean/local structured lookups.
49
+ - `smartbrowser` for explicit approved-browser input/click/filter/wait/screenshot batches; see [action scope and limits](https://github.com/jojo-labs/smart-web/blob/main/docs/operator-actions.md).
590
50
 
591
- All rights reserved. See [LICENSE](LICENSE) for terms.
51
+ Use compact output first; request a larger budget when needed. Partial output is honest evidence, **not task completion**. Never substitute a search snippet, loading screen, or unrelated listing for the requested source.
52
+ The CLI `smart-web fetch <url> --format json` (or `--json`) prints structured output; `--budget full` controls bounded delivery, not source completeness. Invalid budget names fail before retrieval rather than silently using compact output.
53
+ For CLI search, `--max-chars` caps rendered text; use `--json` when the complete projected structured result is needed.
592
54
 
593
- This software is licensed under a proprietary license that permits installation and use as a local MCP server but prohibits modification, redistribution, reverse engineering, or use in competing products.
55
+ ## Documentation
594
56
 
595
- ## References
57
+ - [Product contract and acceptance](https://github.com/jojo-labs/smart-web/blob/main/docs/operator-contract.md)
58
+ - [Host setup](https://github.com/jojo-labs/smart-web/blob/main/docs/host-setup.md) · [Configuration](https://github.com/jojo-labs/smart-web/blob/main/docs/configuration.md)
59
+ - [Architecture and code map](https://github.com/jojo-labs/smart-web/blob/main/docs/architecture.md) · [Verification and maintenance](https://github.com/jojo-labs/smart-web/blob/main/docs/benchmarks.md)
60
+ - [All docs](https://github.com/jojo-labs/smart-web/blob/main/docs/README.md) · [Agent instructions](https://github.com/jojo-labs/smart-web/blob/main/AGENTS.md)
596
61
 
597
- [^1]: Model Context Protocol, "Tools" specification — MCP tools are a retrieval and integration surface, not a requirement to emulate full browser-automation flows inside one tool.
62
+ Report reproducible problems in [GitHub Issues](https://github.com/jojo-labs/smart-web/issues), with sanitized input, expected/observed behavior, and runtime version. Never attach credentials or private browsing content.
598
63
 
599
- [^2]: Model Context Protocol Blog, "Tool Annotations as Risk Vocabulary: What Hints Can and Can't Do" (2026-03-16) — structured tool output and accurate hints improve host-side routing and escalation decisions.
64
+ [Package](https://www.npmjs.com/package/smart-web-mcp) · MCP Registry: `io.github.jojo-labs/smart-web` · [License](https://github.com/jojo-labs/smart-web/blob/main/LICENSE) (proprietary; all rights reserved).