smart-web-mcp 0.44.1 → 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +350 -0
- package/README.md +33 -568
- package/dist/assessment.d.ts +10 -3
- package/dist/assessment.js +21 -6
- package/dist/assessment.js.map +1 -1
- package/dist/body-export.d.ts +10 -0
- package/dist/body-export.js +28 -0
- package/dist/body-export.js.map +1 -0
- package/dist/browser-session.d.ts +32 -1
- package/dist/browser-session.js +373 -3
- package/dist/browser-session.js.map +1 -1
- package/dist/cli-fetch-options.d.ts +7 -0
- package/dist/cli-fetch-options.js +41 -0
- package/dist/cli-fetch-options.js.map +1 -0
- package/dist/cli.js +28 -11
- package/dist/cli.js.map +1 -1
- package/dist/comment-export.d.ts +10 -0
- package/dist/comment-export.js +39 -0
- package/dist/comment-export.js.map +1 -0
- package/dist/composition.d.ts +1 -22
- package/dist/composition.js +0 -16
- package/dist/composition.js.map +1 -1
- package/dist/extraction/budget.d.ts +3 -0
- package/dist/extraction/budget.js +88 -7
- package/dist/extraction/budget.js.map +1 -1
- package/dist/extraction/document.js +134 -9
- package/dist/extraction/document.js.map +1 -1
- package/dist/korea/coupang-filters.d.ts +5 -0
- package/dist/korea/coupang-filters.js +87 -0
- package/dist/korea/coupang-filters.js.map +1 -0
- package/dist/korea/coupang.d.ts +27 -6
- package/dist/korea/coupang.js +202 -82
- package/dist/korea/coupang.js.map +1 -1
- package/dist/korea/proxy-client.d.ts +1 -0
- package/dist/korea/proxy-client.js +6 -2
- package/dist/korea/proxy-client.js.map +1 -1
- package/dist/korea/smartvertical.d.ts +1 -0
- package/dist/korea/smartvertical.js +39 -25
- package/dist/korea/smartvertical.js.map +1 -1
- package/dist/lib.d.ts +4 -2
- package/dist/lib.js +2 -1
- package/dist/lib.js.map +1 -1
- package/dist/mcp-server.js +30 -15
- package/dist/mcp-server.js.map +1 -1
- package/dist/operator-browser-fixture.d.ts +7 -0
- package/dist/operator-browser-fixture.js +19 -0
- package/dist/operator-browser-fixture.js.map +1 -0
- package/dist/operator-canary-cli.d.ts +6 -0
- package/dist/operator-canary-cli.js +190 -0
- package/dist/operator-canary-cli.js.map +1 -0
- package/dist/operator-canary.d.ts +74 -0
- package/dist/operator-canary.js +135 -0
- package/dist/operator-canary.js.map +1 -0
- package/dist/operator-cdp.d.ts +15 -0
- package/dist/operator-cdp.js +195 -0
- package/dist/operator-cdp.js.map +1 -0
- package/dist/operator-frame.d.ts +3 -0
- package/dist/operator-frame.js +55 -0
- package/dist/operator-frame.js.map +1 -0
- package/dist/settings.d.ts +3 -0
- package/dist/settings.js +1 -0
- package/dist/settings.js.map +1 -1
- package/dist/shared.d.ts +33 -8
- package/dist/shared.js +15 -9
- package/dist/shared.js.map +1 -1
- package/dist/smartbrowser.d.ts +169 -0
- package/dist/smartbrowser.js +288 -0
- package/dist/smartbrowser.js.map +1 -0
- package/dist/smartcrawl.js +16 -21
- package/dist/smartcrawl.js.map +1 -1
- package/dist/smartfetch/academic-fallback.js +3 -5
- package/dist/smartfetch/academic-fallback.js.map +1 -1
- package/dist/smartfetch/archive-fallback.js +2 -5
- package/dist/smartfetch/archive-fallback.js.map +1 -1
- package/dist/smartfetch/assets.d.ts +1 -1
- package/dist/smartfetch/assets.js +28 -1
- package/dist/smartfetch/assets.js.map +1 -1
- package/dist/smartfetch/jina-reader.d.ts +1 -0
- package/dist/smartfetch/jina-reader.js +17 -3
- package/dist/smartfetch/jina-reader.js.map +1 -1
- package/dist/smartfetch/pipeline.js +15 -25
- package/dist/smartfetch/pipeline.js.map +1 -1
- package/dist/smartfetch/provider-policy.d.ts +3 -0
- package/dist/smartfetch/provider-policy.js +45 -2
- package/dist/smartfetch/provider-policy.js.map +1 -1
- package/dist/smartfetch/provider-types.d.ts +8 -0
- package/dist/smartfetch/providers/article.js +76 -10
- package/dist/smartfetch/providers/article.js.map +1 -1
- package/dist/smartfetch/providers/blind-post.d.ts +2 -0
- package/dist/smartfetch/providers/blind-post.js +167 -0
- package/dist/smartfetch/providers/blind-post.js.map +1 -0
- package/dist/smartfetch/providers/clien.d.ts +2 -0
- package/dist/smartfetch/providers/clien.js +157 -0
- package/dist/smartfetch/providers/clien.js.map +1 -0
- package/dist/smartfetch/providers/commerce.d.ts +15 -0
- package/dist/smartfetch/providers/commerce.js +123 -7
- package/dist/smartfetch/providers/commerce.js.map +1 -1
- package/dist/smartfetch/providers/dcinside.js +313 -20
- package/dist/smartfetch/providers/dcinside.js.map +1 -1
- package/dist/smartfetch/providers/github-issue.d.ts +3 -0
- package/dist/smartfetch/providers/github-issue.js +101 -0
- package/dist/smartfetch/providers/github-issue.js.map +1 -0
- package/dist/smartfetch/providers/hackernews-html.d.ts +21 -0
- package/dist/smartfetch/providers/hackernews-html.js +148 -0
- package/dist/smartfetch/providers/hackernews-html.js.map +1 -0
- package/dist/smartfetch/providers/hackernews.js +76 -14
- package/dist/smartfetch/providers/hackernews.js.map +1 -1
- package/dist/smartfetch/providers/index.js +9 -1
- package/dist/smartfetch/providers/index.js.map +1 -1
- package/dist/smartfetch/providers/linkedin.js +2 -1
- package/dist/smartfetch/providers/linkedin.js.map +1 -1
- package/dist/smartfetch/providers/naver-blog.d.ts +2 -0
- package/dist/smartfetch/providers/naver-blog.js +87 -45
- package/dist/smartfetch/providers/naver-blog.js.map +1 -1
- package/dist/smartfetch/providers/naver-cafe-native.d.ts +7 -0
- package/dist/smartfetch/providers/naver-cafe-native.js +188 -0
- package/dist/smartfetch/providers/naver-cafe-native.js.map +1 -0
- package/dist/smartfetch/providers/naver-cafe.js +165 -4
- package/dist/smartfetch/providers/naver-cafe.js.map +1 -1
- package/dist/smartfetch/providers/naver-map-search.js +1 -1
- package/dist/smartfetch/providers/naver-map-search.js.map +1 -1
- package/dist/smartfetch/providers/ppomppu.d.ts +2 -0
- package/dist/smartfetch/providers/ppomppu.js +289 -0
- package/dist/smartfetch/providers/ppomppu.js.map +1 -0
- package/dist/smartfetch/providers/reddit.js +24 -3
- package/dist/smartfetch/providers/reddit.js.map +1 -1
- package/dist/smartfetch/providers/threads.js +1 -1
- package/dist/smartfetch/providers/threads.js.map +1 -1
- package/dist/smartfetch/providers/wikipedia.js +138 -0
- package/dist/smartfetch/providers/wikipedia.js.map +1 -1
- package/dist/smartfetch.d.ts +7 -1
- package/dist/smartfetch.js +353 -76
- package/dist/smartfetch.js.map +1 -1
- package/dist/smartsearch-brave.d.ts +2 -0
- package/dist/smartsearch-brave.js +43 -0
- package/dist/smartsearch-brave.js.map +1 -0
- package/dist/smartsearch-browser.d.ts +5 -0
- package/dist/smartsearch-browser.js +116 -0
- package/dist/smartsearch-browser.js.map +1 -0
- package/dist/smartsearch-ddg.d.ts +2 -0
- package/dist/smartsearch-ddg.js +46 -0
- package/dist/smartsearch-ddg.js.map +1 -0
- package/dist/smartsearch-node.d.ts +6 -0
- package/dist/smartsearch-node.js +21 -0
- package/dist/smartsearch-node.js.map +1 -0
- package/dist/smartsearch-postgresql.d.ts +7 -0
- package/dist/smartsearch-postgresql.js +44 -0
- package/dist/smartsearch-postgresql.js.map +1 -0
- package/dist/smartsearch-python.d.ts +8 -0
- package/dist/smartsearch-python.js +87 -0
- package/dist/smartsearch-python.js.map +1 -0
- package/dist/smartsearch-quality.d.ts +6 -0
- package/dist/smartsearch-quality.js +47 -0
- package/dist/smartsearch-quality.js.map +1 -0
- package/dist/smartsearch-youtube.d.ts +7 -0
- package/dist/smartsearch-youtube.js +76 -0
- package/dist/smartsearch-youtube.js.map +1 -0
- package/dist/smartsearch.d.ts +2 -0
- package/dist/smartsearch.js +506 -84
- package/dist/smartsearch.js.map +1 -1
- package/package.json +8 -5
- package/python/requirements-undetected-chromedriver.txt +2 -2
package/README.md
CHANGED
|
@@ -2,598 +2,63 @@
|
|
|
2
2
|
|
|
3
3
|
# smart-web
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
One local MCP entry point for finding, reading, and using the web.
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
**Goal:** finish the user's web task—not merely return a successful request or a browser handoff. Use fast retrieval where it works; continue through the approved local browser-use session when it does not.
|
|
8
8
|
|
|
9
|
-
|
|
10
|
-
- `smartsearch` for discovery
|
|
11
|
-
- `smartcrawl` for same-site multi-page traversal
|
|
12
|
-
- `smartvertical` for Korean/local structured vertical data such as places, weather, transit, finance, law, commerce, and safety lookups
|
|
9
|
+
## Principles
|
|
13
10
|
|
|
14
|
-
|
|
11
|
+
- Judge success by the requested content or action, including comments and exact product identity.
|
|
12
|
+
- Measure search relevance, freshness, query constraints, and usable sources—not result count.
|
|
13
|
+
- Prioritize failures from real workflows across sites, not a single showcase domain.
|
|
14
|
+
- Keep output compact, with source evidence, attempt history, and explicit omissions.
|
|
15
|
+
- Isolate fragile site logic; detect drift and repair it through bounded, verified agent runs.
|
|
15
16
|
|
|
16
|
-
|
|
17
|
+
## Status
|
|
17
18
|
|
|
18
|
-
-
|
|
19
|
-
- MCP Registry: `io.github.jojo-labs/smart-web`
|
|
20
|
-
- Issues: <https://github.com/jojo-labs/smart-web/issues>
|
|
19
|
+
The package exposes four retrieval tools plus the bounded `smartbrowser` action tool. The **operator-first revamp is in progress**: broader interactive coverage and automatic repair remain acceptance goals, not a claim of universal support. The legacy web-task-api backend is retired; browser actions use the approved local session. Cookies persist across batches, but tab state only persists within a batch.
|
|
21
20
|
|
|
22
|
-
|
|
21
|
+
Track shipped evidence and remaining work in [the revamp issue](https://github.com/jojo-labs/smart-web/issues/496). Browser-readable content should not require the user to switch tools; genuine login, permission, or unavailable-content limits remain explicit.
|
|
23
22
|
|
|
24
|
-
##
|
|
23
|
+
## Quick start
|
|
25
24
|
|
|
26
|
-
|
|
27
|
-
- explicit routing: URL → `smartfetch`, query → `smartsearch`, known site → `smartcrawl`
|
|
28
|
-
- up-front smartfetch acquisition lanes instead of broad direct → browser retry chains, with `pipeline.acquire.attempt_map` showing seed/direct/relay/archive/browser phases and skip/failure reasons
|
|
29
|
-
- optional Jina Reader relay for weak generic/article fetches when you explicitly enable it, with JSON mode and alternate-link preservation
|
|
30
|
-
- shared document extraction for article-like pages keeps browser primary content separate from raw shell HTML, preserving headings, links, blocks, and Markdown when requested
|
|
31
|
-
- docs export manifests are versioned and include per-document SHA-256 hashes plus resource-handle manifests for stable local Markdown mirrors; resource-handle manifests include both exported documents and metadata handles for the summary/manifest files, and successful export `structuredContent` embeds the same handle manifest for host follow-up calls
|
|
32
|
-
- budgeted `smartfetch` projection records explicit truncation metadata instead of scattering silent hardcoded caps; MCP calls default to the compact preset for token-sensitive agent hosts
|
|
33
|
-
- weak generic pages can recover through JSON-LD or Next.js payload extraction and same-origin RSS/Atom feed discovery before falling back to a handoff
|
|
34
|
-
- optional local adaptive scraper fallback can call a Scrapling-compatible `scrapling` command for AI-targeted extraction on weak JavaScript shells before spending a full Playwright browser pass
|
|
35
|
-
- paywalled article pages prefer lawful archive recovery and search-indexed reference previews before telling the host to escalate elsewhere
|
|
36
|
-
- LinkedIn authwall handling prefers compliant public fallbacks such as archives and search-indexed reference metadata before giving up
|
|
37
|
-
- legal paper fallback for academic URLs: DOI-aware OpenAlex, Unpaywall, Semantic Scholar, CORE discovery, and Europe PMC enrichment plus bioRxiv/medRxiv API fallback when a direct paper page is thin or blocked
|
|
38
|
-
- site-native public search fallbacks cover Reddit (Arctic Shift, live RSS, PullPush, then JSON), GitHub repository discovery/repo-docs queries, npm search/exact packages, Hacker News, Stack Exchange, Wikipedia, Velog, MDN, and crates.io when a matching `site:` query is available
|
|
39
|
-
- YouTube watch, shorts, embed, live, and youtu.be video URLs get best-effort public caption enrichment by default while non-video YouTube pages stay on the generic fetch path
|
|
40
|
-
- structured output with `assessment`, `pipeline`, and partial-result `research_handoff` fields for host-side decision making
|
|
41
|
-
- local-first defaults with privacy-aware search/fetch behavior
|
|
42
|
-
- useful normalization for common public surfaces instead of raw HTML dumps
|
|
43
|
-
- social unavailable shells such as Threads generic join/login pages stay partial and chrome-free instead of masquerading as verified post content
|
|
44
|
-
- blocked commerce URLs can accept up to five contextual `reference_queries`; public catalog metadata is returned only when the resolved product, item, and vendor identifiers all match, so a title hint cannot silently select a different listing
|
|
45
|
-
|
|
46
|
-
## Tool routing
|
|
47
|
-
|
|
48
|
-
- `smartfetch` first when the user already gave a URL or shortlink
|
|
49
|
-
- `smartsearch` when the user gave a topic, keywords, or a `site:` query
|
|
50
|
-
- `smartcrawl` when you already know the site and need multiple relevant pages
|
|
51
|
-
|
|
52
|
-
| Situation | Tool |
|
|
53
|
-
| ------------------------------------------------- | --------------- |
|
|
54
|
-
| User already gave a URL or shortlink | `smartfetch` |
|
|
55
|
-
| User gave a topic, keywords, or `site:` query | `smartsearch` |
|
|
56
|
-
| You already know the site and need multiple pages | `smartcrawl` |
|
|
57
|
-
| User needs Korean/local structured data | `smartvertical` |
|
|
58
|
-
|
|
59
|
-
`smart-web` is a retrieval layer. It is not a general-purpose click/type/form-fill browser agent.[^1]
|
|
60
|
-
|
|
61
|
-
## What the tools return
|
|
62
|
-
|
|
63
|
-
All public tools are shaped for agent consumption. Common high-signal fields include:
|
|
64
|
-
|
|
65
|
-
- `assessment`: confidence, block/auth hints, recommended handoff
|
|
66
|
-
- `pipeline`: resolve/acquire/normalize/assess stages, plus policy metadata for relay-style usage and authwall likelihood; `pipeline.acquire.attempt_map` describes the default seed, direct/impit, Jina Reader, archive, and browser phases as succeeded, failed, or skipped
|
|
67
|
-
- `research_handoff`: compact metadata on blocked or partial `smartfetch` results telling downstream research workers and editable report renderers which evidence fields to consume
|
|
68
|
-
- `budget`: output projection metadata showing which fields were truncated and by how much, plus final MCP `response_chars` counts for compact-JSON `structuredContent` and rendered text; use `smartfetch`/`smartsearch`/`smartcrawl`/`smartvertical` `output_budget: "compact"` for token-sensitive MCP hosts, `preferred_budget_chars` for host-budget-aware routing, and `"full"` when larger structured payloads are explicitly needed
|
|
69
|
-
- `evidence.selected_output_budget*`: the selected `smartfetch` budget preset, winning control, and resolver reason so hosts can see which v0.x compatibility hint actually took effect
|
|
70
|
-
- `errors`: structured failure reasons
|
|
71
|
-
- `post`, `thread`, `comments`: normalized content surfaces
|
|
72
|
-
- `assets` and `download`: file-like outputs when present
|
|
73
|
-
|
|
74
|
-
This lets hosts decide whether deterministic retrieval was good enough or whether they should escalate to an auth-aware companion flow or a broader browser-task runtime.[^2]
|
|
75
|
-
|
|
76
|
-
For `smartfetch`, MCP `structuredContent` always carries the projected JSON result. MCP calls default to `output_budget: "compact"` so agent hosts do not ingest large page/thread/link payloads unless they ask for them. The projection caps known body/list fields plus arbitrary provider record width, nesting, and aggregate string content, so unknown `post`, thread, or comment fields cannot bypass the preset; `budget.fields` reports omitted nested paths and primary URL/title/author/status/body fields are prioritized. The human-readable `content[0].text` defaults to compact text instead of duplicating that JSON payload, with separate caps for long body text, result, comment, link, and asset sections. Omitted counts are reported in the text view while the projected fields remain available through `structuredContent`; set `format: "json"` only when a host needs JSON repeated in the text channel. Set `output_budget: "balanced"` or `"full"` only for follow-up calls that explicitly need more structured content. For host-side integration, prefer `preferred_budget_chars`; use `context_mode` only as a backward-compatible coarse alias when numeric hints are not supplied. If neither is present, `headroom_tokens` keeps compact output by default and uses `balanced` only when there is enough room. `format: "markdown"` keeps the compact response shape and requests Markdown extraction when available.
|
|
77
|
-
|
|
78
|
-
When a `smartfetch` result is blocked, partial, or low-confidence, `structuredContent.research_handoff` stays small and metadata-only. It points downstream tools at the source URL and the projected evidence bundle: `pipeline`, `assessment`, `errors`, `assets`, `outbound_links`, and the full projected `structuredContent`. Heavy research agents or editable report/rendering tools can consume that bundle downstream, but `smart-web` itself remains the unified retrieval kernel and does not spawn those tools.
|
|
79
|
-
|
|
80
|
-
For `smartsearch`, MCP calls now project both `structuredContent` and rendered text through the same compact/balanced/full budget controls. Compact search keeps enough ranked links for first-pass discovery while capping every normalized string field, including query/engine metadata, result titles and URLs, snippets, notes, assessment, and search-to-crawl handoff guidance; `contextMaxCharacters` remains a final text-channel cap for legacy hosts. Each shortened field is recorded under `budget.fields`. Use `output_budget: "full"` only for a known follow-up when the host really wants every returned snippet/result in `structuredContent`.
|
|
81
|
-
|
|
82
|
-
`smartsearch` reports provider execution separately from search matches under `structuredContent.execution`: the overall `status` is `succeeded` when at least one attempted provider completed, and each entry in `attempts` records its provider, status, result count, and any execution error. If every configured provider fails to execute, MCP returns `isError: true` and the CLI exits non-zero after preserving the structured diagnostics. A successful provider response with zero matching results remains successful even when the compatibility fields stay `engine: "none"` and `results: []`.
|
|
83
|
-
|
|
84
|
-
For `smartcrawl`, MCP summary calls now project both `structuredContent` and rendered text through compact/balanced/full budget controls, so broad docs/forum crawls do not flood agent context with every page, candidate, note, or error by default. Summary projection caps all normalized titles, URLs, discovery/page metadata, notes, errors, assessment, and handoff strings and records each omission under `budget.fields`. `output_dir` export mode remains handle-first: successful export responses embed `resource_handle_manifest` in `structuredContent`, and export summaries include file:// re-open guidance for the summary, manifest, resource-handle JSON, and first Markdown document so agent hosts can resume from handles without parsing the full JSON first. The resource-handle manifest also carries ordered `recommended_resources` for hosts that need a ready-to-open summary, manifest, and first document without inferring role priority. Blocked exports label those paths as planned/unwritten and include retry guidance instead of presenting false file handles. Set `format: "json"` only for hosts that cannot consume `structuredContent` and intentionally need JSON repeated in the text channel.
|
|
85
|
-
|
|
86
|
-
For `smartvertical`, MCP calls use the same budget-control precedence as the web tools. Compact mode caps rendered text plus arbitrary structured `result` arrays and long nested strings, and `structuredContent.budget.fields` reports exactly which vertical fields were shortened. Use `output_budget: "balanced"` or a larger `preferred_budget_chars` value when a Korean/local lookup needs more rows or full legal/safety record text.
|
|
87
|
-
|
|
88
|
-
Use budget presets this way:
|
|
89
|
-
|
|
90
|
-
- `compact`: default first pass for agent hosts, link triage, social/thread previews, and "is this page enough?" checks.
|
|
91
|
-
- `balanced`: follow-up when the first pass found the right page but omitted useful body, result, comment, or link detail.
|
|
92
|
-
- `full`: explicit retrieval pass for a known relevant URL when the host can afford the larger `structuredContent` payload.
|
|
93
|
-
|
|
94
|
-
When compact output omits needed content, make a follow-up call for the same URL/query/start site with a larger budget control instead of asking the model to infer from the compact text. Prefer `preferred_budget_chars` when the host knows its available response budget, or use `output_budget: "balanced"` / `"full"` for manual retries.
|
|
95
|
-
|
|
96
|
-
## Supported high-signal surfaces
|
|
97
|
-
|
|
98
|
-
- communities: Reddit, DCInside, LinkedIn public posts with public fallback recovery, Algumon
|
|
99
|
-
- search-native surfaces: Reddit, GitHub repository discovery and repo-docs path scopes, npm search and exact package path scopes, Hacker News, Stack Exchange, Wikipedia, Velog, MDN, crates.io
|
|
100
|
-
- competitive-programming surfaces: solved.ac problem/search/profile routes, BOJ problem/workbook/user pages, Codeforces problem/profile pages, AtCoder task/contest pages, QOJ problem pages, Jungol problem pages
|
|
101
|
-
- local map place pages: Naver Map, Kakao Map, and redirected `naver.me` place-share links
|
|
102
|
-
- reference/article pages: NamuWiki, Wikipedia, Naver Blog, Tistory, Velog
|
|
103
|
-
- media/social: YouTube video URLs with best-effort public transcripts, X, Threads, Instagram, Telegram
|
|
104
|
-
- commerce pages: Amazon, Coupang, Danawa, Aladin, AliExpress
|
|
105
|
-
- archives/reference wrappers: Wayback snapshots, archive.md lookup pages, arXiv abstract pages with direct paper links
|
|
106
|
-
- academic paper surfaces: arXiv, PubMed, PMC, bioRxiv, medRxiv, plus legal DOI-linked OA copies surfaced from OpenAlex, Unpaywall, Semantic Scholar, CORE, and Europe PMC when a publisher page is thin or paywalled
|
|
107
|
-
|
|
108
|
-
When a source stays blocked or under-specified, `smart-web` prefers partial but honest output over pretending to have verified content.
|
|
109
|
-
|
|
110
|
-
## Install
|
|
25
|
+
Requires Node.js 24 and npm. Start the stdio server:
|
|
111
26
|
|
|
112
27
|
```bash
|
|
113
28
|
npx -y smart-web-mcp
|
|
114
29
|
```
|
|
115
30
|
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
## Host setup
|
|
119
|
-
|
|
120
|
-
### Claude Code
|
|
121
|
-
|
|
122
|
-
```bash
|
|
123
|
-
claude mcp add smart-web -- npx -y smart-web-mcp
|
|
124
|
-
```
|
|
125
|
-
|
|
126
|
-
Suggested tool-use profile: default to `smartfetch` without `format` and without a budget control for compact first-pass retrieval; retry with `output_budget: "balanced"` or `"full"` only when the user asks for deeper extraction. If Claude Code exposes a host response budget, send it as `preferred_budget_chars`.
|
|
127
|
-
|
|
128
|
-
### Codex
|
|
129
|
-
|
|
130
|
-
```bash
|
|
131
|
-
codex mcp add smart-web -- npx -y smart-web-mcp
|
|
132
|
-
```
|
|
133
|
-
|
|
134
|
-
Or via `~/.codex/config.toml` (or project-scoped `.codex/config.toml`):
|
|
135
|
-
|
|
136
|
-
```toml
|
|
137
|
-
[mcp_servers.smart-web]
|
|
138
|
-
command = "npx"
|
|
139
|
-
args = ["-y", "smart-web-mcp"]
|
|
140
|
-
```
|
|
141
|
-
|
|
142
|
-
Suggested tool-use profile: keep default compact text for normal browsing. For targeted follow-up retrieval, pass `preferred_budget_chars` from the remaining response/tool budget when available; otherwise pass one explicit `output_budget` preset.
|
|
143
|
-
|
|
144
|
-
### OpenCode
|
|
145
|
-
|
|
146
|
-
```json
|
|
147
|
-
{
|
|
148
|
-
"$schema": "https://opencode.ai/config.json",
|
|
149
|
-
"mcp": {
|
|
150
|
-
"smart-web": {
|
|
151
|
-
"type": "local",
|
|
152
|
-
"command": ["npx", "-y", "smart-web-mcp"],
|
|
153
|
-
"enabled": true
|
|
154
|
-
}
|
|
155
|
-
}
|
|
156
|
-
}
|
|
157
|
-
```
|
|
158
|
-
|
|
159
|
-
Suggested tool-use profile: call `smartfetch` with no `format` for normal retrieval, `format: "markdown"` only when Markdown extraction is useful, and `format: "json"` only for hosts that intentionally require the full JSON duplicated into `content[0].text`.
|
|
160
|
-
|
|
161
|
-
### Hermes/Raon
|
|
162
|
-
|
|
163
|
-
For Hermes/Raon-style hosts, keep the MCP server command standard and set budget controls per tool call:
|
|
164
|
-
|
|
165
|
-
```json
|
|
166
|
-
{
|
|
167
|
-
"tool": "smartfetch",
|
|
168
|
-
"arguments": {
|
|
169
|
-
"url": "https://example.com/post",
|
|
170
|
-
"preferred_budget_chars": 40000
|
|
171
|
-
}
|
|
172
|
-
}
|
|
173
|
-
```
|
|
174
|
-
|
|
175
|
-
Use `preferred_budget_chars` as the primary host-profile control when Hermes/Raon knows the response budget. Use `headroom_tokens` only when the host has token headroom but no character budget, and use `context_mode` only for older integrations that cannot send numeric budget metadata.
|
|
176
|
-
|
|
177
|
-
Budget-control precedence is deterministic:
|
|
178
|
-
|
|
179
|
-
1. `output_budget`
|
|
180
|
-
2. `preferred_budget_chars`
|
|
181
|
-
3. `headroom_tokens`
|
|
182
|
-
4. `context_mode`
|
|
183
|
-
5. compact default
|
|
184
|
-
|
|
185
|
-
Send only one control when possible. If older host profiles send more than one, `evidence.selected_output_budget_source` reports which control won.
|
|
186
|
-
|
|
187
|
-
### Custom settings file
|
|
188
|
-
|
|
189
|
-
To use a non-default settings path, append `--settings-file` to the command:
|
|
190
|
-
|
|
191
|
-
```bash
|
|
192
|
-
npx -y smart-web-mcp --settings-file /absolute/path/to/smart-web.settings.json
|
|
193
|
-
```
|
|
194
|
-
|
|
195
|
-
For Claude Code:
|
|
196
|
-
|
|
197
|
-
```bash
|
|
198
|
-
claude mcp add smart-web -- npx -y smart-web-mcp --settings-file /absolute/path/to/smart-web.settings.json
|
|
199
|
-
```
|
|
200
|
-
|
|
201
|
-
For OpenCode, add `"args"` after `"command"`:
|
|
202
|
-
|
|
203
|
-
```json
|
|
204
|
-
{
|
|
205
|
-
"mcp": {
|
|
206
|
-
"smart-web": {
|
|
207
|
-
"type": "local",
|
|
208
|
-
"command": ["npx", "-y", "smart-web-mcp", "--settings-file", "/absolute/path/to/smart-web.settings.json"],
|
|
209
|
-
"enabled": true
|
|
210
|
-
}
|
|
211
|
-
}
|
|
212
|
-
}
|
|
213
|
-
```
|
|
214
|
-
|
|
215
|
-
## Configuration
|
|
216
|
-
|
|
217
|
-
### Settings path
|
|
218
|
-
|
|
219
|
-
Default: `~/.config/smart-web/settings.json`
|
|
220
|
-
|
|
221
|
-
Print a template:
|
|
222
|
-
|
|
223
|
-
```bash
|
|
224
|
-
npx -y smart-web-mcp --print-settings-example
|
|
225
|
-
```
|
|
226
|
-
|
|
227
|
-
Initialize the default file:
|
|
228
|
-
|
|
229
|
-
```bash
|
|
230
|
-
npx -y smart-web-mcp --init-settings
|
|
231
|
-
```
|
|
232
|
-
|
|
233
|
-
`smart-web` treats the settings file as the single runtime config surface.
|
|
234
|
-
|
|
235
|
-
### Full settings reference
|
|
236
|
-
|
|
237
|
-
```jsonc
|
|
238
|
-
{
|
|
239
|
-
// "balanced" (default) or "private"
|
|
240
|
-
// "private" disables relay-style providers and public search helpers
|
|
241
|
-
"profile": "balanced",
|
|
242
|
-
|
|
243
|
-
"runtime": {
|
|
244
|
-
// Optional override for staging/export temp files
|
|
245
|
-
// Default: platform cache root (e.g. ~/.cache/smart-web/tmp)
|
|
246
|
-
"tempDir": ""
|
|
247
|
-
},
|
|
248
|
-
|
|
249
|
-
"search": {
|
|
250
|
-
// Self-hosted SearXNG instance URL
|
|
251
|
-
"searxngBaseUrl": "",
|
|
252
|
-
"enableSearxng": true
|
|
253
|
-
},
|
|
254
|
-
|
|
255
|
-
"fetch": {
|
|
256
|
-
// Browser overrides — most users leave these empty
|
|
257
|
-
"chromeChannel": "",
|
|
258
|
-
"chromePath": "",
|
|
259
|
-
// Auto-install Playwright Chromium when missing (default: true)
|
|
260
|
-
"autoInstallPlaywright": true,
|
|
261
|
-
// Jina Reader relay — opt-in third-party path for weak article pages
|
|
262
|
-
"enableJinaReader": false,
|
|
263
|
-
"jinaReaderBaseUrl": "https://r.jina.ai/",
|
|
264
|
-
// Academic fallback — legal OA enrichment for paper URLs
|
|
265
|
-
"enableAcademicFallback": true,
|
|
266
|
-
"enableOpenAlex": true,
|
|
267
|
-
"enableEuropePmc": true,
|
|
268
|
-
"enableBiorxivApi": true,
|
|
269
|
-
// Unpaywall — requires a contact email
|
|
270
|
-
"enableUnpaywall": true,
|
|
271
|
-
"unpaywallEmail": "",
|
|
272
|
-
// Semantic Scholar — optional API key for higher-rate access
|
|
273
|
-
"enableSemanticScholar": true,
|
|
274
|
-
"semanticScholarApiKey": "",
|
|
275
|
-
// CORE — optional API key for richer search-backed enrichment
|
|
276
|
-
"enableCoreDiscovery": true,
|
|
277
|
-
"coreApiKey": "",
|
|
278
|
-
// FxTwitter — transparent x.com → fxtwitter redirect
|
|
279
|
-
"enableFxTwitter": true,
|
|
280
|
-
// Undetected-chromedriver — optional Python Selenium fallback for tough anti-bot pages
|
|
281
|
-
"enableUndetectedChromedriver": true,
|
|
282
|
-
"undetectedChromedriverPython": "python3",
|
|
283
|
-
// Reddit personal retrieval chain (legacy key name retained for compatibility)
|
|
284
|
-
"enableRedditJson": true,
|
|
285
|
-
// YouTube video transcript enrichment for public captions (default: true)
|
|
286
|
-
"enableYoutubeTranscript": true,
|
|
287
|
-
// Archive fallback — Wayback and archive.md recovery
|
|
288
|
-
"enableArchiveFallback": true,
|
|
289
|
-
"enableWayback": true,
|
|
290
|
-
"enableArchiveMd": true,
|
|
291
|
-
// Optional local adaptive scraper fallback. This is disabled by default;
|
|
292
|
-
// install Scrapling separately and opt in only on deployments that want an
|
|
293
|
-
// extra no-secret local shell/bot-challenge extraction attempt before the
|
|
294
|
-
// built-in Playwright fallback.
|
|
295
|
-
"enableAdaptiveScraper": false,
|
|
296
|
-
"adaptiveScraperCommand": "scrapling",
|
|
297
|
-
"adaptiveScraperTimeoutMs": 20000,
|
|
298
|
-
// Output projection budget for smartfetch responses. MCP, CLI, and library
|
|
299
|
-
// calls honor this setting unless the caller passes a per-call budget control.
|
|
300
|
-
// With no setting or per-call control, MCP calls default to compact.
|
|
301
|
-
// preset: "compact" | "balanced" | "full"; numeric fields override the preset.
|
|
302
|
-
"outputBudget": {
|
|
303
|
-
"preset": "balanced",
|
|
304
|
-
"maxPostTextChars": 8000,
|
|
305
|
-
"maxPostMarkdownChars": 10000,
|
|
306
|
-
"maxThreadItems": 10,
|
|
307
|
-
"maxCommentItems": 10,
|
|
308
|
-
"maxOutboundLinks": 60,
|
|
309
|
-
"maxAssets": 15,
|
|
310
|
-
"maxBlocks": 20,
|
|
311
|
-
"maxBlockTextChars": 500,
|
|
312
|
-
"maxItemTextChars": 500
|
|
313
|
-
},
|
|
314
|
-
// Optional default named browser profile for authenticated smartfetch calls.
|
|
315
|
-
// Names are 1-64 characters: letters, numbers, dots, underscores, or hyphens;
|
|
316
|
-
// they must start with a letter or number and cannot end with a dot.
|
|
317
|
-
"browserProfile": ""
|
|
318
|
-
},
|
|
319
|
-
|
|
320
|
-
"network": {
|
|
321
|
-
// Allow localhost/private/reserved fetch targets (default: false)
|
|
322
|
-
// Applies to smartfetch, smartcrawl, and Korean public API helpers.
|
|
323
|
-
"allowPrivateHosts": false
|
|
324
|
-
}
|
|
325
|
-
}
|
|
326
|
-
```
|
|
327
|
-
|
|
328
|
-
### Common setups
|
|
329
|
-
|
|
330
|
-
#### Default — no config file needed
|
|
331
|
-
|
|
332
|
-
Out of the box `smart-web` works with sensible defaults. Create a settings file only when you need to change something.
|
|
333
|
-
|
|
334
|
-
#### Minimal: profile only
|
|
335
|
-
|
|
336
|
-
```json
|
|
337
|
-
{
|
|
338
|
-
"profile": "balanced"
|
|
339
|
-
}
|
|
340
|
-
```
|
|
341
|
-
|
|
342
|
-
#### Private / local-first
|
|
343
|
-
|
|
344
|
-
```json
|
|
345
|
-
{
|
|
346
|
-
"profile": "private",
|
|
347
|
-
"search": {
|
|
348
|
-
"searxngBaseUrl": "http://localhost:8080"
|
|
349
|
-
}
|
|
350
|
-
}
|
|
351
|
-
```
|
|
352
|
-
|
|
353
|
-
In `private` mode, relay-style provider requests are blocked, public no-key search fallbacks are disabled, and SearXNG bases must resolve to localhost or private addresses. Hosted API-key search providers such as Exa, Tavily, and Brave Search are intentionally not part of the default smartsearch runtime; use local/site-native search, DuckDuckGo/Brave HTML fallbacks, or self-hosted SearXNG instead.
|
|
354
|
-
|
|
355
|
-
#### Disable Playwright auto-install
|
|
356
|
-
|
|
357
|
-
```json
|
|
358
|
-
{
|
|
359
|
-
"fetch": {
|
|
360
|
-
"autoInstallPlaywright": false
|
|
361
|
-
}
|
|
362
|
-
}
|
|
363
|
-
```
|
|
364
|
-
|
|
365
|
-
The server will return an actionable `playwright_browsers_missing` or `playwright_browsers_outdated` error instead.
|
|
366
|
-
|
|
367
|
-
#### Opt in to Jina Reader for weak generic article pages
|
|
368
|
-
|
|
369
|
-
```json
|
|
370
|
-
{
|
|
371
|
-
"fetch": {
|
|
372
|
-
"enableJinaReader": true
|
|
373
|
-
}
|
|
374
|
-
}
|
|
375
|
-
```
|
|
376
|
-
|
|
377
|
-
This stays opt-in because it is a third-party relay path. When enabled, it only applies to weak generic/article direct fetches and prefers Jina JSON mode when that improves the page. `profile: "private"` hard-disables relay-style providers.
|
|
378
|
-
|
|
379
|
-
Relay eligibility is provider-policy based. Generic and document-like providers can opt in, while browser/auth/social/commerce providers stay off unless their provider declares support. `pipeline.policy.relay_style_allowed` and `pipeline.policy.relay_style_used` tell hosts whether a relay path was allowed and whether it was used.
|
|
380
|
-
|
|
381
|
-
#### Opt in to a local adaptive scraper command
|
|
382
|
-
|
|
383
|
-
```json
|
|
384
|
-
{
|
|
385
|
-
"fetch": {
|
|
386
|
-
"enableAdaptiveScraper": true,
|
|
387
|
-
"adaptiveScraperCommand": "scrapling",
|
|
388
|
-
"adaptiveScraperTimeoutMs": 20000
|
|
389
|
-
}
|
|
390
|
-
}
|
|
391
|
-
```
|
|
392
|
-
|
|
393
|
-
This lane is for local, no-secret deployments that install Scrapling or a compatible CLI themselves. It runs only after direct retrieval still looks weak, blocked, or shell-like and before the built-in browser fallback. The attempt is recorded as `adaptive_scraper` in `pipeline.acquire.attempt_map`; if the command is missing, times out, returns empty text, or does not improve the active result, smartfetch keeps the existing result and continues through the normal fallback chain.
|
|
394
|
-
|
|
395
|
-
#### Tune smartfetch output budgets
|
|
31
|
+
MCP host configuration:
|
|
396
32
|
|
|
397
33
|
```json
|
|
398
34
|
{
|
|
399
|
-
"
|
|
400
|
-
"
|
|
401
|
-
"maxPostTextChars": 16000,
|
|
402
|
-
"maxThreadItems": 25,
|
|
403
|
-
"maxOutboundLinks": 80
|
|
404
|
-
}
|
|
35
|
+
"mcpServers": {
|
|
36
|
+
"smart-web": { "command": "npx", "args": ["-y", "smart-web-mcp"] }
|
|
405
37
|
}
|
|
406
38
|
}
|
|
407
39
|
```
|
|
408
40
|
|
|
409
|
-
|
|
41
|
+
Use `--init-settings` to create the default settings file, or `--print-settings-example` to inspect it. The published npm package can lag source; verify the runtime you actually use.
|
|
410
42
|
|
|
411
|
-
|
|
412
|
-
|
|
413
|
-
```json
|
|
414
|
-
{
|
|
415
|
-
"fetch": {
|
|
416
|
-
"enableArchiveFallback": false
|
|
417
|
-
}
|
|
418
|
-
}
|
|
419
|
-
```
|
|
420
|
-
|
|
421
|
-
#### Legal OA DOI recovery for paywalled paper pages
|
|
422
|
-
|
|
423
|
-
```json
|
|
424
|
-
{
|
|
425
|
-
"fetch": {
|
|
426
|
-
"enableUnpaywall": true,
|
|
427
|
-
"unpaywallEmail": "research@example.com",
|
|
428
|
-
"enableSemanticScholar": true,
|
|
429
|
-
"enableCoreDiscovery": true
|
|
430
|
-
}
|
|
431
|
-
}
|
|
432
|
-
```
|
|
433
|
-
|
|
434
|
-
#### Enable undetected-chromedriver from a dedicated virtualenv
|
|
435
|
-
|
|
436
|
-
```json
|
|
437
|
-
{
|
|
438
|
-
"fetch": {
|
|
439
|
-
"enableUndetectedChromedriver": true,
|
|
440
|
-
"undetectedChromedriverPython": "/absolute/path/to/venv/bin/python"
|
|
441
|
-
}
|
|
442
|
-
}
|
|
443
|
-
```
|
|
444
|
-
|
|
445
|
-
This helper is optional. `smart-web` still works without it, but when installed and reachable it can act as a second browser engine for stubborn pages where Playwright alone is not enough.
|
|
446
|
-
|
|
447
|
-
#### Custom temp directory
|
|
448
|
-
|
|
449
|
-
```json
|
|
450
|
-
{
|
|
451
|
-
"runtime": {
|
|
452
|
-
"tempDir": "/absolute/path/to/smart-web-tmp"
|
|
453
|
-
}
|
|
454
|
-
}
|
|
455
|
-
```
|
|
456
|
-
|
|
457
|
-
### Advanced overrides
|
|
458
|
-
|
|
459
|
-
These keys live inside `settings.json` when you need to force one provider on or off:
|
|
460
|
-
|
|
461
|
-
**Search**
|
|
462
|
-
|
|
463
|
-
- `search.enableSearxng`
|
|
464
|
-
- `search.enableDuckDuckGo`
|
|
465
|
-
- `search.enableBraveHtml`
|
|
466
|
-
|
|
467
|
-
**Fetch**
|
|
468
|
-
|
|
469
|
-
- `fetch.enableFxTwitter`
|
|
470
|
-
- `fetch.enableXOembed`
|
|
471
|
-
- `fetch.autoInstallPlaywright`
|
|
472
|
-
- `fetch.enableJinaReader`
|
|
473
|
-
- `fetch.jinaReaderBaseUrl`
|
|
474
|
-
- `fetch.enableAcademicFallback`
|
|
475
|
-
- `fetch.enableOpenAlex`
|
|
476
|
-
- `fetch.enableEuropePmc`
|
|
477
|
-
- `fetch.enableBiorxivApi`
|
|
478
|
-
- `fetch.enableUnpaywall`
|
|
479
|
-
- `fetch.unpaywallEmail`
|
|
480
|
-
- `fetch.enableSemanticScholar`
|
|
481
|
-
- `fetch.semanticScholarApiKey`
|
|
482
|
-
- `fetch.enableCoreDiscovery`
|
|
483
|
-
- `fetch.coreApiKey`
|
|
484
|
-
- `fetch.enableUndetectedChromedriver`
|
|
485
|
-
- `fetch.undetectedChromedriverPython`
|
|
486
|
-
- `fetch.enableRedditJson`
|
|
487
|
-
- `fetch.enableYoutubeTranscript`
|
|
488
|
-
- `fetch.enableArchiveFallback`
|
|
489
|
-
- `fetch.enableWayback`
|
|
490
|
-
- `fetch.enableArchiveMd`
|
|
491
|
-
- `fetch.outputBudget.maxPostTextChars`
|
|
492
|
-
- `fetch.outputBudget.maxPostMarkdownChars`
|
|
493
|
-
- `fetch.outputBudget.maxThreadItems`
|
|
494
|
-
- `fetch.outputBudget.maxCommentItems`
|
|
495
|
-
- `fetch.outputBudget.maxOutboundLinks`
|
|
496
|
-
- `fetch.outputBudget.maxAssets`
|
|
497
|
-
- `fetch.outputBudget.maxBlocks`
|
|
498
|
-
- `fetch.outputBudget.maxBlockTextChars`
|
|
499
|
-
- `fetch.outputBudget.maxItemTextChars`
|
|
500
|
-
- `fetch.outputBudget.maxEnvelopeArrayItems`
|
|
501
|
-
- `fetch.outputBudget.maxAssessmentSignals`
|
|
502
|
-
- `fetch.outputBudget.maxTotalEnvelopeEntries`
|
|
503
|
-
- `fetch.outputBudget.maxTotalEnvelopeStringChars`
|
|
504
|
-
|
|
505
|
-
**Compatibility**
|
|
506
|
-
|
|
507
|
-
- `network.localOnly`: advanced override that forces local-only behavior regardless of profile
|
|
508
|
-
|
|
509
|
-
## Provider benchmark
|
|
510
|
-
|
|
511
|
-
Use the provider benchmark when deciding whether to build, wrap, or replace retrieval lanes:
|
|
512
|
-
|
|
513
|
-
```bash
|
|
514
|
-
npm run benchmark:providers
|
|
515
|
-
npm run benchmark:providers:measure
|
|
516
|
-
npm run benchmark:providers:compare
|
|
517
|
-
npm run --silent benchmark:providers:json
|
|
518
|
-
npm run benchmark:providers:json-file
|
|
519
|
-
npm run benchmark:providers:report
|
|
520
|
-
npm run benchmark:providers:gaps
|
|
521
|
-
npm run benchmark:providers:live
|
|
522
|
-
npm run benchmark:providers:live:jina
|
|
523
|
-
npm run benchmark:providers:live:cp
|
|
524
|
-
npm run benchmark:providers:live:browser
|
|
525
|
-
npm run benchmark:providers:live:youtube
|
|
526
|
-
npm run benchmark:providers:live:blocked
|
|
527
|
-
npm run benchmark:providers:live:naver-blog
|
|
528
|
-
npm run benchmark:providers:live:naver-map
|
|
529
|
-
npm run benchmark:providers:live:naver-cafe
|
|
530
|
-
npm run benchmark:providers:live:x-twitter
|
|
531
|
-
npm run benchmark:providers:live:no-secret
|
|
532
|
-
npm run benchmark:providers:live:no-secret:gaps
|
|
533
|
-
npm run watchdog:no-secret:live
|
|
534
|
-
npm run benchmark:providers:trend
|
|
535
|
-
npm run benchmark:smartsearch
|
|
536
|
-
npm run --silent benchmark:smartsearch:json
|
|
537
|
-
npm run watchdog:no-secret
|
|
538
|
-
npm run watchdog:partial-handoff
|
|
539
|
-
```
|
|
540
|
-
|
|
541
|
-
The default report is availability-only and safe for CI. `--run-measurements` (`npm run benchmark:providers:measure`) adds deterministic fixture records for available local/core and optional no-key lanes: success, useful character counts, projected/text character counts, useful-to-projected and useful-to-text yield ratios, latency, extractor/provider id, confidence, partial status, error class, lane label/family/category metadata, provider-family/category lane-status rollups, surface coverage, and recommendation sections grouped by lane, case, and surface.
|
|
542
|
-
|
|
543
|
-
Use `--compare-external` (`npm run benchmark:providers:compare`) to add an `externalComparisons` section that compares current smart-web fixture evidence against optional no-secret/local candidates. The comparison includes the local Crawl4AI command lane, the public Jina Reader relay fixture, and available direct/Playwright baselines. If `crawl4ai` is not installed, the Crawl4AI row is `skipped` with `action: "install-local-prerequisite"` and `prerequisite.command: "crawl4ai"`; this is an expected setup signal, not a benchmark failure or required dependency.
|
|
544
|
-
|
|
545
|
-
Reports carry the `provider-benchmark.v1` schema version plus non-fatal git and runtime metadata so archived JSON, Markdown, and NDJSON trend artifacts are parseable, traceable, and comparable as the harness evolves. Markdown, text, JSON, and NDJSON outputs include aggregate fixture yield averages, compact yield rollups by provider family/category/lane, and a best-current-evidence surface table that ranks measured/skipped candidates so review jobs can compare token efficiency and build-vs-wrap posture between core, browser, relay, and specialist paths without reprocessing the full measurement table.
|
|
546
|
-
|
|
547
|
-
Measurement artifacts also include a live-readiness manifest listing each lane's env/command/settings prerequisites, a deterministic fixture smoke, a safe MCP live-smoke candidate, and surfaces that still lack live evidence. `--run-live-smokes` (`npm run benchmark:providers:live`) is an explicit gated mode that runs only no-secret, core-surface live checks for static articles, long docs/Wikipedia, Hacker News, and docs export; optional no-secret lanes remain gated until explicitly allowlisted: `--allow-jina-live-smoke` (`npm run benchmark:providers:live:jina`) exercises the public Jina Reader relay with a secret-stripped settings file, `--allow-competitive-programming-live-smoke` (`npm run benchmark:providers:live:cp`) adds a public BOJ/Codeforces representative smoke for the competitive-programming specialist surface, `--allow-browser-live-smoke` (`npm run benchmark:providers:live:browser`) runs a Playwright `--force-dynamic` smoke for the JS-rendered article surface only when the local browser lane is available, `--allow-youtube-live-smoke` (`npm run benchmark:providers:live:youtube`) adds the public YouTube seed-lane representative smoke, `--allow-blocked-authwall-live-smoke` (`npm run benchmark:providers:live:blocked`) adds a deterministic HTTP Basic Auth challenge smoke that treats honest blocked/error/handoff signals as positive calibration without private credentials, `--allow-naver-blog-live-smoke` (`npm run benchmark:providers:live:naver-blog`) adds a stable public Naver Blog representative smoke for Korean specialist-lane calibration, `--allow-naver-map-live-smoke` (`npm run benchmark:providers:live:naver-map`) adds a stable public Naver Map place smoke with browser-preflight diagnostics, `--allow-naver-cafe-live-smoke` (`npm run benchmark:providers:live:naver-cafe`) adds a stable public Naver Cafe shared-link smoke that treats useful no-credential partial output as honest calibration, and `--allow-x-twitter-live-smoke` (`npm run benchmark:providers:live:x-twitter`) adds a public X/Twitter post smoke that keeps no-secret social-provider drift visible without private cookies. `--allow-all-no-secret-live-smokes` (`npm run benchmark:providers:live:no-secret`) enables every no-secret gate in one run; hosted API-key search lanes are excluded from the benchmark surface by product policy.
|
|
548
|
-
|
|
549
|
-
Live smoke artifacts record per-case `present`/`missing`/`skipped` evidence next to fixture evidence without printing secret values. When fixture rows exist but all are skipped by missing local commands, comparison and gap artifacts report fixture evidence as `skipped` with a `fixture-skipped-*` comparison instead of implying the fixture rows are absent. Browser-lane live smoke rows also include a secret-safe Playwright preflight diagnostic with executable path, launch timing, `domcontentloaded` navigation timing, HTTP status/final URL, and truncated failure reason so runtime/cache/network failures can be separated from smartfetch normalization failures.
|
|
550
|
-
|
|
551
|
-
The same JSON, text, Markdown, and NDJSON outputs include a compact live-vs-fixture comparison rollup so replacement reviews can distinguish calibrated surfaces from fixture-only, skipped-live, live-disagreeing, and fixture-skipped-by-prerequisite surfaces. Secret-safe provider diagnostics list prerequisite names with satisfied/missing status and never print env values. Use `npm run benchmark:providers:json-file` or add `--json-file <path>` to write the full JSON artifact without relying on stdout capture. Use `npm run benchmark:providers:report` or add `--report-file <path>` with measurement mode to write a Markdown review artifact containing lane availability, diagnostics, surface/corpus coverage, measurement highlights, yield comparisons, per-surface evidence, live readiness, live-smoke diagnostics, and build-vs-wrap recommendations without changing stdout or JSON behavior. Use `npm run benchmark:providers:gaps` or add `--gap-file <path>` with measurement mode to write a compact `provider-benchmark-gaps.v1` JSON artifact containing only the ranked provider evidence gap queue, summary counts, git/runtime metadata, and safe smoke candidates. Use `npm run benchmark:providers:trend` or add `--trend-file <path>` with measurement mode to append compact NDJSON trend records with aggregate lane, prerequisite, measurement, recommendation, surface-evidence, and fixture-yield counts. Unknown `--flag` options fail fast so typoed benchmark jobs do not produce misleading artifacts. The corpus covers static articles, JS-rendered articles, blocked/authwall-like pages, long docs/Wikipedia pages, Naver Blog/Cafe/Map, Hacker News, YouTube, X/Twitter, BOJ/Codeforces, relay, and docs export surfaces. Optional local/no-secret lanes without commands remain explicit skips instead of silently failing; API-key provider candidates are excluded from the default corpus and reports by product policy.
|
|
552
|
-
|
|
553
|
-
Provider benchmark reports also include `providerEvidenceGaps`: a ranked next-action queue built from the live-vs-fixture rollup. Each gap records the surface, priority score, selected and candidate lanes, follow-up action, unlock prerequisites, fixture/live useful chars, fixture skipped-case lanes/reasons, live-smoke failure error classes/reasons, and safe live smoke candidate so replacement reviews can pick the next provider experiment without scanning every measurement row or re-running a timed live calibration. Use `npm run benchmark:providers:live:no-secret:gaps` to refresh only the compact all-no-secret gap artifact at `reports/provider-benchmark-live-no-secret-gaps.json`. When fixture rows exist but were skipped by missing local commands, the live-vs-fixture and gap actions are `unlock-skipped-fixture-prerequisites` rather than fixture-creation actions.
|
|
554
|
-
|
|
555
|
-
`npm run benchmark:smartsearch` is the deterministic quality gate for site-native smartsearch routing. It patches `fetch`, disables generic web-search fallbacks with a temporary settings file, and exercises representative `site:` queries for Reddit, GitHub repo discovery, GitHub repo docs, npm, Hacker News, Stack Exchange, Wikipedia, Velog, MDN, and crates.io without live network access. The human report is compact; `npm run --silent benchmark:smartsearch:json` emits `smartsearch-quality-benchmark.v1` JSON with expected engine, top URL, URL-substring checks, actual engine, top URL, notes, and pass/fail status per case.
|
|
556
|
-
|
|
557
|
-
See `reports/README.md` for the curated report archive index. It identifies `provider-benchmark-live-no-secret.{json,md}` and `provider-benchmark-live-no-secret-gaps.json` as the current maintenance artifacts, and labels older cumulative live-smoke reports as historical calibration snapshots.
|
|
558
|
-
|
|
559
|
-
`npm run watchdog:no-secret` is the deterministic current-surface gate for no-secret maintenance. It fails when `reports/provider-benchmark-live-no-secret-gaps.json` contains current gaps, when runtime/docs/examples/scripts reintroduce hosted API-key provider markers for Exa, Tavily, Firecrawl, Browserless, or Brave Search API, or when the partial/handoff watchdog no longer has a compact text ceiling. It intentionally ignores historical `CHANGELOG.md` entries and archived `reports/` output, and does not run live network smokes.
|
|
560
|
-
|
|
561
|
-
`npm run watchdog:no-secret:live` is the live regression gate for no-secret maintenance. It first refreshes `reports/provider-benchmark-live-no-secret-gaps.json` with `npm run benchmark:providers:live:no-secret:gaps`, then runs `npm run watchdog:no-secret` so non-empty live provider gaps fail locally or in CI. The scheduled/manual GitHub Actions workflow `.github/workflows/no-secret-live-watchdog.yml` runs the same alias weekly with read-only permissions, no secrets, and a short-retention gap artifact.
|
|
562
|
-
|
|
563
|
-
`npm run watchdog:partial-handoff` builds the package and runs a compact no-secret drift watchdog over representative partial/blocked/handoff `smartfetch` outputs (httpbin authwall, LinkedIn profile authwall, Naver Blog unavailable/deleted posts, Threads unavailable/join-login output, DCInside unavailable/deleted posts, Telegram missing/private posts, a non-existent Wikidocs article, Naver Cafe, Naver Map search, and BOJ unavailable/shutdown output). It exits non-zero when expected handoff cases are silently reclassified as non-partial, when partial/handoff `content[0].text` grows beyond the compact watchdog ceiling, or when compact text or structured links/assets leak chrome URLs, script globals, or style-shell tokens, and prints a `partial-handoff-watchdog.v1` JSON report without raw page dumps or credentials. For runner artifacts without JSON stdout, run `node scripts/partial-handoff-watchdog.mjs --report-file <path>` after `npm run build`; the command keeps stdout human-readable and writes the full JSON report to the requested path. Fixture cases can set `expectPartialHandoff: true` to make classifier drift fail even if the output would otherwise be skipped.
|
|
564
|
-
|
|
565
|
-
## Quick verification
|
|
566
|
-
|
|
567
|
-
Sanity-check the server with a few real calls:
|
|
568
|
-
|
|
569
|
-
- `smartfetch` on a normal article URL
|
|
570
|
-
- `smartfetch` on a Medium or other member-only article URL — confirm it returns either archive-backed content or an honest `reference_only` preview
|
|
571
|
-
- `smartfetch` on a Naver Map, Kakao Map, or `naver.me` place-share URL
|
|
572
|
-
- `smartfetch` on an arXiv abstract URL
|
|
573
|
-
- `smartfetch` on a PubMed, PMC, or bioRxiv/medRxiv paper URL
|
|
574
|
-
- `smartfetch` on a known product URL
|
|
575
|
-
- `smartsearch` on a `site:` query
|
|
576
|
-
- `smartcrawl` on a docs site or board you actually use
|
|
577
|
-
|
|
578
|
-
## Reporting issues
|
|
579
|
-
|
|
580
|
-
If `smart-web` returns reproducible incorrect or misleading output, open a GitHub issue with:
|
|
581
|
-
|
|
582
|
-
- exact input and tool arguments
|
|
583
|
-
- observed output
|
|
584
|
-
- expected behavior
|
|
585
|
-
- `smart-web-mcp` version
|
|
586
|
-
|
|
587
|
-
Skip obvious transient network, auth, or rate-limit failures unless the classification itself looks wrong.
|
|
43
|
+
## Tool routing
|
|
588
44
|
|
|
589
|
-
|
|
45
|
+
- `smartfetch` first when the user already gave a URL or shortlink.
|
|
46
|
+
- `smartsearch` when the user gave a topic, keywords, or a `site:` query.
|
|
47
|
+
- `smartcrawl` when you already know the site and need multiple relevant pages.
|
|
48
|
+
- `smartvertical` for Korean/local structured lookups.
|
|
49
|
+
- `smartbrowser` for explicit approved-browser input/click/filter/wait/screenshot batches; see [action scope and limits](https://github.com/jojo-labs/smart-web/blob/main/docs/operator-actions.md).
|
|
590
50
|
|
|
591
|
-
|
|
51
|
+
Use compact output first; request a larger budget when needed. Partial output is honest evidence, **not task completion**. Never substitute a search snippet, loading screen, or unrelated listing for the requested source.
|
|
52
|
+
The CLI `smart-web fetch <url> --format json` (or `--json`) prints structured output; `--budget full` controls bounded delivery, not source completeness. Invalid budget names fail before retrieval rather than silently using compact output.
|
|
53
|
+
For CLI search, `--max-chars` caps rendered text; use `--json` when the complete projected structured result is needed.
|
|
592
54
|
|
|
593
|
-
|
|
55
|
+
## Documentation
|
|
594
56
|
|
|
595
|
-
|
|
57
|
+
- [Product contract and acceptance](https://github.com/jojo-labs/smart-web/blob/main/docs/operator-contract.md)
|
|
58
|
+
- [Host setup](https://github.com/jojo-labs/smart-web/blob/main/docs/host-setup.md) · [Configuration](https://github.com/jojo-labs/smart-web/blob/main/docs/configuration.md)
|
|
59
|
+
- [Architecture and code map](https://github.com/jojo-labs/smart-web/blob/main/docs/architecture.md) · [Verification and maintenance](https://github.com/jojo-labs/smart-web/blob/main/docs/benchmarks.md)
|
|
60
|
+
- [All docs](https://github.com/jojo-labs/smart-web/blob/main/docs/README.md) · [Agent instructions](https://github.com/jojo-labs/smart-web/blob/main/AGENTS.md)
|
|
596
61
|
|
|
597
|
-
|
|
62
|
+
Report reproducible problems in [GitHub Issues](https://github.com/jojo-labs/smart-web/issues), with sanitized input, expected/observed behavior, and runtime version. Never attach credentials or private browsing content.
|
|
598
63
|
|
|
599
|
-
[
|
|
64
|
+
[Package](https://www.npmjs.com/package/smart-web-mcp) · MCP Registry: `io.github.jojo-labs/smart-web` · [License](https://github.com/jojo-labs/smart-web/blob/main/LICENSE) (proprietary; all rights reserved).
|