dsh-search-enhance 0.1.0 → 0.1.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +44 -45
- package/lib/config.d.ts +1 -1
- package/lib/index.js +37 -28
- package/lib/presentation/render.d.ts +2 -2
- package/lib/presentation/render.js +38 -12
- package/lib/prompt/tool-discovery.d.ts +4 -3
- package/lib/prompt/tool-discovery.js +16 -39
- package/lib/providers/direct-fetch.d.ts +4 -2
- package/lib/providers/direct-fetch.js +8 -2
- package/lib/providers/direct-http.js +4 -0
- package/lib/providers/firecrawl-scrape.d.ts +3 -1
- package/lib/providers/firecrawl-scrape.js +12 -5
- package/lib/providers/smart-direct-child.js +5 -2
- package/lib/providers/smart-direct-transport.js +5 -0
- package/lib/providers/smart-direct.d.ts +2 -1
- package/lib/providers/smart-direct.js +5 -1
- package/lib/providers/tavily-extract.d.ts +4 -2
- package/lib/providers/tavily-extract.js +11 -4
- package/lib/providers/web-extract-common.d.ts +7 -0
- package/lib/providers/web-extract-common.js +58 -0
- package/lib/research-plan/index.d.ts +1 -1
- package/lib/research-plan/index.js +8 -6
- package/lib/site-map/types.d.ts +1 -1
- package/lib/tool-discovery/capabilities.d.ts +8 -8
- package/lib/tool-discovery/capabilities.js +24 -23
- package/lib/tool-discovery/fold.d.ts +6 -0
- package/lib/tool-discovery/fold.js +10 -0
- package/lib/tool-discovery/manager.d.ts +7 -22
- package/lib/tool-discovery/manager.js +31 -137
- package/lib/tools/context7.d.ts +1 -1
- package/lib/tools/context7.js +1 -1
- package/lib/tools/docs-search.d.ts +2 -0
- package/lib/tools/docs-search.js +4 -1
- package/lib/tools/index.d.ts +1 -0
- package/lib/tools/index.js +1 -0
- package/lib/tools/research-plan.d.ts +1 -1
- package/lib/tools/research-plan.js +2 -2
- package/lib/tools/schemas.d.ts +6 -5
- package/lib/tools/schemas.js +6 -5
- package/lib/tools/search-call.d.ts +74 -0
- package/lib/tools/search-call.js +334 -0
- package/lib/tools/search-tools.d.ts +46 -23
- package/lib/tools/search-tools.js +43 -36
- package/lib/tools/web-map.d.ts +1 -1
- package/lib/tools/web-map.js +1 -1
- package/lib/tools/web-search.d.ts +2 -0
- package/lib/tools/web-search.js +4 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -12,47 +12,51 @@
|
|
|
12
12
|
▼
|
|
13
13
|
DSH Agent
|
|
14
14
|
│
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
15
|
+
└─ 固定模型工具 surface(schema 与顺序不随披露状态变化)
|
|
16
|
+
├─ web_search ──────> Grok 主搜索
|
|
17
|
+
│ ├─ 按需要补充 Context7 / Exa
|
|
18
|
+
│ ├─ 按需要补充 Tavily / Firecrawl
|
|
19
|
+
│ └─ 返回回答、来源;有 source_ref 时追加 search_sources manifest
|
|
20
|
+
│
|
|
21
|
+
├─ docs_search ─────> Context7 / Exa 文档检索
|
|
22
|
+
│ └─ 返回文档片段、来源;有 source_ref 时追加 search_sources manifest
|
|
23
|
+
│
|
|
24
|
+
├─ web_extract ─────> Tavily → Firecrawl → smart_direct → direct
|
|
25
|
+
│ └─ 读取选中网页的正文
|
|
26
|
+
│
|
|
27
|
+
├─ search_tools ────> 按需返回 capability / operation manifest
|
|
28
|
+
│
|
|
29
|
+
└─ search_call ─────> 调用已经激活的延迟 operation
|
|
30
|
+
├─ Context7 精细查询
|
|
31
|
+
├─ 完整来源分页
|
|
32
|
+
├─ 站点页面发现
|
|
33
|
+
├─ 研究计划
|
|
34
|
+
└─ 配置诊断
|
|
32
35
|
```
|
|
33
36
|
|
|
34
37
|
插件继续使用 DSH 原有的 `web_search` 名称,不会再增加第二个普通搜索入口。在本来可以使用 `web_search` 的 Agent 中,插件提供增强后的搜索;如果某个 Agent 已经禁用网页搜索,插件不会强行重新开启。
|
|
35
38
|
|
|
36
39
|
普通搜索以 Grok 为主。其他服务只负责补充文档、来源或网页内容,不会替代 Grok 的主搜索位置。
|
|
37
40
|
|
|
38
|
-
###
|
|
41
|
+
### 固定模型工具 surface
|
|
39
42
|
|
|
40
|
-
|
|
43
|
+
默认使用 `progressive`。在 DSH 未另行限制的 Agent 中,插件在初始步骤和后续步骤提供的五个固定搜索入口(Native tool / Code Mode SDK)是:
|
|
41
44
|
|
|
42
|
-
| 工具 |
|
|
45
|
+
| 工具 | 调用方式与用途 |
|
|
43
46
|
| --- | --- |
|
|
44
|
-
| `web_search` |
|
|
45
|
-
| `docs_search` |
|
|
46
|
-
| `web_extract` |
|
|
47
|
-
| `search_tools` |
|
|
47
|
+
| `web_search` | 直接调用。使用 Grok 生成通用搜索的主要回答,并按搜索类型补充其他来源 |
|
|
48
|
+
| `docs_search` | 直接调用。检索库、框架、SDK、API 和源码仓库文档 |
|
|
49
|
+
| `web_extract` | 直接调用。读取指定网页正文,用于核对搜索摘要中的重要内容 |
|
|
50
|
+
| `search_tools` | 直接调用。按需返回延迟能力的 operation manifest,不注册新的模型工具 |
|
|
51
|
+
| `search_call` | 固定网关。通过 `search_call({ operation, arguments })` 调用已经激活的延迟 operation |
|
|
48
52
|
|
|
49
|
-
|
|
53
|
+
这里的“延迟”指 operation 是否处于 active 状态,而不是工具是否出现在列表中。延迟 operation 的参数 schema 通过工具结果中的 manifest 披露;输出 schema 只保存在内部 registry,用于校验规范执行结果,不会追加到模型历史,也不会作为独立工具加入模型 surface。
|
|
50
54
|
|
|
51
55
|
### 渐进式披露
|
|
52
56
|
|
|
53
57
|
`search_tools` 一次可以选择一到五组能力:
|
|
54
58
|
|
|
55
|
-
| 能力组 |
|
|
59
|
+
| 能力组 | 按需返回 manifest 的 operation | 适用场景 |
|
|
56
60
|
| --- | --- | --- |
|
|
57
61
|
| `context7` | `context7_resolve_library_id`、`context7_query_docs`、`context7_get_library_docs`、`context7_get_cached_doc_raw` | 需要精确选择库版本或进一步读取 Context7 文档 |
|
|
58
62
|
| `sources` | `search_sources` | 搜索结果中的来源较多,需要继续分页读取完整来源 |
|
|
@@ -60,23 +64,24 @@ DSH Agent
|
|
|
60
64
|
| `planning` | `research_plan` | 明确要求深度研究、多来源核对或复杂比较时先制定计划 |
|
|
61
65
|
| `diagnostics` | `search_diagnostics` | 用户明确要求检查搜索配置或连接状态 |
|
|
62
66
|
|
|
63
|
-
|
|
67
|
+
披露与调用遵循以下规则:
|
|
64
68
|
|
|
65
|
-
1.
|
|
66
|
-
2.
|
|
67
|
-
3. `
|
|
68
|
-
4.
|
|
69
|
-
5.
|
|
69
|
+
1. `search_tools` 返回所请求能力组的 operation manifest,其中包含真实的参数 schema 和 `search_call` 路由,但不包含内部输出 schema;它不会增加、删除或改写模型工具 schema。`search_call` 仍使用 registry 保存的输出 schema 校验规范结果。
|
|
70
|
+
2. 在 `progressive` 模式下,新披露的能力组从下一模型 step 开始 active;同一步内提前调用会失败。激活范围属于当前 Agent,重复请求会再次返回同一 manifest,但不会创建第二套状态或入口。
|
|
71
|
+
3. 在 `all` 模式下,唯一变化是所有延迟 operation 从一开始就 active;`search_tools` 仍按需返回 manifest。`all` 不会“显示全部 12 个工具”,两种模式的五个模型工具及其 schema 完全相同。
|
|
72
|
+
4. 延迟 operation 只能通过 `search_call({ operation, arguments })` 调用,不能直接调用 `search_sources`、`web_map` 等名称;resident 的 `web_search`、`docs_search` 和 `web_extract` 仍然直接调用。
|
|
73
|
+
5. `web_search` 或 `docs_search` 成功返回 `source_ref` 时,插件会自动激活 `sources`,并在结果中追加 `search_sources` manifest;在 `progressive` 模式下可从下一 step 通过 `search_call` 使用它。
|
|
74
|
+
6. 固定 surface 仍受 DSH 原有 Preset、guard 和工具限制约束,插件不会绕过这些限制。
|
|
70
75
|
|
|
71
|
-
|
|
76
|
+
这种固定网关设计保留了按需披露,同时避免插件因披露状态变化而改写发送给 DeepSeek 的 system 文本、tool schema/顺序或 Code Mode SDK 前缀,从而消除插件自身造成的前缀变化。
|
|
72
77
|
|
|
73
78
|
### 一次完整搜索如何进行
|
|
74
79
|
|
|
75
|
-
1.
|
|
76
|
-
2.
|
|
77
|
-
3.
|
|
78
|
-
4.
|
|
79
|
-
5.
|
|
80
|
+
1. 通用问题直接调用 `web_search`,文档问题直接调用 `docs_search`。
|
|
81
|
+
2. 当搜索产生来源时,结果会包含可见来源、`source_ref` 和追加的 `search_sources` manifest,同时自动激活 `sources`。
|
|
82
|
+
3. 在下一 step 需要更多来源时,调用 `search_call({ operation: 'search_sources', arguments: { source_ref, offset: 0, limit: 20, format: 'compact' } })` 分页读取,而不是直接调用 `search_sources`。
|
|
83
|
+
4. 对重要结论,选择权威链接并直接调用 `web_extract` 获取网页正文。
|
|
84
|
+
5. 如果任务需要站点内发现、研究计划、精细 Context7 查询或连接检查,先调用例如 `search_tools({ capabilities: ['site_map'] })` 取得 manifest;`progressive` 模式从下一 step、`all` 模式立即通过 `search_call({ operation: 'web_map', arguments: { url: 'https://example.com' } })` 调用相应 operation。
|
|
80
85
|
6. 最终回答综合主搜索、补充来源和已经读取的网页正文,并保留来源链接。
|
|
81
86
|
|
|
82
87
|
`source_ref` 只是完整来源列表的引用,不等同于网页正文;重要事实仍应通过 `web_extract` 读取原页面后再下结论。
|
|
@@ -147,7 +152,7 @@ dsh web
|
|
|
147
152
|
|
|
148
153
|
不确定时保持默认值即可,也可以在单次搜索中临时选择其他设置。
|
|
149
154
|
|
|
150
|
-
“工具披露模式”建议保持 `progressive
|
|
155
|
+
“工具披露模式”建议保持 `progressive`,让延迟 operation 按需披露并从下一模型 step 激活;`all` 只让所有延迟 operation 从一开始处于 active 状态。两种模式都保留同一组五个模型工具及相同 schema,不会显示额外的独立工具。
|
|
151
156
|
|
|
152
157
|
### 4. 配置网页代理
|
|
153
158
|
|
|
@@ -210,9 +215,3 @@ dsh plugin --profile web add dsh-search-enhance@latest
|
|
|
210
215
|
```bash
|
|
211
216
|
dsh plugin --profile web remove dsh-search-enhance
|
|
212
217
|
```
|
|
213
|
-
|
|
214
|
-
## 使用提示
|
|
215
|
-
|
|
216
|
-
- 网页正文读取不会执行页面 JavaScript,也不能处理登录、验证码或浏览器会话。
|
|
217
|
-
- 网页读取可以访问 DSH 所在机器能够访问的地址;在敏感网络中只处理可信链接,并配合网络访问限制。
|
|
218
|
-
- 如果配置页面没有出现,先用 `dsh plugin --profile web list --depth 0` 确认 npm 包已经安装,然后重启 DSH 并强制刷新浏览器。
|
package/lib/config.d.ts
CHANGED
|
@@ -8,7 +8,7 @@ export type SearchDepth = (typeof SEARCH_DEPTHS)[number];
|
|
|
8
8
|
export declare const TOOL_DISCOVERY_MODES: readonly ["progressive", "all"];
|
|
9
9
|
export type ToolDiscoveryMode = (typeof TOOL_DISCOVERY_MODES)[number];
|
|
10
10
|
export interface ToolDiscoveryConfig {
|
|
11
|
-
/** progressive
|
|
11
|
+
/** progressive gates deferred operations per Agent; all activates every operation. */
|
|
12
12
|
readonly mode: ToolDiscoveryMode;
|
|
13
13
|
}
|
|
14
14
|
/** Stable high-level docs_search default; the Settings cap is independently bounded. */
|
package/lib/index.js
CHANGED
|
@@ -14,8 +14,8 @@ import { TavilySearchProvider } from './providers/tavily.js';
|
|
|
14
14
|
import { TavilyExtractProvider } from './providers/tavily-extract.js';
|
|
15
15
|
import { TavilyMapProvider } from './providers/tavily-map.js';
|
|
16
16
|
import { SOURCE_RECORD_DOMAIN_SPEC, SOURCE_RECORD_TABLE_NAME, SearchEnhanceSourceService, SourceRecordStore, } from './source-storage/index.js';
|
|
17
|
-
import {
|
|
18
|
-
import { ForegroundOperationScope, createContext7Tools, createDocsSearchTool,
|
|
17
|
+
import { foldEffectiveToolDisclosureEvents, installAgentToolDisclosure, } from './tool-discovery/index.js';
|
|
18
|
+
import { ForegroundOperationScope, DeferredOperationRegistry, createContext7Tools, createDocsSearchTool, createResearchPlanTool, createSearchCallTool, createSearchDiagnosticsTool, createSearchSourcesTool, createSearchToolsTool, createWebExtractTool, createWebMapTool, createWebSearchTool, } from './tools/index.js';
|
|
19
19
|
import { WebExtractOrchestrator } from './web-extract/orchestrator.js';
|
|
20
20
|
import { installWebConfigBridge } from './web-config/host.js';
|
|
21
21
|
export const name = 'search-enhance';
|
|
@@ -94,25 +94,12 @@ export async function apply(ctx, config) {
|
|
|
94
94
|
direct: new DirectFetchProvider(),
|
|
95
95
|
getConfig,
|
|
96
96
|
});
|
|
97
|
-
const
|
|
98
|
-
|
|
99
|
-
operations,
|
|
100
|
-
orchestrator,
|
|
101
|
-
sources: ctx.searchEnhanceSources,
|
|
102
|
-
});
|
|
103
|
-
const globalToolDefinitions = [
|
|
104
|
-
createDocsSearchTool({
|
|
97
|
+
const deferredOperationDefinitions = [
|
|
98
|
+
...createContext7Tools({
|
|
105
99
|
documentation,
|
|
106
100
|
getConfig,
|
|
107
101
|
operations,
|
|
108
|
-
sources: ctx.searchEnhanceSources,
|
|
109
|
-
}),
|
|
110
|
-
createWebExtractTool({
|
|
111
|
-
getConfig,
|
|
112
|
-
operations,
|
|
113
|
-
orchestrator: webExtract,
|
|
114
102
|
}),
|
|
115
|
-
createSearchToolsTool({ mode: effective.toolDiscovery.mode }),
|
|
116
103
|
createSearchSourcesTool({
|
|
117
104
|
getConfig,
|
|
118
105
|
operations,
|
|
@@ -125,7 +112,9 @@ export async function apply(ctx, config) {
|
|
|
125
112
|
}),
|
|
126
113
|
createResearchPlanTool({
|
|
127
114
|
getConfig,
|
|
128
|
-
isWebMapAvailable: agent =>
|
|
115
|
+
isWebMapAvailable: agent => agent !== undefined && (effective.toolDiscovery.mode === 'all'
|
|
116
|
+
|| foldEffectiveToolDisclosureEvents(agent.session.events).activeGroups
|
|
117
|
+
.includes('site_map')),
|
|
129
118
|
operations,
|
|
130
119
|
}),
|
|
131
120
|
createSearchDiagnosticsTool({
|
|
@@ -133,22 +122,42 @@ export async function apply(ctx, config) {
|
|
|
133
122
|
operations,
|
|
134
123
|
reporter: diagnostics,
|
|
135
124
|
}),
|
|
136
|
-
|
|
125
|
+
];
|
|
126
|
+
const deferredOperations = new DeferredOperationRegistry(deferredOperationDefinitions);
|
|
127
|
+
const sourceOperationNotice = deferredOperations.renderCapabilityDisclosure('sources');
|
|
128
|
+
const webSearchDefinition = createWebSearchTool({
|
|
129
|
+
getConfig,
|
|
130
|
+
operations,
|
|
131
|
+
orchestrator,
|
|
132
|
+
sourceOperationNotice,
|
|
133
|
+
sources: ctx.searchEnhanceSources,
|
|
134
|
+
});
|
|
135
|
+
const residentToolDefinitions = [
|
|
136
|
+
createDocsSearchTool({
|
|
137
137
|
documentation,
|
|
138
138
|
getConfig,
|
|
139
139
|
operations,
|
|
140
|
+
sourceOperationNotice,
|
|
141
|
+
sources: ctx.searchEnhanceSources,
|
|
142
|
+
}),
|
|
143
|
+
createWebExtractTool({
|
|
144
|
+
getConfig,
|
|
145
|
+
operations,
|
|
146
|
+
orchestrator: webExtract,
|
|
147
|
+
}),
|
|
148
|
+
createSearchToolsTool({
|
|
149
|
+
mode: effective.toolDiscovery.mode,
|
|
150
|
+
registry: deferredOperations,
|
|
151
|
+
}),
|
|
152
|
+
createSearchCallTool({
|
|
153
|
+
mode: effective.toolDiscovery.mode,
|
|
154
|
+
registry: deferredOperations,
|
|
140
155
|
}),
|
|
141
156
|
];
|
|
142
|
-
for (const definition of
|
|
157
|
+
for (const definition of residentToolDefinitions)
|
|
143
158
|
ctx.tools.register(definition);
|
|
144
|
-
installAgentToolDisclosure(ctx, {
|
|
145
|
-
|
|
146
|
-
deferredToolNames: globalToolDefinitions
|
|
147
|
-
.map(definition => definition.name)
|
|
148
|
-
.filter(isDeferredToolName),
|
|
149
|
-
webSearchDefinition,
|
|
150
|
-
});
|
|
151
|
-
registerToolDiscoveryGuidance(ctx, effective.toolDiscovery.mode);
|
|
159
|
+
installAgentToolDisclosure(ctx, { webSearchDefinition });
|
|
160
|
+
registerToolDiscoveryGuidance(ctx);
|
|
152
161
|
installWebConfigBridge(ctx);
|
|
153
162
|
}
|
|
154
163
|
//# sourceMappingURL=index.js.map
|
|
@@ -4,9 +4,9 @@ import type { DocsSearchOutput, WebSearchOutput, SearchDiagnosticsOutput, Search
|
|
|
4
4
|
* limit is carried in the canonical value, so replay never consults Settings,
|
|
5
5
|
* a cache, the clock, or network state.
|
|
6
6
|
*/
|
|
7
|
-
export declare function renderWebSearchText(value: WebSearchOutput): string;
|
|
7
|
+
export declare function renderWebSearchText(value: WebSearchOutput, sourceOperationNotice?: string): string;
|
|
8
8
|
/** Pure Native projection for docs_search; it never reads Settings, cache, storage, or the network. */
|
|
9
|
-
export declare function renderDocsSearchText(value: DocsSearchOutput): string;
|
|
9
|
+
export declare function renderDocsSearchText(value: DocsSearchOutput, sourceOperationNotice?: string): string;
|
|
10
10
|
/** Pure model-text projection for one private-storage source page. */
|
|
11
11
|
export declare function renderSearchSourcesText(value: SearchSourcesOutput): string;
|
|
12
12
|
/**
|
|
@@ -3,6 +3,37 @@ const DISCOVERY_NOTICE = 'Evidence level: discovery. Snippets are discovery meta
|
|
|
3
3
|
function inline(value) {
|
|
4
4
|
return value.replace(/\s+/gu, ' ').trim();
|
|
5
5
|
}
|
|
6
|
+
function renderWithBoundedNotice(text, notice, maximumBytes, fallbackNotice) {
|
|
7
|
+
if (notice === undefined || notice.length === 0) {
|
|
8
|
+
return truncateUtf8(text, maximumBytes).text;
|
|
9
|
+
}
|
|
10
|
+
const separator = '\n\n';
|
|
11
|
+
const complete = `${text}${separator}${notice}`;
|
|
12
|
+
if (utf8ByteLength(complete) <= maximumBytes)
|
|
13
|
+
return complete;
|
|
14
|
+
const noticeBytes = utf8ByteLength(notice);
|
|
15
|
+
if (noticeBytes > maximumBytes) {
|
|
16
|
+
// Never expose a partial JSON capability manifest. A short plain-text
|
|
17
|
+
// fallback can still point the model at the durable source reference.
|
|
18
|
+
return truncateUtf8(fallbackNotice ?? notice, maximumBytes).text;
|
|
19
|
+
}
|
|
20
|
+
const prefixBudget = maximumBytes - noticeBytes - utf8ByteLength(separator);
|
|
21
|
+
if (prefixBudget <= 0)
|
|
22
|
+
return notice;
|
|
23
|
+
const prefix = truncateUtf8(text, prefixBudget).text;
|
|
24
|
+
return prefix.length === 0 ? notice : `${prefix}${separator}${notice}`;
|
|
25
|
+
}
|
|
26
|
+
function sourceDisclosureTail(sourceRef, operationNotice) {
|
|
27
|
+
if (sourceRef === undefined)
|
|
28
|
+
return undefined;
|
|
29
|
+
return [
|
|
30
|
+
`Source reference: ${sourceRef}`,
|
|
31
|
+
operationNotice,
|
|
32
|
+
].filter((line) => line !== undefined && line.length > 0).join('\n\n');
|
|
33
|
+
}
|
|
34
|
+
function sourceDisclosureFallback(sourceRef) {
|
|
35
|
+
return `Source reference: ${sourceRef}\n\nCall search_tools({ capabilities: ["sources"] }) to retrieve the search_sources manifest.`;
|
|
36
|
+
}
|
|
6
37
|
function warningText(warning) {
|
|
7
38
|
const subject = [warning.provider, warning.capability].filter(Boolean).join('/');
|
|
8
39
|
const suffix = [subject, warning.error_kind].filter(Boolean).join(', ');
|
|
@@ -58,11 +89,7 @@ function sourceSection(value) {
|
|
|
58
89
|
return lines.join('\n');
|
|
59
90
|
}
|
|
60
91
|
function sourceSummary(value) {
|
|
61
|
-
|
|
62
|
-
if (value.source_ref !== undefined) {
|
|
63
|
-
lines.push(`Source reference: ${value.source_ref}`);
|
|
64
|
-
}
|
|
65
|
-
return lines.join('\n');
|
|
92
|
+
return `Sources shown: ${value.returned_sources}/${value.total_sources}`;
|
|
66
93
|
}
|
|
67
94
|
function limitationsSection(value) {
|
|
68
95
|
const lines = [];
|
|
@@ -80,7 +107,7 @@ function limitationsSection(value) {
|
|
|
80
107
|
* limit is carried in the canonical value, so replay never consults Settings,
|
|
81
108
|
* a cache, the clock, or network state.
|
|
82
109
|
*/
|
|
83
|
-
export function renderWebSearchText(value) {
|
|
110
|
+
export function renderWebSearchText(value, sourceOperationNotice) {
|
|
84
111
|
const sections = [
|
|
85
112
|
answerSection(value),
|
|
86
113
|
sourceSection(value),
|
|
@@ -93,7 +120,7 @@ export function renderWebSearchText(value) {
|
|
|
93
120
|
&& value.model_text_max_bytes > 0
|
|
94
121
|
? value.model_text_max_bytes
|
|
95
122
|
: 0;
|
|
96
|
-
return
|
|
123
|
+
return renderWithBoundedNotice(complete, sourceDisclosureTail(value.source_ref, sourceOperationNotice), maximumBytes, value.source_ref === undefined ? undefined : sourceDisclosureFallback(value.source_ref));
|
|
97
124
|
}
|
|
98
125
|
function docsWarningText(warning) {
|
|
99
126
|
const subject = [warning.provider, warning.path].filter(Boolean).join('/');
|
|
@@ -173,7 +200,7 @@ function docsSourceSection(value) {
|
|
|
173
200
|
return lines.join('\n');
|
|
174
201
|
}
|
|
175
202
|
/** Pure Native projection for docs_search; it never reads Settings, cache, storage, or the network. */
|
|
176
|
-
export function renderDocsSearchText(value) {
|
|
203
|
+
export function renderDocsSearchText(value, sourceOperationNotice) {
|
|
177
204
|
const status = value.state === 'partial' ? 'partial' : 'complete';
|
|
178
205
|
const explanation = value.state === 'partial'
|
|
179
206
|
? 'Some documentation paths failed or stale cache data was used; available discovery results follow.'
|
|
@@ -188,8 +215,7 @@ export function renderDocsSearchText(value) {
|
|
|
188
215
|
[
|
|
189
216
|
`Sources shown: ${value.returned_sources}/${value.total_sources}`,
|
|
190
217
|
`Snippets shown: ${value.returned_snippets}/${value.total_snippets}`,
|
|
191
|
-
|
|
192
|
-
].filter((line) => line !== undefined).join('\n'),
|
|
218
|
+
].join('\n'),
|
|
193
219
|
[
|
|
194
220
|
docsCachePathText('Context7 resolve cache', value.cache.resolve),
|
|
195
221
|
docsCachePathText('Context7 docs cache', value.cache.docs),
|
|
@@ -211,7 +237,7 @@ export function renderDocsSearchText(value) {
|
|
|
211
237
|
&& value.model_text_max_bytes > 0
|
|
212
238
|
? value.model_text_max_bytes
|
|
213
239
|
: 0;
|
|
214
|
-
return
|
|
240
|
+
return renderWithBoundedNotice(complete, sourceDisclosureTail(value.source_ref, sourceOperationNotice), maximumBytes, value.source_ref === undefined ? undefined : sourceDisclosureFallback(value.source_ref));
|
|
215
241
|
}
|
|
216
242
|
function renderPageSources(page) {
|
|
217
243
|
if (page.sources.length === 0)
|
|
@@ -490,7 +516,7 @@ function completeSearchDiagnosticsText(value) {
|
|
|
490
516
|
`Search API protocol: ${value.configuration.search_api_protocol}`,
|
|
491
517
|
`Search model configured: ${value.configuration.search_model_configured}`,
|
|
492
518
|
`Thinking/fallback: ${value.configuration.thinking_level}/${value.configuration.fallback_mode}`,
|
|
493
|
-
`
|
|
519
|
+
`Configured deferred operations: web_map=${value.configuration.web_map_enabled}, research_plan=${value.configuration.research_plan_enabled}, diagnostics=${value.configuration.diagnostics_enabled} (invoke active operations through search_call)`,
|
|
494
520
|
`Search routes enabled: tavily=${value.configuration.tavily_search_enabled}, firecrawl=${value.configuration.firecrawl_search_enabled}`,
|
|
495
521
|
`Extract routes enabled: tavily=${value.configuration.tavily_extract_enabled}, firecrawl=${value.configuration.firecrawl_scrape_enabled}, smart_direct=${value.configuration.smart_direct_enabled}, direct=${value.configuration.direct_enabled}`,
|
|
496
522
|
].join('\n');
|
|
@@ -1,5 +1,6 @@
|
|
|
1
1
|
import type { Context } from '@deepseek-ai/cordis';
|
|
2
|
-
|
|
3
|
-
|
|
4
|
-
|
|
2
|
+
export declare const TOOL_DISCOVERY_GUIDANCE: string;
|
|
3
|
+
export declare const EVIDENCE_DISCIPLINE_GUIDANCE: string;
|
|
4
|
+
/** Register deterministic guidance that never reads Agent state or tool visibility. */
|
|
5
|
+
export declare function registerToolDiscoveryGuidance(ctx: Context): void;
|
|
5
6
|
//# sourceMappingURL=tool-discovery.d.ts.map
|
|
@@ -1,49 +1,26 @@
|
|
|
1
|
-
|
|
2
|
-
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
webExtractVisible
|
|
16
|
-
? 'Before asserting decisive factual or causal conclusions, inspect selected authoritative URLs with web_extract; never present an inferred mechanism as source-stated fact, and label unestablished mechanisms as inference or unconfirmed.'
|
|
17
|
-
: 'Never present an inferred mechanism as source-stated fact; label unestablished mechanisms as inference or unconfirmed.',
|
|
18
|
-
].join('\n');
|
|
19
|
-
}
|
|
20
|
-
/** Register scope-aware disclosure and evidence guidance for Search Enhance tools. */
|
|
21
|
-
export function registerToolDiscoveryGuidance(ctx, mode) {
|
|
1
|
+
export const TOOL_DISCOVERY_GUIDANCE = [
|
|
2
|
+
'Search Enhance keeps a fixed model-facing surface: web_search, docs_search, web_extract, search_tools, and search_call.',
|
|
3
|
+
'Use search_tools only when the resident search tools cannot complete the task. It returns append-only capability and operation manifests; do not activate every capability preemptively.',
|
|
4
|
+
'Run a manifested deferred operation with search_call({ operation, arguments }). In progressive mode, a newly disclosed capability is callable on the next model step; in all mode, deferred operations are active immediately. search_call fails closed while an operation is inactive.',
|
|
5
|
+
'Activate planning only for explicit deep research, multi-source verification, or complex comparison. Activate diagnostics only when the user asks about Provider configuration or connectivity.',
|
|
6
|
+
'A successful web_search or docs_search result with source_ref activates the sources capability. Use the appended search_sources manifest, or call search_tools for sources to replay that manifest before search_call.',
|
|
7
|
+
].join('\n');
|
|
8
|
+
export const EVIDENCE_DISCIPLINE_GUIDANCE = [
|
|
9
|
+
'For current or external factual questions, start with one focused web_search (use docs_search for SDK/API documentation); do not inspect local files, settings, sessions, or credentials unless the user explicitly asks about local state.',
|
|
10
|
+
'Treat web_search/docs_search answers, snippets, and source metadata as discovery, not claim-level evidence.',
|
|
11
|
+
'Before asserting decisive factual or causal conclusions, inspect selected authoritative URLs with web_extract; never present an inferred mechanism as source-stated fact, and label unestablished mechanisms as inference or unconfirmed.',
|
|
12
|
+
].join('\n');
|
|
13
|
+
/** Register deterministic guidance that never reads Agent state or tool visibility. */
|
|
14
|
+
export function registerToolDiscoveryGuidance(ctx) {
|
|
22
15
|
ctx.systemPrompt.section({
|
|
23
16
|
name: 'search-enhance:tool-discovery',
|
|
24
17
|
order: 121,
|
|
25
|
-
text:
|
|
26
|
-
if (mode === 'all'
|
|
27
|
-
|| context.agent === undefined
|
|
28
|
-
|| ctx.tools.get('search_tools', context.scope) === undefined)
|
|
29
|
-
return '';
|
|
30
|
-
const active = new Set(foldToolDisclosureEvents(context.agent.session.events).activeGroups);
|
|
31
|
-
const remaining = CAPABILITY_GROUPS.filter(group => !active.has(group));
|
|
32
|
-
if (remaining.length === 0)
|
|
33
|
-
return '';
|
|
34
|
-
return [
|
|
35
|
-
`Additional Search Enhance capabilities are deferred for this Agent: ${remaining.join(', ')}.`,
|
|
36
|
-
'Use search_tools only when web_search, docs_search, and web_extract cannot complete the task; do not activate every group preemptively.',
|
|
37
|
-
'Activate planning only for explicit deep research, multi-source verification, or complex comparison. Activate diagnostics only when the user asks about Provider configuration or connectivity.',
|
|
38
|
-
'A successful disclosure applies on the next model step and cannot bypass another Preset restriction. In Code Mode, call search_tools through the current run_code SDK and wait for the next step before using newly disclosed bindings.',
|
|
39
|
-
'Successful source-producing searches disclose sources automatically, so do not request it again when already active.',
|
|
40
|
-
].join('\n');
|
|
41
|
-
},
|
|
18
|
+
text: TOOL_DISCOVERY_GUIDANCE,
|
|
42
19
|
});
|
|
43
20
|
ctx.systemPrompt.section({
|
|
44
21
|
name: 'search-enhance:evidence-discipline',
|
|
45
22
|
order: 122,
|
|
46
|
-
text:
|
|
23
|
+
text: EVIDENCE_DISCIPLINE_GUIDANCE,
|
|
47
24
|
});
|
|
48
25
|
}
|
|
49
26
|
//# sourceMappingURL=tool-discovery.js.map
|
|
@@ -9,8 +9,10 @@ export interface DirectFetchProviderDependencies extends DirectHttpDependencies
|
|
|
9
9
|
/**
|
|
10
10
|
* Production `direct` route. It performs bounded Node HTTP(S) requests without
|
|
11
11
|
* JavaScript, cookies, login flows, CAPTCHA handling, browser emulation, or any
|
|
12
|
-
* destination-network classification.
|
|
13
|
-
*
|
|
12
|
+
* destination-network classification. Recognizable anti-bot interstitials are
|
|
13
|
+
* unavailable rather than returned as direct page evidence. Localhost, private,
|
|
14
|
+
* metadata, reserved, and DNS-rebinding targets are deliberately not blocked
|
|
15
|
+
* by this adapter.
|
|
14
16
|
*/
|
|
15
17
|
export declare class DirectFetchProvider implements WebExtractAdapter {
|
|
16
18
|
readonly route: "direct";
|
|
@@ -1,5 +1,6 @@
|
|
|
1
1
|
import { abortableDelay, exponentialBackoffMs, isProviderError, ProviderError, runWithTimeout, throwIfAborted, } from '../provider-runtime/index.js';
|
|
2
2
|
import { normalizeWebExtractUrl } from '../web-extract/url.js';
|
|
3
|
+
import { isLikelyAntiBotChallenge } from './web-extract-common.js';
|
|
3
4
|
import { isWebExtractFormat } from '../web-extract/types.js';
|
|
4
5
|
import { inspectDirectHtml, projectDirectContent, } from './direct-content.js';
|
|
5
6
|
import { fetchDirectHttpHop, } from './direct-http.js';
|
|
@@ -35,8 +36,10 @@ function delayForRetry(failedRetryNumber, error, policy, random) {
|
|
|
35
36
|
/**
|
|
36
37
|
* Production `direct` route. It performs bounded Node HTTP(S) requests without
|
|
37
38
|
* JavaScript, cookies, login flows, CAPTCHA handling, browser emulation, or any
|
|
38
|
-
* destination-network classification.
|
|
39
|
-
*
|
|
39
|
+
* destination-network classification. Recognizable anti-bot interstitials are
|
|
40
|
+
* unavailable rather than returned as direct page evidence. Localhost, private,
|
|
41
|
+
* metadata, reserved, and DNS-rebinding targets are deliberately not blocked
|
|
42
|
+
* by this adapter.
|
|
40
43
|
*/
|
|
41
44
|
export class DirectFetchProvider {
|
|
42
45
|
route = PROVIDER;
|
|
@@ -89,6 +92,9 @@ export class DirectFetchProvider {
|
|
|
89
92
|
}
|
|
90
93
|
const contentInput = this.contentInput(response, input);
|
|
91
94
|
const inspection = inspectDirectHtml(contentInput);
|
|
95
|
+
if (isLikelyAntiBotChallenge(inspection.scanText, response.statusCode)) {
|
|
96
|
+
throw new ProviderError({ capability: CAPABILITY, kind: 'unavailable', provider: PROVIDER });
|
|
97
|
+
}
|
|
92
98
|
const navigationTarget = inspection.metaRefreshUrl ?? inspection.alternateUrl;
|
|
93
99
|
if (navigationTarget !== undefined) {
|
|
94
100
|
target = this.nextTarget(navigationTarget, response.url, input.config.webExtract.maxUrlCharacters, directConfig.maxRedirects, seen, budget);
|
|
@@ -5,6 +5,7 @@ import { pipeline } from 'node:stream/promises';
|
|
|
5
5
|
import { createBrotliDecompress, createGunzip, createInflate, } from 'node:zlib';
|
|
6
6
|
import { isProviderError, parseRetryAfterMs, ProviderError, providerHttpError, RETRYABLE_HTTP_STATUSES, throwIfAborted, } from '../provider-runtime/index.js';
|
|
7
7
|
import { isDirectTextLikeContentType, } from './direct-content.js';
|
|
8
|
+
import { isLikelyAntiBotChallenge } from './web-extract-common.js';
|
|
8
9
|
const PROVIDER = 'direct';
|
|
9
10
|
const CAPABILITY = 'web_extract';
|
|
10
11
|
const HTTP_REDIRECT_STATUSES = new Set([301, 302, 303, 307, 308]);
|
|
@@ -468,6 +469,9 @@ export async function fetchDirectHttpHop(input, dependencies = {}) {
|
|
|
468
469
|
if (input.config.proxyUrl !== undefined && statusCode === 407) {
|
|
469
470
|
throw providerHttpError({ capability: CAPABILITY, provider: PROVIDER, status: statusCode });
|
|
470
471
|
}
|
|
472
|
+
if (isLikelyAntiBotChallenge('', statusCode, scalarHeader(response, 'cf-mitigated'))) {
|
|
473
|
+
throw new ProviderError({ capability: CAPABILITY, kind: 'unavailable', provider: PROVIDER });
|
|
474
|
+
}
|
|
471
475
|
const contentType = scalarHeader(response, 'content-type');
|
|
472
476
|
const contentLength = parseContentLength(scalarHeader(response, 'content-length'));
|
|
473
477
|
const contentDisposition = scalarHeader(response, 'content-disposition');
|
|
@@ -11,7 +11,9 @@ export interface FirecrawlScrapeProviderDependencies extends ProviderHttpDepende
|
|
|
11
11
|
export declare function firecrawlFormatForWebExtract(format: WebExtractFormat): FirecrawlFormat | undefined;
|
|
12
12
|
/**
|
|
13
13
|
* Parse one Firecrawl v2 scrape envelope. The requested format is selected from
|
|
14
|
-
* `data[format]`;
|
|
14
|
+
* `data[format]`; recognizable anti-bot challenge content throws a fixed
|
|
15
|
+
* unavailable error instead of consuming empty-content retries, and `metadata`
|
|
16
|
+
* is projected only for explicit scalar fields.
|
|
15
17
|
*/
|
|
16
18
|
export declare function parseFirecrawlScrapeResponse(body: string, format: Extract<WebExtractFormat, 'markdown' | 'html' | 'raw'>, maximumUrlCharacters: number, maximumContentCharacters: number): WebExtractAdapterResult | undefined;
|
|
17
19
|
/** Registration-free Firecrawl v2 `POST /scrape` adapter. */
|
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
import { boundedExtractedContent, remoteMetadata, responseRecord, } from './web-extract-common.js';
|
|
1
|
+
import { boundedExtractedContent, isLikelyAntiBotChallenge, remoteMetadata, responseRecord, } from './web-extract-common.js';
|
|
2
2
|
import { parseProviderJson, providerEndpoint, resolveOptionalCredential, } from './helpers.js';
|
|
3
3
|
import { ProviderError, throwIfAborted, } from '../provider-runtime/index.js';
|
|
4
4
|
import { ProviderHttpClient, } from './http.js';
|
|
@@ -15,7 +15,9 @@ export function firecrawlFormatForWebExtract(format) {
|
|
|
15
15
|
}
|
|
16
16
|
/**
|
|
17
17
|
* Parse one Firecrawl v2 scrape envelope. The requested format is selected from
|
|
18
|
-
* `data[format]`;
|
|
18
|
+
* `data[format]`; recognizable anti-bot challenge content throws a fixed
|
|
19
|
+
* unavailable error instead of consuming empty-content retries, and `metadata`
|
|
20
|
+
* is projected only for explicit scalar fields.
|
|
19
21
|
*/
|
|
20
22
|
export function parseFirecrawlScrapeResponse(body, format, maximumUrlCharacters, maximumContentCharacters) {
|
|
21
23
|
const root = responseRecord(parseProviderJson(body, PROVIDER, CAPABILITY), PROVIDER);
|
|
@@ -31,14 +33,19 @@ export function parseFirecrawlScrapeResponse(body, format, maximumUrlCharacters,
|
|
|
31
33
|
if (firecrawlFormat === undefined)
|
|
32
34
|
return undefined;
|
|
33
35
|
const data = root.data;
|
|
34
|
-
const
|
|
36
|
+
const metadata = remoteMetadata(isRecord(data.metadata) ? data.metadata : undefined, maximumUrlCharacters);
|
|
37
|
+
const extracted = data[firecrawlFormat];
|
|
38
|
+
if (typeof extracted === 'string'
|
|
39
|
+
&& isLikelyAntiBotChallenge(extracted, metadata.statusCode)) {
|
|
40
|
+
throw new ProviderError({ capability: CAPABILITY, kind: 'unavailable', provider: PROVIDER });
|
|
41
|
+
}
|
|
42
|
+
const content = boundedExtractedContent(extracted, maximumContentCharacters);
|
|
35
43
|
if (content === undefined)
|
|
36
44
|
return undefined;
|
|
37
|
-
const metadata = isRecord(data.metadata) ? data.metadata : undefined;
|
|
38
45
|
return {
|
|
39
46
|
content: content.content,
|
|
40
47
|
truncated: content.truncated,
|
|
41
|
-
...
|
|
48
|
+
...metadata,
|
|
42
49
|
};
|
|
43
50
|
}
|
|
44
51
|
function isRecord(value) {
|
|
@@ -109,6 +109,7 @@ async function smartDirectChildMain() {
|
|
|
109
109
|
'content-length',
|
|
110
110
|
'content-disposition',
|
|
111
111
|
'content-encoding',
|
|
112
|
+
'cf-mitigated',
|
|
112
113
|
'location',
|
|
113
114
|
'retry-after',
|
|
114
115
|
]);
|
|
@@ -128,8 +129,9 @@ async function smartDirectChildMain() {
|
|
|
128
129
|
finish({ error: 'invalid_response' }, 1);
|
|
129
130
|
return;
|
|
130
131
|
}
|
|
132
|
+
const challengeMitigated = selected['cf-mitigated']?.trim().toLowerCase() === 'challenge';
|
|
131
133
|
const contentLengthHeader = selected['content-length'];
|
|
132
|
-
if (contentLengthHeader !== undefined) {
|
|
134
|
+
if (!challengeMitigated && contentLengthHeader !== undefined) {
|
|
133
135
|
if (!/^\d+$/.test(contentLengthHeader)) {
|
|
134
136
|
await response.body?.cancel();
|
|
135
137
|
finish({ error: 'invalid_response' }, 1);
|
|
@@ -165,7 +167,8 @@ async function smartDirectChildMain() {
|
|
|
165
167
|
'text/x-markdown',
|
|
166
168
|
'application/markdown',
|
|
167
169
|
].includes(mime ?? '');
|
|
168
|
-
const shouldRead =
|
|
170
|
+
const shouldRead = !challengeMitigated
|
|
171
|
+
&& status >= 200
|
|
169
172
|
&& status <= 299
|
|
170
173
|
&& disposition !== 'attachment'
|
|
171
174
|
&& mimeSupported
|
|
@@ -2,6 +2,7 @@ import { PassThrough, Transform, Writable, } from 'node:stream';
|
|
|
2
2
|
import { pipeline } from 'node:stream/promises';
|
|
3
3
|
import { createBrotliDecompress, createGunzip, createInflate, } from 'node:zlib';
|
|
4
4
|
import { isProviderError, parseRetryAfterMs, ProviderError, providerHttpError, RETRYABLE_HTTP_STATUSES, throwIfAborted, } from '../provider-runtime/index.js';
|
|
5
|
+
import { isLikelyAntiBotChallenge } from './web-extract-common.js';
|
|
5
6
|
const PROVIDER = 'smart_direct';
|
|
6
7
|
const CAPABILITY = 'web_extract';
|
|
7
8
|
const HTTP_REDIRECT_STATUSES = new Set([301, 302, 303, 307, 308]);
|
|
@@ -223,6 +224,10 @@ export async function settleSmartDirectWreqResponse(response, input, dependencie
|
|
|
223
224
|
if (!Number.isInteger(statusCode) || statusCode < 100 || statusCode > 599) {
|
|
224
225
|
throw new ProviderError({ capability: CAPABILITY, kind: 'invalid_response', provider: PROVIDER });
|
|
225
226
|
}
|
|
227
|
+
if (isLikelyAntiBotChallenge('', statusCode, scalarHeader(response, 'cf-mitigated'))) {
|
|
228
|
+
await discardBody(response, input.signal);
|
|
229
|
+
return { kind: 'unavailable', statusCode, url: response.url };
|
|
230
|
+
}
|
|
226
231
|
const contentType = scalarHeader(response, 'content-type');
|
|
227
232
|
const contentLength = parseContentLength(scalarHeader(response, 'content-length'));
|
|
228
233
|
const contentDisposition = scalarHeader(response, 'content-disposition');
|
|
@@ -39,7 +39,8 @@ export declare function smartDirectMarkdownToText(markdown: string): string;
|
|
|
39
39
|
/**
|
|
40
40
|
* Production smart_direct route: wreq browser TLS/HTTP fingerprint transport,
|
|
41
41
|
* bounded linkedom DOM construction, and Defuddle readable-content cleaning.
|
|
42
|
-
*
|
|
42
|
+
* Recognizable anti-bot interstitials are unavailable rather than extracted as
|
|
43
|
+
* evidence. This is not browser automation and executes no page JavaScript.
|
|
43
44
|
*/
|
|
44
45
|
export declare class SmartDirectProvider implements WebExtractAdapter {
|
|
45
46
|
readonly route: "smart_direct";
|