@puwenhui/dsh-rag-kb 0.1.0 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +59 -60
- package/USAGE.zh.md +3 -0
- package/lib/client.js +60 -4
- package/lib/plugin.mjs +208 -48
- package/package.json +2 -2
package/README.md
CHANGED
|
@@ -1,21 +1,24 @@
|
|
|
1
1
|
# @puwenhui/dsh-rag-kb
|
|
2
2
|
|
|
3
|
-
DSH RAG 知识库插件:**文档导入 → 智能切块 → 本地向量化 →
|
|
3
|
+
DSH RAG 知识库插件:**文档导入 → 智能切块 → 本地向量化 → 混合检索(向量+FTS5)→ Reranker 精排 → 引用溯源**。
|
|
4
4
|
|
|
5
5
|
> 面向公司非技术人员的使用帮助请阅读 **[USAGE.zh.md](./USAGE.zh.md)**。
|
|
6
6
|
|
|
7
7
|
## 核心特性
|
|
8
8
|
|
|
9
|
-
-
|
|
10
|
-
-
|
|
11
|
-
-
|
|
9
|
+
- **混合检索**:语义向量(bge-small-zh)+ FTS5 trigram 关键词双路合并,同义/改述和精确词/错误码都能找到
|
|
10
|
+
- **Reranker 精排**(可选):bge-reranker-base cross-encoder,对混合检索结果做语义精排,噪音过滤更准
|
|
11
|
+
- **GPU 加速**:DirectML(RTX 4060 实测批量索引快 4 倍),无 GPU 自动回退 CPU
|
|
12
|
+
- **引用溯源**:脚注式引用(正文上标¹²³ + 末尾参考来源清单),agent 回答自带出处
|
|
12
13
|
- **会话语义管理**:所有操作(添加/删除/保存/检索)直接在对话里用大白话完成
|
|
13
|
-
-
|
|
14
|
+
- **多格式支持**:pdf / docx / md / txt / csv / json
|
|
15
|
+
- **完全本地**:模型本地运行,数据不出机器
|
|
16
|
+
- **目录自动索引**:配置目录后文件放入/变更自动入库
|
|
14
17
|
- **对话内容保存**:把对话中的结论/总结一键存入知识库
|
|
15
18
|
|
|
16
19
|
## 版本要求
|
|
17
20
|
|
|
18
|
-
- **DSH ≥ 0.2.0
|
|
21
|
+
- **DSH ≥ 0.2.0**
|
|
19
22
|
|
|
20
23
|
## 安装
|
|
21
24
|
|
|
@@ -23,25 +26,19 @@ DSH RAG 知识库插件:**文档导入 → 智能切块 → 本地向量化
|
|
|
23
26
|
dsh plugin --profile web add @puwenhui/dsh-rag-kb
|
|
24
27
|
```
|
|
25
28
|
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
## 配置
|
|
29
|
+
## 配置(设置 → 插件 →「知识库」标签页)
|
|
29
30
|
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
|
33
|
-
|
|
34
|
-
|
|
|
35
|
-
|
|
|
36
|
-
|
|
|
37
|
-
|
|
|
38
|
-
|
|
39
|
-
**首次使用**会自动下载中文向量化模型(约 100MB,下载后完全离线,模型缓存在 `$DSH_HOME/rag-kb/hf-cache/`)。
|
|
31
|
+
| 字段 | 默认 | 说明 |
|
|
32
|
+
|---|---|---|
|
|
33
|
+
| 知识库描述 | 空 | agent 检索前阅读,知道库里有什么 |
|
|
34
|
+
| 自动索引目录 | 空 | 放入文档自动入库(可选) |
|
|
35
|
+
| 检索返回条数 | 5 | 混合检索召回 top-K |
|
|
36
|
+
| 相似度阈值 | 0.3 | 0~1,过低噪音多 |
|
|
37
|
+
| 启用 Reranker | 关 | 开启后需下载 ~1GB 模型 |
|
|
38
|
+
| 重排后返回条数 | 3 | Reranker 精排后返回 top-N |
|
|
40
39
|
|
|
41
40
|
## 会话语义操作
|
|
42
41
|
|
|
43
|
-
所有知识库操作直接在对话里说,agent 自动调用对应工具:
|
|
44
|
-
|
|
45
42
|
| 你说 | agent 调用 |
|
|
46
43
|
|---|---|
|
|
47
44
|
| 「把 D:\docs 加进知识库」 | `knowledge_manage add` |
|
|
@@ -50,51 +47,53 @@ dsh plugin --profile web add @puwenhui/dsh-rag-kb
|
|
|
50
47
|
| 「删掉旧版手册」 | `knowledge_manage remove` |
|
|
51
48
|
| 「根据知识库回答:XXX」 | `knowledge_search` |
|
|
52
49
|
|
|
53
|
-
##
|
|
54
|
-
|
|
55
|
-
| 工具 | 说明 |
|
|
56
|
-
|---|---|
|
|
57
|
-
| `knowledge_search` | 语义检索:query + top_k → 带来源/相似度/原文的结果 |
|
|
58
|
-
| `knowledge_status` | 知识库状态:文档列表/配置/描述 |
|
|
59
|
-
| `knowledge_manage` | 管理操作:save / add / remove / reindex / describe |
|
|
60
|
-
|
|
61
|
-
## 技术架构
|
|
50
|
+
## 检索架构
|
|
62
51
|
|
|
63
52
|
```
|
|
64
|
-
文档 →
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
→ node:sqlite 存储(chunk + 向量 BLOB)
|
|
68
|
-
→ 检索时内存余弦相似 top-K(万级 chunk 毫秒级)
|
|
69
|
-
```
|
|
70
|
-
|
|
71
|
-
- **向量库**:`node:sqlite`(Node 22.5+ 内置,零 npm 依赖,万级 chunk 内存检索 <50ms)
|
|
72
|
-
- **Embedding**:`@huggingface/transformers` + `Xenova/bge-small-zh-v1.5`(纯 JS ONNX,本地 CPU 3.6ms/chunk)
|
|
73
|
-
- **PDF 解析**:`unpdf`(纯 JS)| **DOCX**:`mammoth`(纯 JS)
|
|
74
|
-
- **数据目录**:`$DSH_HOME/rag-kb/`(index.db + hf-cache/)
|
|
75
|
-
|
|
76
|
-
## 已知限制
|
|
77
|
-
|
|
78
|
-
- 单库设计(不分知识库),适合 <1000 篇文档
|
|
79
|
-
- 512 维向量(bge-small-zh),如需多语言可换 bge-m3(1024 维)
|
|
80
|
-
- 附件出站(图片/文件回传 QQ)走路径文本,暂不支持直接推送
|
|
53
|
+
文档 → 解析(pdf/docx/md) → 标题感知切块(节标题强制断块 + 400 token + 10% 重叠)
|
|
54
|
+
→ 向量化(transformers.js bge-small-zh-v1.5, 512维, GPU/CPU)
|
|
55
|
+
→ FTS5 trigram 关键词索引(与向量索引同步写入)
|
|
81
56
|
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
src/ # 宿主半边(工具 + 索引 + 检索 + systemPrompt 注入 + SSE + 目录监视)
|
|
86
|
-
client-src/ # 浏览器半边(设置标签页)
|
|
87
|
-
lib/plugin.mjs # 宿主半边构建产物
|
|
88
|
-
lib/client.js # 浏览器半边构建产物
|
|
89
|
-
cordis.patch.yml
|
|
57
|
+
检索: 混合检索(向量余弦 + FTS5 bm25 双路合并, via=vector/keyword/both)
|
|
58
|
+
→ [可选] Reranker 精排(bge-reranker-base cross-encoder, 召回×3 → 精排 → top-N)
|
|
59
|
+
→ 引用溯源(citation + position 字段)
|
|
90
60
|
```
|
|
91
61
|
|
|
92
|
-
|
|
62
|
+
### 分块策略(v0.3.0 起)
|
|
93
63
|
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
64
|
+
面向 Markdown 手册/知识库文档优化,保证**一节 = 一块**,引用溯源可精确到节:
|
|
65
|
+
|
|
66
|
+
| 规则 | 说明 |
|
|
67
|
+
|---|---|
|
|
68
|
+
| 节标题强制断块 | 遇 `##` ~ `######` 立即断块,短节不再被合并;刻意不含单级 `#`——配置/代码注释普遍用它开头,误判会把命令序列切碎 |
|
|
69
|
+
| 续块补回节标题 | 长节切成多块时,每块都带节标题,检索命中时能看清出处 |
|
|
70
|
+
| 软下限 + 硬上限 | 块过短(如仅「节标题 + 来源行」)宁可略微超长也不单独成块;合并超过 1.2 倍上限则强制断开 |
|
|
71
|
+
| 滑窗按换行对齐 | 超长无结构段落按换行位置切分,不把命令行从中间截断 |
|
|
72
|
+
| 兜底宽度 | 块长控制在约 600 字符(≈400 token),实测上限 ~920 字符,避免超出 embedding 序列长度被截断 |
|
|
73
|
+
|
|
74
|
+
实测效果(13 篇技术/管理文档,114 节):切成 605 块,**96% 的块带节标题**。
|
|
75
|
+
|
|
76
|
+
## 技术栈
|
|
77
|
+
|
|
78
|
+
| 层 | 选型 | 说明 |
|
|
79
|
+
|---|---|---|
|
|
80
|
+
| Embedding | transformers.js + bge-small-zh-v1.5 ONNX | 512 维中文,本地 CPU 3.6ms/条 |
|
|
81
|
+
| Reranker | transformers.js + bge-reranker-base ONNX | cross-encoder 精排 |
|
|
82
|
+
| GPU 加速 | DirectML (device: 'dml') | 批量索引快 4 倍,无 GPU 回退 CPU |
|
|
83
|
+
| 向量库 | node:sqlite BLOB + 内存余弦 | 万级 chunk <50ms |
|
|
84
|
+
| FTS | SQLite FTS5 + trigram tokenizer | 中文 3 字滑窗精确匹配 |
|
|
85
|
+
| PDF 解析 | unpdf(纯 JS) | 无原生依赖 |
|
|
86
|
+
| DOCX 解析 | mammoth(纯 JS) | 无原生依赖 |
|
|
87
|
+
| 数据目录 | $DSH_HOME/rag-kb/ | index.db + hf-cache/ |
|
|
88
|
+
|
|
89
|
+
## 模型缓存
|
|
90
|
+
|
|
91
|
+
| 模型 | 大小 | 位置 |
|
|
92
|
+
|---|---|---|
|
|
93
|
+
| bge-small-zh-v1.5 | ~91MB | $DSH_HOME/rag-kb/hf-cache/ |
|
|
94
|
+
| bge-reranker-base | ~1GB | 同上(启用 Reranker 时自动下载) |
|
|
95
|
+
|
|
96
|
+
离线部署:管理员把整个 `hf-cache/` 目录拷贝到同事机器的相同位置即可。
|
|
98
97
|
|
|
99
98
|
## License
|
|
100
99
|
|
package/USAGE.zh.md
CHANGED
|
@@ -93,6 +93,9 @@ AI 会自动提取内容、起标题、入库。
|
|
|
93
93
|
- 把「相似度阈值」调高(如 0.5)——结果更精准但可能漏掉
|
|
94
94
|
- 在「知识库描述」里写得更具体,帮助 AI 判断什么时候该检索
|
|
95
95
|
|
|
96
|
+
**Q:什么是 Reranker?要不要开?**
|
|
97
|
+
Reranker 是对检索结果做二次精排的技术,能过滤噪音让结果更准。如果知识库内容多且检索有噪音,建议开启(需下载约 1GB 模型)。文档少(<50篇)时不开也够用。
|
|
98
|
+
|
|
96
99
|
**Q:能同时存多少文档?**
|
|
97
100
|
建议 1000 篇以内。超过后检索仍可用,但不同主题的文档可能互相干扰。
|
|
98
101
|
|
package/lib/client.js
CHANGED
|
@@ -38,6 +38,18 @@ window.__ModuleLoader__.load({ id: "@puwenhui/dsh-rag-kb", factory: (require) =>
|
|
|
38
38
|
label: "相似度阈值",
|
|
39
39
|
type: "number",
|
|
40
40
|
hint: "0~1,默认 0.3,过低噪音多"
|
|
41
|
+
},
|
|
42
|
+
{
|
|
43
|
+
key: "enableReranker",
|
|
44
|
+
label: "启用 Reranker 重排序",
|
|
45
|
+
type: "boolean",
|
|
46
|
+
hint: "首次使用会下载约 500MB 模型"
|
|
47
|
+
},
|
|
48
|
+
{
|
|
49
|
+
key: "rerankTopN",
|
|
50
|
+
label: "重排后返回条数",
|
|
51
|
+
type: "number",
|
|
52
|
+
hint: "默认 3"
|
|
41
53
|
}
|
|
42
54
|
];
|
|
43
55
|
function KbSettingsTab(props) {
|
|
@@ -169,18 +181,23 @@ window.__ModuleLoader__.load({ id: "@puwenhui/dsh-rag-kb", factory: (require) =>
|
|
|
169
181
|
}), /* @__PURE__ */ (0, react_jsx_runtime.jsxs)("div", {
|
|
170
182
|
style: {
|
|
171
183
|
display: "grid",
|
|
172
|
-
gridTemplateColumns: "1fr",
|
|
173
|
-
gap: "10px",
|
|
184
|
+
gridTemplateColumns: "1fr 1fr",
|
|
185
|
+
gap: "10px 14px",
|
|
174
186
|
marginTop: "12px"
|
|
175
187
|
},
|
|
176
188
|
children: [FIELDS.map((f) => {
|
|
177
189
|
const overridden = Object.hasOwn(userLayer, f.key);
|
|
190
|
+
const isBoolean = f.type === "boolean";
|
|
178
191
|
return /* @__PURE__ */ (0, react_jsx_runtime.jsxs)("label", {
|
|
179
192
|
style: {
|
|
180
193
|
display: "flex",
|
|
181
194
|
flexDirection: "column",
|
|
182
195
|
gap: "4px",
|
|
183
|
-
fontSize: "13px"
|
|
196
|
+
fontSize: "13px",
|
|
197
|
+
...isBoolean ? {
|
|
198
|
+
gridColumn: "span 1",
|
|
199
|
+
justifyContent: "center"
|
|
200
|
+
} : {}
|
|
184
201
|
},
|
|
185
202
|
children: [
|
|
186
203
|
/* @__PURE__ */ (0, react_jsx_runtime.jsxs)("span", {
|
|
@@ -193,7 +210,45 @@ window.__ModuleLoader__.load({ id: "@puwenhui/dsh-rag-kb", factory: (require) =>
|
|
|
193
210
|
children: "·已覆盖"
|
|
194
211
|
})]
|
|
195
212
|
}),
|
|
196
|
-
/* @__PURE__ */ (0, react_jsx_runtime.
|
|
213
|
+
isBoolean ? /* @__PURE__ */ (0, react_jsx_runtime.jsxs)("div", {
|
|
214
|
+
style: {
|
|
215
|
+
display: "flex",
|
|
216
|
+
alignItems: "center",
|
|
217
|
+
gap: "12px",
|
|
218
|
+
padding: "4px 0"
|
|
219
|
+
},
|
|
220
|
+
children: [/* @__PURE__ */ (0, react_jsx_runtime.jsxs)("label", {
|
|
221
|
+
style: {
|
|
222
|
+
display: "flex",
|
|
223
|
+
alignItems: "center",
|
|
224
|
+
gap: "4px",
|
|
225
|
+
cursor: "pointer"
|
|
226
|
+
},
|
|
227
|
+
children: [/* @__PURE__ */ (0, react_jsx_runtime.jsx)("input", {
|
|
228
|
+
type: "radio",
|
|
229
|
+
name: f.key,
|
|
230
|
+
checked: shown[f.key] === true,
|
|
231
|
+
onChange: () => {
|
|
232
|
+
edit(f.key, true);
|
|
233
|
+
}
|
|
234
|
+
}), /* @__PURE__ */ (0, react_jsx_runtime.jsx)("span", { children: "开" })]
|
|
235
|
+
}), /* @__PURE__ */ (0, react_jsx_runtime.jsxs)("label", {
|
|
236
|
+
style: {
|
|
237
|
+
display: "flex",
|
|
238
|
+
alignItems: "center",
|
|
239
|
+
gap: "4px",
|
|
240
|
+
cursor: "pointer"
|
|
241
|
+
},
|
|
242
|
+
children: [/* @__PURE__ */ (0, react_jsx_runtime.jsx)("input", {
|
|
243
|
+
type: "radio",
|
|
244
|
+
name: f.key,
|
|
245
|
+
checked: shown[f.key] !== true,
|
|
246
|
+
onChange: () => {
|
|
247
|
+
edit(f.key, false);
|
|
248
|
+
}
|
|
249
|
+
}), /* @__PURE__ */ (0, react_jsx_runtime.jsx)("span", { children: "关" })]
|
|
250
|
+
})]
|
|
251
|
+
}) : /* @__PURE__ */ (0, react_jsx_runtime.jsx)("input", {
|
|
197
252
|
value: draft !== null && draft[f.key] !== void 0 ? String(draft[f.key]) : String(shown[f.key] ?? ""),
|
|
198
253
|
placeholder: f.placeholder ?? "",
|
|
199
254
|
onChange: (e) => {
|
|
@@ -212,6 +267,7 @@ window.__ModuleLoader__.load({ id: "@puwenhui/dsh-rag-kb", factory: (require) =>
|
|
|
212
267
|
}, f.key);
|
|
213
268
|
}), /* @__PURE__ */ (0, react_jsx_runtime.jsxs)("div", {
|
|
214
269
|
style: {
|
|
270
|
+
gridColumn: "1 / -1",
|
|
215
271
|
display: "flex",
|
|
216
272
|
alignItems: "center",
|
|
217
273
|
gap: "10px",
|
package/lib/plugin.mjs
CHANGED
|
@@ -25,18 +25,22 @@ var __require = /* #__PURE__ */ (() => createRequire(import.meta.url))();
|
|
|
25
25
|
/** 用户可见配置(设置面板「知识库」标签页) */
|
|
26
26
|
const PublicConfig = z.object({
|
|
27
27
|
watchDir: z.string().default("").description("本地文档目录(放入即自动索引)"),
|
|
28
|
-
topK: z.number().default(5).description("
|
|
28
|
+
topK: z.number().default(5).description("混合检索召回条数(默认 5)"),
|
|
29
29
|
minScore: z.number().default(.3).description("相似度阈值(0~1,默认 0.3)"),
|
|
30
|
-
description: z.string().default("").description("知识库描述(agent 检索前会阅读)")
|
|
30
|
+
description: z.string().default("").description("知识库描述(agent 检索前会阅读)"),
|
|
31
|
+
enableReranker: z.boolean().default(false).description("启用 reranker 重排序(首次下载约 500MB 模型)"),
|
|
32
|
+
rerankTopN: z.number().default(3).description("重排后返回条数(默认 3)")
|
|
31
33
|
});
|
|
32
34
|
/** 完整配置(含高级项) */
|
|
33
35
|
const Config = z.object({
|
|
34
36
|
watchDir: z.string().default("").description("本地文档目录"),
|
|
35
|
-
topK: z.number().default(5).description("
|
|
37
|
+
topK: z.number().default(5).description("混合检索召回条数"),
|
|
36
38
|
minScore: z.number().default(.3).description("相似度阈值"),
|
|
37
39
|
description: z.string().default("").description("知识库描述"),
|
|
38
40
|
embeddingModel: z.string().default("Xenova/bge-small-zh-v1.5").description("embedding 模型"),
|
|
39
|
-
batchSize: z.number().default(32).description("批量推理大小")
|
|
41
|
+
batchSize: z.number().default(32).description("批量推理大小"),
|
|
42
|
+
enableReranker: z.boolean().default(false).description("启用 reranker 重排序"),
|
|
43
|
+
rerankTopN: z.number().default(3).description("重排后返回条数")
|
|
40
44
|
});
|
|
41
45
|
/** 全字段 volatile 注入(npm 版 schemastery 无 .volatile(),0.2.0 设置面板只收录 volatile 字段) */
|
|
42
46
|
for (const schema of [PublicConfig, Config]) for (const child of Object.values(schema.dict ?? {})) child.meta.volatile = true;
|
|
@@ -48,7 +52,9 @@ function resolveKbConfig(c = {}) {
|
|
|
48
52
|
minScore: Math.max(0, Math.min(1, c.minScore ?? .3)),
|
|
49
53
|
description: c.description ?? "",
|
|
50
54
|
embeddingModel: c.embeddingModel ?? "Xenova/bge-small-zh-v1.5",
|
|
51
|
-
batchSize: Math.max(1, Math.min(128, c.batchSize ?? 32))
|
|
55
|
+
batchSize: Math.max(1, Math.min(128, c.batchSize ?? 32)),
|
|
56
|
+
enableReranker: c.enableReranker === true,
|
|
57
|
+
rerankTopN: Math.max(1, Math.min(20, c.rerankTopN ?? 3))
|
|
52
58
|
};
|
|
53
59
|
}
|
|
54
60
|
//#endregion
|
|
@@ -95,6 +101,10 @@ var KbStore = class {
|
|
|
95
101
|
FOREIGN KEY(doc_id) REFERENCES documents(doc_id) ON DELETE CASCADE
|
|
96
102
|
);
|
|
97
103
|
CREATE INDEX IF NOT EXISTS idx_chunks_doc ON chunks(doc_id);
|
|
104
|
+
CREATE VIRTUAL TABLE IF NOT EXISTS chunk_fts USING fts5(
|
|
105
|
+
text,
|
|
106
|
+
tokenize='trigram'
|
|
107
|
+
);
|
|
98
108
|
`);
|
|
99
109
|
}
|
|
100
110
|
/** 文档内容 hash 作 doc_id(同名不同内容=不同文档) */
|
|
@@ -118,6 +128,9 @@ var KbStore = class {
|
|
|
118
128
|
this.loaded = false;
|
|
119
129
|
}
|
|
120
130
|
deleteDoc(docId) {
|
|
131
|
+
const ids = this.db.prepare("SELECT id FROM chunks WHERE doc_id=?").all(docId);
|
|
132
|
+
const ftsDel = this.db.prepare("DELETE FROM chunk_fts WHERE rowid=?");
|
|
133
|
+
for (const { id } of ids) ftsDel.run(id);
|
|
121
134
|
this.db.prepare("DELETE FROM chunks WHERE doc_id=?").run(docId);
|
|
122
135
|
this.db.prepare("DELETE FROM documents WHERE doc_id=?").run(docId);
|
|
123
136
|
this.loaded = false;
|
|
@@ -128,13 +141,22 @@ var KbStore = class {
|
|
|
128
141
|
getDoc(docId) {
|
|
129
142
|
return this.db.prepare("SELECT * FROM documents WHERE doc_id=?").get(docId);
|
|
130
143
|
}
|
|
131
|
-
/**
|
|
144
|
+
/** 批量写入切块+向量+FTS 索引(一个事务,失败回滚) */
|
|
132
145
|
insertChunks(docId, chunks) {
|
|
133
146
|
this.db.exec("BEGIN");
|
|
134
147
|
try {
|
|
135
148
|
this.db.prepare("DELETE FROM chunks WHERE doc_id=?").run(docId);
|
|
149
|
+
const oldIds = this.db.prepare("SELECT id FROM chunks WHERE doc_id=?").all(docId);
|
|
150
|
+
const ftsDel = this.db.prepare("DELETE FROM chunk_fts WHERE rowid=?");
|
|
151
|
+
for (const { id } of oldIds) ftsDel.run(id);
|
|
136
152
|
const stmt = this.db.prepare("INSERT INTO chunks(doc_id, seq, text, embedding) VALUES(?,?,?,?)");
|
|
137
|
-
|
|
153
|
+
const ftsIns = this.db.prepare("INSERT INTO chunk_fts(rowid, text) VALUES(?,?)");
|
|
154
|
+
for (const c of chunks) {
|
|
155
|
+
const buf = Buffer.from(c.embedding.buffer, c.embedding.byteOffset, c.embedding.byteLength);
|
|
156
|
+
const info = stmt.run(docId, c.seq, c.text, buf);
|
|
157
|
+
const rowid = Number(info.lastInsertRowid);
|
|
158
|
+
ftsIns.run(rowid, c.text);
|
|
159
|
+
}
|
|
138
160
|
this.db.exec("COMMIT");
|
|
139
161
|
} catch (e) {
|
|
140
162
|
this.db.exec("ROLLBACK");
|
|
@@ -157,22 +179,69 @@ var KbStore = class {
|
|
|
157
179
|
}
|
|
158
180
|
this.loaded = true;
|
|
159
181
|
}
|
|
160
|
-
/**
|
|
161
|
-
|
|
182
|
+
/** 混合检索:向量余弦 + FTS5 关键词双路合并(加权排序) */
|
|
183
|
+
hybridSearch(queryVec, queryText, topK, minScore) {
|
|
162
184
|
this.ensureLoaded();
|
|
163
|
-
const
|
|
185
|
+
const vectorScores = /* @__PURE__ */ new Map();
|
|
164
186
|
for (let i = 0; i < this.vectors.length; i++) {
|
|
165
187
|
const v = this.vectors[i];
|
|
166
188
|
let dot = 0;
|
|
167
189
|
for (let j = 0; j < queryVec.length; j++) dot += queryVec[j] * v[j];
|
|
168
|
-
if (dot >= minScore)
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
|
|
190
|
+
if (dot >= minScore) {
|
|
191
|
+
const key = `${this.vectorRows[i].docId}:${this.vectorRows[i].seq}`;
|
|
192
|
+
vectorScores.set(key, {
|
|
193
|
+
docId: this.vectorRows[i].docId,
|
|
194
|
+
seq: this.vectorRows[i].seq,
|
|
195
|
+
score: dot
|
|
196
|
+
});
|
|
197
|
+
}
|
|
173
198
|
}
|
|
174
|
-
|
|
175
|
-
|
|
199
|
+
const keywordScores = /* @__PURE__ */ new Map();
|
|
200
|
+
if (queryText.trim().length >= 3) try {
|
|
201
|
+
const safe = queryText.replace(/["'*()\-:]/g, " ").trim().slice(0, 50);
|
|
202
|
+
if (safe.length >= 3) {
|
|
203
|
+
const rows = this.db.prepare(`
|
|
204
|
+
SELECT c.doc_id, c.seq, bm25(chunk_fts) AS rank
|
|
205
|
+
FROM chunk_fts f
|
|
206
|
+
JOIN chunks c ON c.id = f.rowid
|
|
207
|
+
WHERE chunk_fts MATCH ?
|
|
208
|
+
ORDER BY rank
|
|
209
|
+
LIMIT ?
|
|
210
|
+
`).all(`"${safe}"`, topK * 2);
|
|
211
|
+
for (const row of rows) {
|
|
212
|
+
const key = `${row.doc_id}:${row.seq}`;
|
|
213
|
+
const kwScore = 1 / (1 + Math.abs(row.rank));
|
|
214
|
+
keywordScores.set(key, {
|
|
215
|
+
docId: row.doc_id,
|
|
216
|
+
seq: row.seq,
|
|
217
|
+
score: kwScore
|
|
218
|
+
});
|
|
219
|
+
}
|
|
220
|
+
}
|
|
221
|
+
} catch {}
|
|
222
|
+
const merged = /* @__PURE__ */ new Map();
|
|
223
|
+
for (const [key, v] of vectorScores) merged.set(key, {
|
|
224
|
+
...v,
|
|
225
|
+
score: v.score,
|
|
226
|
+
via: keywordScores.has(key) ? "both" : "vector"
|
|
227
|
+
});
|
|
228
|
+
for (const [key, k] of keywordScores) if (merged.has(key)) {
|
|
229
|
+
const existing = merged.get(key);
|
|
230
|
+
existing.score = existing.score + k.score * .3;
|
|
231
|
+
existing.via = "both";
|
|
232
|
+
} else merged.set(key, {
|
|
233
|
+
...k,
|
|
234
|
+
via: "keyword"
|
|
235
|
+
});
|
|
236
|
+
return [...merged.values()].sort((a, b) => b.score - a.score).slice(0, topK);
|
|
237
|
+
}
|
|
238
|
+
/** 向后兼容:纯向量检索(内部调 hybridSearch 空 queryText) */
|
|
239
|
+
search(queryVec, topK, minScore) {
|
|
240
|
+
return this.hybridSearch(queryVec, "", topK, minScore).map((r) => ({
|
|
241
|
+
docId: r.docId,
|
|
242
|
+
seq: r.seq,
|
|
243
|
+
score: r.score
|
|
244
|
+
}));
|
|
176
245
|
}
|
|
177
246
|
/** 按 doc_id + seq 取 chunk 原文 */
|
|
178
247
|
chunkText(docId, seq) {
|
|
@@ -192,11 +261,17 @@ var indexer_exports = /* @__PURE__ */ __exportAll({
|
|
|
192
261
|
chunkText: () => chunkText,
|
|
193
262
|
extractText: () => extractText,
|
|
194
263
|
getEmbedder: () => getEmbedder,
|
|
195
|
-
indexFile: () => indexFile
|
|
264
|
+
indexFile: () => indexFile,
|
|
265
|
+
rerank: () => rerank
|
|
196
266
|
});
|
|
197
267
|
/** 中文友好的切块参数 */
|
|
198
268
|
const CHUNK_TOKENS = 400;
|
|
199
|
-
|
|
269
|
+
/**
|
|
270
|
+
* 节标题行:`##` ~ `######` 视为节标题,遇之强制断块(保证「一节 = 一块」)。
|
|
271
|
+
* 刻意不含单级 `#`——nginx/shell/mysql 等配置与代码的注释普遍用单 `#` 开头,
|
|
272
|
+
* 误判会把命令序列切碎(实测语音平台文档 27 节被切成 231 块)。
|
|
273
|
+
*/
|
|
274
|
+
const HEADING_RE = /^#{2,6} /;
|
|
200
275
|
async function extractText(filePath) {
|
|
201
276
|
const ext = extname(filePath).toLowerCase();
|
|
202
277
|
const buf = await readFile(filePath);
|
|
@@ -215,58 +290,86 @@ async function extractText(filePath) {
|
|
|
215
290
|
if (ext === ".md" || ext === ".txt" || ext === ".log" || ext === ".csv" || ext === ".json") return { text: buf.toString("utf8") };
|
|
216
291
|
throw new Error(`不支持的文件类型 ${ext}(支持 pdf/docx/md/txt/csv/json)`);
|
|
217
292
|
}
|
|
218
|
-
/**
|
|
293
|
+
/** 标题感知切块:节标题强制断块(一节 = 一块)+ 空行分段 + 过长滑窗兜底 */
|
|
219
294
|
function chunkText(text) {
|
|
220
295
|
const normalized = text.replace(/\r\n/g, "\n").replace(/\n{3,}/g, "\n\n").trim();
|
|
221
296
|
if (normalized === "") return [];
|
|
222
|
-
const paragraphs = normalized.split(/\n(?=#{
|
|
297
|
+
const paragraphs = normalized.split(/\n(?=#{2,6} )|\n\n+/).map((p) => p.trim()).filter((p) => p !== "");
|
|
223
298
|
const chunks = [];
|
|
224
299
|
let current = "";
|
|
300
|
+
let sectionHeading = "";
|
|
225
301
|
const maxChars = CHUNK_TOKENS * 1.5;
|
|
226
|
-
const overlap = Math.floor(maxChars * OVERLAP_RATIO);
|
|
227
302
|
for (const para of paragraphs) {
|
|
228
|
-
|
|
303
|
+
const isHeading = HEADING_RE.test(para);
|
|
304
|
+
if (isHeading) sectionHeading = para.split("\n")[0].trim();
|
|
305
|
+
if (isHeading && current !== "") {
|
|
229
306
|
chunks.push(current);
|
|
230
|
-
current =
|
|
307
|
+
current = "";
|
|
231
308
|
}
|
|
232
|
-
if (para.length > maxChars * 2) {
|
|
233
|
-
|
|
234
|
-
|
|
235
|
-
|
|
309
|
+
if (para.length > maxChars * 1.2) {
|
|
310
|
+
const prefix = current;
|
|
311
|
+
current = "";
|
|
312
|
+
const mergePrefix = prefix !== "" && prefix.length <= maxChars * .6;
|
|
313
|
+
if (prefix !== "" && !mergePrefix) chunks.push(prefix);
|
|
314
|
+
for (let i = 0; i < para.length;) {
|
|
315
|
+
let end = Math.min(i + maxChars, para.length);
|
|
316
|
+
if (end < para.length) {
|
|
317
|
+
const nl = para.lastIndexOf("\n", end);
|
|
318
|
+
if (nl > i + maxChars * .5) end = nl + 1;
|
|
319
|
+
}
|
|
320
|
+
const piece = para.slice(i, end);
|
|
321
|
+
const isFirst = i === 0;
|
|
322
|
+
i = end;
|
|
323
|
+
chunks.push(isFirst ? mergePrefix ? `${prefix}\n\n${piece}` : piece : sectionHeading !== "" ? `${sectionHeading}\n\n${piece}` : piece);
|
|
236
324
|
}
|
|
237
|
-
|
|
325
|
+
current = sectionHeading;
|
|
238
326
|
continue;
|
|
239
327
|
}
|
|
328
|
+
const merged = current + "\n" + para;
|
|
329
|
+
if (current !== "" && merged.length > maxChars && (current.length >= maxChars * .5 || merged.length > maxChars * 1.2)) {
|
|
330
|
+
chunks.push(current);
|
|
331
|
+
const tail = current.slice(-60);
|
|
332
|
+
current = sectionHeading !== "" ? `${sectionHeading}\n\n${tail}` : tail;
|
|
333
|
+
}
|
|
240
334
|
current = current === "" ? para : current + "\n" + para;
|
|
241
335
|
}
|
|
242
336
|
if (current !== "") chunks.push(current);
|
|
243
337
|
return chunks.filter((c) => c.trim().length >= 10);
|
|
244
338
|
}
|
|
245
|
-
let
|
|
246
|
-
|
|
339
|
+
let gpuPipeline;
|
|
340
|
+
let cpuPipeline;
|
|
341
|
+
async function loadPipeline(modelName, device) {
|
|
247
342
|
const { pipeline, env } = await import("@huggingface/transformers");
|
|
248
343
|
const { dshHomePath } = await import("@deepseek-ai/dsh-home-paths");
|
|
249
344
|
env.cacheDir = dshHomePath("rag-kb", "hf-cache");
|
|
250
|
-
|
|
345
|
+
try {
|
|
346
|
+
return await pipeline("feature-extraction", modelName, { device });
|
|
347
|
+
} catch (e) {
|
|
348
|
+
if (device === "dml") {
|
|
349
|
+
console.log("[rag-kb] DML GPU 不可用,回退 CPU");
|
|
350
|
+
return await pipeline("feature-extraction", modelName, { device: "cpu" });
|
|
351
|
+
}
|
|
352
|
+
throw e;
|
|
353
|
+
}
|
|
251
354
|
}
|
|
252
355
|
async function getEmbedder(modelName) {
|
|
253
|
-
pipelinePromise ??= loadPipeline(modelName);
|
|
254
|
-
const embedder = await pipelinePromise;
|
|
255
356
|
return {
|
|
256
|
-
/**
|
|
357
|
+
/** 单条向量化(检索查询用,CPU 单条比 GPU 快) */
|
|
257
358
|
async embed(text) {
|
|
258
|
-
|
|
359
|
+
cpuPipeline ??= await loadPipeline(modelName, "cpu");
|
|
360
|
+
const out = await cpuPipeline(text, {
|
|
259
361
|
pooling: "cls",
|
|
260
362
|
normalize: true
|
|
261
363
|
});
|
|
262
364
|
return new Float32Array(out.data);
|
|
263
365
|
},
|
|
264
|
-
/** 批量向量化(索引用,
|
|
366
|
+
/** 批量向量化(索引用,GPU/DML 批量快 4 倍,不可用自动回退 CPU) */
|
|
265
367
|
async embedBatch(texts, batchSize) {
|
|
368
|
+
gpuPipeline ??= await loadPipeline(modelName, "dml");
|
|
266
369
|
const results = [];
|
|
267
370
|
for (let i = 0; i < texts.length; i += batchSize) {
|
|
268
371
|
const batch = texts.slice(i, i + batchSize);
|
|
269
|
-
const out = await
|
|
372
|
+
const out = await gpuPipeline(batch, {
|
|
270
373
|
pooling: "cls",
|
|
271
374
|
normalize: true
|
|
272
375
|
});
|
|
@@ -277,6 +380,31 @@ async function getEmbedder(modelName) {
|
|
|
277
380
|
}
|
|
278
381
|
};
|
|
279
382
|
}
|
|
383
|
+
let rerankerPipeline;
|
|
384
|
+
async function loadReranker() {
|
|
385
|
+
const { pipeline, env } = await import("@huggingface/transformers");
|
|
386
|
+
const { dshHomePath } = await import("@deepseek-ai/dsh-home-paths");
|
|
387
|
+
env.cacheDir = dshHomePath("rag-kb", "hf-cache");
|
|
388
|
+
return pipeline("text-classification", "Xenova/bge-reranker-base");
|
|
389
|
+
}
|
|
390
|
+
/**
|
|
391
|
+
* 对混合检索结果做 cross-encoder 精排。
|
|
392
|
+
* 输入:query + 候选列表(含原文),输出按相关性重排序。
|
|
393
|
+
*/
|
|
394
|
+
async function rerank(query, candidates) {
|
|
395
|
+
rerankerPipeline ??= await loadReranker();
|
|
396
|
+
const pairs = candidates.map((c) => ({
|
|
397
|
+
text: query,
|
|
398
|
+
text_pair: c.text
|
|
399
|
+
}));
|
|
400
|
+
const scores = await rerankerPipeline(pairs);
|
|
401
|
+
const scored = candidates.map((c, i) => ({
|
|
402
|
+
...c,
|
|
403
|
+
rerankScore: scores[i].score
|
|
404
|
+
}));
|
|
405
|
+
scored.sort((a, b) => b.rerankScore - a.rerankScore);
|
|
406
|
+
return scored;
|
|
407
|
+
}
|
|
280
408
|
/** 完整索引管线:文件路径 → 切块+向量数组 */
|
|
281
409
|
async function indexFile(filePath, embedder, batchSize, onProgress) {
|
|
282
410
|
const { text } = await extractText(filePath);
|
|
@@ -355,16 +483,46 @@ async function indexOne(ctx, filePath, source) {
|
|
|
355
483
|
};
|
|
356
484
|
}
|
|
357
485
|
}
|
|
358
|
-
/**
|
|
486
|
+
/** 混合语义检索(向量+关键词)+ 可选 reranker 精排 + 引用元数据 */
|
|
359
487
|
async function search(query, topK) {
|
|
360
488
|
const s = await ensureStore();
|
|
361
489
|
const queryVec = await (await getEmbedder(cfg?.embeddingModel ?? "Xenova/bge-small-zh-v1.5")).embed(query);
|
|
362
|
-
|
|
363
|
-
|
|
364
|
-
|
|
365
|
-
|
|
366
|
-
|
|
490
|
+
const effectiveTopK = topK ?? cfg?.topK ?? 5;
|
|
491
|
+
const recallK = cfg?.enableReranker === true ? effectiveTopK * 3 : effectiveTopK;
|
|
492
|
+
let results = s.hybridSearch(queryVec, query, recallK, cfg?.minScore ?? .3).map((h) => {
|
|
493
|
+
const docName = s.getDoc(h.docId)?.name ?? h.docId;
|
|
494
|
+
const fullText = s.chunkText(h.docId, h.seq);
|
|
495
|
+
return {
|
|
496
|
+
docId: h.docId,
|
|
497
|
+
doc: docName,
|
|
498
|
+
seq: h.seq,
|
|
499
|
+
score: Number(h.score.toFixed(3)),
|
|
500
|
+
via: h.via,
|
|
501
|
+
text: fullText,
|
|
502
|
+
citation: `${docName} 第${h.seq + 1}段`,
|
|
503
|
+
position: `文档「${docName}」第 ${h.seq + 1} 段`
|
|
504
|
+
};
|
|
505
|
+
});
|
|
506
|
+
if (cfg?.enableReranker === true && results.length > 0) try {
|
|
507
|
+
const reranked = await rerank(query, results);
|
|
508
|
+
const topN = cfg?.rerankTopN ?? 3;
|
|
509
|
+
results = reranked.slice(0, topN).map((r) => ({
|
|
510
|
+
...r,
|
|
511
|
+
text: r.text.slice(0, 800),
|
|
512
|
+
rerankScore: Number(r.rerankScore.toFixed(4))
|
|
513
|
+
}));
|
|
514
|
+
} catch (e) {
|
|
515
|
+
console.warn("[rag-kb] reranker 失败,使用混合检索结果:", String(e.message ?? e));
|
|
516
|
+
results = results.slice(0, effectiveTopK).map((r) => ({
|
|
517
|
+
...r,
|
|
518
|
+
text: r.text.slice(0, 800)
|
|
519
|
+
}));
|
|
520
|
+
}
|
|
521
|
+
else results = results.slice(0, effectiveTopK).map((r) => ({
|
|
522
|
+
...r,
|
|
523
|
+
text: r.text.slice(0, 800)
|
|
367
524
|
}));
|
|
525
|
+
return results;
|
|
368
526
|
}
|
|
369
527
|
function apply(ctx, config = {}) {
|
|
370
528
|
const unwrap = (v) => v !== null && typeof v === "object" && "get" in v ? v.get() : v;
|
|
@@ -432,12 +590,12 @@ function apply(ctx, config = {}) {
|
|
|
432
590
|
const disposers = [
|
|
433
591
|
ctx.tools.register({
|
|
434
592
|
name: "knowledge_search",
|
|
435
|
-
description: "
|
|
593
|
+
description: "在本地知识库中混合检索(语义向量 + 关键词)。返回最相关的知识块,每条带 citation 引用标注。回答知识库相关问题时:先检索 → 组织回答 → 在答案中标注来源 citation。query 支持自然语言和精确关键词(错误码/型号/人名等)。",
|
|
436
594
|
parameters: compileParams({
|
|
437
595
|
query: {
|
|
438
596
|
type: "string",
|
|
439
597
|
required: true,
|
|
440
|
-
description: "
|
|
598
|
+
description: "检索问题(自然语言或精确关键词)"
|
|
441
599
|
},
|
|
442
600
|
top_k: {
|
|
443
601
|
type: "number",
|
|
@@ -635,11 +793,13 @@ function apply(ctx, config = {}) {
|
|
|
635
793
|
const docs = store?.listDocs().filter((d) => d.status === "ready").length ?? 0;
|
|
636
794
|
return [
|
|
637
795
|
"You have a local knowledge base accessible via these tools:",
|
|
638
|
-
"- knowledge_search:
|
|
796
|
+
"- knowledge_search: hybrid search (semantic + keyword). Use BEFORE answering questions about KB content.",
|
|
639
797
|
"- knowledge_status: list documents and config",
|
|
640
798
|
"- knowledge_manage: add/save/remove/reindex documents",
|
|
799
|
+
"IMPORTANT: When the user says 搜一下/查一下/搜索, FIRST try knowledge_search on the local KB. Only do a web search if KB has no relevant results AND the topic is clearly not about local documents. Do NOT search both simultaneously — this confuses the answer.",
|
|
641
800
|
desc !== "" ? `KB description: ${desc}` : "",
|
|
642
801
|
`Currently ${docs} document(s) indexed.`,
|
|
802
|
+
"CITATION STYLE (footnote format): In the body of your answer, mark sources with superscript numbers like ¹ ² ³ (Unicode superscripts). At the END, add a \"参考来源\" section. IMPORTANT: MERGE citations from the same document into ONE entry — e.g. if citing 文档A 第1段/第2段/第3段, list as \"1. 文档A 第1-3段\" (not 3 separate lines). Only list each document once with all its paragraph numbers combined. Keep the answer body clean.",
|
|
643
803
|
"When the user says \"save this to the knowledge base\" (保存到知识库/记下来), extract the valuable content from the conversation and call knowledge_manage with action=\"save\", providing a concise title and the text.",
|
|
644
804
|
"For file/directory indexing use action=\"add\" with a path."
|
|
645
805
|
].filter(Boolean).join("\n");
|
|
@@ -652,6 +812,6 @@ function apply(ctx, config = {}) {
|
|
|
652
812
|
});
|
|
653
813
|
}
|
|
654
814
|
//#endregion
|
|
655
|
-
export { Config, KbStore, apply, chunkText, extractText, getEmbedder, indexFile, inject, name, resolveKbConfig };
|
|
815
|
+
export { Config, KbStore, apply, chunkText, extractText, getEmbedder, indexFile, inject, name, rerank, resolveKbConfig };
|
|
656
816
|
|
|
657
817
|
//# sourceMappingURL=plugin.mjs.map
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@puwenhui/dsh-rag-kb",
|
|
3
|
-
"version": "0.
|
|
4
|
-
"description": "DSH RAG
|
|
3
|
+
"version": "0.3.0",
|
|
4
|
+
"description": "DSH RAG 知识库插件:混合检索(向量+FTS5)+Reranker 重排+GPU 加速+引用溯源+会话语义管理",
|
|
5
5
|
"publishConfig": {
|
|
6
6
|
"registry": "https://registry.npmjs.org/",
|
|
7
7
|
"access": "public"
|