@puwenhui/dsh-rag-kb 0.1.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,21 +1,24 @@
1
1
  # @puwenhui/dsh-rag-kb
2
2
 
3
- DSH RAG 知识库插件:**文档导入 → 智能切块 → 本地向量化 → 语义检索**,让 DSH 对话拥有私域知识。
3
+ DSH RAG 知识库插件:**文档导入 → 智能切块 → 本地向量化 → 混合检索(向量+FTS5)→ Reranker 精排 → 引用溯源**。
4
4
 
5
5
  > 面向公司非技术人员的使用帮助请阅读 **[USAGE.zh.md](./USAGE.zh.md)**。
6
6
 
7
7
  ## 核心特性
8
8
 
9
- - **语义检索**:自然语言提问,agent 自动从知识库找答案(中文/英文均可)
10
- - **多格式支持**:pdf / docx / md / txt / csv / json
11
- - **完全本地**:向量化模型本地运行(bge-small-zh ONNX),数据不出机器
9
+ - **混合检索**:语义向量(bge-small-zh)+ FTS5 trigram 关键词双路合并,同义/改述和精确词/错误码都能找到
10
+ - **Reranker 精排**(可选):bge-reranker-base cross-encoder,对混合检索结果做语义精排,噪音过滤更准
11
+ - **GPU 加速**:DirectML(RTX 4060 实测批量索引快 4 倍),无 GPU 自动回退 CPU
12
+ - **引用溯源**:脚注式引用(正文上标¹²³ + 末尾参考来源清单),agent 回答自带出处
12
13
  - **会话语义管理**:所有操作(添加/删除/保存/检索)直接在对话里用大白话完成
13
- - **目录自动索引**:配置一个目录,文件放入/变更自动入库
14
+ - **多格式支持**:pdf / docx / md / txt / csv / json
15
+ - **完全本地**:模型本地运行,数据不出机器
16
+ - **目录自动索引**:配置目录后文件放入/变更自动入库
14
17
  - **对话内容保存**:把对话中的结论/总结一键存入知识库
15
18
 
16
19
  ## 版本要求
17
20
 
18
- - **DSH ≥ 0.2.0**(依赖 settings 标签页体系与 volatile 字段机制)
21
+ - **DSH ≥ 0.2.0**
19
22
 
20
23
  ## 安装
21
24
 
@@ -23,25 +26,19 @@ DSH RAG 知识库插件:**文档导入 → 智能切块 → 本地向量化
23
26
  dsh plugin --profile web add @puwenhui/dsh-rag-kb
24
27
  ```
25
28
 
26
- (源码宿主用 `pnpm dsh plugin --profile web add @puwenhui/dsh-rag-kb`)
27
-
28
- ## 配置
29
+ ## 配置(设置 → 插件 →「知识库」标签页)
29
30
 
30
- 安装后打开 `dsh web`,**设置 → 插件 →「知识库」**标签页:
31
-
32
- | 字段 | 说明 |
33
- |---|---|
34
- | 知识库描述 | agent 检索前阅读,知道库里有什么 |
35
- | 自动索引目录 | 放入文档自动入库(可选) |
36
- | 检索返回条数 | 默认 5 |
37
- | 相似度阈值 | 0~1,默认 0.3 |
38
-
39
- **首次使用**会自动下载中文向量化模型(约 100MB,下载后完全离线,模型缓存在 `$DSH_HOME/rag-kb/hf-cache/`)。
31
+ | 字段 | 默认 | 说明 |
32
+ |---|---|---|
33
+ | 知识库描述 | 空 | agent 检索前阅读,知道库里有什么 |
34
+ | 自动索引目录 | 空 | 放入文档自动入库(可选) |
35
+ | 检索返回条数 | 5 | 混合检索召回 top-K |
36
+ | 相似度阈值 | 0.3 | 0~1,过低噪音多 |
37
+ | 启用 Reranker | 关 | 开启后需下载 ~1GB 模型 |
38
+ | 重排后返回条数 | 3 | Reranker 精排后返回 top-N |
40
39
 
41
40
  ## 会话语义操作
42
41
 
43
- 所有知识库操作直接在对话里说,agent 自动调用对应工具:
44
-
45
42
  | 你说 | agent 调用 |
46
43
  |---|---|
47
44
  | 「把 D:\docs 加进知识库」 | `knowledge_manage add` |
@@ -50,51 +47,53 @@ dsh plugin --profile web add @puwenhui/dsh-rag-kb
50
47
  | 「删掉旧版手册」 | `knowledge_manage remove` |
51
48
  | 「根据知识库回答:XXX」 | `knowledge_search` |
52
49
 
53
- ## 模型工具(3 个)
54
-
55
- | 工具 | 说明 |
56
- |---|---|
57
- | `knowledge_search` | 语义检索:query + top_k → 带来源/相似度/原文的结果 |
58
- | `knowledge_status` | 知识库状态:文档列表/配置/描述 |
59
- | `knowledge_manage` | 管理操作:save / add / remove / reindex / describe |
60
-
61
- ## 技术架构
50
+ ## 检索架构
62
51
 
63
52
  ```
64
- 文档 → unpdf/mammoth 提取纯文本
65
- → 标题感知切块(400 token + 10% 重叠)
66
- → transformers.js 向量化(bge-small-zh-v1.5, 512 维, GPU 自动检测/CPU fallback)
67
- → node:sqlite 存储(chunk + 向量 BLOB)
68
- → 检索时内存余弦相似 top-K(万级 chunk 毫秒级)
69
- ```
70
-
71
- - **向量库**:`node:sqlite`(Node 22.5+ 内置,零 npm 依赖,万级 chunk 内存检索 <50ms)
72
- - **Embedding**:`@huggingface/transformers` + `Xenova/bge-small-zh-v1.5`(纯 JS ONNX,本地 CPU 3.6ms/chunk)
73
- - **PDF 解析**:`unpdf`(纯 JS)| **DOCX**:`mammoth`(纯 JS)
74
- - **数据目录**:`$DSH_HOME/rag-kb/`(index.db + hf-cache/)
75
-
76
- ## 已知限制
77
-
78
- - 单库设计(不分知识库),适合 <1000 篇文档
79
- - 512 维向量(bge-small-zh),如需多语言可换 bge-m3(1024 维)
80
- - 附件出站(图片/文件回传 QQ)走路径文本,暂不支持直接推送
53
+ 文档 → 解析(pdf/docx/md) → 标题感知切块(节标题强制断块 + 400 token + 10% 重叠)
54
+ → 向量化(transformers.js bge-small-zh-v1.5, 512维, GPU/CPU)
55
+ → FTS5 trigram 关键词索引(与向量索引同步写入)
81
56
 
82
- ## 目录结构
83
-
84
- ```
85
- src/ # 宿主半边(工具 + 索引 + 检索 + systemPrompt 注入 + SSE + 目录监视)
86
- client-src/ # 浏览器半边(设置标签页)
87
- lib/plugin.mjs # 宿主半边构建产物
88
- lib/client.js # 浏览器半边构建产物
89
- cordis.patch.yml
57
+ 检索: 混合检索(向量余弦 + FTS5 bm25 双路合并, via=vector/keyword/both)
58
+ → [可选] Reranker 精排(bge-reranker-base cross-encoder, 召回×3 → 精排 → top-N)
59
+ → 引用溯源(citation + position 字段)
90
60
  ```
91
61
 
92
- ## 开发
62
+ ### 分块策略(v0.3.0 起)
93
63
 
94
- - 宿主半边:`npx tsdown`(产出 lib/plugin.mjs)
95
- - 浏览器半边:`cd client-src && npm install && npm run build`
96
- - 部署到本地 profile:构建后复制 `lib/plugin.mjs` 和 `lib/client.js` 到 profile 的 node_modules 目录
97
- - 发布:bump version → `npm publish`
64
+ 面向 Markdown 手册/知识库文档优化,保证**一节 = 一块**,引用溯源可精确到节:
65
+
66
+ | 规则 | 说明 |
67
+ |---|---|
68
+ | 节标题强制断块 | 遇 `##` ~ `######` 立即断块,短节不再被合并;刻意不含单级 `#`——配置/代码注释普遍用它开头,误判会把命令序列切碎 |
69
+ | 续块补回节标题 | 长节切成多块时,每块都带节标题,检索命中时能看清出处 |
70
+ | 软下限 + 硬上限 | 块过短(如仅「节标题 + 来源行」)宁可略微超长也不单独成块;合并超过 1.2 倍上限则强制断开 |
71
+ | 滑窗按换行对齐 | 超长无结构段落按换行位置切分,不把命令行从中间截断 |
72
+ | 兜底宽度 | 块长控制在约 600 字符(≈400 token),实测上限 ~920 字符,避免超出 embedding 序列长度被截断 |
73
+
74
+ 实测效果(13 篇技术/管理文档,114 节):切成 605 块,**96% 的块带节标题**。
75
+
76
+ ## 技术栈
77
+
78
+ | 层 | 选型 | 说明 |
79
+ |---|---|---|
80
+ | Embedding | transformers.js + bge-small-zh-v1.5 ONNX | 512 维中文,本地 CPU 3.6ms/条 |
81
+ | Reranker | transformers.js + bge-reranker-base ONNX | cross-encoder 精排 |
82
+ | GPU 加速 | DirectML (device: 'dml') | 批量索引快 4 倍,无 GPU 回退 CPU |
83
+ | 向量库 | node:sqlite BLOB + 内存余弦 | 万级 chunk <50ms |
84
+ | FTS | SQLite FTS5 + trigram tokenizer | 中文 3 字滑窗精确匹配 |
85
+ | PDF 解析 | unpdf(纯 JS) | 无原生依赖 |
86
+ | DOCX 解析 | mammoth(纯 JS) | 无原生依赖 |
87
+ | 数据目录 | $DSH_HOME/rag-kb/ | index.db + hf-cache/ |
88
+
89
+ ## 模型缓存
90
+
91
+ | 模型 | 大小 | 位置 |
92
+ |---|---|---|
93
+ | bge-small-zh-v1.5 | ~91MB | $DSH_HOME/rag-kb/hf-cache/ |
94
+ | bge-reranker-base | ~1GB | 同上(启用 Reranker 时自动下载) |
95
+
96
+ 离线部署:管理员把整个 `hf-cache/` 目录拷贝到同事机器的相同位置即可。
98
97
 
99
98
  ## License
100
99
 
package/USAGE.zh.md CHANGED
@@ -93,6 +93,9 @@ AI 会自动提取内容、起标题、入库。
93
93
  - 把「相似度阈值」调高(如 0.5)——结果更精准但可能漏掉
94
94
  - 在「知识库描述」里写得更具体,帮助 AI 判断什么时候该检索
95
95
 
96
+ **Q:什么是 Reranker?要不要开?**
97
+ Reranker 是对检索结果做二次精排的技术,能过滤噪音让结果更准。如果知识库内容多且检索有噪音,建议开启(需下载约 1GB 模型)。文档少(<50篇)时不开也够用。
98
+
96
99
  **Q:能同时存多少文档?**
97
100
  建议 1000 篇以内。超过后检索仍可用,但不同主题的文档可能互相干扰。
98
101
 
package/lib/client.js CHANGED
@@ -38,6 +38,18 @@ window.__ModuleLoader__.load({ id: "@puwenhui/dsh-rag-kb", factory: (require) =>
38
38
  label: "相似度阈值",
39
39
  type: "number",
40
40
  hint: "0~1,默认 0.3,过低噪音多"
41
+ },
42
+ {
43
+ key: "enableReranker",
44
+ label: "启用 Reranker 重排序",
45
+ type: "boolean",
46
+ hint: "首次使用会下载约 500MB 模型"
47
+ },
48
+ {
49
+ key: "rerankTopN",
50
+ label: "重排后返回条数",
51
+ type: "number",
52
+ hint: "默认 3"
41
53
  }
42
54
  ];
43
55
  function KbSettingsTab(props) {
@@ -169,18 +181,23 @@ window.__ModuleLoader__.load({ id: "@puwenhui/dsh-rag-kb", factory: (require) =>
169
181
  }), /* @__PURE__ */ (0, react_jsx_runtime.jsxs)("div", {
170
182
  style: {
171
183
  display: "grid",
172
- gridTemplateColumns: "1fr",
173
- gap: "10px",
184
+ gridTemplateColumns: "1fr 1fr",
185
+ gap: "10px 14px",
174
186
  marginTop: "12px"
175
187
  },
176
188
  children: [FIELDS.map((f) => {
177
189
  const overridden = Object.hasOwn(userLayer, f.key);
190
+ const isBoolean = f.type === "boolean";
178
191
  return /* @__PURE__ */ (0, react_jsx_runtime.jsxs)("label", {
179
192
  style: {
180
193
  display: "flex",
181
194
  flexDirection: "column",
182
195
  gap: "4px",
183
- fontSize: "13px"
196
+ fontSize: "13px",
197
+ ...isBoolean ? {
198
+ gridColumn: "span 1",
199
+ justifyContent: "center"
200
+ } : {}
184
201
  },
185
202
  children: [
186
203
  /* @__PURE__ */ (0, react_jsx_runtime.jsxs)("span", {
@@ -193,7 +210,45 @@ window.__ModuleLoader__.load({ id: "@puwenhui/dsh-rag-kb", factory: (require) =>
193
210
  children: "·已覆盖"
194
211
  })]
195
212
  }),
196
- /* @__PURE__ */ (0, react_jsx_runtime.jsx)("input", {
213
+ isBoolean ? /* @__PURE__ */ (0, react_jsx_runtime.jsxs)("div", {
214
+ style: {
215
+ display: "flex",
216
+ alignItems: "center",
217
+ gap: "12px",
218
+ padding: "4px 0"
219
+ },
220
+ children: [/* @__PURE__ */ (0, react_jsx_runtime.jsxs)("label", {
221
+ style: {
222
+ display: "flex",
223
+ alignItems: "center",
224
+ gap: "4px",
225
+ cursor: "pointer"
226
+ },
227
+ children: [/* @__PURE__ */ (0, react_jsx_runtime.jsx)("input", {
228
+ type: "radio",
229
+ name: f.key,
230
+ checked: shown[f.key] === true,
231
+ onChange: () => {
232
+ edit(f.key, true);
233
+ }
234
+ }), /* @__PURE__ */ (0, react_jsx_runtime.jsx)("span", { children: "开" })]
235
+ }), /* @__PURE__ */ (0, react_jsx_runtime.jsxs)("label", {
236
+ style: {
237
+ display: "flex",
238
+ alignItems: "center",
239
+ gap: "4px",
240
+ cursor: "pointer"
241
+ },
242
+ children: [/* @__PURE__ */ (0, react_jsx_runtime.jsx)("input", {
243
+ type: "radio",
244
+ name: f.key,
245
+ checked: shown[f.key] !== true,
246
+ onChange: () => {
247
+ edit(f.key, false);
248
+ }
249
+ }), /* @__PURE__ */ (0, react_jsx_runtime.jsx)("span", { children: "关" })]
250
+ })]
251
+ }) : /* @__PURE__ */ (0, react_jsx_runtime.jsx)("input", {
197
252
  value: draft !== null && draft[f.key] !== void 0 ? String(draft[f.key]) : String(shown[f.key] ?? ""),
198
253
  placeholder: f.placeholder ?? "",
199
254
  onChange: (e) => {
@@ -212,6 +267,7 @@ window.__ModuleLoader__.load({ id: "@puwenhui/dsh-rag-kb", factory: (require) =>
212
267
  }, f.key);
213
268
  }), /* @__PURE__ */ (0, react_jsx_runtime.jsxs)("div", {
214
269
  style: {
270
+ gridColumn: "1 / -1",
215
271
  display: "flex",
216
272
  alignItems: "center",
217
273
  gap: "10px",
package/lib/plugin.mjs CHANGED
@@ -25,18 +25,22 @@ var __require = /* #__PURE__ */ (() => createRequire(import.meta.url))();
25
25
  /** 用户可见配置(设置面板「知识库」标签页) */
26
26
  const PublicConfig = z.object({
27
27
  watchDir: z.string().default("").description("本地文档目录(放入即自动索引)"),
28
- topK: z.number().default(5).description("检索返回条数(默认 5)"),
28
+ topK: z.number().default(5).description("混合检索召回条数(默认 5)"),
29
29
  minScore: z.number().default(.3).description("相似度阈值(0~1,默认 0.3)"),
30
- description: z.string().default("").description("知识库描述(agent 检索前会阅读)")
30
+ description: z.string().default("").description("知识库描述(agent 检索前会阅读)"),
31
+ enableReranker: z.boolean().default(false).description("启用 reranker 重排序(首次下载约 500MB 模型)"),
32
+ rerankTopN: z.number().default(3).description("重排后返回条数(默认 3)")
31
33
  });
32
34
  /** 完整配置(含高级项) */
33
35
  const Config = z.object({
34
36
  watchDir: z.string().default("").description("本地文档目录"),
35
- topK: z.number().default(5).description("检索返回条数"),
37
+ topK: z.number().default(5).description("混合检索召回条数"),
36
38
  minScore: z.number().default(.3).description("相似度阈值"),
37
39
  description: z.string().default("").description("知识库描述"),
38
40
  embeddingModel: z.string().default("Xenova/bge-small-zh-v1.5").description("embedding 模型"),
39
- batchSize: z.number().default(32).description("批量推理大小")
41
+ batchSize: z.number().default(32).description("批量推理大小"),
42
+ enableReranker: z.boolean().default(false).description("启用 reranker 重排序"),
43
+ rerankTopN: z.number().default(3).description("重排后返回条数")
40
44
  });
41
45
  /** 全字段 volatile 注入(npm 版 schemastery 无 .volatile(),0.2.0 设置面板只收录 volatile 字段) */
42
46
  for (const schema of [PublicConfig, Config]) for (const child of Object.values(schema.dict ?? {})) child.meta.volatile = true;
@@ -48,7 +52,9 @@ function resolveKbConfig(c = {}) {
48
52
  minScore: Math.max(0, Math.min(1, c.minScore ?? .3)),
49
53
  description: c.description ?? "",
50
54
  embeddingModel: c.embeddingModel ?? "Xenova/bge-small-zh-v1.5",
51
- batchSize: Math.max(1, Math.min(128, c.batchSize ?? 32))
55
+ batchSize: Math.max(1, Math.min(128, c.batchSize ?? 32)),
56
+ enableReranker: c.enableReranker === true,
57
+ rerankTopN: Math.max(1, Math.min(20, c.rerankTopN ?? 3))
52
58
  };
53
59
  }
54
60
  //#endregion
@@ -95,6 +101,10 @@ var KbStore = class {
95
101
  FOREIGN KEY(doc_id) REFERENCES documents(doc_id) ON DELETE CASCADE
96
102
  );
97
103
  CREATE INDEX IF NOT EXISTS idx_chunks_doc ON chunks(doc_id);
104
+ CREATE VIRTUAL TABLE IF NOT EXISTS chunk_fts USING fts5(
105
+ text,
106
+ tokenize='trigram'
107
+ );
98
108
  `);
99
109
  }
100
110
  /** 文档内容 hash 作 doc_id(同名不同内容=不同文档) */
@@ -118,6 +128,9 @@ var KbStore = class {
118
128
  this.loaded = false;
119
129
  }
120
130
  deleteDoc(docId) {
131
+ const ids = this.db.prepare("SELECT id FROM chunks WHERE doc_id=?").all(docId);
132
+ const ftsDel = this.db.prepare("DELETE FROM chunk_fts WHERE rowid=?");
133
+ for (const { id } of ids) ftsDel.run(id);
121
134
  this.db.prepare("DELETE FROM chunks WHERE doc_id=?").run(docId);
122
135
  this.db.prepare("DELETE FROM documents WHERE doc_id=?").run(docId);
123
136
  this.loaded = false;
@@ -128,13 +141,22 @@ var KbStore = class {
128
141
  getDoc(docId) {
129
142
  return this.db.prepare("SELECT * FROM documents WHERE doc_id=?").get(docId);
130
143
  }
131
- /** 批量写入切块+向量(一个事务,失败回滚) */
144
+ /** 批量写入切块+向量+FTS 索引(一个事务,失败回滚) */
132
145
  insertChunks(docId, chunks) {
133
146
  this.db.exec("BEGIN");
134
147
  try {
135
148
  this.db.prepare("DELETE FROM chunks WHERE doc_id=?").run(docId);
149
+ const oldIds = this.db.prepare("SELECT id FROM chunks WHERE doc_id=?").all(docId);
150
+ const ftsDel = this.db.prepare("DELETE FROM chunk_fts WHERE rowid=?");
151
+ for (const { id } of oldIds) ftsDel.run(id);
136
152
  const stmt = this.db.prepare("INSERT INTO chunks(doc_id, seq, text, embedding) VALUES(?,?,?,?)");
137
- for (const c of chunks) stmt.run(docId, c.seq, c.text, Buffer.from(c.embedding.buffer, c.embedding.byteOffset, c.embedding.byteLength));
153
+ const ftsIns = this.db.prepare("INSERT INTO chunk_fts(rowid, text) VALUES(?,?)");
154
+ for (const c of chunks) {
155
+ const buf = Buffer.from(c.embedding.buffer, c.embedding.byteOffset, c.embedding.byteLength);
156
+ const info = stmt.run(docId, c.seq, c.text, buf);
157
+ const rowid = Number(info.lastInsertRowid);
158
+ ftsIns.run(rowid, c.text);
159
+ }
138
160
  this.db.exec("COMMIT");
139
161
  } catch (e) {
140
162
  this.db.exec("ROLLBACK");
@@ -157,22 +179,69 @@ var KbStore = class {
157
179
  }
158
180
  this.loaded = true;
159
181
  }
160
- /** 余弦相似检索(向量已归一化,等价点积) */
161
- search(queryVec, topK, minScore) {
182
+ /** 混合检索:向量余弦 + FTS5 关键词双路合并(加权排序) */
183
+ hybridSearch(queryVec, queryText, topK, minScore) {
162
184
  this.ensureLoaded();
163
- const scored = [];
185
+ const vectorScores = /* @__PURE__ */ new Map();
164
186
  for (let i = 0; i < this.vectors.length; i++) {
165
187
  const v = this.vectors[i];
166
188
  let dot = 0;
167
189
  for (let j = 0; j < queryVec.length; j++) dot += queryVec[j] * v[j];
168
- if (dot >= minScore) scored.push({
169
- docId: this.vectorRows[i].docId,
170
- seq: this.vectorRows[i].seq,
171
- score: dot
172
- });
190
+ if (dot >= minScore) {
191
+ const key = `${this.vectorRows[i].docId}:${this.vectorRows[i].seq}`;
192
+ vectorScores.set(key, {
193
+ docId: this.vectorRows[i].docId,
194
+ seq: this.vectorRows[i].seq,
195
+ score: dot
196
+ });
197
+ }
173
198
  }
174
- scored.sort((a, b) => b.score - a.score);
175
- return scored.slice(0, topK);
199
+ const keywordScores = /* @__PURE__ */ new Map();
200
+ if (queryText.trim().length >= 3) try {
201
+ const safe = queryText.replace(/["'*()\-:]/g, " ").trim().slice(0, 50);
202
+ if (safe.length >= 3) {
203
+ const rows = this.db.prepare(`
204
+ SELECT c.doc_id, c.seq, bm25(chunk_fts) AS rank
205
+ FROM chunk_fts f
206
+ JOIN chunks c ON c.id = f.rowid
207
+ WHERE chunk_fts MATCH ?
208
+ ORDER BY rank
209
+ LIMIT ?
210
+ `).all(`"${safe}"`, topK * 2);
211
+ for (const row of rows) {
212
+ const key = `${row.doc_id}:${row.seq}`;
213
+ const kwScore = 1 / (1 + Math.abs(row.rank));
214
+ keywordScores.set(key, {
215
+ docId: row.doc_id,
216
+ seq: row.seq,
217
+ score: kwScore
218
+ });
219
+ }
220
+ }
221
+ } catch {}
222
+ const merged = /* @__PURE__ */ new Map();
223
+ for (const [key, v] of vectorScores) merged.set(key, {
224
+ ...v,
225
+ score: v.score,
226
+ via: keywordScores.has(key) ? "both" : "vector"
227
+ });
228
+ for (const [key, k] of keywordScores) if (merged.has(key)) {
229
+ const existing = merged.get(key);
230
+ existing.score = existing.score + k.score * .3;
231
+ existing.via = "both";
232
+ } else merged.set(key, {
233
+ ...k,
234
+ via: "keyword"
235
+ });
236
+ return [...merged.values()].sort((a, b) => b.score - a.score).slice(0, topK);
237
+ }
238
+ /** 向后兼容:纯向量检索(内部调 hybridSearch 空 queryText) */
239
+ search(queryVec, topK, minScore) {
240
+ return this.hybridSearch(queryVec, "", topK, minScore).map((r) => ({
241
+ docId: r.docId,
242
+ seq: r.seq,
243
+ score: r.score
244
+ }));
176
245
  }
177
246
  /** 按 doc_id + seq 取 chunk 原文 */
178
247
  chunkText(docId, seq) {
@@ -192,11 +261,17 @@ var indexer_exports = /* @__PURE__ */ __exportAll({
192
261
  chunkText: () => chunkText,
193
262
  extractText: () => extractText,
194
263
  getEmbedder: () => getEmbedder,
195
- indexFile: () => indexFile
264
+ indexFile: () => indexFile,
265
+ rerank: () => rerank
196
266
  });
197
267
  /** 中文友好的切块参数 */
198
268
  const CHUNK_TOKENS = 400;
199
- const OVERLAP_RATIO = .1;
269
+ /**
270
+ * 节标题行:`##` ~ `######` 视为节标题,遇之强制断块(保证「一节 = 一块」)。
271
+ * 刻意不含单级 `#`——nginx/shell/mysql 等配置与代码的注释普遍用单 `#` 开头,
272
+ * 误判会把命令序列切碎(实测语音平台文档 27 节被切成 231 块)。
273
+ */
274
+ const HEADING_RE = /^#{2,6} /;
200
275
  async function extractText(filePath) {
201
276
  const ext = extname(filePath).toLowerCase();
202
277
  const buf = await readFile(filePath);
@@ -215,58 +290,86 @@ async function extractText(filePath) {
215
290
  if (ext === ".md" || ext === ".txt" || ext === ".log" || ext === ".csv" || ext === ".json") return { text: buf.toString("utf8") };
216
291
  throw new Error(`不支持的文件类型 ${ext}(支持 pdf/docx/md/txt/csv/json)`);
217
292
  }
218
- /** 标题感知切块:按标题/空行分段,过长段再固定长度切 */
293
+ /** 标题感知切块:节标题强制断块(一节 = 一块)+ 空行分段 + 过长滑窗兜底 */
219
294
  function chunkText(text) {
220
295
  const normalized = text.replace(/\r\n/g, "\n").replace(/\n{3,}/g, "\n\n").trim();
221
296
  if (normalized === "") return [];
222
- const paragraphs = normalized.split(/\n(?=#{1,4} )|\n\n+/).map((p) => p.trim()).filter((p) => p !== "");
297
+ const paragraphs = normalized.split(/\n(?=#{2,6} )|\n\n+/).map((p) => p.trim()).filter((p) => p !== "");
223
298
  const chunks = [];
224
299
  let current = "";
300
+ let sectionHeading = "";
225
301
  const maxChars = CHUNK_TOKENS * 1.5;
226
- const overlap = Math.floor(maxChars * OVERLAP_RATIO);
227
302
  for (const para of paragraphs) {
228
- if (current !== "" && (current + "\n" + para).length > maxChars) {
303
+ const isHeading = HEADING_RE.test(para);
304
+ if (isHeading) sectionHeading = para.split("\n")[0].trim();
305
+ if (isHeading && current !== "") {
229
306
  chunks.push(current);
230
- current = current.length > overlap ? current.slice(-60) : "";
307
+ current = "";
231
308
  }
232
- if (para.length > maxChars * 2) {
233
- if (current !== "") {
234
- chunks.push(current);
235
- current = "";
309
+ if (para.length > maxChars * 1.2) {
310
+ const prefix = current;
311
+ current = "";
312
+ const mergePrefix = prefix !== "" && prefix.length <= maxChars * .6;
313
+ if (prefix !== "" && !mergePrefix) chunks.push(prefix);
314
+ for (let i = 0; i < para.length;) {
315
+ let end = Math.min(i + maxChars, para.length);
316
+ if (end < para.length) {
317
+ const nl = para.lastIndexOf("\n", end);
318
+ if (nl > i + maxChars * .5) end = nl + 1;
319
+ }
320
+ const piece = para.slice(i, end);
321
+ const isFirst = i === 0;
322
+ i = end;
323
+ chunks.push(isFirst ? mergePrefix ? `${prefix}\n\n${piece}` : piece : sectionHeading !== "" ? `${sectionHeading}\n\n${piece}` : piece);
236
324
  }
237
- for (let i = 0; i < para.length; i += 540) chunks.push(para.slice(i, i + maxChars));
325
+ current = sectionHeading;
238
326
  continue;
239
327
  }
328
+ const merged = current + "\n" + para;
329
+ if (current !== "" && merged.length > maxChars && (current.length >= maxChars * .5 || merged.length > maxChars * 1.2)) {
330
+ chunks.push(current);
331
+ const tail = current.slice(-60);
332
+ current = sectionHeading !== "" ? `${sectionHeading}\n\n${tail}` : tail;
333
+ }
240
334
  current = current === "" ? para : current + "\n" + para;
241
335
  }
242
336
  if (current !== "") chunks.push(current);
243
337
  return chunks.filter((c) => c.trim().length >= 10);
244
338
  }
245
- let pipelinePromise;
246
- async function loadPipeline(modelName) {
339
+ let gpuPipeline;
340
+ let cpuPipeline;
341
+ async function loadPipeline(modelName, device) {
247
342
  const { pipeline, env } = await import("@huggingface/transformers");
248
343
  const { dshHomePath } = await import("@deepseek-ai/dsh-home-paths");
249
344
  env.cacheDir = dshHomePath("rag-kb", "hf-cache");
250
- return pipeline("feature-extraction", modelName);
345
+ try {
346
+ return await pipeline("feature-extraction", modelName, { device });
347
+ } catch (e) {
348
+ if (device === "dml") {
349
+ console.log("[rag-kb] DML GPU 不可用,回退 CPU");
350
+ return await pipeline("feature-extraction", modelName, { device: "cpu" });
351
+ }
352
+ throw e;
353
+ }
251
354
  }
252
355
  async function getEmbedder(modelName) {
253
- pipelinePromise ??= loadPipeline(modelName);
254
- const embedder = await pipelinePromise;
255
356
  return {
256
- /** 单条向量化(检索查询用) */
357
+ /** 单条向量化(检索查询用,CPU 单条比 GPU 快) */
257
358
  async embed(text) {
258
- const out = await embedder(text, {
359
+ cpuPipeline ??= await loadPipeline(modelName, "cpu");
360
+ const out = await cpuPipeline(text, {
259
361
  pooling: "cls",
260
362
  normalize: true
261
363
  });
262
364
  return new Float32Array(out.data);
263
365
  },
264
- /** 批量向量化(索引用,batch 推理吃满多核/GPU) */
366
+ /** 批量向量化(索引用,GPU/DML 批量快 4 倍,不可用自动回退 CPU) */
265
367
  async embedBatch(texts, batchSize) {
368
+ gpuPipeline ??= await loadPipeline(modelName, "dml");
266
369
  const results = [];
267
370
  for (let i = 0; i < texts.length; i += batchSize) {
268
371
  const batch = texts.slice(i, i + batchSize);
269
- const out = await embedder(batch, {
372
+ const out = await gpuPipeline(batch, {
270
373
  pooling: "cls",
271
374
  normalize: true
272
375
  });
@@ -277,6 +380,31 @@ async function getEmbedder(modelName) {
277
380
  }
278
381
  };
279
382
  }
383
+ let rerankerPipeline;
384
+ async function loadReranker() {
385
+ const { pipeline, env } = await import("@huggingface/transformers");
386
+ const { dshHomePath } = await import("@deepseek-ai/dsh-home-paths");
387
+ env.cacheDir = dshHomePath("rag-kb", "hf-cache");
388
+ return pipeline("text-classification", "Xenova/bge-reranker-base");
389
+ }
390
+ /**
391
+ * 对混合检索结果做 cross-encoder 精排。
392
+ * 输入:query + 候选列表(含原文),输出按相关性重排序。
393
+ */
394
+ async function rerank(query, candidates) {
395
+ rerankerPipeline ??= await loadReranker();
396
+ const pairs = candidates.map((c) => ({
397
+ text: query,
398
+ text_pair: c.text
399
+ }));
400
+ const scores = await rerankerPipeline(pairs);
401
+ const scored = candidates.map((c, i) => ({
402
+ ...c,
403
+ rerankScore: scores[i].score
404
+ }));
405
+ scored.sort((a, b) => b.rerankScore - a.rerankScore);
406
+ return scored;
407
+ }
280
408
  /** 完整索引管线:文件路径 → 切块+向量数组 */
281
409
  async function indexFile(filePath, embedder, batchSize, onProgress) {
282
410
  const { text } = await extractText(filePath);
@@ -355,16 +483,46 @@ async function indexOne(ctx, filePath, source) {
355
483
  };
356
484
  }
357
485
  }
358
- /** 语义检索 */
486
+ /** 混合语义检索(向量+关键词)+ 可选 reranker 精排 + 引用元数据 */
359
487
  async function search(query, topK) {
360
488
  const s = await ensureStore();
361
489
  const queryVec = await (await getEmbedder(cfg?.embeddingModel ?? "Xenova/bge-small-zh-v1.5")).embed(query);
362
- return s.search(queryVec, topK ?? cfg?.topK ?? 5, cfg?.minScore ?? .3).map((h) => ({
363
- doc: s.getDoc(h.docId)?.name ?? h.docId,
364
- seq: h.seq,
365
- score: Number(h.score.toFixed(3)),
366
- text: s.chunkText(h.docId, h.seq).slice(0, 800)
490
+ const effectiveTopK = topK ?? cfg?.topK ?? 5;
491
+ const recallK = cfg?.enableReranker === true ? effectiveTopK * 3 : effectiveTopK;
492
+ let results = s.hybridSearch(queryVec, query, recallK, cfg?.minScore ?? .3).map((h) => {
493
+ const docName = s.getDoc(h.docId)?.name ?? h.docId;
494
+ const fullText = s.chunkText(h.docId, h.seq);
495
+ return {
496
+ docId: h.docId,
497
+ doc: docName,
498
+ seq: h.seq,
499
+ score: Number(h.score.toFixed(3)),
500
+ via: h.via,
501
+ text: fullText,
502
+ citation: `${docName} 第${h.seq + 1}段`,
503
+ position: `文档「${docName}」第 ${h.seq + 1} 段`
504
+ };
505
+ });
506
+ if (cfg?.enableReranker === true && results.length > 0) try {
507
+ const reranked = await rerank(query, results);
508
+ const topN = cfg?.rerankTopN ?? 3;
509
+ results = reranked.slice(0, topN).map((r) => ({
510
+ ...r,
511
+ text: r.text.slice(0, 800),
512
+ rerankScore: Number(r.rerankScore.toFixed(4))
513
+ }));
514
+ } catch (e) {
515
+ console.warn("[rag-kb] reranker 失败,使用混合检索结果:", String(e.message ?? e));
516
+ results = results.slice(0, effectiveTopK).map((r) => ({
517
+ ...r,
518
+ text: r.text.slice(0, 800)
519
+ }));
520
+ }
521
+ else results = results.slice(0, effectiveTopK).map((r) => ({
522
+ ...r,
523
+ text: r.text.slice(0, 800)
367
524
  }));
525
+ return results;
368
526
  }
369
527
  function apply(ctx, config = {}) {
370
528
  const unwrap = (v) => v !== null && typeof v === "object" && "get" in v ? v.get() : v;
@@ -432,12 +590,12 @@ function apply(ctx, config = {}) {
432
590
  const disposers = [
433
591
  ctx.tools.register({
434
592
  name: "knowledge_search",
435
- description: "在本地知识库中语义检索。返回最相关的知识块(带来源文档名、相似度、原文片段)。回答与知识库内容相关的问题前应先检索。query 支持自然语言。",
593
+ description: "在本地知识库中混合检索(语义向量 + 关键词)。返回最相关的知识块,每条带 citation 引用标注。回答知识库相关问题时:先检索 → 组织回答 → 在答案中标注来源 citation。query 支持自然语言和精确关键词(错误码/型号/人名等)。",
436
594
  parameters: compileParams({
437
595
  query: {
438
596
  type: "string",
439
597
  required: true,
440
- description: "检索问题(自然语言,中英文均可)"
598
+ description: "检索问题(自然语言或精确关键词)"
441
599
  },
442
600
  top_k: {
443
601
  type: "number",
@@ -635,11 +793,13 @@ function apply(ctx, config = {}) {
635
793
  const docs = store?.listDocs().filter((d) => d.status === "ready").length ?? 0;
636
794
  return [
637
795
  "You have a local knowledge base accessible via these tools:",
638
- "- knowledge_search: semantic search the KB (use BEFORE answering questions about its content)",
796
+ "- knowledge_search: hybrid search (semantic + keyword). Use BEFORE answering questions about KB content.",
639
797
  "- knowledge_status: list documents and config",
640
798
  "- knowledge_manage: add/save/remove/reindex documents",
799
+ "IMPORTANT: When the user says 搜一下/查一下/搜索, FIRST try knowledge_search on the local KB. Only do a web search if KB has no relevant results AND the topic is clearly not about local documents. Do NOT search both simultaneously — this confuses the answer.",
641
800
  desc !== "" ? `KB description: ${desc}` : "",
642
801
  `Currently ${docs} document(s) indexed.`,
802
+ "CITATION STYLE (footnote format): In the body of your answer, mark sources with superscript numbers like ¹ ² ³ (Unicode superscripts). At the END, add a \"参考来源\" section. IMPORTANT: MERGE citations from the same document into ONE entry — e.g. if citing 文档A 第1段/第2段/第3段, list as \"1. 文档A 第1-3段\" (not 3 separate lines). Only list each document once with all its paragraph numbers combined. Keep the answer body clean.",
643
803
  "When the user says \"save this to the knowledge base\" (保存到知识库/记下来), extract the valuable content from the conversation and call knowledge_manage with action=\"save\", providing a concise title and the text.",
644
804
  "For file/directory indexing use action=\"add\" with a path."
645
805
  ].filter(Boolean).join("\n");
@@ -652,6 +812,6 @@ function apply(ctx, config = {}) {
652
812
  });
653
813
  }
654
814
  //#endregion
655
- export { Config, KbStore, apply, chunkText, extractText, getEmbedder, indexFile, inject, name, resolveKbConfig };
815
+ export { Config, KbStore, apply, chunkText, extractText, getEmbedder, indexFile, inject, name, rerank, resolveKbConfig };
656
816
 
657
817
  //# sourceMappingURL=plugin.mjs.map
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "@puwenhui/dsh-rag-kb",
3
- "version": "0.1.0",
4
- "description": "DSH RAG 知识库插件:文档导入→切块→本地向量化→语义检索,knowledge_search 工具 + 会话语义管理 + 设置面板",
3
+ "version": "0.3.0",
4
+ "description": "DSH RAG 知识库插件:混合检索(向量+FTS5)+Reranker 重排+GPU 加速+引用溯源+会话语义管理",
5
5
  "publishConfig": {
6
6
  "registry": "https://registry.npmjs.org/",
7
7
  "access": "public"