dsh-novel-writer 1.5.0 → 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.en.md +20 -0
- package/README.md +46 -27
- package/cordis.patch.yml +1 -0
- package/lib/analysis.js +301 -10
- package/lib/client.js +371 -158
- package/lib/embedding.js +252 -0
- package/lib/index.js +280 -12
- package/lib/models/config.json +31 -0
- package/lib/models/onnx/model_quantized.onnx +0 -0
- package/lib/models/special_tokens_map.json +7 -0
- package/lib/models/tokenizer.json +21278 -0
- package/lib/models/tokenizer_config.json +15 -0
- package/package.json +8 -6
- package/skills/novel-writing/SKILL.md +31 -2
- package/INSTALL.md +0 -36
- package/test/client-test.mjs +0 -34
- package/test/e2e-test.mjs +0 -155
- package/test/pattern-test.mjs +0 -48
package/README.en.md
ADDED
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
# dsh-novel-writer — Novel Writing Assistant Plugin (v2.0.0)
|
|
2
|
+
|
|
3
|
+
A DSH (DeepSeek Harness) bundle plugin. **v2.0.0 ships a local semantic embedding engine (bge-small-zh-v1.5 ONNX, bundled, CPU inference, 0 token cost)** — semantic capabilities without any API fees.
|
|
4
|
+
|
|
5
|
+
## Tools (14, each with an independent UI switch)
|
|
6
|
+
novel_books / novel_chapters / novel_read / novel_keywords / novel_new_chapter / novel_import / novel_sentence_analysis / novel_sentence_config / novel_style_check / novel_plot / novel_settings / novel_summary / novel_continuity_check / **novel_semantic_search (new in v2.0.0)**
|
|
7
|
+
|
|
8
|
+
## v2.0.0 — Semantic Layer
|
|
9
|
+
- **Local engine** (lib/embedding.js): Xenova/bge-small-zh-v1.5 (quantized), 512-dim Chinese vectors, ~23MB, bundled; @huggingface/tokenizers + onnxruntime-web (WASM), local CPU, 0 token / 0 API cost;
|
|
10
|
+
- **Lazy load** (~0.3s first call); automatic fallback to pure-rule mode if unavailable;
|
|
11
|
+
- **Index cache**: per-book vectors at `.novel-writer/embedding/<book>.json` — repeated searches are instant;
|
|
12
|
+
- **novel_semantic_search**: natural-language search across the whole book (foreshadowing clues, emotional scenes, setting mentions) — matches by meaning even when keywords differ;
|
|
13
|
+
- **novel_style_check upgrade**: adds semantic similarity (target chapter vs rest of book) alongside rule-based fingerprint similarity;
|
|
14
|
+
- **Switch: semanticEmbedding** — default ON (probe-based: auto-enabled when the model is present); turn it off in the sidebar「写作助手功能」panel to fully return to rule-only mode.
|
|
15
|
+
|
|
16
|
+
## History (v0.3 → v1.6)
|
|
17
|
+
Sentence-pattern analysis; emotion purification + AI review (strong/weak lexicon, pollution caveat); emotion quantification (Valence sliding window: variance/slope/conflict + implicit imagery + complexity score); worldview/genre/theme detection (modern/western/eastern + 15 genres + 35 themes); pragmatic-style review; settings tables (5) + foreshadowing registry + summaries + continuity audit + style check; analysis cache & report export (`.novel-writer/analysis/`).
|
|
18
|
+
|
|
19
|
+
## Size
|
|
20
|
+
~95MB total (23MB quantized model + 70MB runtime); zip 31MB. If disk is tight, disable the semantic switch and delete `lib/models/` + `node_modules/` to go back to rule-only.
|
package/README.md
CHANGED
|
@@ -1,36 +1,55 @@
|
|
|
1
|
-
# dsh-novel-writer — 小说写作助手插件(
|
|
1
|
+
# dsh-novel-writer — 小说写作助手插件(v2.0.0)
|
|
2
2
|
|
|
3
|
-
DSH(DeepSeek Harness)小说写作助手 bundle
|
|
4
|
-
浏览器端仅依赖 Web GUI 自带的 react。
|
|
5
|
-
**v0.9.0 在 v0.8.0 基础上补全「世界观层」:第五张设定表「世界观/用语规范」、文化基准自动判断(detect)、用语风格扫描、续写用语一致性检查**:分析缓存与报告导出、风格自检、伏笔登记表、
|
|
6
|
-
关键词三字组/疑似人名、情感词典去噪、祈使句与环境分类改进、局域网开关可选放行。
|
|
3
|
+
DSH(DeepSeek Harness)小说写作助手 bundle 插件。**v2.0.0 集成本地语义嵌入引擎(bge-small-zh-v1.5 ONNX,随插件分发,本地 CPU 推理,0 token 成本)**,在不增加任何 API 费用的前提下获得语义级能力。
|
|
7
4
|
|
|
8
|
-
## 工具清单(
|
|
5
|
+
## 工具清单(14 个,全部带独立 UI 开关)
|
|
9
6
|
|
|
10
|
-
novel_books / novel_chapters / novel_read / novel_keywords / novel_new_chapter / novel_import / novel_sentence_analysis / novel_sentence_config / novel_style_check / novel_plot /
|
|
7
|
+
novel_books / novel_chapters / novel_read / novel_keywords / novel_new_chapter / novel_import / novel_sentence_analysis / novel_sentence_config / novel_style_check / novel_plot / novel_settings(设定管理)/ novel_summary(章节摘要)/ novel_continuity_check(连贯性审计)/ **novel_semantic_search(语义检索,v2.0.0 新增)**
|
|
11
8
|
|
|
12
|
-
##
|
|
13
|
-
- **novel_settings 第五张表「世界观/用语规范」(category=worldview)**:
|
|
14
|
-
- `detect`:自动判断文化基准(western/eastern/mixed/unknown)+ 置信度 + 证据(中西词表按时代错置分类:器物/称谓/计量/宗教仪式/市井风貌/服饰/食物/制度);
|
|
15
|
-
- 登记:name / basis / bannedWords(禁用词)/ recommended(替代词映射)/ ritual(仪式规范);
|
|
16
|
-
- 文化基准由 AI 自行判断,不预设东西方。
|
|
17
|
-
- **novel_continuity_check 用语风格扫描**:对照 worldview 禁用词表输出「用语冲突」候选(带建议替换词,未登记时用默认基准并提示 detect);
|
|
18
|
-
- **续写提示词**:新增世界观一致性检查步骤。
|
|
9
|
+
## v2.0.0 新增能力(语义层)
|
|
19
10
|
|
|
20
|
-
|
|
11
|
+
### 本地语义嵌入引擎(embedding.js)
|
|
12
|
+
- 模型:Xenova/bge-small-zh-v1.5(512 维中文向量,约 91MB),随插件分发,开箱即用;
|
|
13
|
+
- 运行时:**@huggingface/tokenizers**(官方轻量分词,294KB 零依赖)+ onnxruntime-web(WASM 推理),本地 CPU,0 token、0 API 费用;
|
|
14
|
+
- 懒加载:首次调用才加载(约 0.3s);加载失败自动回退纯规则,不影响任何既有功能;
|
|
15
|
+
- 索引缓存:每书向量落盘 `.novel-writer/embedding/<书>.json`,重复检索秒开(47 万字约 3100 段,首次建索引约 20-30s,之后命中缓存)。
|
|
21
16
|
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
- **novel_continuity_check**:对照设定表扫描全书 → 数字口径/人物缺场/别名/重复条目矛盾候选;
|
|
25
|
-
- **novel_plot 字段化**:type/priority/relatedCharacters/locations/payoffCondition/mentionedIn/lastMentioned + scan 自动提及追踪;
|
|
26
|
-
- **细节密度指标**:动作链密度/物件名词/感官词(千字归一)进 novel_sentence_analysis 与 novel_style_check;
|
|
27
|
-
- **统一数据目录**:§BT§<书库根>/.novel-writer/§BT§ 下 plots/ settings/ summaries/ analysis/ audits/(伏笔旧位置自动迁移;分析缓存与关键词报告落盘到 analysis/);
|
|
28
|
-
- **UI 路径修复**:cordis 空字符串 root 回退 lastRoot、面板刷新按钮、打开/复制按钮不再禁用、显示 dataDir。
|
|
17
|
+
### 新工具:novel_semantic_search(语义检索)
|
|
18
|
+
用自然语言描述在全书中检索语义相关段落——找伏笔线索、情感场景、设定提及,**即使原文没有相同关键词也能命中**(如搜「压抑克制的时刻」能命中没有"压抑"二字的段落)。
|
|
29
19
|
|
|
30
|
-
|
|
20
|
+
### novel_style_check 升级:语义级风格对比
|
|
21
|
+
规则指纹相似度(句式/句长/情绪)之外,新增**语义相似度**(目标章 vs 全书其余部分向量余弦),双维度判定风格一致性。
|
|
31
22
|
|
|
32
|
-
|
|
33
|
-
-
|
|
34
|
-
-
|
|
35
|
-
|
|
23
|
+
### 功能开关:语义增强(semanticEmbedding)
|
|
24
|
+
- **默认开启(探测式)**:模型存在即自动启用;用户可在侧边栏「写作助手功能」面板关闭,关闭后语义检索返回提示、风格检查跳过语义维度;
|
|
25
|
+
- 关闭后插件完整回退到 v1.6.0 纯规则能力。
|
|
26
|
+
|
|
27
|
+
## 历史能力(v0.3 → v1.6 累积)
|
|
28
|
+
- 句式模式分析(九类句式分布/排列/段落/句长/情感曲线/风格指纹/密度);
|
|
29
|
+
- 情感净化 + AI 复核(强/弱情绪词分级、污染预警 caveat、强制抽查);
|
|
30
|
+
- 情感量化(Valence 滑动窗口:方差/斜率/矛盾指数 + 隐性意象 + 复杂度评分);
|
|
31
|
+
- 世界观/流派/题材检测(modern/western/eastern + 15 流派 + 35 题材,主副题材);
|
|
32
|
+
- 语用级审查(称谓/客套/仪式/语气,中西/现代禁用词);
|
|
33
|
+
- 设定管理五张表 + 伏笔登记 + 章节摘要 + 连贯性审计 + 风格自检;
|
|
34
|
+
- 分析缓存与报告导出(`.novel-writer/analysis/`)。
|
|
35
|
+
|
|
36
|
+
## 功能开关(侧边栏「写作助手功能」面板)
|
|
37
|
+
- 总开关 enabled;autoAnalyze;14 个工具级开关;
|
|
38
|
+
- 功能级:emotionCaveat(情感净化预警)/ genreTheme(题材与流派检测)/ emotionComplexity(情感量化)/ **semanticEmbedding(语义增强,v2.0.0)**。
|
|
39
|
+
|
|
40
|
+
## 目录结构
|
|
36
41
|
```
|
|
42
|
+
dsh-novel-writer-v2.0.0/
|
|
43
|
+
├── lib/
|
|
44
|
+
│ ├── index.js # 宿主端(14 工具 + state 路由 + 开关门禁)
|
|
45
|
+
│ ├── analysis.js # 规则引擎(句式/情感净化/情感量化/题材)
|
|
46
|
+
│ ├── embedding.js # 本地语义嵌入引擎(bge-small-zh ONNX)
|
|
47
|
+
│ ├── models/ # 模型文件(onnx/model.onnx 91MB + tokenizer)
|
|
48
|
+
│ └── client.js # 浏览器端(侧边栏面板 + 全部开关)
|
|
49
|
+
├── node_modules/ # @huggingface/tokenizers + onnxruntime-web(随插件分发)
|
|
50
|
+
├── test/ # 引擎 + client + e2e 测试
|
|
51
|
+
└── skills/novel-writing/SKILL.md # 模型使用指南
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
## 体积说明
|
|
55
|
+
**约 95MB**(模型 23MB + 运行时 70MB),zip 31MB;已移除 transformers.js/onnxruntime-node/sharp 全家桶(省 85%)。如磁盘紧张可在「语义增强」开关关闭后删除 `lib/models/` 与 `node_modules/`,插件回到纯规则模式。
|
package/cordis.patch.yml
CHANGED
|
@@ -1,3 +1,4 @@
|
|
|
1
|
+
# dsh-novel-writer v2.0.0 bundle patch: 14 tools (incl. novel_semantic_search) + local embedding engine.
|
|
1
2
|
# dsh-novel-writer bundle patch: mounts the novel-writing assistant plugin
|
|
2
3
|
# into the profile's loader tree. The entry's name is the plugin package
|
|
3
4
|
# itself, resolved from the profile's node_modules.
|
package/lib/analysis.js
CHANGED
|
@@ -1,13 +1,11 @@
|
|
|
1
1
|
/**
|
|
2
|
-
* dsh-novel-writer — 句式模式分析引擎 (v1.
|
|
2
|
+
* dsh-novel-writer — 句式模式分析引擎 (v1.6.0 合并版)
|
|
3
3
|
*
|
|
4
4
|
* 融合 v0.3.0(深度分析)与 v0.4.0(轻量节奏参考)两版引擎:
|
|
5
5
|
* - 九类句式:陈述 / 环境 / 心理 / 对话 / 疑问 / 反问 / 感叹 / 祈使 / 省略留白
|
|
6
6
|
* - 排列规律:句式转移、2/3 连句模板、段首段尾句式、按章节的压缩排列序列(跑长编码)
|
|
7
7
|
* - 风格特征:句长分布、短长句占比、对话/心理/环境密度、主观性指数、风格指纹
|
|
8
8
|
* - 主观情感:轻量情感词典(喜/怒/哀/惧/惊)+ 强度副词加权,输出情感曲线
|
|
9
|
-
* - v1.5.0 情感净化:强/弱情绪词分级(clean 只计强情绪词)、污染源检测(R18/战斗/恐怖
|
|
10
|
-
* 密度阈值)、caveat 预警与 aiAction AI 复核指令、confidence 可信度
|
|
11
9
|
* - 节奏建议:给模型的 guidance 文本(句式分布 + 高频组合 + 章节节奏序列)
|
|
12
10
|
* - 采样上限:maxSentences 保护超长文本(默认 20000 句)
|
|
13
11
|
*
|
|
@@ -70,10 +68,6 @@ const EMOTION_WORDS = Object.freeze({
|
|
|
70
68
|
fear: [...STRONG_EMOTION_WORDS.fear, ...WEAK_EMOTION_WORDS.fear],
|
|
71
69
|
surprise: [...STRONG_EMOTION_WORDS.surprise, ...WEAK_EMOTION_WORDS.surprise]
|
|
72
70
|
});
|
|
73
|
-
/** 去重后的情感词表(模块级缓存,避免每次分析重复 Set 构建)。 */
|
|
74
|
-
const EMOTION_WORDS_UNIQUE = Object.freeze(Object.fromEntries(
|
|
75
|
-
Object.entries(EMOTION_WORDS).map(([emotion, words]) => [emotion, [...new Set(words)]])
|
|
76
|
-
));
|
|
77
71
|
|
|
78
72
|
/** v1.5.0 情绪污染源词表:检测到高密度时降低情感可信度并提示 AI 复核。 */
|
|
79
73
|
const EMOTION_POLLUTION = Object.freeze({
|
|
@@ -411,7 +405,302 @@ export function buildGuidance(ratios, topPatterns, chapterSequences) {
|
|
|
411
405
|
}
|
|
412
406
|
|
|
413
407
|
/**
|
|
414
|
-
*
|
|
408
|
+
* v1.6.0 Valence 效价映射表(基于大连理工中文情感词汇本体库框架):
|
|
409
|
+
* 七维(乐/好/怒/哀/惧/恶/惊)× 强度五档(1/3/4/7/9 → ±0.1/0.3/0.5/0.7/0.9)。
|
|
410
|
+
* 乐=正向,好=正向,怒/哀/惧/恶/惊=负向(惊=中性偏负)。纯规则查表,0 token。
|
|
411
|
+
*/
|
|
412
|
+
const VALENCE_WORDS = Object.freeze({
|
|
413
|
+
"欣喜": 0.9, "狂喜": 0.9, "欢天喜地": 0.9, "雀跃": 0.7, "高兴": 0.7, "开心": 0.7, "快乐": 0.7,
|
|
414
|
+
"喜悦": 0.7, "愉快": 0.5, "欢喜": 0.7, "兴奋": 0.7, "愉悦": 0.5, "欢快": 0.7, "乐": 0.5,
|
|
415
|
+
"欢": 0.5, "喜": 0.5, "笑": 0.3, "微笑": 0.3, "笑容": 0.3, "笑眯眯": 0.3, "哈哈": 0.3,
|
|
416
|
+
"欣慰": 0.5, "满足": 0.3, "痛快": 0.5, "爽快": 0.3, "甜": 0.5, "甜蜜": 0.5, "幸福": 0.9,
|
|
417
|
+
"美满": 0.9, "温馨": 0.5, "温暖": 0.3, "踏实": 0.1, "平静": 0.1, "释然": 0.4, "解脱": 0.4,
|
|
418
|
+
"喜欢": 0.7, "喜爱": 0.7, "欣赏": 0.5, "爱": 0.9, "疼爱": 0.7, "宠爱": 0.7, "仰慕": 0.7,
|
|
419
|
+
"尊敬": 0.5, "敬仰": 0.7, "崇拜": 0.7, "心动": 0.5, "眷恋": 0.7, "依恋": 0.7, "思念": 0.3,
|
|
420
|
+
"心疼": 0.3, "怜爱": 0.5, "温柔": 0.3, "珍惜": 0.5, "信赖": 0.5, "感恩": 0.7,
|
|
421
|
+
"愤怒": -0.9, "暴怒": -0.9, "怒发冲冠": -0.9, "火冒三丈": -0.9, "恼羞成怒": -0.9,
|
|
422
|
+
"生气": -0.7, "恼火": -0.7, "气愤": -0.7, "怒": -0.7, "恨": -0.9, "怨恨": -0.9, "憎恨": -0.9,
|
|
423
|
+
"厌恶": -0.7, "憎恶": -0.9, "不满": -0.5, "恼": -0.5, "气": -0.3, "发火": -0.7, "咬牙": -0.3,
|
|
424
|
+
"怒意": -0.7, "怒火": -0.9, "愤恨": -0.9, "恼羞": -0.7, "气冲冲": -0.7, "咬牙切齿": -0.5,
|
|
425
|
+
"悲伤": -0.7, "悲痛": -0.9, "悲痛欲绝": -0.9, "悲哀": -0.7, "哀伤": -0.7, "难过": -0.5,
|
|
426
|
+
"伤心": -0.7, "痛苦": -0.7, "心碎": -0.9, "绝望": -0.9, "哭泣": -0.7, "哭": -0.5, "泪": -0.3,
|
|
427
|
+
"眼泪": -0.3, "流泪": -0.5, "哽咽": -0.5, "抽泣": -0.5, "叹息": -0.3, "叹气": -0.3,
|
|
428
|
+
"惆怅": -0.5, "失落": -0.5, "忧伤": -0.5, "黯然": -0.5, "心酸": -0.7, "辛酸": -0.7,
|
|
429
|
+
"悲": -0.7, "凄凉": -0.7, "苦涩": -0.7, "苦闷": -0.5, "沮丧": -0.7, "消沉": -0.7,
|
|
430
|
+
"落寞": -0.5, "孤寂": -0.5, "郁闷": -0.3, "低落": -0.3, "压抑": -0.5, "心灰意冷": -0.9,
|
|
431
|
+
"恐惧": -0.9, "害怕": -0.7, "惊慌": -0.7, "不安": -0.3, "紧张": -0.3, "担心": -0.3,
|
|
432
|
+
"畏惧": -0.7, "惊恐": -0.9, "胆怯": -0.5, "发抖": -0.3, "哆嗦": -0.3, "心慌": -0.5,
|
|
433
|
+
"毛骨悚然": -0.9, "冷汗": -0.5, "忐忑": -0.5, "惶恐": -0.9, "心悸": -0.5, "惊惶": -0.7,
|
|
434
|
+
"胆战心惊": -0.9, "心虚": -0.5, "焦虑": -0.5, "恐慌": -0.9,
|
|
435
|
+
"恶心": -0.7, "厌恶": -0.7, "鄙视": -0.7, "轻蔑": -0.5, "嫌弃": -0.7, "反感": -0.5,
|
|
436
|
+
"憎恶": -0.9, "作呕": -0.7,
|
|
437
|
+
"惊讶": -0.1, "震惊": -0.5, "意外": -0.1, "吃惊": -0.3, "诧异": -0.3, "愕然": -0.3,
|
|
438
|
+
"愣住": -0.1, "目瞪口呆": -0.5, "难以置信": -0.5, "不可思议": -0.3, "惊愕": -0.5,
|
|
439
|
+
"惊奇": 0.1, "震撼": -0.3, "傻眼": -0.3, "呆住": -0.1, "惊呆": -0.5, "骇然": -0.5
|
|
440
|
+
});
|
|
441
|
+
|
|
442
|
+
/**
|
|
443
|
+
* v1.6.0 隐性情感载体映射表(意象/动作 → 效价 + 标签 + 脆弱标记)。
|
|
444
|
+
* 参考中国古典诗歌意象体系 + 现代微动作意象,规则查表 0 token。
|
|
445
|
+
*/
|
|
446
|
+
const IMPLICIT_CARRIERS = Object.freeze([
|
|
447
|
+
{ words: ["雨", "阴雨", "细雨", "冷雨", "秋雨"], valence: -0.3, label: "压抑·萧瑟" },
|
|
448
|
+
{ words: ["黄昏", "暮色", "残阳", "夕阳西下"], valence: -0.3, label: "迟暮·萧瑟" },
|
|
449
|
+
{ words: ["枯枝", "落叶", "枯叶", "败叶"], valence: -0.3, label: "凋零·萧瑟" },
|
|
450
|
+
{ words: ["冷风", "寒风", "北风", "秋风"], valence: -0.3, label: "寒冷·孤寂" },
|
|
451
|
+
{ words: ["昏暗", "阴影", "灰暗", "幽暗"], valence: -0.3, label: "压抑" },
|
|
452
|
+
{ words: ["孤雁", "孤鸿", "寒鸦"], valence: -0.25, label: "孤独" },
|
|
453
|
+
{ words: ["残月", "冷月", "孤月"], valence: -0.25, label: "孤独·凄清" },
|
|
454
|
+
{ words: ["梧桐", "芭蕉"], valence: -0.25, label: "愁绪" },
|
|
455
|
+
{ words: ["寒蝉", "秋虫"], valence: -0.2, label: "凄切" },
|
|
456
|
+
{ words: ["荒芜", "废墟", "断壁", "残垣"], valence: -0.4, label: "荒凉·衰败" },
|
|
457
|
+
{ words: ["空荡", "空旷", "空落落"], valence: -0.3, label: "空虚" },
|
|
458
|
+
{ words: ["暖光", "暖阳", "炉火", "烛火", "灯火"], valence: 0.2, label: "短暂温暖", fragile: true },
|
|
459
|
+
{ words: ["茶烟", "炊烟", "轻烟"], valence: 0.2, label: "短暂温暖", fragile: true },
|
|
460
|
+
{ words: ["余晖", "黄昏的光"], valence: 0.15, label: "短暂温暖", fragile: true },
|
|
461
|
+
{ words: ["摩挲", "摩挲杯沿"], valence: -0.2, label: "焦虑·克制" },
|
|
462
|
+
{ words: ["攥紧衣角", "攥着衣角", "握紧衣角"], valence: -0.2, label: "焦虑·克制" },
|
|
463
|
+
{ words: ["咬唇", "咬住嘴唇", "咬着下唇"], valence: -0.2, label: "隐忍·克制" },
|
|
464
|
+
{ words: ["低垂眼帘", "垂下眼帘", "垂下眼"], valence: -0.2, label: "隐忍·欲言又止" },
|
|
465
|
+
{ words: ["沉默良久", "久久沉默"], valence: -0.2, label: "隐忍" },
|
|
466
|
+
{ words: ["指尖发白", "指节发白", "攥紧拳头"], valence: -0.3, label: "压抑·愤怒" },
|
|
467
|
+
{ words: ["颤抖的手", "手在抖"], valence: -0.3, label: "紧张·恐惧" },
|
|
468
|
+
{ words: ["转身", "转过身"], valence: -0.3, label: "疏离·决绝" },
|
|
469
|
+
{ words: ["走出", "推门而出", "大步离开"], valence: -0.3, label: "决绝" },
|
|
470
|
+
{ words: ["背影", "远去的背影"], valence: -0.3, label: "疏离·失落" },
|
|
471
|
+
{ words: ["回头", "回望"], valence: -0.2, label: "不舍·眷恋" },
|
|
472
|
+
{ words: ["轻笑", "苦笑", "扯了扯嘴角"], valence: -0.15, label: "无奈·强颜" },
|
|
473
|
+
{ words: ["摇头", "摇了摇头", "垂下头"], valence: -0.15, label: "无奈·妥协" },
|
|
474
|
+
{ words: ["垂手", "放下手", "手垂落"], valence: -0.15, label: "无力·放弃" }
|
|
475
|
+
]);
|
|
476
|
+
|
|
477
|
+
/** v1.6.0 隐性载体扫描:意象/动作 → 负向/正向占比 + top 载体 + 脆弱标记。 */
|
|
478
|
+
export function implicitEmotionScan(text) {
|
|
479
|
+
let negHits = 0, posHits = 0, fragileHits = 0;
|
|
480
|
+
const carrierCounts = new Map();
|
|
481
|
+
for (const carrier of IMPLICIT_CARRIERS) {
|
|
482
|
+
for (const word of carrier.words) {
|
|
483
|
+
const n = text.split(word).length - 1;
|
|
484
|
+
if (n > 0) {
|
|
485
|
+
if (carrier.valence < 0) negHits += n * -carrier.valence;
|
|
486
|
+
else { posHits += n * carrier.valence; if (carrier.fragile) fragileHits += n; }
|
|
487
|
+
carrierCounts.set(carrier.label + ":" + word, (carrierCounts.get(carrier.label + ":" + word) ?? 0) + n);
|
|
488
|
+
}
|
|
489
|
+
}
|
|
490
|
+
}
|
|
491
|
+
const total = negHits + posHits;
|
|
492
|
+
const top = [...carrierCounts.entries()].sort((x, y) => y[1] - x[1]).slice(0, 8).map(([k, n]) => ({ carrier: k.split(":")[1], label: k.split(":")[0], count: n }));
|
|
493
|
+
return {
|
|
494
|
+
negative: total === 0 ? 0 : Math.round((negHits / total) * 100) / 100,
|
|
495
|
+
positive: total === 0 ? 0 : Math.round((posHits / total) * 100) / 100,
|
|
496
|
+
fragile: fragileHits > 0,
|
|
497
|
+
totalHits: negHits + posHits,
|
|
498
|
+
topCarriers: top
|
|
499
|
+
};
|
|
500
|
+
}
|
|
501
|
+
|
|
502
|
+
/**
|
|
503
|
+
* v1.6.0 情感量化:Valence 滑动窗口时间序列 → 方差/斜率/矛盾指数。
|
|
504
|
+
*/
|
|
505
|
+
const EMOTION_CATS = ["joy", "anger", "sorrow", "fear", "surprise"];
|
|
506
|
+
const EMOTION_MAX_ENTROPY = Math.log(EMOTION_CATS.length);
|
|
507
|
+
|
|
508
|
+
/** v1.6.0 滑动窗口:每 100 字算平均效价 → 时间序列 + 正负词计数。 */
|
|
509
|
+
export function valenceSeries(text, winChars = 100) {
|
|
510
|
+
const series = [];
|
|
511
|
+
const windowPosNeg = [];
|
|
512
|
+
const entries = Object.entries(VALENCE_WORDS).sort((x, y) => y[0].length - x[0].length);
|
|
513
|
+
for (let start = 0; start < text.length; start += winChars) {
|
|
514
|
+
const slice = text.slice(start, start + winChars);
|
|
515
|
+
let sum = 0, n = 0, pos = 0, neg = 0;
|
|
516
|
+
for (const [word, val] of entries) {
|
|
517
|
+
let from = 0;
|
|
518
|
+
while (from < slice.length) {
|
|
519
|
+
const idx = slice.indexOf(word, from);
|
|
520
|
+
if (idx === -1) break;
|
|
521
|
+
sum += val; n += 1;
|
|
522
|
+
if (val > 0) pos += 1; else if (val < 0) neg += 1;
|
|
523
|
+
from = idx + word.length;
|
|
524
|
+
}
|
|
525
|
+
}
|
|
526
|
+
series.push(n === 0 ? 0 : Math.round((sum / n) * 1000) / 1000);
|
|
527
|
+
windowPosNeg.push({ pos, neg });
|
|
528
|
+
}
|
|
529
|
+
const posWords = windowPosNeg.reduce((s, w) => s + w.pos, 0);
|
|
530
|
+
const negWords = windowPosNeg.reduce((s, w) => s + w.neg, 0);
|
|
531
|
+
return { series, posWords, negWords, windowCount: series.length, windowPosNeg };
|
|
532
|
+
}
|
|
533
|
+
|
|
534
|
+
/** v1.6.0 三指标:方差 V + 相邻撕裂 V_adj + 斜率 Δ(最小二乘+鲁棒版)+ 矛盾指数 C。 */
|
|
535
|
+
export function valenceStats(text) {
|
|
536
|
+
const { series, posWords, negWords } = valenceSeries(text);
|
|
537
|
+
const n = series.length;
|
|
538
|
+
if (n === 0) return { windows: 0 };
|
|
539
|
+
const mean = series.reduce((x, y) => x + y, 0) / n;
|
|
540
|
+
const variance = series.reduce((s, x) => s + (x - mean) ** 2, 0) / n;
|
|
541
|
+
let adjSum = 0;
|
|
542
|
+
for (let i = 1; i < n; i += 1) adjSum += Math.abs(series[i] - series[i - 1]);
|
|
543
|
+
const adjVariance = n > 1 ? adjSum / (n - 1) : 0;
|
|
544
|
+
const iMean = (n - 1) / 2;
|
|
545
|
+
let num = 0, den = 0;
|
|
546
|
+
for (let i = 0; i < n; i += 1) {
|
|
547
|
+
num += (i - iMean) * (series[i] - mean);
|
|
548
|
+
den += (i - iMean) ** 2;
|
|
549
|
+
}
|
|
550
|
+
const slope = den === 0 ? 0 : num / den;
|
|
551
|
+
const delta = slope * (n - 1);
|
|
552
|
+
const third = Math.max(1, Math.floor(n / 3));
|
|
553
|
+
const headMean = series.slice(0, third).reduce((x, y) => x + y, 0) / third;
|
|
554
|
+
const tailMean = series.slice(-third).reduce((x, y) => x + y, 0) / third;
|
|
555
|
+
const { windowPosNeg } = valenceSeries(text);
|
|
556
|
+
const totalWords = posWords + negWords;
|
|
557
|
+
const posRatio = totalWords === 0 ? 0 : posWords / totalWords;
|
|
558
|
+
const negRatio = totalWords === 0 ? 0 : negWords / totalWords;
|
|
559
|
+
// 坑1方案:矛盾指数按"窗口内原始词"算再平均(避免全书平均掩盖"同窗交织"vs"分段喜悲")
|
|
560
|
+
const windowConflicts = windowPosNeg
|
|
561
|
+
.filter((w) => w.pos + w.neg > 0)
|
|
562
|
+
.map((w) => 2 * Math.min(w.pos / (w.pos + w.neg), w.neg / (w.pos + w.neg)));
|
|
563
|
+
const conflict = windowConflicts.length === 0 ? 0 : windowConflicts.reduce((x, y) => x + y, 0) / windowConflicts.length;
|
|
564
|
+
return {
|
|
565
|
+
windows: n,
|
|
566
|
+
variance: Math.round(variance * 1000) / 1000,
|
|
567
|
+
adjVariance: Math.round(adjVariance * 1000) / 1000,
|
|
568
|
+
delta: Math.round(delta * 1000) / 1000,
|
|
569
|
+
deltaRobust: Math.round((tailMean - headMean) * 1000) / 1000,
|
|
570
|
+
conflict: Math.round(conflict * 1000) / 1000,
|
|
571
|
+
posRatio: Math.round(posRatio * 1000) / 1000,
|
|
572
|
+
negRatio: Math.round(negRatio * 1000) / 1000,
|
|
573
|
+
meanValence: Math.round(mean * 1000) / 1000
|
|
574
|
+
};
|
|
575
|
+
}
|
|
576
|
+
|
|
577
|
+
/** v1.6.0 显隐对比:显性均值 vs 隐性方向 → 表里不一。 */
|
|
578
|
+
export function explicitImplicitCompare(explicitMean, implicit) {
|
|
579
|
+
if (!implicit || implicit.totalHits === 0) return { explicitImplicitConflict: false, explicitSign: "neutral", implicitSign: "neutral" };
|
|
580
|
+
const explicitSign = explicitMean > 0.15 ? "positive" : explicitMean < -0.15 ? "negative" : "neutral";
|
|
581
|
+
const implicitSign = implicit.negative >= 0.6 ? "negative" : implicit.positive >= 0.6 ? "positive" : "neutral";
|
|
582
|
+
return {
|
|
583
|
+
explicitImplicitConflict: explicitSign === "positive" && implicitSign === "negative",
|
|
584
|
+
explicitSign,
|
|
585
|
+
implicitSign
|
|
586
|
+
};
|
|
587
|
+
}
|
|
588
|
+
|
|
589
|
+
/** v1.6.0 复合情感共现(规则):同段多情感类别 → 高频矛盾对。 */
|
|
590
|
+
export function compositeEmotionPairs(blocks) {
|
|
591
|
+
const pairCounts = new Map();
|
|
592
|
+
for (const block of blocks) {
|
|
593
|
+
const present = new Set();
|
|
594
|
+
for (const sentence of block) {
|
|
595
|
+
const e = sentence.emotion;
|
|
596
|
+
for (const k of EMOTION_CATS) if ((e.cleanScores?.[k] ?? 0) > 0) present.add(k);
|
|
597
|
+
}
|
|
598
|
+
const list = [...present];
|
|
599
|
+
for (let i = 0; i < list.length; i += 1) {
|
|
600
|
+
for (let j = i + 1; j < list.length; j += 1) {
|
|
601
|
+
const key = [list[i], list[j]].sort().join("+");
|
|
602
|
+
pairCounts.set(key, (pairCounts.get(key) ?? 0) + 1);
|
|
603
|
+
}
|
|
604
|
+
}
|
|
605
|
+
}
|
|
606
|
+
const labels = {
|
|
607
|
+
"joy+sorrow": "悲喜交加", "joy+anger": "又爱又恨/喜怒交织", "sorrow+anger": "哀怒交加/愤懑",
|
|
608
|
+
"fear+sorrow": "悲伤恐惧", "joy+fear": "惊喜交加", "anger+fear": "惊惧愤怒", "sorrow+surprise": "愕然悲伤"
|
|
609
|
+
};
|
|
610
|
+
return [...pairCounts.entries()].map(([pair, count]) => ({ pair, count, label: labels[pair] ?? "复合情感" }))
|
|
611
|
+
.sort((x, y) => y.count - x.count).slice(0, 5);
|
|
612
|
+
}
|
|
613
|
+
|
|
614
|
+
/** v1.6.0 单章分布(五维 clean)→ 熵/多样性/主次。 */
|
|
615
|
+
export function chapterEmotionStats(counts) {
|
|
616
|
+
const total = EMOTION_CATS.reduce((s, k) => s + (counts[k] ?? 0), 0);
|
|
617
|
+
if (total === 0) return null;
|
|
618
|
+
const p = EMOTION_CATS.map((k) => (counts[k] ?? 0) / total);
|
|
619
|
+
const entropy = -p.filter((x) => x > 0).reduce((s, x) => s + x * Math.log(x), 0);
|
|
620
|
+
const diversity = p.filter((x) => x >= 0.15).length;
|
|
621
|
+
const sorted = EMOTION_CATS.map((k, i) => ({ emotion: k, ratio: p[i] })).sort((x, y) => y.ratio - x.ratio);
|
|
622
|
+
return {
|
|
623
|
+
entropy: Math.round(entropy * 1000) / 1000,
|
|
624
|
+
diversity,
|
|
625
|
+
dominant: sorted[0].emotion,
|
|
626
|
+
dominantRatio: Math.round(sorted[0].ratio * 1000) / 1000,
|
|
627
|
+
secondary: sorted[1].emotion,
|
|
628
|
+
secondaryRatio: Math.round(sorted[1].ratio * 1000) / 1000
|
|
629
|
+
};
|
|
630
|
+
}
|
|
631
|
+
|
|
632
|
+
/** v1.6.0 全书聚合:复杂度评分 0-1(熵归一化0.5 + 多样性0.25 + 主次冲突0.25)+ 章间漂移。 */
|
|
633
|
+
export function emotionComplexity(perChapter) {
|
|
634
|
+
const stats = perChapter.map((c) => ({ chapter: c.chapter, ...(chapterEmotionStats(c.counts) ?? {}) })).filter((s) => s.entropy !== void 0);
|
|
635
|
+
if (stats.length === 0) {
|
|
636
|
+
// 无情感词命中:不复杂(low),恒有值供模型读取
|
|
637
|
+
return {
|
|
638
|
+
score: 0, level: "low", entropy: 0, maxEntropy: Math.round(EMOTION_MAX_ENTROPY * 1000) / 1000,
|
|
639
|
+
diversity: 0, dominant: "neutral", dominantRatio: 0, secondary: "neutral", secondaryRatio: 0,
|
|
640
|
+
conflict: "", conflictStrength: 0,
|
|
641
|
+
chapterDrift: { meanEntropy: 0, entropyVariance: 0, swinging: false }
|
|
642
|
+
};
|
|
643
|
+
}
|
|
644
|
+
const totalCounts = {};
|
|
645
|
+
for (const c of perChapter) for (const k of EMOTION_CATS) totalCounts[k] = (totalCounts[k] ?? 0) + (c.counts[k] ?? 0);
|
|
646
|
+
const global = chapterEmotionStats(totalCounts);
|
|
647
|
+
const entropies = stats.map((s) => s.entropy);
|
|
648
|
+
const meanEntropy = entropies.reduce((x, y) => x + y, 0) / entropies.length;
|
|
649
|
+
const entropyVariance = entropies.reduce((s, e) => s + (e - meanEntropy) ** 2, 0) / entropies.length;
|
|
650
|
+
const meanDiversity = stats.reduce((s, x) => s + x.diversity, 0) / stats.length;
|
|
651
|
+
const entropyNorm = global ? global.entropy / EMOTION_MAX_ENTROPY : 0;
|
|
652
|
+
const conflictStrength = global ? global.secondaryRatio / Math.max(global.dominantRatio, 0.0001) : 0;
|
|
653
|
+
const score = Math.min(1, Math.max(0, entropyNorm * 0.5 + (meanDiversity / 5) * 0.25 + conflictStrength * 0.25));
|
|
654
|
+
const level = score >= 0.6 ? "high" : score >= 0.4 ? "medium" : "low";
|
|
655
|
+
return {
|
|
656
|
+
score: Math.round(score * 100) / 100,
|
|
657
|
+
level,
|
|
658
|
+
entropy: global ? global.entropy : 0,
|
|
659
|
+
maxEntropy: Math.round(EMOTION_MAX_ENTROPY * 1000) / 1000,
|
|
660
|
+
diversity: global ? global.diversity : 0,
|
|
661
|
+
dominant: global?.dominant ?? "neutral",
|
|
662
|
+
dominantRatio: global?.dominantRatio ?? 0,
|
|
663
|
+
secondary: global?.secondary ?? "neutral",
|
|
664
|
+
secondaryRatio: global?.secondaryRatio ?? 0,
|
|
665
|
+
conflict: global ? global.dominant + "↔" + global.secondary : "",
|
|
666
|
+
conflictStrength: Math.round(conflictStrength * 100) / 100,
|
|
667
|
+
chapterDrift: {
|
|
668
|
+
meanEntropy: Math.round(meanEntropy * 1000) / 1000,
|
|
669
|
+
entropyVariance: Math.round(entropyVariance * 10000) / 10000,
|
|
670
|
+
swinging: entropyVariance > 0.03
|
|
671
|
+
}
|
|
672
|
+
};
|
|
673
|
+
}
|
|
674
|
+
|
|
675
|
+
/**
|
|
676
|
+
* v1.6.0 情感量化入口:三指标 + 显隐对比 + 复杂度 + 复合共现(纯规则 0 token)。
|
|
677
|
+
*/
|
|
678
|
+
export function emotionalQuantification(text, perChapter, blocks) {
|
|
679
|
+
const stats = valenceStats(text);
|
|
680
|
+
const implicit = implicitEmotionScan(text);
|
|
681
|
+
const meanValence = stats.meanValence ?? 0;
|
|
682
|
+
// perChapter 为空(单章/无分章输入)时,用当前文本自身 clean 计数聚合,保证 complexity 恒有值
|
|
683
|
+
if (!Array.isArray(perChapter) || perChapter.length === 0) {
|
|
684
|
+
const counts = { joy: 0, anger: 0, sorrow: 0, fear: 0, surprise: 0 };
|
|
685
|
+
for (const block of blocks) {
|
|
686
|
+
for (const sentence of block) {
|
|
687
|
+
const cs = sentence.emotion.cleanScores ?? {};
|
|
688
|
+
for (const k of EMOTION_CATS) counts[k] += cs[k] ?? 0;
|
|
689
|
+
}
|
|
690
|
+
}
|
|
691
|
+
perChapter = [{ chapter: "全书", counts }];
|
|
692
|
+
}
|
|
693
|
+
return {
|
|
694
|
+
stats,
|
|
695
|
+
implicit,
|
|
696
|
+
compare: explicitImplicitCompare(meanValence, implicit),
|
|
697
|
+
complexity: emotionComplexity(perChapter),
|
|
698
|
+
composites: compositeEmotionPairs(blocks)
|
|
699
|
+
};
|
|
700
|
+
}
|
|
701
|
+
|
|
702
|
+
/**
|
|
703
|
+
* 全书/单章句式模式分析主入口(v1.6.0 合并版)。
|
|
415
704
|
* @param text 正文文本。
|
|
416
705
|
* @param options { top: 句式模板条数(默认 8), maxSentences: 采样上限(默认 20000), chapterTexts: [{chapter, text}] 可选分章输入 }
|
|
417
706
|
* @returns 结构化分析结果(与 novel_sentence_analysis 输出 schema 一致)。
|
|
@@ -633,7 +922,7 @@ export function analyzeText(text, options = {}) {
|
|
|
633
922
|
emotion: emotionName,
|
|
634
923
|
label: EMOTION_LABELS[emotionName],
|
|
635
924
|
count: round(emotionCounts[emotionName], 2),
|
|
636
|
-
words:
|
|
925
|
+
words: [...new Set(EMOTION_WORDS[emotionName])].filter((word) => emotionWordCounts.has(word)).slice(0, 10)
|
|
637
926
|
})),
|
|
638
927
|
cleanScores: ["joy", "anger", "sorrow", "fear", "surprise"].map((emotionName) => ({
|
|
639
928
|
emotion: emotionName,
|
|
@@ -651,7 +940,9 @@ export function analyzeText(text, options = {}) {
|
|
|
651
940
|
})
|
|
652
941
|
.sort((a, b) => b.count - a.count || a.word.localeCompare(b.word))
|
|
653
942
|
.slice(0, 12),
|
|
654
|
-
curve: emotionCurve(blockMeta, curveSegments)
|
|
943
|
+
curve: emotionCurve(blockMeta, curveSegments),
|
|
944
|
+
// v1.6.0:情感量化(Valence 三指标 + 显隐对比 + 复杂度 + 复合共现)
|
|
945
|
+
quantification: emotionalQuantification(text, [], blockMeta.map((m) => m.sentences).filter((s) => s.length > 0))
|
|
655
946
|
};
|
|
656
947
|
|
|
657
948
|
// 主观性指数(启发式 0-100)
|