ppxans-harness 3.2.0 → 3.2.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,10 +1,54 @@
1
+ <div align="center">
2
+
1
3
  # 🦐 PPXANS-Harness
2
4
 
3
- **皮皮虾神经系(ANS)+ Harness 一体化智能体内核** —— 纯 Node.js、**零运行时依赖**的可审计自主智能体。
5
+ ### 皮皮虾(PPXANS)—— 一个能自己记住、自己修复、自己学习,而且每一步都留痕的 AI 智能体内核
6
+
7
+ **纯 Node.js · 零运行时依赖 · 下载即跑**
8
+
9
+ [![License](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE)
10
+ [![Node](https://img.shields.io/badge/node-%3E%3D20-brightgreen.svg)](package.json)
11
+ [![Runtime deps](https://img.shields.io/badge/runtime_dependencies-0-brightgreen.svg)](package.json)
12
+ [![Tests](https://img.shields.io/badge/tests-1029_passing-brightgreen.svg)](#-测试--评测--ci)
13
+ [![Self-heal](https://img.shields.io/badge/self--heal-7%2F7-brightgreen.svg)](scripts/selfheal-bench.js)
14
+ [![MCP](https://img.shields.io/badge/MCP-server_%2B_client-blueviolet.svg)](#mcp-标准端点-streamable-http)
15
+
16
+ <img src="docs/demo/terminal.svg" width="760" alt="PPXANS-Harness demo">
17
+
18
+ </div>
19
+
20
+ 接上任意 **OpenAI 兼容大模型**,它就变成一个**有记性、会成长、可审计**的助手:有自己的五层记忆(记得你是谁、忘掉无关的)、启动自愈、从失败里学、每次工具调用都写进 SHA-256 防篡改账本,还自带标准 **MCP 服务端** —— Claude Desktop / Cursor / 任何 MCP 客户端**开箱即用**。
21
+
22
+ > **English —** PPXANS-Harness is a self-contained AI **agent kernel in pure Node.js, with zero runtime dependencies**. Point it at any OpenAI-compatible model and you get an agent with a 5-layer memory, startup self-healing, failure-driven self-learning, a tamper-evident tool-call audit chain, multi-agent orchestration, and a built-in MCP server. `npm start` and go — there is no `npm install` step.
23
+
24
+ ### ⚡ 30 秒跑起来
25
+
26
+ ```bash
27
+ git clone https://github.com/chen6896qqwee/PPXANS-Harness.git
28
+ cd PPXANS-Harness
29
+
30
+ # 任选一家模型,或本地 LM Studio(默认 http://127.0.0.1:1234/v1)
31
+ export ZHIPU_API_KEY=xxx # 或 OPENAI_API_KEY / DEEPSEEK_API_KEY / DASHSCOPE_API_KEY ...
32
+
33
+ npm start # → http://127.0.0.1:8899 (内核 + Web 界面,同进程同端口,自动开浏览器)
34
+ ```
35
+
36
+ **没有 `npm install`。** 主包的 `package.json` 里根本没有 `dependencies` 字段 —— 只用 Node 内置模块,`node bin/ppx-web.js` 就能起。
37
+
38
+ ### 这东西到底是什么?(说人话)
39
+
40
+ 不是框架,不是 SDK,是**一个完整能跑的产品**。你可以把它理解成:给大模型装上**记忆、免疫系统和体检报告**的底座。
4
41
 
5
- 一个会自我修复、自我学习、可审计验证的超级 Agent:**85 内置工具 · L0–L4 五层记忆 + 事实有效期 · SHA-256 审计哈希链 · 标准 MCP 服务端+客户端 · 多模型路由 · 本地向量/ASR 可选 · 自愈 7/7 · 1014 测试全绿 · Web UI**。
42
+ | 常见 Agent 的毛病 | PPXANS-Harness 怎么做 |
43
+ |---|---|
44
+ | 一关窗口就失忆 | 五层记忆 L0–L4 跨会话留存;软删可回滚、事实带有效期、装不下才裁剪 |
45
+ | 一崩就全没了 | 启动体检 + 损坏文件修复 + 崩溃恢复,自愈基准 **7/7** |
46
+ | 同一个错反复犯 | 失败沉淀成经验(refine),成功沉淀成技能(refineSkill),还会自动升级 |
47
+ | 干了啥说不清 | 每次工具调用 append-only 写进 **SHA-256 链式账本**,改一行全链校验失败并定位到行 |
48
+ | 生态孤岛 | 标准 **MCP 服务端**(`POST /mcp`)+ 客户端,外部工具与客户端双向接入 |
49
+ | 依赖地狱 | 主包**零运行时依赖**,`node bin/ppx-web.js` 直接起 |
6
50
 
7
- 自带标准 MCP 服务(`POST /mcp`),Claude Desktop / Cursor / 任何 MCP 客户端**开箱即用**;支持各大模型 API + 本地模型,自由回退。
51
+ **运行时实测**:63 内置工具 + 22 个 `ppx.*` 管理工具(MCP 共暴露 **85**)· 自愈 **7/7 100%** · 全量测试 **1029 项(1025 通过 / 0 失败 / 4 skip)** · 渐进披露把固定开销从 8172 降到 ~3357 tok/请求(**-59%**)。
8
52
 
9
53
  > 🛡️ **自愈基准**:`node scripts/selfheal-bench.js` → **7/7 100%**(发布前门禁,`PPX_MIN_SELFHEAL` 可设阈值)
10
54
  > 🔗 **审计哈希链**:`npm run audit:verify` —— append-only + SHA-256 链式防篡改,篡改/删除可定位到行
@@ -112,7 +156,7 @@ npm run selfheal
112
156
  npm run chat # 终端对话 CLI (ppx / ppxans)
113
157
  npm run serve # 仅 HTTP 接口 (无界面): http://127.0.0.1:8899
114
158
  npm run web:check # Web UI 静态自检 (图标/DOM id/语法解析/静态资源)
115
- npm test # 全量测试 (1014 项)
159
+ npm test # 全量测试 (1029 项)
116
160
  ```
117
161
 
118
162
  ### MCP 标准端点 (Streamable HTTP)
@@ -148,7 +192,7 @@ npm test # 全量测试 (1014 项)
148
192
  ## 🧪 测试 / 评测 / CI
149
193
 
150
194
  ```bash
151
- npm test # 全量 1014 项 (0 失败)
195
+ npm test # 全量 1029 项 (1025 通过 / 0 失败 / 4 skip)
152
196
  npm run eval # 本地能力评测 (7 项, 无需 LLM)
153
197
  npm run eval -- --llm # LLM 端到端评测 (需 provider)
154
198
  npm run bench # 并发/长会话吞吐压测
@@ -199,7 +243,7 @@ PPXANS-Harness/
199
243
  ├── bin/ ppx / ppxans / ppx-web / ppx-serve / ppx-channels 入口
200
244
  ├── data/ 运行时数据 (不进 git)
201
245
  ├── references/ 第三方项目来源登记
202
- ├── test/ 测试 (1014 项, v3 新模块全覆盖)
246
+ ├── test/ 测试 (1029 项, v3 新模块全覆盖)
203
247
  └── docs/ 文档 (ARCHITECTURE-V3 / QUICKSTART / web-launch 等)
204
248
  ```
205
249
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ppxans-harness",
3
- "version": "3.2.0",
3
+ "version": "3.2.2",
4
4
  "description": "PPXANS-Harness - 皮皮虾神经系 (ANS) + Harness 一体化智能体内核. 零运行时依赖纯Node. 自愈 + 自学习 + 五层记忆 + SHA-256 审计哈希链 + 多 Agent 军团 + 治理内核(deny-wins/熔断/seam) + 进化内核(Playbook/故障记忆) + 符号画布 + supervisor. 自研底座: 仅 OpenAI 兼容 HTTP 直连, 无外部引擎依赖. Agent Nervous System harness in pure Node, zero runtime dependencies.",
5
5
  "type": "module",
6
6
  "license": "Apache-2.0",
@@ -41,7 +41,7 @@
41
41
  "bench:audit": "node scripts/audit-bench.js",
42
42
  "bench:ctx": "node scripts/ctx-profile.js",
43
43
  "test": "node --test --test-force-exit test/*.test.js",
44
- "prepublishOnly": "npm run selfheal && npm test",
44
+ "prepublishOnly": "npm run selfheal && npm test && npm run eval",
45
45
  "package": "node scripts/package.js",
46
46
  "package:portable": "node scripts/package.js portable",
47
47
  "package:installer": "node scripts/package.js installer",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "ppx-memory": {
3
- "uses": 43,
4
- "lastUsed": "2026-10-03T10:20:18.436Z"
3
+ "uses": 51,
4
+ "lastUsed": "2026-10-03T11:44:04.018Z"
5
5
  }
6
6
  }
@@ -0,0 +1,2 @@
1
+ {"conversations": [{"from": "human", "value": "把 OpenClaw 从 9.3 升级到最新 9.4 版本"}, {"from": "gpt", "value": "先做完整备份 (openclaw backup create --verify 通过),然后直接执行 npm install -g openclaw@2026.9.4。装完验证:openclaw --version 显示 2026.9.4 (3a9d69d),插件列表显示 2026.9.4,vec0.dll 存在,Gateway 正常重启运行。绕开官方更新器的 canary 机制,用 npm 直装跳过有 bug 的彩排路径,更新成功且无数据丢失。"}], "chosen": {"from": "gpt", "value": "先做完整备份 (openclaw backup create --verify 通过),然后直接执行 npm install -g openclaw@2026.9.4。装完验证:openclaw --version 显示 2026.9.4 (3a9d69d),插件列表显示 2026.9.4,vec0.dll 存在,Gateway 正常重启运行。绕开官方更新器的 canary 机制,用 npm 直装跳过有 bug 的彩排路径,更新成功且无数据丢失。"}, "rejected": {"from": "gpt", "value": "运行 openclaw update --yes --channel stable。更新器在 validating 阶段失败:candidate snapshot 报 ENOENT,错误路径是 ~\\Temp\\openclaw-update-canary-xxx\\?\\$OPENCLAW_STATE_DIR\\agents\\main\\agent —— $OPENCLAW_STATE_DIR 环境变量在 Windows 路径拼接时未被展开成实际路径,导致 mkdir 失败。结果 runtime-verification-failed,版本仍停留在 9.3,更新彻底失败。"}, "meta": {"task_type": "package-upgrade", "reward_chosen": 1.0, "reward_rejected": 0.0, "source": "real-session-20260911"}}
2
+ {"conversations": [{"from": "human", "value": "OpenClaw 更新后一直卡在 verifying 状态,doctor/repair 都拒绝执行,怎么修?"}, {"from": "gpt", "value": "先查 update_runs 表找到那条卡死的记录:7df8bfa1 的所有步骤实际都已 completed,verification_json 显示 booted=true, versionMatch=true, pluginErrors=[],只是最后一步 verifying 的完成信号在 gateway restart 时丢失。备份状态库后用 SQL 把该 run 标记为 phase=finished, status=succeeded,再清掉另一条 9-07 的陈旧 requested-only 占位记录。结果 update_runs 无 running 残留,doctor 恢复正常,更新确认成功。根因是状态机死记录,不是重复重试能解决的。"}], "chosen": {"from": "gpt", "value": "先查 update_runs 表找到那条卡死的记录:7df8bfa1 的所有步骤实际都已 completed,verification_json 显示 booted=true, versionMatch=true, pluginErrors=[],只是最后一步 verifying 的完成信号在 gateway restart 时丢失。备份状态库后用 SQL 把该 run 标记为 phase=finished, status=succeeded,再清掉另一条 9-07 的陈旧 requested-only 占位记录。结果 update_runs 无 running 残留,doctor 恢复正常,更新确认成功。根因是状态机死记录,不是重复重试能解决的。"}, "rejected": {"from": "gpt", "value": "反复运行 openclaw update repair 和 openclaw doctor --fix。每次都报 'Doctor could not enter maintenance: another OpenClaw process owns gateway-lifecycle',因为状态库里有一条 status=running 的 update_runs 记录,repair 认为更新仍在进行中,永远拒绝进入维护模式。重复了 5 次全部失败,纯死循环。"}, "meta": {"task_type": "update-state-recovery", "reward_chosen": 1.0, "reward_rejected": 0.0, "source": "real-session-20260911"}}
@@ -0,0 +1,15 @@
1
+ {
2
+ "task": "按 agent-data-flywheel 技能第5节生成 DPO 偏好对数据集并验证",
3
+ "observations": [
4
+ {"modality": "text", "content": "第二轮技能测试:需从真实会话事件生成偏好对(成功vs失败),按第5节飞轮流水线跑通 收集轨迹->验证器打分->过滤->生成偏好对"}
5
+ ],
6
+ "memory": [
7
+ "素材来源:本次会话真实事件——OpenClaw 9.3->9.4 升级(官方更新器失败 vs npm直装成功)、更新卡verifying修复(死循环重试 vs 状态库手动收尾)",
8
+ "技能第5节要求:成功轨迹与失败轨迹分开归档生成偏好对;奖励必须可验证(规则验证器优先)"
9
+ ],
10
+ "thought": "用真实事件构造两对 DPO 偏好对,转 LLaMA-Factory 格式,再写规则验证器打分,验证整个流水线前段可跑通",
11
+ "action": {"tool": "python", "args": {"files": ["gen_dpo_pairs.py", "validate_dpo.py"], "result": "flywheel-dpo-pairs.jsonl"}},
12
+ "result": {"ok": true, "summary": "生成2对DPO偏好对并归档到技能 trajectories/;规则验证器打分 0.80/1.00 全过,平均0.90,合格率2/2;格式为 LLaMA-Factory dpo stage 兼容 JSONL"},
13
+ "reward": 1.0,
14
+ "reflection": "数据飞轮流水线前段(轨迹->偏好对->验证器)真实可跑通。关键发现:①从真实会话事件取材效率最高,chosen/rejected 天然清晰可验证;②规则验证器打分标准(结构40%+可验证信号40%+反思20%)能有效区分好坏数据;③Windows 下 Python 写文件必须显式 encoding='utf-8' 且 ensure_ascii=False,否则中文损坏;④验证器可验证信号词表是关键资产,后续可扩展成独立配置。下一步可跑 LLaMA-Factory 真实 DPO 训练(需GPU+训练环境),本机无GPU故验证到数据层为止"
15
+ }
@@ -0,0 +1,17 @@
1
+ {
2
+ "task": "按 agent-professional-training 技能首次真实运行「感知—记忆—推理—行动—反思」闭环,产出第 1 条训练轨迹并通过校验器",
3
+ "observations": [
4
+ {"modality": "text", "content": "用户指令:'你试试'。上下文:SKILL.md 7.6 节已新增 NexRL/Nex-N2.5 框架并映射到 L1/L2 行为层动作;技能要求每次任务第 5 步强制写轨迹。环境:Windows / PowerShell / OpenClaw 工作区,桌面 repo 与已安装技能同步于 git main=d697231"}
5
+ ],
6
+ "memory": [
7
+ "技能第 2 节闭环五步:感知→记忆→推理→行动→反思,第 5 步强制写轨迹",
8
+ "技能第 4 节:轨迹 8 必填字段 task/observations/memory/thought/action/result/reward/reflection",
9
+ "技能 4.2 校验器:8 字段齐全 + reward 数值 + modality 合法",
10
+ "先例:2026-09-11 已有 round2-trajectory.json(真实事件取材)+ dpo-pairs.jsonl"
11
+ ],
12
+ "thought": "任务'试试'=验证闭环能真实跑通。规划:①感知输入(用户指令+环境);②检索记忆(技能规程+先例格式);③推理拆解(写轨迹→写校验器→跑校验→归档→同步 git);④行动(写入第 4 节格式轨迹、运行 4.2 校验器);⑤反思(对照 9.0 行为层自检表)+ 强制写本条轨迹。核心认知:'行动'步本身就是'写轨迹+跑校验'——训练即做事,做事即训练",
13
+ "action": {"tool": "write+exec", "args": {"files": ["trajectories/2026-09-13-first-run-trajectory.json", "validate_trajectory.py"], "commands": ["python validate_trajectory.py trajectories/2026-09-13-first-run-trajectory.json", "同步已安装技能 + git commit"]}},
14
+ "result": {"ok": true, "summary": "闭环五步真实跑通:首条轨迹按第 4 节 8 字段格式写入并通过 4.2 校验器;归档 trajectories/;同步已安装技能目录并 git 提交;无 GPU 纯行为层进化,符合技能主线 L1/L2"},
15
+ "reward": 1.0,
16
+ "reflection": "首次真实运行闭环验证可行。关键发现:①闭环'行动'步本身就是'写轨迹+跑校验'——训练即做事、做事即训练,不需任何额外训练代码;②记忆检索要先例:参照 2026-09-11 轨迹格式避免重造格式;③校验器是闭环的'环境反馈',reward 从校验结果来,不依赖 GPU;④本轮'教训沉淀':轨迹 JSON 中命令路径要写相对路径,避免换机失效;⑤下一步:持续按闭环干活攒轨迹,满 50 条跑第 5 节飞轮生成偏好对(对照 14 节自检清单)"
17
+ }
@@ -154,6 +154,8 @@ export class PPXAgent {
154
154
  // 2026-10-03l: 加 cost (USD) 维度 + budget.usd 支出上限 (超限后 chat/chatStream 拒绝继续烧钱)
155
155
  this.usageStats = { calls: 0, tokens: 0, cost: 0, byModel: {} };
156
156
  this._budgetExceeded = false;
157
+ this._usageLastFlush = 0;
158
+ this._healthCache = null;
157
159
  this._installUsageTracking();
158
160
  // 待审批映射 (codex approval flow): id -> { req, resolve, timer }
159
161
  this._pendingApprovals = new Map();
@@ -602,13 +604,26 @@ export class PPXAgent {
602
604
  if (native.length) clients = [...native, ...fence];
603
605
  }
604
606
  if (clients.length > 1) {
605
- try {
606
- const states = await Promise.all(clients.map((c) => c.health ? c.health() : Promise.resolve(true)));
607
- const healthy = clients.filter((_, i) => states[i]);
607
+ // 健康探测 TTL 缓存 (2026-10-03m): 每轮 chat 都全量探活 = 高频对话下每次多一段串行探活延迟。
608
+ // TTL 内复用上次结果 (默认 30s, config.agent.health_cache_ms 可调, 0 = 关闭缓存);
609
+ // provider 集合变化 (键不匹配) 时重新探测。
610
+ const ttl = Number(this.config?.agent?.health_cache_ms ?? 30000);
611
+ const cacheKey = clients.map((c) => c.model || c.name).join("|");
612
+ const cached = ttl > 0 && this._healthCache && this._healthCache.key === cacheKey
613
+ && Date.now() - this._healthCache.ts < ttl ? this._healthCache.states : null;
614
+ if (cached) {
615
+ const healthy = clients.filter((_, i) => cached[i]);
608
616
  if (healthy.length) clients = healthy;
609
- else info("所有 provider 健康探测失败, 按原配置顺序尝试兜底");
610
- } catch (e) {
611
- warn("health 探测异常, 按原顺序回退:", e.message);
617
+ } else {
618
+ try {
619
+ const states = await Promise.all(clients.map((c) => c.health ? c.health() : Promise.resolve(true)));
620
+ this._healthCache = { ts: Date.now(), key: cacheKey, states };
621
+ const healthy = clients.filter((_, i) => states[i]);
622
+ if (healthy.length) clients = healthy;
623
+ else info("所有 provider 健康探测失败, 按原配置顺序尝试兜底");
624
+ } catch (e) {
625
+ warn("health 探测异常, 按原顺序回退:", e.message);
626
+ }
612
627
  }
613
628
  }
614
629
  let lastErr = null;
@@ -1273,6 +1288,11 @@ export class PPXAgent {
1273
1288
  this.usageStats.byModel[m].cost = Math.round(((this.usageStats.byModel[m].cost || 0) + cost) * 1e6) / 1e6;
1274
1289
  this._checkBudget();
1275
1290
  }
1291
+ // 周期落盘 (2026-10-03m): 崩溃不丢账 —— 每满 10 次调用落一次,
1292
+ // 长跑进程被 kill 时 usage-stats.json 最多丢 9 笔而非全量 (退出另有兑底)
1293
+ if (this.usageStats.calls % 10 === 0) {
1294
+ this._flushUsageStats();
1295
+ }
1276
1296
  } catch { /* 统计失败不阻塞调用 */ }
1277
1297
  return r;
1278
1298
  };
@@ -1298,16 +1318,22 @@ export class PPXAgent {
1298
1318
  return `[预算耗尽] 本进程累计支出 $${(this.usageStats.cost || 0).toFixed(4)} 已达上限 $${(Number.isFinite(cap) ? cap : 0).toFixed(2)},已停止继续调用模型。调整 config/ppx.json 的 budget.usd(0 或删除 = 不限)后重启生效。`;
1299
1319
  }
1300
1320
 
1301
- shutdown() {
1302
- this.stopProactiveTicker();
1303
- // 使用统计落盘 (ZCode 使用统计对齐): data/usage-stats.json (含 cost 金额维度 + 预算上限快照)
1321
+ // 使用统计落盘 (ZCode 使用统计对齐): data/usage-stats.json (含 cost 金额维度 + 预算上限快照)
1322
+ // 2026-10-03m: 从 shutdown 抽出供周期落盘复用 (周期性 + 退出兑底双保险)
1323
+ _flushUsageStats() {
1304
1324
  try {
1305
1325
  const usagePayload = { updated: new Date().toISOString(), ...this.usageStats };
1306
1326
  const cap = Number(this.config?.budget?.usd);
1307
1327
  if (Number.isFinite(cap) && cap > 0) usagePayload.budget_usd = cap;
1308
- fs.writeFileSync(path.join(this.dataDir, "usage-stats.json"),
1309
- JSON.stringify(usagePayload, null, 2));
1310
- } catch { /* 落盘失败不阻塞退出 */ }
1328
+ fs.writeFileSync(path.join(this.dataDir, "usage-stats.json"), JSON.stringify(usagePayload, null, 2));
1329
+ this._usageLastFlush = Date.now();
1330
+ } catch { /* 落盘失败不阻塞主链 */ }
1331
+ }
1332
+
1333
+ shutdown() {
1334
+ this.stopProactiveTicker();
1335
+ // 使用统计落盘 (退出兑底; 长跑期间已有周期落盘, 此处只补尾部增量)
1336
+ this._flushUsageStats();
1311
1337
  this._mcp?.close?.();
1312
1338
  // 释放内嵌数据库句柄 (SQLite 后端必需: 不关会导致文件被占用, 无法迁移/清理)
1313
1339
  try { this.facts?.close?.(); } catch { /* JSON 后端无 close, 静默跳过 */ }