@dsh-bio/dsh-bio-gem 0.1.14 → 0.1.16
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +10 -8
- package/docs/ARCHITECTURE.md +18 -18
- package/docs/releases/v0.1.12.md +2 -2
- package/docs/releases/v0.1.16.md +31 -0
- package/package.json +4 -4
- package/python/annotate.py +1 -1
- package/python/benchmark.py +3 -3
- package/python/coherence.py +2 -2
- package/python/double_knockout.py +21 -1
- package/python/enrichment.py +2 -0
- package/python/fluxscan.py +6 -4
- package/python/fsutil.py +44 -0
- package/python/gapseq_wsl.py +28 -6
- package/python/gem_ops.py +2 -2
- package/python/l3_fix.py +2 -2
- package/python/ledger.py +2 -2
- package/python/model_card.py +1 -1
- package/python/precursor_scan.py +1 -1
- package/python/quality.py +2 -0
- package/python/roundtrip_check.py +1 -1
- package/python/sampling.py +2 -0
- package/python/secretion.py +18 -1
- package/python/sensitivity.py +3 -1
- package/python/targets.py +3 -0
- package/python/validate.py +3 -3
- package/skills/gem-expert.md +4 -4
- package/src/capabilities.js +18 -5
- package/src/integration.js +1 -1
- package/src/python.js +2 -2
- package/src/tools.js +7 -6
- package/src/workdir.js +67 -0
- package/docs/DECISIONS-2026-08-29.md +0 -37
- package/docs/DECISIONS-2026-09-21.md +0 -76
- package/docs/DECISIONS-/351/230/266/346/256/265A.md +0 -67
- package/docs/DECISIONS-/351/230/266/346/256/265E.md +0 -56
package/README.md
CHANGED
|
@@ -17,7 +17,7 @@ Genome-scale metabolic model builder for dsh: genome in, validated SBML out.
|
|
|
17
17
|
| Python | 3.10+,且装 **`cobra`** | 分析/验证/补洞/账本/基准/导出 —— 除 `gem_build` 外的 22 个工具 |
|
|
18
18
|
| `pyrodigal` | 装在同一个 Python 环境(可选但建议) | 裸基因组自动注释兜底(`gem_annotate` / fna 输入) |
|
|
19
19
|
| **CarveMe** | 独立 venv `~/.dsh/dsh-bio-gem/venv-carveme`,含 `carve.exe` + **`diamond.exe`** | `gem_build`(carveme 引擎) |
|
|
20
|
-
| WSL2 + gapseq |
|
|
20
|
+
| WSL2 + gapseq | 可选,按 WSL2 环境拓扑(见第 4 节) | `gem_build`(gapseq 引擎) |
|
|
21
21
|
|
|
22
22
|
> **Python 环境从哪来(v0.1.4 起)**:解释器探测顺序为
|
|
23
23
|
> `GEM_PYTHON` → **宿主 `dsh-bio-genie` 的自举环境** → `CONDA_PREFIX` → `PATH` 中的 `python`,
|
|
@@ -38,7 +38,7 @@ npx -y @deepseek-ai/dsh plugin --profile web add @dsh-bio/dsh-bio-gem
|
|
|
38
38
|
# 方式二:从 GitHub 安装(拉源码;本插件纯 ESM 无构建步骤,可直接加载)
|
|
39
39
|
npx -y @deepseek-ai/dsh plugin --profile web add github:moonbowterfly/dsh-bio-gem
|
|
40
40
|
|
|
41
|
-
#
|
|
41
|
+
# 方式三:从本地目录安装(本地源码)
|
|
42
42
|
npx -y @deepseek-ai/dsh plugin --profile web add ./dsh-bio-gem
|
|
43
43
|
```
|
|
44
44
|
|
|
@@ -53,10 +53,10 @@ npx -y @deepseek-ai/dsh plugin --profile web add ./dsh-bio-gem
|
|
|
53
53
|
|
|
54
54
|
3. 重新打开桌面端生效(也可直接在桌面端内 **「插件」页**输入包名安装,无需退出应用)。
|
|
55
55
|
|
|
56
|
-
-
|
|
56
|
+
- 若已全局安装 dsh CLI,把 `npx -y @deepseek-ai/dsh` 换成 `dsh` 即可。
|
|
57
57
|
- `--profile <name>` 是**必填选项**(不传报 `required option '--profile <name>' not specified`);Web 端固定用 `web`,桌面端固定用 `desktop`。
|
|
58
58
|
- 安装完**重启对应的 dsh**(web:重新双击启动入口;桌面端:重开应用)。
|
|
59
|
-
- **引擎兼容**:0.1.x 侧经 dsh 0.1.5-rc.2
|
|
59
|
+
- **引擎兼容**:0.1.x 侧经 dsh 0.1.5-rc.2 完整兼容核验与实机验证;**0.2.0+(含官方桌面端)已于 2026-10-01 实测**(工具全量注册 + `gem_media_resolve` 真实执行)。
|
|
60
60
|
- 版本刚发布时可能短时间拉不到:registry 首次分发有几分钟延迟,`pnpm` 还可能缓存住 404。遇到 `ERR_PNPM_FETCH_404 ... is not in the npm registry` 时等几分钟重试,或在命令末尾追加 `--registry https://registry.npmjs.org/` 绕过缓存。
|
|
61
61
|
|
|
62
62
|
验证插件层已生效(不用启动服务):
|
|
@@ -121,13 +121,15 @@ unzip -o diamond.zip diamond.exe -d "$HOME/.dsh/dsh-bio-gem/venv-carveme/Scripts
|
|
|
121
121
|
|
|
122
122
|
### 4.(可选)gapseq 引擎(WSL2)
|
|
123
123
|
|
|
124
|
-
`gem_build` 的 `engine=gapseq` 走 WSL2 桥(`gem_gapseq` 原子四步:setup / launch / status / fetch),质量档耗时 30-60 分钟/模型,非必需——默认的 `engine=carveme`
|
|
124
|
+
`gem_build` 的 `engine=gapseq` 走 WSL2 桥(`gem_gapseq` 原子四步:setup / launch / status / fetch),质量档耗时 30-60 分钟/模型,非必需——默认的 `engine=carveme` 已能出可验证模型。桥按 WSL2 + conda 环境拓扑实现(默认路径见 `python/gapseq_wsl.py` 顶部常量,换机器需相应调整),故目前**视为实验性可选能力**。没有 WSL2 不影响其余 22 个工具与 carveme 构建。
|
|
125
|
+
|
|
126
|
+
能力探针只读检查 `doall` 使用的环境内 `seq/Bacteria` 元数据和 `rev/rxn/unrev` 文件,不请求 Zenodo。检查路径与当前固定的 conda 环境路径一致。
|
|
125
127
|
|
|
126
128
|
### 5. 自检(仓库源码目录内)
|
|
127
129
|
|
|
128
130
|
```sh
|
|
129
131
|
# 冒烟:不依赖 dsh,直测 Python 层 + 工具注册表
|
|
130
|
-
#
|
|
132
|
+
# 断言锚定固定夹具路径(见 test/smoke.js 顶部常量),换机器先改路径
|
|
131
133
|
GEM_PYTHON=<你的-cobra-python> node test/smoke.js --skip-build # 跳过 ~70s 的 build 单测
|
|
132
134
|
GEM_PYTHON=<你的-cobra-python> node test/smoke.js # 含 build 单测
|
|
133
135
|
# 托管领域扩展的只读 integration 协议(无需 dsh 实例)
|
|
@@ -210,7 +212,7 @@ CarveMe + diamond are required only by `gem_build`; the other 20 tools need noth
|
|
|
210
212
|
| `gem_essentiality` | 全量必需基因扫描(FVA 预筛 + 手工敲除)| ✅ C58: 必需 155 |
|
|
211
213
|
| `gem_annotate` | 基因组→蛋白(官方优先 + pyrodigal 兜底,纯 Windows)| ✅ pyrodigal 5330 |
|
|
212
214
|
| `gem_build` | CarveMe/gapseq 双引擎构建(fna/faa;后台 job + 进度;M9 或目标介质验证闭环)| ✅ carveme 70s / fna 全链 63.5s / gapseq 实测中 |
|
|
213
|
-
| `gem_gapseq` | gapseq 原子四步(WSL 可选:setup/launch/status/fetch)| ✅
|
|
215
|
+
| `gem_gapseq` | gapseq 原子四步(WSL 可选:setup/launch/status/fetch)| ✅ 实测全通 |
|
|
214
216
|
| `gem_media_resolve` | 跨引擎介质解析 RPC(自然名→EX ID;消费侧统一入口)| ✅ AB→20 EX |
|
|
215
217
|
| `gem_fluxscan` | 通量区间制(FVA 区间+pFBA 点值;条件对比区间分离判定,overlap=伪影禁止引用)| ✅ C58 AB 0.519981 / 蔗糖 supplement 0.97077 |
|
|
216
218
|
| `gem_sensitivity` | 结构性灵敏度(GAM×biomass 22 组合全量+稳定性三分类+单组分漂移+模型卡鲁棒性 v3)| ✅ C58 基准复现 155 |
|
|
@@ -224,7 +226,7 @@ CarveMe + diamond are required only by `gem_build`; the other 20 tools need noth
|
|
|
224
226
|
| `gem_quality` | 模型质量报告(gem-qi-v1:blocked/环路(fastcc 方向锥)/元素平衡/孤儿与死端/覆盖/连通性 + 启发式聚合分;分项与 failed_checks 为准)| ✅ C58+AB: qi 67.23 / blocked 1032 / cyclic 289 |
|
|
225
227
|
| `gem_sample` | 通量空间采样(ACHR 默认 / OptGP 大样本;growth_floor_fraction 受限空间;全空间 vs 受限边界声明)| ✅ C58+AB: 全空间 median 0.019 / floor-0.9 min 0.468 |
|
|
226
228
|
|
|
227
|
-
|
|
229
|
+
架构见 [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md)。
|
|
228
230
|
|
|
229
231
|
## 开发速查
|
|
230
232
|
|
package/docs/ARCHITECTURE.md
CHANGED
|
@@ -4,22 +4,22 @@
|
|
|
4
4
|
|
|
5
5
|
dsh 平台的 **GEM 构建侧插件**:输入细菌全基因组(支持多质粒/多染色体),自动构建→验证→补洞→出报告(标准 SBML + 模型卡),产出后可被 dsh-bio-genie 现有消费工具(FBA/必需性/生产包络线/模型面板)直接加载使用。
|
|
6
6
|
|
|
7
|
-
硬性原则(沿袭 bio-genie
|
|
7
|
+
硬性原则(沿袭 bio-genie):**用户零手动安装、零自愈、通用化(不针对特定机器特化)、结论可溯源**。
|
|
8
8
|
|
|
9
9
|
## 2. 决策记录(为什么这么设计)
|
|
10
10
|
|
|
11
11
|
| 日期 | 决策 | 依据 |
|
|
12
12
|
|---|---|---|
|
|
13
13
|
| 08-28 | 插件名 dsh-bio-gem;资产盘点:消费侧已就绪、补构建侧闭环 | 用户拍板 |
|
|
14
|
-
| 08-29 | 引擎路线:**任务门槛路由**(不是简单 auto);落地节奏 **M1 CarveMe+补洞 → M2 gapseq WSL 桥 → M3 双引擎交叉** |
|
|
14
|
+
| 08-29 | 引擎路线:**任务门槛路由**(不是简单 auto);落地节奏 **M1 CarveMe+补洞 → M2 gapseq WSL 桥 → M3 双引擎交叉** | 独立设计评估 + 实测(CarveMe AB 不生长=补洞是生存线;WSL 桥显著降级交付风险;Docker 非 WSL 替代)|
|
|
15
15
|
| 08-29 | MVP 工具集:gem_build / gem_validate(G1G2G3 必做,G4 条件、G5 抽检)/ gem_gapfind(L1L2L3)/ gem_gapfill(L1L2 规则自动)/ gem_report(薄版模型卡);**gem_essentiality 不进首版** | 消费侧 bio_gene_knockout 已存在,避免重复实现 |
|
|
16
|
-
| 08-29 |
|
|
16
|
+
| 08-29 | 修正建议:弃 μ 判据用 FBA 通量判据;pyrodigal 注释前端降 backlog;测试矩阵首版收敛 C58+2 公开株 | 输出口径为 objective_value;默认输入是带注释基因组 |
|
|
17
17
|
|
|
18
|
-
|
|
18
|
+
**采纳原则**:外部分析缺实际环境上下文时,凡冲突处以实测与产品原则为准。
|
|
19
19
|
|
|
20
20
|
## 3. 工具契约(21 工具 ↔ Python 层;21 op + build CLI)
|
|
21
21
|
|
|
22
|
-
| 工具 | Python 层 |
|
|
22
|
+
| 工具 | Python 层 | 状态 |
|
|
23
23
|
|---|---|---|
|
|
24
24
|
| gem_build | build.py CLI(CarveMe M9 gapfill;fna 自动注释)| ✅ M1+模块 DONE(C58 63-70s)|
|
|
25
25
|
| gem_validate | op validate(G1-G6 + GATE_REGISTRY)| ✅ M1 DONE |
|
|
@@ -33,15 +33,15 @@ dsh 平台的 **GEM 构建侧插件**:输入细菌全基因组(支持多质
|
|
|
33
33
|
| gem_report | op model_info(+ ledger_summary 基率摘要)| ✅ DONE |
|
|
34
34
|
| gem_media_resolve | op media_resolve(介质解析 RPC,消费侧统一入口)| ✅ DONE |
|
|
35
35
|
| gem_biomass | op biomass_inspect / biomass_apply(inspect 组分+对照参考;apply 覆盖表+三联对照+原文件不动回滚)| ✅ Q2 DONE(复位 delta 0.0)|
|
|
36
|
-
| gem_fluxscan | op fluxscan(通量区间制:FVA 区间+pFBA 点值+条件对区间分离判定,overlap=伪影禁止引用)| ✅
|
|
37
|
-
| gem_sensitivity | op sensitivity(GAM×biomass 22 组合全量+稳定性三分类+单组分漂移;模型卡 robustness v3)| ✅
|
|
38
|
-
| gem_ledger | op ledger(prediction ledger:list/query/update;幂等追加式账本)| ✅
|
|
39
|
-
| gem_benchmark | op benchmark(通用基准对比:六关并列/生长[介质层两级策略]/biomass 探针/必需性对比含退化护栏/表型/账本回填/md 落盘;model 参数支持 bigg:<id> 下载)| ✅
|
|
40
|
-
| gem_secretion | op secretion(可分泌谱:production envelope;边界声明内置;wt<=EPS 退化护栏不登记)| ✅
|
|
41
|
-
| gem_double_knockout | op double_knockout(双敲 v1:GPR 穷尽先验+全扫 max_pairs 预算;假设声明内置)| ✅
|
|
42
|
-
| gem_enrichment | op enrichment(必需基因通路富集:超几何+BH FDR;无注释 annotation_unavailable 兜底)| ✅
|
|
43
|
-
| gem_targets | op targets(靶点规范导出:11 字段锁定 schema;账本计数闭合;引物设计不做)| ✅
|
|
44
|
-
| gem_precursor_scan | op precursor_scan(阻塞前体分析:基线通量→可生长即返「无阻塞」;不生长则逐前体移除测试定位阻塞点)| ✅ 2026-09-11
|
|
36
|
+
| gem_fluxscan | op fluxscan(通量区间制:FVA 区间+pFBA 点值+条件对区间分离判定,overlap=伪影禁止引用)| ✅ 已完成(C58 AB 0.519981 / 蔗糖 supplement 0.97077)|
|
|
37
|
+
| gem_sensitivity | op sensitivity(GAM×biomass 22 组合全量+稳定性三分类+单组分漂移;模型卡 robustness v3)| ✅ 已完成(基准复现 155)|
|
|
38
|
+
| gem_ledger | op ledger(prediction ledger:list/query/update;幂等追加式账本)| ✅ 已完成(C58 155+19 条幂等复跑)|
|
|
39
|
+
| gem_benchmark | op benchmark(通用基准对比:六关并列/生长[介质层两级策略]/biomass 探针/必需性对比含退化护栏/表型/账本回填/md 落盘;model 参数支持 bigg:<id> 下载)| ✅ 已完成 |
|
|
40
|
+
| gem_secretion | op secretion(可分泌谱:production envelope;边界声明内置;wt<=EPS 退化护栏不登记)| ✅ 已完成(C58 85 可分泌)|
|
|
41
|
+
| gem_double_knockout | op double_knockout(双敲 v1:GPR 穷尽先验+全扫 max_pairs 预算;假设声明内置)| ✅ 已完成(Atu3364↔Atu4682 对应命中)|
|
|
42
|
+
| gem_enrichment | op enrichment(必需基因通路富集:超几何+BH FDR;无注释 annotation_unavailable 兜底)| ✅ 已完成(C58 55 条 FDR 显著)|
|
|
43
|
+
| gem_targets | op targets(靶点规范导出:11 字段锁定 schema;账本计数闭合;引物设计不做)| ✅ 已完成(258 行三类闭合)|
|
|
44
|
+
| gem_precursor_scan | op precursor_scan(阻塞前体分析:基线通量→可生长即返「无阻塞」;不生长则逐前体移除测试定位阻塞点)| ✅ 2026-09-11(实测归因产出)|
|
|
45
45
|
|
|
46
46
|
> Python 分发器 `gem_ops.py` 共 **23 个 op**(annotate/benchmark/biomass_apply/biomass_inspect/double_knockout/enrichment/essential_scan/fluxscan/gapfill/gapfind/gapseq/l3_fix/ledger/media_resolve/model_info/phenotype_fix/precursor_scan/quality/sample/secretion/sensitivity/targets/validate);`gem_build` 不经分发器,由 `build.py` CLI 直接调用(长任务,jobs.js 拉起)。
|
|
47
47
|
>
|
|
@@ -49,7 +49,7 @@ dsh 平台的 **GEM 构建侧插件**:输入细菌全基因组(支持多质
|
|
|
49
49
|
|
|
50
50
|
> **precursor_scan 的判据取舍(勿回退)**:初版曾用「全开交换下逐前体 demand 能否净生产」的**绝对可达性**判据,在教科书模型 e_coli_core 上把 atp_c/accoa_c/nad_c/nadph_c 误报为「结构缺失」(辅因子有循环补给路径,稳态下不净生产 ≠ 网络不能供给),故否决。现行判据为**相对判断**:先测基线通量,可生长即直接返回「无阻塞」;不生长才逐前体做移除测试,由「移除后是否恢复通量」直接定义阻塞点。验证锚:toy 单点阻塞模型(精确命中)、e_coli_core(growable,零误报)、iNX1344_v3(infeasible_or_constrained,与 agent 手工探索结论一致)。
|
|
51
51
|
|
|
52
|
-
> 其余工具层约定:附模型卡统一写入 `python/model_card.py`(lineage/verified_phenotypes/essential_genes/robustness v3)与往返保真自检 `python/roundtrip_check.py`;预测账本 `python/ledger.py`(一个模型一个账本:`~/.dsh/dsh-bio-gem/ledger/<模型名>.jsonl`,按模型 basename 分,显式 ledger_path 可覆盖;无参查询=聚合全局视图;旧全局 predictions.jsonl 已迁移为 legacy
|
|
52
|
+
> 其余工具层约定:附模型卡统一写入 `python/model_card.py`(lineage/verified_phenotypes/essential_genes/robustness v3)与往返保真自检 `python/roundtrip_check.py`;预测账本 `python/ledger.py`(一个模型一个账本:`~/.dsh/dsh-bio-gem/ledger/<模型名>.jsonl`,按模型 basename 分,显式 ledger_path 可覆盖;无参查询=聚合全局视图;旧全局 predictions.jsonl 已迁移为 legacy)。**生长/通量数值口径**:所有产出生长/通量数值的工具输出均带 `units: mmol/gDW/h` 与单点 FBA 声明;条件间通量对比一律走 gem_fluxscan 区间分离判定(overlap=伪影禁止引用)。
|
|
53
53
|
|
|
54
54
|
## 4. 引擎路线(M1→M2→M3)
|
|
55
55
|
|
|
@@ -69,7 +69,7 @@ dsh 平台的 **GEM 构建侧插件**:输入细菌全基因组(支持多质
|
|
|
69
69
|
| G5 | 必需基因抽检(≤30 基因)| 条件 | 有参照集才跑;映射覆盖 <80% 时 SKIP(WARN) |
|
|
70
70
|
| G6 | ATP 泄漏检测(全关交换后 ATP demand 应≈0)| ✅ | leak ≤0.01 判 PASS;ATP 解析走 id→name→formula 三级回退(跨 ID 体系)|
|
|
71
71
|
|
|
72
|
-
**G0 的由来(2026-09-10
|
|
72
|
+
**G0 的由来(2026-09-10 实测)**:MetaCyc 风格 id 的公开模型(iNX1344_v3)上,
|
|
73
73
|
`gem_gapfind` 报 5 个 L3「内部通路缺口」,实为 biomass 前体未映射所致——agent 为逐个
|
|
74
74
|
证伪手写 cobra 代码 18 次。现 `gem_validate` 在 G1 之前输出 `g0`,`gem_gapfind` 返回
|
|
75
75
|
`coherence_warning` + `interpretation_guard`,把该结论前置给 agent。
|
|
@@ -126,7 +126,7 @@ job 化 + 进度事件(粒度 ≤5s)+ 分步 checkpoint(每步落盘,可
|
|
|
126
126
|
|
|
127
127
|
零手动干预下:**基因组进 → 四个消费工具(FBA/必需性/包络线/面板)不经修改即可用的 SBML 出**,且模型在声明培养基上生长为正;C58 端到端演示通过(build→面板可见→FBA 可跑→必需性可跑);模型卡齐全(引擎/版本/补洞记录/验证结果,同输入重跑一致);5-6 Mb 基因组 p95 ≤ 20 min。
|
|
128
128
|
|
|
129
|
-
## 附录 A
|
|
129
|
+
## 附录 A:性能基准(2026-08-30 实测,独占运行)
|
|
130
130
|
|
|
131
131
|
分析 Python 3.13.13 / cobra 0.32.1 / GLPK;C58=gapseq 2485 反应/1084 基因;iNX1344_v4=1441 反应/1344 基因。
|
|
132
132
|
|
|
@@ -142,4 +142,4 @@ job 化 + 进度事件(粒度 ≤5s)+ 分步 checkpoint(每步落盘,可
|
|
|
142
142
|
| 单组分 ±25% 灵敏度 | 75 组分×2=150 次 FBA,54.4s | 47 组分×2=94 次 FBA,7.9s |
|
|
143
143
|
| 必需性漂移 top10(含生长探针) | 522.0s(含 7 刚性对跳过探针) | 33.8s(20/20 全部"不生长跳过") |
|
|
144
144
|
|
|
145
|
-
> 注:FVA 占单条件耗时 ~75%;sensitivity 线性于组合数(每组合 fresh 读模+FVA+敲除循环)。GLPK 对个别扰动 LP 有病态停摆前科,sensitivity 内置 LP_TIMEOUT_S=30
|
|
145
|
+
> 注:FVA 占单条件耗时 ~75%;sensitivity 线性于组合数(每组合 fresh 读模+FVA+敲除循环)。GLPK 对个别扰动 LP 有病态停摆前科,sensitivity 内置 LP_TIMEOUT_S=30 护栏。
|
package/docs/releases/v0.1.12.md
CHANGED
|
@@ -17,7 +17,7 @@
|
|
|
17
17
|
- `test/smoke.js` 重构:模型资产**多候选解析**(`--assets-root` / `DSH_BIO_GEM_ASSETS` / 新旧位置)——资产迁移后不再硬编码路径崩溃;**缺资产 / 缺账本标记 SKIP 而非失败**(`--require-assets` 供 CI 严格模式);CarveMe 跳过判断移至调用前;子进程非零退出 fail-closed。
|
|
18
18
|
- `npm test` 接线:smoke + integration + optional-injection(真注册验证,此前未纳入任何 npm 入口)。
|
|
19
19
|
|
|
20
|
-
**引擎兼容(2026-09-19)**:经 dsh **0.1.5-rc.2**
|
|
20
|
+
**引擎兼容(2026-09-19)**:经 dsh **0.1.5-rc.2** 完整兼容核验(全量变更:零适配命中)与实机验证(工具注册 / 代谢建模面板 / 账本与模型数据)。
|
|
21
21
|
|
|
22
22
|
## 安装
|
|
23
23
|
|
|
@@ -38,6 +38,6 @@ dsh plugin --profile web add @dsh-bio/dsh-bio-gem
|
|
|
38
38
|
|
|
39
39
|
## 已知边界(诚实清单)
|
|
40
40
|
|
|
41
|
-
- `gem_build` 的 gapseq 引擎为实验性(WSL2
|
|
41
|
+
- `gem_build` 的 gapseq 引擎为实验性(WSL2 拓扑绑定,30–60 分钟/模型);默认 carveme 引擎已可出可验证模型。
|
|
42
42
|
- CarveMe 运行时需手动准备一次(README 第 3 节)——「装完插件」不等于「构建可用」,分析/验证类能力不受影响。
|
|
43
43
|
- 注释依赖 SBML groups(gapseq 系模型自带);无注释模型按契约返回 `annotation_unavailable` 兜底,不伪造通路。
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
# @dsh-bio/dsh-bio-gem v0.1.16
|
|
2
|
+
|
|
3
|
+
> **主题**:导出与工作区修复批次 + 公开文档整理。
|
|
4
|
+
> 基因组尺度代谢模型(GEM)构建插件:基因组输入 → 构建 → 验证 → 补洞 → 出报告(标准 SBML + 模型卡)。
|
|
5
|
+
|
|
6
|
+
## 修复与改进
|
|
7
|
+
|
|
8
|
+
**导出产物规范化**
|
|
9
|
+
- CSV 导出保持**纯数据表**(标准表头 + 行);声明/参数/免责信息移入旁车文件 `<path>.meta.json`(返回体附 `export_csv_meta` 路径)——标准解析器(pandas 等)可直接读取。
|
|
10
|
+
- 导出路径父目录不存在时**自动创建**(mkdir -p 语义),不再报 FileNotFoundError。
|
|
11
|
+
|
|
12
|
+
**工作区与运行环境**
|
|
13
|
+
- 工具相对路径基于会话工作区解析(python 子进程 cwd 对齐会话目录)。
|
|
14
|
+
- gapseq 能力探针改为只读检查本地序列库元数据,不再依赖在线检查(离线环境不再误报)。
|
|
15
|
+
|
|
16
|
+
**公开文档整理**
|
|
17
|
+
- 清理内部流程信息;架构文档措辞规范化。
|
|
18
|
+
|
|
19
|
+
## 安装
|
|
20
|
+
|
|
21
|
+
```sh
|
|
22
|
+
npx -y @deepseek-ai/dsh plugin --profile web add @dsh-bio/dsh-bio-gem
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
桌面端(0.2.0+):用桌面端自带 CLI(`--profile desktop`)或应用内「插件」页安装。
|
|
26
|
+
分析用 Python(cobra)优先复用 dsh-bio-genie 自举环境;未装 genie 时见 README 第 2 节。
|
|
27
|
+
|
|
28
|
+
## 已知边界(诚实清单)
|
|
29
|
+
|
|
30
|
+
- `gem_build` 的 gapseq 引擎为实验性(WSL2 环境绑定,30–60 分钟/模型);默认 carveme 引擎已可出可验证模型。
|
|
31
|
+
- 其余边界见 v0.1.12 清单。
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@dsh-bio/dsh-bio-gem",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.16",
|
|
4
4
|
"description": "基因组尺度代谢模型(GEM)构建插件:输入细菌全基因组(蛋白FASTA,支持多质粒/多染色体),自动构建+验证+补洞+出报告(SBML + 模型卡),供 dsh-bio-genie 消费工具加载使用 | Genome-scale metabolic model builder for dsh",
|
|
5
5
|
"repository": {
|
|
6
6
|
"type": "git",
|
|
@@ -52,8 +52,8 @@
|
|
|
52
52
|
},
|
|
53
53
|
"scripts": {
|
|
54
54
|
"smoke": "node test/smoke.js",
|
|
55
|
-
"test": "node test/check-version-source.mjs && node test/smoke.js --skip-build && node test/integration.js && node --import ./test/register-dsh-tools.mjs test/optional-injection.js && node --import ./test/register-dsh-tools.mjs test/check-capabilities.mjs && node test/check-counts.mjs",
|
|
56
|
-
"test:full": "node test/smoke.js && node test/integration.js && node --import ./test/register-dsh-tools.mjs test/optional-injection.js && node --import ./test/register-dsh-tools.mjs test/check-capabilities.mjs && node test/check-counts.mjs"
|
|
55
|
+
"test": "node test/check-version-source.mjs && node test/smoke.js --skip-build && python -I test/test_gapseq_probe.py && python -I test/test_fsutil.py && node test/integration.js && node test/capabilities.js && node --import ./test/register-dsh-tools.mjs test/optional-injection.js && node --import ./test/register-dsh-tools.mjs test/check-capabilities.mjs && node test/check-counts.mjs && node test/capabilities-warn-state.mjs && node test/workdir.mjs",
|
|
56
|
+
"test:full": "node test/smoke.js && node test/integration.js && node test/capabilities.js && node --import ./test/register-dsh-tools.mjs test/optional-injection.js && node --import ./test/register-dsh-tools.mjs test/check-capabilities.mjs && node test/check-counts.mjs"
|
|
57
57
|
},
|
|
58
58
|
"peerDependencies": {
|
|
59
59
|
"@deepseek-ai/dsh-tools": "*"
|
|
@@ -63,4 +63,4 @@
|
|
|
63
63
|
"optional": true
|
|
64
64
|
}
|
|
65
65
|
}
|
|
66
|
-
}
|
|
66
|
+
}
|
package/python/annotate.py
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
# annotate.py — 路线 P0 注释步骤(纯 Windows)
|
|
2
|
-
#
|
|
2
|
+
# 策略(设计评审采纳 + 验证协议):官方注释优先 + pyrodigal 兜底
|
|
3
3
|
# 1) 同目录 *_protein.faa 存在(NCBI dataset 常见)→ 直接用
|
|
4
4
|
# 2) 同目录 *.gff(含 CDS)→ 解析坐标从 fna 提取 + 翻译(transl_table 11)
|
|
5
5
|
# 3) 仅 .fna → pyrodigal(多序列模式;总长 <100kb 时 meta 模式)
|
package/python/benchmark.py
CHANGED
|
@@ -125,7 +125,7 @@ def _essential_full_scan(model_path, medium, log, tag):
|
|
|
125
125
|
|
|
126
126
|
def map_genes(genes, model_b):
|
|
127
127
|
"""基因映射尽力而为(与物种无关):策略1 identity(同 id);策略2 gene.name 匹配。
|
|
128
|
-
反应桥(EC
|
|
128
|
+
反应桥(EC/名字)需两侧注释充分;两侧命名空间注释层不足时不启用,如实报告。"""
|
|
129
129
|
b_by_id = {g.id for g in model_b.genes}
|
|
130
130
|
b_by_name = {}
|
|
131
131
|
for g in model_b.genes:
|
|
@@ -387,9 +387,9 @@ def write_md(path, out):
|
|
|
387
387
|
def fetch_bigg_model(model_id, dest_dir=None):
|
|
388
388
|
"""从 BiGG 下载模型 SBML(B3 最小版)。
|
|
389
389
|
URL 策略(实测 2026-08-30):静态库 http://bigg.ucsd.edu/static/models/<id>.xml 返回标准 SBML;
|
|
390
|
-
|
|
390
|
+
`/api/v2/universal/models/<id>/download` 实为 404(universal 是 reactions 命名空间),
|
|
391
391
|
/api/v2/models/<id>/download 返回 200 但内容是 BiGG JSON(非 SBML)——两者均不采用。
|
|
392
|
-
|
|
392
|
+
直连失败走系统代理;都失败抛错(调用方如实报告,不阻塞本地对比)。"""
|
|
393
393
|
import urllib.request
|
|
394
394
|
import shutil
|
|
395
395
|
urls = [f"http://bigg.ucsd.edu/static/models/{model_id}.xml"]
|
package/python/coherence.py
CHANGED
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
# 目的:在 G3/G4/G5 与 gapfind 之前,先判断**模型自身数据质量是否允许下结论**,
|
|
4
4
|
# 避免把「未映射代谢物」这类数据问题,误报成「通路缺口 / 必需基因异常」。
|
|
5
5
|
#
|
|
6
|
-
# 实测来源(2026-09-10
|
|
6
|
+
# 实测来源(2026-09-10,iNX1344_v3 —— MetaCyc 风格 id 的公开模型):
|
|
7
7
|
# - gem_gapfind 报 5 个 L3「内部通路缺口」,根因实为 biomass 前体未映射;
|
|
8
8
|
# - agent 为证伪这些假阳性,手写 cobra 代码 18 次(占该轮调用的一半)。
|
|
9
9
|
#
|
|
@@ -16,7 +16,7 @@
|
|
|
16
16
|
# 2. 「前体可达性(demand 逐前体 FBA)」——初版实现同样在 e_coli_core 上把
|
|
17
17
|
# atp_c / accoa_c / nad_c / nadph_c 误报为「既不能合成也不能摄取」。全开交换下
|
|
18
18
|
# 的 demand 语义与胞内辅因子/能量货币的循环补给路径纠缠,判据未成熟。
|
|
19
|
-
# 正确方法(agent
|
|
19
|
+
# 正确方法(agent 曾手工探索过)待重新设计后引入。
|
|
20
20
|
#
|
|
21
21
|
# 保留的判据都经过「问题模型报出真问题 + 标准模型零误报」双向验证:
|
|
22
22
|
# - id 体系识别(信息性,决定下游名称映射口径)
|
|
@@ -18,6 +18,7 @@ sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
|
|
18
18
|
|
|
19
19
|
from silentio import silent_read_sbml
|
|
20
20
|
from essential_scan import setup_model_medium, scan_essentiality, EPS
|
|
21
|
+
from fsutil import ensure_parent_dir, write_meta_sidecar
|
|
21
22
|
|
|
22
23
|
ASSUMPTION_NOTE = "细菌双敲验证率无大规模实验数据支撑,本结果=假设生成,供实验设计参考非结论"
|
|
23
24
|
|
|
@@ -174,16 +175,35 @@ def double_knockout(model_path, medium=None, max_pairs=5000, export_csv=None,
|
|
|
174
175
|
|
|
175
176
|
|
|
176
177
|
def _export_csv(path, results, out):
|
|
178
|
+
"""纯数据 CSV(标准表头,机器可读)+ 旁车 meta。
|
|
179
|
+
|
|
180
|
+
2026-10-05 修:此前把 assumption_note 写成首行 `"# assumption", note` 双字段行——
|
|
181
|
+
pandas/自动解析会把它当表头或脏行(实测 agent 需专门跳过 '#' 行)。
|
|
182
|
+
现改为:CSV 保持纯表;说明与参数写入 <path>.meta.json(返回体 export_csv_meta)。
|
|
183
|
+
"""
|
|
184
|
+
path = ensure_parent_dir(path)
|
|
177
185
|
n = 0
|
|
178
186
|
with open(path, "w", newline="", encoding="utf-8-sig") as f:
|
|
179
187
|
w = csv.writer(f)
|
|
180
|
-
w.writerow(["# assumption", out["assumption_note"]])
|
|
181
188
|
w.writerow(["gene_a", "gene_b", "single_a_growth", "single_b_growth",
|
|
182
189
|
"double_growth", "rationale", "source"])
|
|
183
190
|
for r in results:
|
|
184
191
|
w.writerow([r["pair"][0], r["pair"][1], r["single_a_growth"],
|
|
185
192
|
r["single_b_growth"], r["double_growth"], r["rationale"], r["source"]])
|
|
186
193
|
n += 1
|
|
194
|
+
meta = write_meta_sidecar(path, {
|
|
195
|
+
"assumption_note": out.get("assumption_note"),
|
|
196
|
+
"model": out.get("model"),
|
|
197
|
+
"medium": out.get("medium"),
|
|
198
|
+
"medium_preset": out.get("medium_preset"),
|
|
199
|
+
"wt_growth": out.get("wt_growth"),
|
|
200
|
+
"units": out.get("units"),
|
|
201
|
+
"eps": out.get("eps"),
|
|
202
|
+
"max_pairs": out.get("max_pairs"),
|
|
203
|
+
"pairs_found": out.get("pairs_found"),
|
|
204
|
+
})
|
|
205
|
+
if meta:
|
|
206
|
+
out["export_csv_meta"] = meta
|
|
187
207
|
return n
|
|
188
208
|
|
|
189
209
|
|
package/python/enrichment.py
CHANGED
|
@@ -13,6 +13,7 @@ from math import comb
|
|
|
13
13
|
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
|
14
14
|
|
|
15
15
|
from silentio import silent_read_sbml
|
|
16
|
+
from fsutil import ensure_parent_dir
|
|
16
17
|
|
|
17
18
|
NOTE_UNAVAILABLE = ("模型无 SBML groups(通路)注释。可补途径:①用 gapseq 重建(自带 MetaCyc pathway "
|
|
18
19
|
"groups);②从 BiGG API 取模型 subsystem 后注入 groups;③按 ec-code 注释做粗分类。")
|
|
@@ -153,6 +154,7 @@ def enrichment(model_path, gene_list=None, pathway_source=None, ledger_path=None
|
|
|
153
154
|
"timing_seconds": round(time.time() - t0, 1),
|
|
154
155
|
}
|
|
155
156
|
if export_csv:
|
|
157
|
+
export_csv = ensure_parent_dir(export_csv)
|
|
156
158
|
with open(export_csv, "w", newline="", encoding="utf-8-sig") as f:
|
|
157
159
|
w = csv.writer(f)
|
|
158
160
|
w.writerow(["pathway", "genes_hit_count", "genes_hit", "background_hit",
|
package/python/fluxscan.py
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
# fluxscan.py — M1 通量区间制(阶段 A 可信度内核第一件)
|
|
2
|
-
#
|
|
2
|
+
# 语义(已锁定):每反应输出 fva_min/fva_max/pfba;条件对比消费区间分离判定;
|
|
3
3
|
# overlap = 点值差异是求解器伪影,禁止引用。
|
|
4
4
|
# 计算口径:每 condition 独立 silent_read_sbml 重读模型(勿深拷贝);介质 setup 对齐 validate G3
|
|
5
5
|
# (expand_medium -> 全交换清零 EX_/DM_/SK_/boundary -> resolve_medium 设 bounds);
|
|
@@ -21,6 +21,7 @@ sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
|
|
21
21
|
from silentio import silent_read_sbml
|
|
22
22
|
from gapfind import expand_medium, resolve_medium, build_ex_index, match_ex
|
|
23
23
|
from validate import parse_formula
|
|
24
|
+
from fsutil import ensure_parent_dir
|
|
24
25
|
|
|
25
26
|
EX_PREFIX = ("EX_", "DM_", "SK_")
|
|
26
27
|
DEFAULT_FRACTION = 0.9999
|
|
@@ -168,6 +169,7 @@ def _pair_result(a_name, b_name, data_a, data_b, scope, tol, only_diff):
|
|
|
168
169
|
|
|
169
170
|
def _export_csv(path, model_reactions, pair_results):
|
|
170
171
|
"""全量 CSV:每行 = 条件对 × 反应(双侧区间/点值/判定)。返回 (rows, bytes)。"""
|
|
172
|
+
path = ensure_parent_dir(path)
|
|
171
173
|
names = {r.id: r for r in model_reactions}
|
|
172
174
|
rows = 0
|
|
173
175
|
with open(path, "w", newline="", encoding="utf-8-sig") as f:
|
|
@@ -267,15 +269,15 @@ if __name__ == "__main__":
|
|
|
267
269
|
# 双协议(历史坑:stdin 与 argv-file 两派并存,新脚本必须都支持)
|
|
268
270
|
if "--selftest" in sys.argv:
|
|
269
271
|
cases = [
|
|
270
|
-
# (la, ua, lb, ub, tol, direction, hard) ——
|
|
272
|
+
# (la, ua, lb, ub, tol, direction, hard) —— 5 组判定样例
|
|
271
273
|
(1.0, 2.0, 5.0, 6.0, 1e-6, "b_higher", True), # ① 分离:b 高
|
|
272
274
|
(5.0, 6.0, 1.0, 2.0, 1e-6, "a_higher", True), # ① 分离:a 高
|
|
273
275
|
(1.0, 3.0, 2.0, 4.0, 1e-6, None, False), # ② 重叠(伪影)
|
|
274
276
|
(1.0, 3.0, 3.0, 5.0, 1e-6, None, False), # ② 接触容差内仍 overlap(边界)
|
|
275
277
|
(-5.0, -2.0, 0.0, 0.0, 1e-6, "b_higher", True), # ③ 零通量/负向
|
|
276
278
|
(-5.0, -2.0, 0.0, 0.0, 0.0, "b_higher", True), # ④ 精确边界(tol=0)
|
|
277
|
-
# ⑤ 容差:锁定公式下 gap=5e-7 < tol=1e-6 -> overlap
|
|
278
|
-
# 与锁定公式 ua+tol<lb 矛盾:5e-7 不大于 1e-6
|
|
279
|
+
# ⑤ 容差:锁定公式下 gap=5e-7 < tol=1e-6 -> overlap(该样例期望"分离",
|
|
280
|
+
# 与锁定公式 ua+tol<lb 矛盾:5e-7 不大于 1e-6。按公式实现)
|
|
279
281
|
(1.0, 1.0000005, 1.000001, 2.0, 1e-6, None, False),
|
|
280
282
|
# 同数字、tol(1e-7)<gap(5e-7) -> 分离:证明容差机制本身生效
|
|
281
283
|
(1.0, 1.0000005, 1.000001, 2.0, 1e-7, "b_higher", True),
|
package/python/fsutil.py
ADDED
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
"""fsutil.py — 文件系统小工具(导出路径处理,2026-10-05 增补)。
|
|
2
|
+
|
|
3
|
+
背景(round1 TC4 观察 GEM-2,round2 复核):export_csv/export_path 传相对路径
|
|
4
|
+
且父目录不存在时,直接 FileNotFoundError ——对用户是零价值的报错(需手动 mkdir)。
|
|
5
|
+
统一改为自动创建父目录(mkdir -p 语义),返回绝对路径。
|
|
6
|
+
- 部分导出曾把声明文本写成 CSV 首行 `"# xxx", note` 双字段行——标准 CSV 解析器
|
|
7
|
+
(pandas 默认)会把它当表头或脏行(实测 agent 需专门跳过 '#' 行)。
|
|
8
|
+
统一改为:CSV 保持纯数据表,声明与参数写入旁车文件 <path>.meta.json。
|
|
9
|
+
"""
|
|
10
|
+
from __future__ import annotations
|
|
11
|
+
|
|
12
|
+
import json
|
|
13
|
+
import os
|
|
14
|
+
import time
|
|
15
|
+
|
|
16
|
+
|
|
17
|
+
def ensure_parent_dir(path: str) -> str:
|
|
18
|
+
"""确保导出路径的父目录存在,返回绝对路径。
|
|
19
|
+
|
|
20
|
+
空 dirname(纯文件名 → 当前目录)时只返回绝对路径;
|
|
21
|
+
目录已存在时为幂等 no-op。
|
|
22
|
+
"""
|
|
23
|
+
abspath = os.path.abspath(path)
|
|
24
|
+
parent = os.path.dirname(abspath)
|
|
25
|
+
if parent:
|
|
26
|
+
os.makedirs(parent, exist_ok=True)
|
|
27
|
+
return abspath
|
|
28
|
+
|
|
29
|
+
|
|
30
|
+
def write_meta_sidecar(path: str, payload: dict) -> str | None:
|
|
31
|
+
"""在导出文件旁写 <path>.meta.json(声明/参数/时间戳)。失败仅告警返回 None。
|
|
32
|
+
|
|
33
|
+
调用方约定:CSV 本体只放纯数据;任何注释性说明(assumption/boundary 等)
|
|
34
|
+
放进 payload,由本函数落盘。
|
|
35
|
+
"""
|
|
36
|
+
meta_path = path + ".meta.json"
|
|
37
|
+
try:
|
|
38
|
+
data = dict(payload)
|
|
39
|
+
data.setdefault("generated_at", time.strftime("%Y-%m-%dT%H:%M:%S%z"))
|
|
40
|
+
with open(meta_path, "w", encoding="utf-8") as fh:
|
|
41
|
+
json.dump(data, fh, ensure_ascii=False, indent=1)
|
|
42
|
+
return meta_path
|
|
43
|
+
except Exception:
|
|
44
|
+
return None
|
package/python/gapseq_wsl.py
CHANGED
|
@@ -25,6 +25,7 @@ WSL_DISTRO = os.environ.get("GEM_GAPSEQ_DISTRO", "Ubuntu-22.04")
|
|
|
25
25
|
CONDA_SH = "/opt/miniforge3/etc/profile.d/conda.sh"
|
|
26
26
|
GAPSEQ_ENV = "gapseq"
|
|
27
27
|
WSL_WORK = "/opt/gem-gapseq-work"
|
|
28
|
+
SEQDB_DIR = f"/opt/miniforge3/envs/{GAPSEQ_ENV}/share/gapseq/dat/seq/Bacteria"
|
|
28
29
|
|
|
29
30
|
|
|
30
31
|
def wsl_run(bash_cmd, timeout=300):
|
|
@@ -79,13 +80,34 @@ def probe():
|
|
|
79
80
|
import re
|
|
80
81
|
mm = re.search(r"gapseq version:\s*([\d.]+)", out)
|
|
81
82
|
res["gapseq_version"] = mm.group(1) if mm else out[:80]
|
|
82
|
-
|
|
83
|
+
# doall 实际读取 conda 环境内的 seq/Bacteria;旧探针却查 F: 备份并调用
|
|
84
|
+
# update-sequences -c,后者依赖 Zenodo 在线记录。网络解析失败会误报后端缺失。
|
|
85
|
+
# 本地元数据 + 三类序列文件同时检查,缺任何一项都不宣称可用。
|
|
83
86
|
rc, out, err = wsl_run(
|
|
84
|
-
f"
|
|
85
|
-
f"
|
|
86
|
-
|
|
87
|
+
f"set -o pipefail; cat {SEQDB_DIR}/version_seqDB.json && "
|
|
88
|
+
f"find {SEQDB_DIR}/rev -type f | wc -l && "
|
|
89
|
+
f"find {SEQDB_DIR}/rxn -type f | wc -l && "
|
|
90
|
+
f"find {SEQDB_DIR}/unrev -type f | wc -l", 120)
|
|
91
|
+
try:
|
|
92
|
+
meta, offset = json.JSONDecoder().raw_decode(out)
|
|
93
|
+
counts = [int(value) for value in out[offset:].split()]
|
|
94
|
+
metadata_ok = (isinstance(meta, dict)
|
|
95
|
+
and isinstance(meta.get("version"), list)
|
|
96
|
+
and len(meta["version"]) == 1
|
|
97
|
+
and isinstance(meta["version"][0], str)
|
|
98
|
+
and bool(meta["version"][0])
|
|
99
|
+
and isinstance(meta.get("zenodoID"), list)
|
|
100
|
+
and len(meta["zenodoID"]) == 1
|
|
101
|
+
and isinstance(meta["zenodoID"][0], int)
|
|
102
|
+
and meta["zenodoID"][0] > 0)
|
|
103
|
+
res["checks"]["seqdb"] = rc == 0 and metadata_ok and len(counts) == 3 and all(n > 0 for n in counts)
|
|
104
|
+
if res["checks"]["seqdb"]:
|
|
105
|
+
res["seqdb_version"] = meta["version"][0]
|
|
106
|
+
res["seqdb_counts"] = dict(zip(("rev", "rxn", "unrev"), counts))
|
|
107
|
+
except (IndexError, ValueError, TypeError, json.JSONDecodeError):
|
|
108
|
+
res["checks"]["seqdb"] = False
|
|
87
109
|
if not res["checks"]["seqdb"]:
|
|
88
|
-
res["detail"] = f"
|
|
110
|
+
res["detail"] = f"gapseq 实际序列库缺失或不完整: {(out or err)[:400]}"
|
|
89
111
|
res["level"] = "DEGRADED"
|
|
90
112
|
return res
|
|
91
113
|
res["capable"] = True
|
|
@@ -248,4 +270,4 @@ if __name__ == "__main__":
|
|
|
248
270
|
f"source {CONDA_SH} && conda activate {GAPSEQ_ENV} && gapseq -v 2>&1", 120)
|
|
249
271
|
print(out[:300] or err[:300])
|
|
250
272
|
else:
|
|
251
|
-
print("unknown cmd")
|
|
273
|
+
print("unknown cmd")
|
package/python/gem_ops.py
CHANGED
|
@@ -264,7 +264,7 @@ def op_l3_fix(args):
|
|
|
264
264
|
|
|
265
265
|
# ---------------------------------------------------------------------------
|
|
266
266
|
# op: fluxscan — 阶段A-M1 通量区间制(FVA 区间 + pFBA 点值 + 条件对区间分离判定)
|
|
267
|
-
#
|
|
267
|
+
# 语义锁定:overlap = 求解器伪影禁止引用;判定公式见 fluxscan.judge_interval(单测锁定)
|
|
268
268
|
# ---------------------------------------------------------------------------
|
|
269
269
|
def op_fluxscan(args):
|
|
270
270
|
from fluxscan import fluxscan, DEFAULT_FRACTION, DEFAULT_TOL
|
|
@@ -533,7 +533,7 @@ OPS["targets"] = op_targets
|
|
|
533
533
|
# ---------------------------------------------------------------------------
|
|
534
534
|
# op: precursor_scan — 阻塞前体分析
|
|
535
535
|
# 「模型为什么不长」的结构级定位:逐前体做移除测试,找出卡住生长的前体。
|
|
536
|
-
# 来源:2026-09-11
|
|
536
|
+
# 来源:2026-09-11 实测归因(agent 手写该逻辑 6+ 次,无工具可用)。
|
|
537
537
|
# 判据刻意用「相对判断」而非绝对可达性 —— 对可生长模型天然零误报。
|
|
538
538
|
# ---------------------------------------------------------------------------
|
|
539
539
|
def op_precursor_scan(args):
|
package/python/l3_fix.py
CHANGED
|
@@ -28,7 +28,7 @@ EX_PREFIX = ("EX_", "DM_", "SK_")
|
|
|
28
28
|
SOURCE_TAG = "gem-l3fix"
|
|
29
29
|
NEW_RXN_SUFFIX = "_l3fix"
|
|
30
30
|
|
|
31
|
-
# BiGG 基名 -> ModelSEED cpd 号(2026-08-29
|
|
31
|
+
# BiGG 基名 -> ModelSEED cpd 号(2026-08-29 以 C58 名字+公式+电荷逐一验证;
|
|
32
32
|
# 映射时仍做公式/电荷校验,不一致即弃用防静默污染化学计量)
|
|
33
33
|
COFACTOR_BRIDGE = {
|
|
34
34
|
"h2o": "00001", "atp": "00002", "nad": "00003", "nadh": "00004",
|
|
@@ -42,7 +42,7 @@ COFACTOR_BRIDGE = {
|
|
|
42
42
|
COMP_MAP = {"c": "c0", "e": "e0", "p": "p0"} # BiGG 区室后缀 -> gapseq 区室 id
|
|
43
43
|
|
|
44
44
|
WHITELIST_DIR = os.path.join(os.path.expanduser("~"), ".dsh", "dsh-bio-gem", "whitelist")
|
|
45
|
-
# 本地白名单数据库(license 守则:
|
|
45
|
+
# 本地白名单数据库(license 守则: 仅本地留存,不进 git/发布包;GEM_WHITELIST_DB_DIR 可覆盖)
|
|
46
46
|
RXN_DB_DIR = os.environ.get("GEM_WHITELIST_DB_DIR", r"D:\Program\hermes\temp\gem_whitelist")
|
|
47
47
|
DEFAULT_UNIVERSAL = r"D:\Program\hermes\temp\gem_universal\iML1515.xml"
|
|
48
48
|
|
package/python/ledger.py
CHANGED
|
@@ -7,7 +7,7 @@
|
|
|
7
7
|
# 幂等: 同 model+condition+type+content 哈希去重,重复运行不追加。
|
|
8
8
|
# 完整性: 逐行 JSON 解析校验,损坏行跳过并在返回里报 corrupt_rows + 行号(不阻塞);
|
|
9
9
|
# 写入失败只 WARN 不使主流程失败。update 重写文件但保留损坏行原样(不删行)。
|
|
10
|
-
#
|
|
10
|
+
# 证据分级优先级(口径锁定): EVIDENCE_literature > EVIDENCE_sequence > EVIDENCE_rule > EVIDENCE_math
|
|
11
11
|
import os
|
|
12
12
|
import re
|
|
13
13
|
import sys
|
|
@@ -29,7 +29,7 @@ def _now_iso():
|
|
|
29
29
|
|
|
30
30
|
|
|
31
31
|
def _content_hash(model, condition, rtype, content):
|
|
32
|
-
#
|
|
32
|
+
# 实测发现:Windows 路径斜杠/大小写差异("F:/a" vs "F:\a")会使同一模型的
|
|
33
33
|
# 重复登记漏过去重——hash 前做 normcase+normpath 归一化(首登记的存储格式不变)。
|
|
34
34
|
m = os.path.normcase(os.path.normpath(model)) if model else ""
|
|
35
35
|
raw = "\x1f".join([m, condition or "", rtype or "", content or ""])
|
package/python/model_card.py
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
# 纪律: 各工具完成后**仅当产物模型旁已有 card** 才向后追加;无卡不动(不凭空造卡)。
|
|
5
5
|
# 兼容: build.py 旧卡(无 schema 字段)读取时即时迁移到 v2(新增字段缺失不报错)。
|
|
6
6
|
# units: growth_rate 一律 1/h(比生长速率;biomass 反应 gDW 归一化口径,数值 = μ)。
|
|
7
|
-
# 2026-09-21 由 mmol/gDW/h 演进为 1/h
|
|
7
|
+
# 2026-09-21 由 mmol/gDW/h 演进为 1/h:数值不变、生物语义更准(schema 向后兼容,旧卡照读)。
|
|
8
8
|
import os
|
|
9
9
|
import json
|
|
10
10
|
import time
|
package/python/precursor_scan.py
CHANGED
package/python/quality.py
CHANGED
|
@@ -14,6 +14,7 @@ from cobra.flux_analysis import fastcc, find_blocked_reactions
|
|
|
14
14
|
|
|
15
15
|
from gapfind import expand_medium, resolve_medium
|
|
16
16
|
from silentio import silent_read_sbml
|
|
17
|
+
from fsutil import ensure_parent_dir
|
|
17
18
|
|
|
18
19
|
|
|
19
20
|
EX_PREFIXES = ("EX_", "DM_", "SK_")
|
|
@@ -357,6 +358,7 @@ def _connectivity(model):
|
|
|
357
358
|
|
|
358
359
|
def _write_export_csv(path, records):
|
|
359
360
|
"""Write long-form complete lists; JSON samples remain deliberately capped."""
|
|
361
|
+
path = ensure_parent_dir(path)
|
|
360
362
|
with open(path, "w", encoding="utf-8", newline="") as handle:
|
|
361
363
|
writer = csv.DictWriter(handle, fieldnames=("check", "item_id", "detail"))
|
|
362
364
|
writer.writeheader()
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# roundtrip_check.py — SBML 往返保真自检(Q2 工程质量件 A)
|
|
2
2
|
# cobra 读 → write_sbml_model → 读回:断言反应/代谢物/基因数一致 + GPR 字符串精确一致
|
|
3
|
-
#
|
|
3
|
+
# (经典暗坑:序列化静默丢 GPR——fbc v2 写入路径回归护栏;≥5 个复合 GPR 反应样本必查)
|
|
4
4
|
import os
|
|
5
5
|
import sys
|
|
6
6
|
import json
|
package/python/sampling.py
CHANGED
|
@@ -11,6 +11,7 @@ from cobra.util.solver import linear_reaction_coefficients
|
|
|
11
11
|
|
|
12
12
|
from gapfind import expand_medium, resolve_medium
|
|
13
13
|
from silentio import silent_read_sbml
|
|
14
|
+
from fsutil import ensure_parent_dir
|
|
14
15
|
|
|
15
16
|
|
|
16
17
|
EX_PREFIXES = ("EX_", "DM_", "SK_")
|
|
@@ -219,6 +220,7 @@ def sample_fluxes(
|
|
|
219
220
|
growth_summary = _describe(samples[growth_reaction.id].to_numpy())
|
|
220
221
|
|
|
221
222
|
if export_csv:
|
|
223
|
+
export_csv = ensure_parent_dir(export_csv)
|
|
222
224
|
samples.to_csv(export_csv, index=False)
|
|
223
225
|
|
|
224
226
|
if growth_floor_fraction is None:
|