renkin 1.0.0 → 1.0.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -171,9 +171,16 @@ Use `--format mermaid` for GitHub/Notion-compatible flowcharts.
171
171
 
172
172
  ## Current Limitations
173
173
 
174
- ⚠️ Benchmark numbers are under active re-measurement after a validator-accuracy
175
- fix — historical 78.0%/95.9%/81.8%(ChEMBL) figures elsewhere in this repo predate
176
- that fix and are invalidated. RENKIN does not predict yields, calibrated
174
+ Current release: **v1.0.2**.
175
+
176
+ ⚠️ The corrected 4,903-target v1.0.1 shared-stock comparison is complete. RENKIN
177
+ recorded 591/4,903 primary successes (12.05%) versus AiZynthFinder 4.4.1's
178
+ 200/4,903 (4.08%); the paired difference was +7.975 percentage points with 95%
179
+ CI [+7.098, +8.852]. The full v1.0.1 arm passed integrity verification, so both
180
+ the statistical and formal publication gates pass. This is a result for the
181
+ declared shared-stock endpoint, not universal CASP superiority. Historical
182
+ 78.0%/95.9%/81.8%(ChEMBL) figures
183
+ elsewhere in this repo predate validator fixes and are invalidated. RENKIN does not predict yields, calibrated
177
184
  experimental success probabilities, or side reactions, and does not search
178
185
  the literature automatically (`success_probability` is a template-frequency
179
186
  search-ranking score, not a calibrated prediction — see
@@ -410,16 +417,17 @@ for the full acceptance criteria and licensing split.
410
417
  | **Ring-context safety guard** | `--ring-context-policy conservative --ring-context-sidecar <path>` — opt-in match-level filter that rejects an extracted template's ring-opening/closing disconnection when its historical training data never observed that bond as ring-forming/-breaking; default `disabled` (unchanged legacy behavior) — see [Issue #72](https://github.com/kent-tokyo/renkin/issues/72) |
411
418
  | **LightGBM candidate reranker** | `--reranker-model`/`--reranker-freq-table` (CLI) or `reranker_model_path`/`reranker_freq_table_path` (Python) — opt-in, ordering-only re-ranking via a frozen LightGBM model; never changes which candidates are generated, only their order, and reproduces legacy ordering byte-for-byte when off. Paired 100-target route-search gate: `route_to_configured_stock` 16→20 (+4/-0). `python3 scripts/fetch_reranker_model.py` fetches the frozen model (SHA-256-verified, not bundled in any package — see [Roadmap](#roadmap)) |
412
419
  | **Coverage mode** (opt-in) | `--search-mode coverage --coverage-templates <path>` (CLI) or `search_mode="coverage"`, `coverage_templates_path=...` (Python) — if the default template set finds no route, automatically escalates to a larger, separately loaded template set, cooperatively cancellable via `--coverage-timeout-secs`. Standard-mode output is byte-for-byte unchanged when not used. `python3 scripts/fetch_coverage_templates.py` fetches the frozen 2,000-template Stage-2 set (SHA-256-verified, not bundled in any package, same reasoning as the reranker model — see [Roadmap](#roadmap)) |
420
+ | **Staged recovery** (opt-in, native) | `--search-mode recovery --beam-diversity-slots N` preserves a successful baseline and conditionally escalates element gating, diversity, depth, and caller-supplied narrow-to-broad coverage tiers. Every attempt is audit-visible; standard mode and its defaults are unchanged. See [staged recovery mode](docs/guides/staged-recovery.md) |
413
421
  | **RENKIN Bridge / `audit-route`** | `renkin audit-route route.json [--format auto\|renkin\|aizynthfinder\|syntheseus\|synplanner] [--stock stock.smi] [--output human\|json]` — tool-neutral route audit: structural integrity, stock, and declared-reaction forward-replay validation, each reported independently as `pass`/`fail`/`not_evaluable`, rolled up into a route-level `pass`/`fail`/`partial` verdict. Reads RENKIN-native route JSON (v0.25.0), real AiZynthFinder route JSON — single-target and gzip-compressed batch output, verified against AiZynthFinder 4.3.2, 4.4.0, and 4.4.1 specifically, not claimed for every version (v0.26.0, version matrix widened v0.32.0) — Syntheseus routes via the optional `renkin.syntheseus_exporter`'s `syntheseus-route-v1` interchange schema, since Syntheseus itself has no native route export (v0.30.0) — and real SynPlanner 1.6.0 `write_routes_json` exports directly, no exporter package needed (v0.34.0); `--format auto` detects the input shape and hard-errors rather than guessing on anything ambiguous. [AiZynthFinder walkthrough →](https://kent-tokyo.github.io/renkin/guides/aizynthfinder-audit-demo/) · [Syntheseus walkthrough →](https://kent-tokyo.github.io/renkin/guides/syntheseus-audit-demo/) · [SynPlanner walkthrough →](https://kent-tokyo.github.io/renkin/guides/synplanner-audit-demo/) |
414
- | **Route scoring** | `confidence`, `step_confidence`, `success_probability` (Retro-prob style), `convergency`, `atom_economy`, `route_cost` (`Σ BB cost + steps×0.5`, or actual prices via `--bb-prices`/`--stock`) per step/route — see caveat below the table |
422
+ | **Route scoring & diagnostics** | Separate `confidence`, `success_probability`, cost, feasibility findings, building-block diversity, and template-proxy chemical-idea diversity; no aggregate laboratory-feasibility score is fabricated — see the caveat below and the [diagnostics](docs/guides/route-feasibility-diagnostics.md) / [diversity](docs/guides/route-set-diversity.md) guides |
415
423
  | **Step metadata provenance** | Each step reports `metadata_source`/`metadata_scope` so it's machine-readable whether `conditions`/`reaction_family` came from a rule-author default vs. something more grounded; absent (not fabricated) for extracted templates |
416
424
  | **Pareto multi-objective search** | `--format pareto` returns a Pareto front across `route_cost`/`success_probability`/`steps`; objectives configurable via `--objectives` |
417
425
  | **Constraint DSL** | `--constraints constraints.json` — element/building-block filters, step/cost limits, confidence thresholds, required/avoided/preferred reaction families; enables LLM → RENKIN pipelines |
418
426
  | **Output formats & diagnostics** | `--format json\|tree\|mermaid\|explain\|compare\|compare-json\|pareto`; zero-route JSON includes a `diagnostics` block with `likely_causes`/`suggestions` |
419
427
  | **`renkin-forward` toolkit** | `predict` (rank forward products), `enumerate` (bounded products from one reactant + partner library), `hints` (partner-free retrieval hints, no concrete product), `validate` (forward-verify each retro step) — see the [Forward guides](docs/guides/forward-retrieval-hints.md#predict--enumerate--hints-at-a-glance) |
420
428
  | **`renkin-bench`** | USPTO-50k/PaRoutes evaluation with `--plausibility` (forward-validated composite score), `--failure-taxonomy`, atom-balance checks (`target_MW > Σ precursor_MW`), and multi-stage `cascade` re-runs on unsolved targets — see [Benchmark](#benchmark) |
421
- | **Stock management** | `renkin stock stats\|validate\|coverage` for legacy CSV plus the `vendor_stock` library API for v0.38 CSV/TSV vendor records (SMILES, ID, vendor, price, lead time, availability), explicit exact/parent/stereo/tautomer match modes, and an InChIKey candidate index |
422
- | **MCP server** | `renkin-mcp` exposes 6 tools: `find_routes`, `validate_route`, `explain_route`, `find_pareto_routes`, `plan_with_constraints`, `estimate_diversity` |
429
+ | **Stock management** | `renkin stock stats\|validate\|coverage\|compile`; integrity-checked `.rstock` snapshots avoid reparsing large stocks, while the `vendor_stock` API supports source, price, lead-time, availability, and exact/parent/stereo/tautomer policy |
430
+ | **MCP server** | `renkin-mcp` exposes 7 tools over stdio (`find_routes`, `validate_route`, `explain_route`, `find_pareto_routes`, `plan_with_constraints`, `estimate_diversity`, `diagnose_failure`) and supports the legacy `2024-11-05` and modern `2026-07-28` protocol revisions; see the [MCP guide](docs/guides/mcp.md) |
423
431
  | **`renkin-doctor`** | Environment diagnostic binary — templates, building blocks, Python import, tool versions, data integrity |
424
432
  | **`renkin-kg`** | Reaction knowledge graph builder — bipartite mol↔reaction graphs from routes, GraphML/Cypher export |
425
433
  | **Multi-target** | `pip install renkin` (pre-built wheels, Linux/macOS/Windows) · `npm install renkin` (~500 KB WASM, near-native browser speed) |
@@ -469,6 +477,23 @@ USPTO-50k test set (4,907 molecules, full evaluation):
469
477
 
470
478
  > **Evaluation definition**: A molecule is *solved* if `find_routes` returns at least one route whose leaf precursors are all in the building block set, within depth=5 and beam=100. Ground-truth reactants from USPTO-50k are **not** checked — any commercially accessible route counts.
471
479
 
480
+ ### Formal v1.0.1 shared-stock comparison (4,903 paired targets)
481
+
482
+ | Arm | Primary route-to-shared-stock successes | Rate |
483
+ |---|---:|---:|
484
+ | RENKIN v1.0.1 | 591 / 4,903 | 12.05% |
485
+ | AiZynthFinder 4.4.1 | 200 / 4,903 | 4.08% |
486
+
487
+ The paired RENKIN-minus-AiZynthFinder difference is **+7.975 percentage
488
+ points**, with paired-bootstrap 95% CI **[+7.098, +8.852]**. The statistical
489
+ and formal publication gates both **PASS**. The complete v1.0.1 arm passed
490
+ target-set, manifest, schema, and route-hash integrity verification; all 591
491
+ reported routes had parseable normalized trees terminating in the configured
492
+ shared stock. The frozen v1.0.0 HOLD artifact remains preserved separately.
493
+ This is a shared-stock route endpoint, not experimental yield or universal CASP
494
+ superiority. [Protocol and status](docs/benchmark/formal-v1.0-competitor-comparison.md)
495
+ · [corrected report](data/comparison/formal_v1.0.1/formal_report.md)
496
+
472
497
  ### Corrected baseline (commit `e20dc8c`, 2026-07-22)
473
498
 
474
499
  | Public label | Internal metric | Value |
@@ -557,7 +582,7 @@ The JSON output includes `avg_nodes_expanded`, `avg_confidence`, `avg_convergenc
557
582
  }
558
583
  ```
559
584
 
560
- **Tools** (6):
585
+ **Tools** (7):
561
586
 
562
587
  | Tool | Description |
563
588
  |---|---|
@@ -567,13 +592,18 @@ The JSON output includes `avg_nodes_expanded`, `avg_confidence`, `avg_convergenc
567
592
  | `find_pareto_routes` | Pareto-front multi-objective route search |
568
593
  | `plan_with_constraints` | Constraint-DSL planning (element/building-block filters, step/cost limits, confidence thresholds, required/avoided/preferred reaction families) |
569
594
  | `estimate_diversity` | Route diversity and coverage metrics |
595
+ | `diagnose_failure` | Structured explanation of why a search found no route |
570
596
 
571
597
  `find_routes` also accepts `search_mode: "coverage"` with a required
572
598
  `coverage_templates` path. It runs the standard Stage 1 first and escalates
573
599
  to Stage 2 only when Stage 1 finds no route; the response reports the selected
574
600
  stage, timeout status, and per-stage elapsed time.
575
601
 
576
- The server auto-detects `data/building_blocks.smi` and `data/templates_extracted_5000.smi` in the working directory. Falls back to the embedded `DEFAULT_BUILDING_BLOCKS` / `default_rules()` defaults if not found (152 unique building blocks per `ChemEnv::bb_count()`, 22 handcrafted rules, including the new graph-based `carbamate_cleavage` rule).
602
+ The server auto-detects `data/building_blocks.smi` and the optional, locally
603
+ generated `data/templates_extracted_5000.smi` in the working directory. It
604
+ falls back to the embedded `DEFAULT_BUILDING_BLOCKS` / `default_rules()`
605
+ defaults if they are not found (152 unique building blocks per
606
+ `ChemEnv::bb_count()`, 23 handcrafted rules).
577
607
 
578
608
  ```bash
579
609
  cargo build --release
@@ -612,7 +642,7 @@ Target SMILES
612
642
  ┌─────────────────────────┐
613
643
  │ chem_env.rs │ ← chematic wrapper
614
644
  │ - SMILES parse │ canonical-SMILES FxHashSet BB lookup (O(1))
615
- │ - 21 built-in + up to 50k via --templates │ fragment sanitization + ring-leak filter
645
+ │ - 23 built-in + up to 50k via --templates │ fragment sanitization + ring-leak filter
616
646
  │ - Building block check │ apply_retro memoization cache
617
647
  └────────────┬────────────┘
618
648
  │ par_iter (rayon / sequential on WASM)
@@ -655,7 +685,8 @@ renkin/ ← Cargo workspace root
655
685
  │ ├── bin/benchmark.rs # renkin-bench binary (--plausibility flag)
656
686
  │ ├── bin/doctor.rs # renkin-doctor diagnostic binary
657
687
  │ ├── bin/fp.rs # renkin-fp ECFP4 fingerprint (nn-scoring feature)
658
- │ ├── bin/mcp.rs # renkin-mcp MCP server (6 tools)
688
+ │ ├── bin/mcp.rs # renkin-mcp stdio launcher
689
+ │ ├── mcp/ # Dual-era protocol + 7 tool handlers
659
690
  │ ├── chem_env.rs # retro rules + BB lookup + template loader
660
691
  │ ├── score.rs # SA Score heuristic + step cost
661
692
  │ ├── search.rs # A* / AND-OR tree engine + beam pruning
@@ -669,7 +700,7 @@ renkin/ ← Cargo workspace root
669
700
  │ └── renkin-kg/ # reaction knowledge graph builder (GraphML / Cypher export)
670
701
  ├── data/
671
702
  │ ├── building_blocks.smi # 402 curated commercial starting materials (loaded/deduplicated count)
672
- │ ├── templates_extracted_5000.smi # 5,000 auto-extracted SMIRKS templates
703
+ │ ├── templates_extracted_500.smi # 500 checked-in auto-extracted SMIRKS templates
673
704
  │ ├── benchmark_targets.smi # internal benchmark set
674
705
  │ └── bench_chunks/ # USPTO-50k per-chunk results
675
706
  ├── scripts/
@@ -706,7 +737,7 @@ see "Earlier milestones" below for older shipped work.
706
737
 
707
738
  ### In progress
708
739
 
709
- - [ ] Candidate-generation coverage gap — 33.0% (1,618/4,903) of the formal TEST corpus has zero positive candidates in-pool, a ceiling reranking cannot fix by construction; template-diversity-scaling confirmed as a strong mechanism (Phase A.5/B.2, see coverage mode above), higher-level-template research direction not yet started
740
+ - [ ] Candidate-generation coverage gap — 33.0% (1,618/4,903) of the formal TEST corpus has zero positive candidates in-pool, a ceiling reranking cannot fix by construction. Template-diversity scaling remains a strong mechanism; a first provenance-bounded radius-zero abstraction track is now implemented and disjoint-VAL gated ([#240](https://github.com/kent-tokyo/renkin/issues/240)), but its 4/22 targeted residual recovery is not a full-corpus remeasurement or a shipped default
710
741
  - [ ] Template retrieval index (element bitmask + bond-center prefilter) for the 50k template set
711
742
  - [ ] Calibrated route confidence (map `success_probability` to empirical solve rate)
712
743
 
@@ -752,12 +783,12 @@ see [Benchmark](#benchmark) for the corrected historical baseline.
752
783
  - [x] `renkin-doctor` — environment diagnostic binary (templates, BBs, Python, binaries)
753
784
  - [x] Failure diagnostics — zero-route output includes `likely_causes` + `suggestions` JSON block
754
785
  - [x] `--format explain|compare|compare-json` — human-readable and tabular route output
755
- - [x] `renkin stock stats|validate|coverage` — stock CSV management subcommand
786
+ - [x] `renkin stock stats|validate|coverage|compile` — stock inspection plus integrity-checked compiled `.rstock` snapshots
756
787
  - [x] Pareto multi-objective search — `--format pareto`, `--objectives`, `find_pareto_routes` MCP
757
788
  - [x] Constraint DSL — `--constraints JSON`, `plan_with_constraints` MCP tool
758
789
  - [x] `renkin template stats|validate|dedup|explain|coverage` — template quality tools
759
790
  - [x] `renkin-kg` — reaction knowledge graph (bipartite mol↔reaction, GraphML/Cypher export)
760
- - [x] MCP server (`renkin-mcp`) — expanded to 6 tools (`explain_route`, `find_pareto_routes`, `plan_with_constraints`, ...)
791
+ - [x] MCP server (`renkin-mcp`) — 7 tools with legacy `2024-11-05` and modern `2026-07-28` stdio protocol support
761
792
  - [x] Core search engine foundation — SMIRKS retro-reaction rules + fragment sanitization, A\*/AND-OR tree search with closed list + degenerate-route filter, SA Score heuristic + beam search, `rayon` parallel rule application (sequential fallback on WASM), FxHashMap/SmallVec beam frontier/SA-Score-memoization/`Arc<PathNode>` path-sharing perf work
762
793
  - [x] Multi-target packaging — Python bindings (PyO3 + maturin, `pip install renkin`), WASM build (`npm install renkin`), published to crates.io/PyPI/npm with GitHub Actions CI/CD, WASM browser playground + i18n (EN/JA/ZH)
763
794
  - [x] Benchmark CLI (`renkin-bench`) + USPTO-50k evaluation, `--format tree|mermaid` visualization, MkDocs documentation site + GitHub Pages playground
package/package.json CHANGED
@@ -5,7 +5,7 @@
5
5
  "kent-tokyo <kent-tokyo@users.noreply.github.com>"
6
6
  ],
7
7
  "description": "Ultra-fast retrosynthesis engine for computer-aided synthesis planning (CASP) — pure Rust, WASM-ready, Python bindings via PyO3",
8
- "version": "1.0.0",
8
+ "version": "1.0.2",
9
9
  "license": "MIT",
10
10
  "repository": {
11
11
  "type": "git",
package/renkin_bg.wasm CHANGED
Binary file