asciichem 0.28.2 → 0.29.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.github/workflows/ci.yml +9 -5
- data/.github/workflows/release.yml +1 -1
- data/CHANGELOG.md +29 -0
- data/benchmarks/README.md +34 -0
- data/benchmarks/parsanol_recheck.rb +71 -34
- data/lib/asciichem/cli.rb +93 -73
- data/lib/asciichem/engine/parsanol_engine.rb +118 -0
- data/lib/asciichem/engine/parslet_engine.rb +34 -0
- data/lib/asciichem/engine.rb +59 -0
- data/lib/asciichem/grammar.rb +2 -457
- data/lib/asciichem/grammar_rules.rb +476 -0
- data/lib/asciichem/parser.rb +8 -10
- data/lib/asciichem/transform.rb +1 -851
- data/lib/asciichem/transform_rules.rb +855 -0
- data/lib/asciichem/version.rb +1 -1
- data/lib/asciichem.rb +3 -0
- metadata +6 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 9579311feee35b70f9d15879379582cb67a8b6b9fe004e7c714e12448835e53c
|
|
4
|
+
data.tar.gz: 6f732f510f6d8923f2d1c855113ec34db33910a258135ee36a954035c90a252b
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 509996644be4e271c9432f8c5cc7baee5120f58b75af97f3f0bc6050b776dca6c567b1f69586bd87258237f12a56eeabccd4085e45b41772506ae099230d261c
|
|
7
|
+
data.tar.gz: e4b197451e71cc877398d6ce00130d1ac8dc0f2d73878b05dae4f849f145d99ac7de3e4e8e098db574eeb0c528f67030005a3538ed219f2178fb3ee7bc94c6cc
|
data/.github/workflows/ci.yml
CHANGED
|
@@ -13,13 +13,17 @@ jobs:
|
|
|
13
13
|
matrix:
|
|
14
14
|
ruby: ["3.3", "3.4"]
|
|
15
15
|
steps:
|
|
16
|
-
- uses: actions/checkout@
|
|
16
|
+
- uses: actions/checkout@v7
|
|
17
17
|
- name: Clone conformance corpus (asciichem-tests)
|
|
18
|
-
|
|
19
|
-
|
|
18
|
+
uses: actions/checkout@v7
|
|
19
|
+
with:
|
|
20
|
+
repository: asciichem/asciichem-tests
|
|
21
|
+
path: .asciichem-tests
|
|
22
|
+
- name: Point the suite at the corpus and record its version
|
|
20
23
|
run: |
|
|
21
|
-
|
|
22
|
-
|
|
24
|
+
echo "ASCIICHEM_CORPUS=$(pwd)/.asciichem-tests/corpus/fixtures" >> "$GITHUB_ENV"
|
|
25
|
+
git -C .asciichem-tests fetch --depth 1 --tags --quiet
|
|
26
|
+
echo "ASCIICHEM_CORPUS_VERSION=$(git -C .asciichem-tests describe --tags --abbrev=0 2>/dev/null || echo main)" >> "$GITHUB_ENV"
|
|
23
27
|
- uses: ruby/setup-ruby@v1
|
|
24
28
|
with:
|
|
25
29
|
ruby-version: ${{ matrix.ruby }}
|
data/CHANGELOG.md
CHANGED
|
@@ -3,6 +3,35 @@
|
|
|
3
3
|
All notable changes to AsciiChem are documented here.
|
|
4
4
|
This project follows [Semantic Versioning](https://semver.org/).
|
|
5
5
|
|
|
6
|
+
## [0.29.1] - 2026-09-16
|
|
7
|
+
|
|
8
|
+
### Added
|
|
9
|
+
- CLI `convert --engine parslet|parsanol` selects the parsing engine
|
|
10
|
+
per invocation (parslet stays the default); absent parsanol gem
|
|
11
|
+
exits 6 with install guidance. The ASCIICHEM_ENGINE env var keeps
|
|
12
|
+
working for programmatic selection.
|
|
13
|
+
|
|
14
|
+
## [0.29.0] - 2026-09-15
|
|
15
|
+
|
|
16
|
+
### Added
|
|
17
|
+
- Opt-in Parsanol parsing engine (TODO.impl 63/64; parsanol-ruby#25):
|
|
18
|
+
`AsciiChem::Engine.use(:parsanol)` runs the SAME grammar and
|
|
19
|
+
transform (extracted into backend-neutral GrammarRules /
|
|
20
|
+
TransformRules modules) over Parsanol's Rust-backed
|
|
21
|
+
parslet-compat layer - the full suite (1981 examples) passes
|
|
22
|
+
identically under either engine, at ~2.4x parse speed
|
|
23
|
+
(219 vs 90 i/s on the benchmark workload). Parsanol is a soft
|
|
24
|
+
dependency (gemspec unchanged); absent gem raises guidance.
|
|
25
|
+
ASCIICHEM_ENGINE=parsanol selects it for test runs.
|
|
26
|
+
|
|
27
|
+
### Changed
|
|
28
|
+
- Cascade legs are captured as one :segments repeat (the
|
|
29
|
+
electron-config pattern) so every engine arrays them; parsanol
|
|
30
|
+
merges - and overwrites - bare repeated sibling captures,
|
|
31
|
+
silently dropping legs (reported upstream). Transform
|
|
32
|
+
canonicaliser consumes the segments shape; spec'd for scalar and
|
|
33
|
+
array forms.
|
|
34
|
+
|
|
6
35
|
## [0.28.2] - 2026-09-14
|
|
7
36
|
|
|
8
37
|
### Fixed
|
data/benchmarks/README.md
CHANGED
|
@@ -85,3 +85,37 @@ a meaningful re-measure.** Corpus correctness is already there; the
|
|
|
85
85
|
native path is the whole point and remains unmeasurable until
|
|
86
86
|
serialization survives a multi-rule grammar.
|
|
87
87
|
|
|
88
|
+
### Re-check 3 (2026-09-15, parsanol 1.3.15)
|
|
89
|
+
|
|
90
|
+
The mode-routing/VM rework landed; native now engages for the full
|
|
91
|
+
grammar (`PARSANOL_MODE=native` in `benchmarks/parsanol_recheck.rb`,
|
|
92
|
+
fork-per-case gate so Rust aborts are reported, not fatal):
|
|
93
|
+
|
|
94
|
+
- **219/221 corpus cases green under native** — every accept case
|
|
95
|
+
except the two embedded-math inputs, and all 51 rejects clean
|
|
96
|
+
- **3.2x faster than parslet** on the workload (4.28 ms vs 13.64 ms
|
|
97
|
+
per 10-input pass, same session, ±3.0%)
|
|
98
|
+
- The two failures are the embedded-math grammar paths hitting
|
|
99
|
+
`serialize_dynamic` — the still-unfixed `@next_id` collision from
|
|
100
|
+
the re-check above (manifests as the Rust panic or a Ruby-side
|
|
101
|
+
`NoMethodError` on the native error path). Upstream thread:
|
|
102
|
+
parsanol-ruby#25 (third comment).
|
|
103
|
+
- Separately noted upstream: `H2` / `_2O` are accepted under native
|
|
104
|
+
but rejected under parslet (optimizer Str/Re run-merging semantics;
|
|
105
|
+
no corpus case covers these spellings today).
|
|
106
|
+
|
|
107
|
+
### Re-check 5 (2026-09-15, parsanol 1.3.17)
|
|
108
|
+
|
|
109
|
+
The ffi-gem cdylib tier (Rust engine on every runtime) changes
|
|
110
|
+
nothing for us: gate still **221/221**, 4.40 ms/i on the recheck
|
|
111
|
+
workload. The opt-in engine shipped in asciichem 0.29.0 is
|
|
112
|
+
unaffected; `@next_id` remains unfixed upstream but no longer fires
|
|
113
|
+
on corpus inputs. The re-check-3 verdict below is superseded — the
|
|
114
|
+
engine IS switchable and shipped (TODO.impl 64).
|
|
115
|
+
|
|
116
|
+
**Verdict: one upstream one-liner from adoption evaluation.** With
|
|
117
|
+
`@next_id += 1` fixed, the entire corpus passes under native at
|
|
118
|
+
3x parslet speed — at that point the decision is whether to make the
|
|
119
|
+
engine switchable (opt-in, soft dependency) in the gem.
|
|
120
|
+
|
|
121
|
+
|
|
@@ -1,26 +1,35 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
|
-
# Parsanol
|
|
4
|
-
#
|
|
5
|
-
#
|
|
6
|
-
# on the
|
|
3
|
+
# Parsanol gate + benchmark against the SHIPPED opt-in engine
|
|
4
|
+
# (asciichem 0.29.0+): AsciiChem::Engine.use(:parsanol) runs the same
|
|
5
|
+
# GrammarRules/TransformRules over Parsanol's Rust-backed compat
|
|
6
|
+
# layer. (1) gates on the shared corpus, forked per case so a Rust
|
|
7
|
+
# panic aborts the child — reported, not fatal; (2) gates on the
|
|
8
|
+
# issue-25 EOF repro; (3) measures against the parslet path
|
|
9
|
+
# (benchmarks/engines.rb).
|
|
7
10
|
#
|
|
8
|
-
#
|
|
9
|
-
#
|
|
10
|
-
# increments @next_id so the second callback panics the Rust core —
|
|
11
|
-
# parsanol-ruby#25). The measurement therefore forces :ruby, the
|
|
12
|
-
# only working mode for full parslet grammars via the shim.
|
|
11
|
+
# Requires the parsanol gem (add to the Gemfile, or point -I at a
|
|
12
|
+
# local checkout):
|
|
13
13
|
#
|
|
14
|
-
#
|
|
15
|
-
#
|
|
14
|
+
# bundle exec ruby -I ../parsanol/parsanol-ruby/lib benchmarks/parsanol_recheck.rb
|
|
15
|
+
#
|
|
16
|
+
# PARSANOL_MODE=ruby forces Parsanol's pure-Ruby backend (no Rust
|
|
17
|
+
# core) for comparison.
|
|
16
18
|
require "benchmark/ips"
|
|
17
19
|
require "asciichem"
|
|
20
|
+
require "asciichem/engine/parsanol_engine"
|
|
18
21
|
require "json"
|
|
19
22
|
|
|
20
|
-
|
|
23
|
+
AsciiChem::Engine.use(:parsanol)
|
|
24
|
+
|
|
25
|
+
engine = AsciiChem::Engine.current
|
|
26
|
+
puts "parsanol #{Parsanol::VERSION} | engine: #{engine}"
|
|
27
|
+
puts "grammar superclass: #{engine.grammar.superclass}"
|
|
28
|
+
puts "mode: #{ENV.fetch("PARSANOL_MODE", "native")}"
|
|
21
29
|
|
|
22
|
-
|
|
23
|
-
|
|
30
|
+
if ENV.fetch("PARSANOL_MODE", "native") == "ruby"
|
|
31
|
+
Parsanol::Native.singleton_class.define_method(:available?) { false }
|
|
32
|
+
end
|
|
24
33
|
|
|
25
34
|
# -- 1. Issue-25 repro: repeat-of-maybe at end of input --------------
|
|
26
35
|
begin
|
|
@@ -31,38 +40,66 @@ rescue AsciiChem::ParseError => e
|
|
|
31
40
|
end
|
|
32
41
|
|
|
33
42
|
# -- 2. Shared-corpus gate --------------------------------------------
|
|
43
|
+
# Each case runs in a forked child: a Rust panic aborts the child
|
|
44
|
+
# process (unrescuable in Ruby), and the parent reports it by name
|
|
45
|
+
# instead of dying.
|
|
34
46
|
corpus_dir = File.expand_path("../../asciichem-tests/corpus/fixtures", __dir__)
|
|
35
47
|
cases = Dir[File.join(corpus_dir, "*.json")].sort.flat_map { |p| JSON.parse(File.read(p)) }
|
|
36
48
|
parser_cases = cases.select { |c| c.key?("input") && !c.key?("lint") && !c.key?("convention") }
|
|
37
49
|
|
|
38
|
-
|
|
50
|
+
def run_in_child
|
|
51
|
+
reader, writer = IO.pipe
|
|
52
|
+
pid = fork do
|
|
53
|
+
reader.close
|
|
54
|
+
Marshal.dump(yield, writer)
|
|
55
|
+
rescue StandardError => e
|
|
56
|
+
Marshal.dump({ exception: e.class.name, message: e.message }, writer)
|
|
57
|
+
ensure
|
|
58
|
+
writer.close
|
|
59
|
+
end
|
|
60
|
+
writer.close
|
|
61
|
+
payload = Marshal.load(reader)
|
|
62
|
+
reader.close
|
|
63
|
+
_, status = Process.waitpid2(pid)
|
|
64
|
+
[payload, status]
|
|
65
|
+
end
|
|
66
|
+
|
|
67
|
+
pass = fail_parse = fail_reject = fail_roundtrip = fatal = 0
|
|
39
68
|
parser_cases.each do |fixture|
|
|
40
69
|
input = fixture.fetch("input")
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
70
|
+
payload, status = run_in_child do
|
|
71
|
+
formula = AsciiChem.parse(input)
|
|
72
|
+
{ text: (formula.to_text if fixture["roundTrip"]) }
|
|
73
|
+
end
|
|
74
|
+
if status.signaled? || !status.success?
|
|
75
|
+
fatal += 1
|
|
76
|
+
puts " FATAL (child #{status.exitstatus ? "exit #{status.exitstatus}" : "aborted"}): #{fixture["id"]} #{input.inspect}" if fatal <= 8
|
|
77
|
+
next
|
|
78
|
+
end
|
|
79
|
+
if payload.key?(:exception)
|
|
80
|
+
if fixture.fetch("parses")
|
|
50
81
|
fail_parse += 1
|
|
51
|
-
puts " PARSE FAIL: #{input.inspect} -> #{
|
|
52
|
-
|
|
53
|
-
else
|
|
54
|
-
begin
|
|
55
|
-
AsciiChem.parse(input)
|
|
56
|
-
fail_reject += 1
|
|
57
|
-
puts " SHOULD REJECT: #{input.inspect}" if fail_reject <= 8
|
|
58
|
-
rescue AsciiChem::ParseError, Parslet::ParseFailed
|
|
82
|
+
puts " PARSE FAIL: #{input.inspect} -> #{payload[:message][0, 90]}" if fail_parse <= 8
|
|
83
|
+
else
|
|
59
84
|
pass += 1
|
|
60
85
|
end
|
|
86
|
+
next
|
|
87
|
+
end
|
|
88
|
+
unless fixture.fetch("parses")
|
|
89
|
+
fail_reject += 1
|
|
90
|
+
puts " SHOULD REJECT: #{input.inspect}" if fail_reject <= 8
|
|
91
|
+
next
|
|
92
|
+
end
|
|
93
|
+
if fixture["roundTrip"] && payload[:text] != input
|
|
94
|
+
fail_roundtrip += 1
|
|
95
|
+
puts " ROUNDTRIP DIFF: #{input.inspect} -> #{payload[:text].inspect}" if fail_roundtrip <= 5
|
|
96
|
+
next
|
|
61
97
|
end
|
|
98
|
+
pass += 1
|
|
62
99
|
end
|
|
63
100
|
total = parser_cases.length
|
|
64
|
-
puts format("corpus gate: %d/%d ok (parse-fails %d, should-reject %d, roundtrip-diffs %d)",
|
|
65
|
-
pass, total, fail_parse, fail_reject, fail_roundtrip)
|
|
101
|
+
puts format("corpus gate: %d/%d ok (parse-fails %d, should-reject %d, roundtrip-diffs %d, fatal %d)",
|
|
102
|
+
pass, total, fail_parse, fail_reject, fail_roundtrip, fatal)
|
|
66
103
|
|
|
67
104
|
# -- 3. Performance ----------------------------------------------------
|
|
68
105
|
WORKLOAD = [
|
data/lib/asciichem/cli.rb
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
|
-
require
|
|
3
|
+
require 'thor'
|
|
4
4
|
|
|
5
5
|
module AsciiChem
|
|
6
6
|
# Thor-based command line interface. Invoked via the `asciichem`
|
|
@@ -8,25 +8,30 @@ module AsciiChem
|
|
|
8
8
|
class Cli < Thor
|
|
9
9
|
# Use lowercase 'asciichem' as the program name in help output
|
|
10
10
|
# and command banners, matching the executable name.
|
|
11
|
-
package_name
|
|
11
|
+
package_name 'asciichem'
|
|
12
12
|
|
|
13
|
-
desc
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
13
|
+
desc 'convert -i INPUT -t FORMAT',
|
|
14
|
+
'Convert INPUT to FORMAT (mathml|text|html|latex|svg|structural-svg|model-json|cml|smiles|molfile)'
|
|
15
|
+
method_option :input, aliases: '-i', type: :string,
|
|
16
|
+
desc: "Source text (or '-' for stdin)"
|
|
17
|
+
method_option :file, aliases: '-f', type: :string,
|
|
18
|
+
desc: 'Read source from a file'
|
|
19
|
+
method_option :from, type: :string, default: 'asciichem',
|
|
20
|
+
desc: 'Input grammar: asciichem|smiles|molfile'
|
|
21
|
+
method_option :format, aliases: '-t', type: :string, default: 'mathml',
|
|
22
|
+
desc: 'Output format'
|
|
23
|
+
method_option :engine, type: :string, default: 'parslet',
|
|
24
|
+
desc: 'Parsing engine: parslet (default) | parsanol (opt-in, needs the parsanol gem)'
|
|
22
25
|
def convert
|
|
23
|
-
unless options[
|
|
24
|
-
raise AsciiChem::ParseError, "provide -i INPUT or -f FILE"
|
|
25
|
-
end
|
|
26
|
+
raise AsciiChem::ParseError, 'provide -i INPUT or -f FILE' unless options['input'] || options['file']
|
|
26
27
|
|
|
28
|
+
select_engine(options[:engine])
|
|
27
29
|
source = read_source
|
|
28
30
|
formula = ingest(source, options[:from])
|
|
29
31
|
puts render(formula, options[:format])
|
|
32
|
+
rescue AsciiChem::Engine::Error => e
|
|
33
|
+
warn "Engine error: #{e.message}"
|
|
34
|
+
exit 6
|
|
30
35
|
rescue AsciiChem::ParseError => e
|
|
31
36
|
warn "Parse error: #{e.message}"
|
|
32
37
|
exit 1
|
|
@@ -35,9 +40,9 @@ module AsciiChem
|
|
|
35
40
|
exit 2
|
|
36
41
|
end
|
|
37
42
|
|
|
38
|
-
desc
|
|
39
|
-
method_option :input, aliases:
|
|
40
|
-
|
|
43
|
+
desc 'parse-cml -i INPUT', 'Parse CML XML and emit AsciiChem text'
|
|
44
|
+
method_option :input, aliases: '-i', type: :string, required: true,
|
|
45
|
+
desc: 'CML XML source'
|
|
41
46
|
def parse_cml
|
|
42
47
|
formula = AsciiChem::Cml.parse(options[:input])
|
|
43
48
|
puts formula.to_text
|
|
@@ -46,8 +51,8 @@ module AsciiChem
|
|
|
46
51
|
exit 1
|
|
47
52
|
end
|
|
48
53
|
|
|
49
|
-
desc
|
|
50
|
-
method_option :input, aliases:
|
|
54
|
+
desc 'roundtrip -i INPUT', 'Parse and re-emit; exit non-zero if not equal'
|
|
55
|
+
method_option :input, aliases: '-i', type: :string, required: true
|
|
51
56
|
def roundtrip
|
|
52
57
|
original = options[:input]
|
|
53
58
|
rendered = AsciiChem.parse(original).to_text
|
|
@@ -60,11 +65,11 @@ module AsciiChem
|
|
|
60
65
|
end
|
|
61
66
|
end
|
|
62
67
|
|
|
63
|
-
desc
|
|
64
|
-
method_option :input, aliases:
|
|
65
|
-
|
|
66
|
-
method_option :format, aliases:
|
|
67
|
-
|
|
68
|
+
desc 'lint -i INPUT', 'Run chemistry checks; exit 1 on error, 0 if clean'
|
|
69
|
+
method_option :input, aliases: '-i', type: :string, required: true,
|
|
70
|
+
desc: 'AsciiChem source text'
|
|
71
|
+
method_option :format, aliases: '-f', type: :string, default: 'text',
|
|
72
|
+
desc: 'Output format: text or json'
|
|
68
73
|
def lint
|
|
69
74
|
formula = AsciiChem.parse(options[:input])
|
|
70
75
|
diagnostics = AsciiChem::Linter.run(formula)
|
|
@@ -76,7 +81,7 @@ module AsciiChem
|
|
|
76
81
|
end
|
|
77
82
|
|
|
78
83
|
map %w[--version -v] => :version
|
|
79
|
-
desc
|
|
84
|
+
desc 'version', 'Print the asciichem version'
|
|
80
85
|
def version
|
|
81
86
|
puts "asciichem #{AsciiChem::VERSION}"
|
|
82
87
|
end
|
|
@@ -86,32 +91,32 @@ module AsciiChem
|
|
|
86
91
|
"asciichem #{command.usage}"
|
|
87
92
|
end
|
|
88
93
|
|
|
89
|
-
desc
|
|
90
|
-
method_option :cas, type: :string, desc:
|
|
91
|
-
method_option :name, type: :string, desc:
|
|
92
|
-
method_option :cid, type: :string, desc:
|
|
93
|
-
method_option :inchikey, type: :string, desc:
|
|
94
|
-
method_option :smiles, type: :string, desc:
|
|
95
|
-
method_option :source, type: :string, default:
|
|
96
|
-
method_option :refresh, type: :boolean, default: false, desc:
|
|
97
|
-
method_option :format, aliases:
|
|
98
|
-
|
|
94
|
+
desc 'resolve --cas X | --name X | ...', 'Resolve a substance from a source (network; cached)'
|
|
95
|
+
method_option :cas, type: :string, desc: 'CAS registry number'
|
|
96
|
+
method_option :name, type: :string, desc: 'Substance name'
|
|
97
|
+
method_option :cid, type: :string, desc: 'PubChem CID'
|
|
98
|
+
method_option :inchikey, type: :string, desc: 'InChIKey'
|
|
99
|
+
method_option :smiles, type: :string, desc: 'SMILES'
|
|
100
|
+
method_option :source, type: :string, default: 'pubchem', desc: 'Resolver source'
|
|
101
|
+
method_option :refresh, type: :boolean, default: false, desc: 'Bypass the cache'
|
|
102
|
+
method_option :format, aliases: '-t', type: :string, default: 'model-json',
|
|
103
|
+
desc: 'Output: model-json | text | smiles'
|
|
99
104
|
def resolve
|
|
100
105
|
convention, value = %i[cas name cid inchikey smiles]
|
|
101
106
|
.filter_map { |k| [k, options[k.to_s]] if options[k.to_s] }
|
|
102
107
|
.first
|
|
103
|
-
raise AsciiChem::Error,
|
|
108
|
+
raise AsciiChem::Error, 'give one of --cas/--name/--cid/--inchikey/--smiles' unless value
|
|
104
109
|
|
|
105
|
-
convention = { cas:
|
|
106
|
-
inchikey:
|
|
110
|
+
convention = { cas: 'cas', name: 'name', cid: 'pubchem-cid',
|
|
111
|
+
inchikey: 'inchikey', smiles: 'smiles' }.fetch(convention)
|
|
107
112
|
substance = AsciiChem::Resolver[options[:source]].new.resolve(
|
|
108
113
|
value: value, convention: convention, refresh: options[:refresh]
|
|
109
114
|
)
|
|
110
115
|
raise AsciiChem::Error, "#{options[:source]} does not know #{value.inspect}" unless substance
|
|
111
116
|
|
|
112
117
|
puts case options[:format].to_s
|
|
113
|
-
when
|
|
114
|
-
when
|
|
118
|
+
when 'text' then substance.preferred_name.to_s
|
|
119
|
+
when 'smiles' then substance.identifier_value('canonical-smiles').to_s
|
|
115
120
|
else substance.to_model_json
|
|
116
121
|
end
|
|
117
122
|
rescue AsciiChem::Error => e
|
|
@@ -119,22 +124,22 @@ module AsciiChem
|
|
|
119
124
|
exit 3
|
|
120
125
|
end
|
|
121
126
|
|
|
122
|
-
desc
|
|
123
|
-
method_option :cas, type: :string, desc:
|
|
124
|
-
method_option :name, type: :string, desc:
|
|
125
|
-
method_option :cid, type: :string, desc:
|
|
126
|
-
method_option :inchikey, type: :string, desc:
|
|
127
|
-
method_option :smiles, type: :string, desc:
|
|
128
|
-
method_option :source, type: :string, default:
|
|
129
|
-
method_option :refresh, type: :boolean, default: false, desc:
|
|
127
|
+
desc 'cite --cas X | --name X | ...', 'Resolve a substance and emit a dataset-type Relaton bibitem (XML)'
|
|
128
|
+
method_option :cas, type: :string, desc: 'CAS registry number'
|
|
129
|
+
method_option :name, type: :string, desc: 'Substance name'
|
|
130
|
+
method_option :cid, type: :string, desc: 'PubChem CID'
|
|
131
|
+
method_option :inchikey, type: :string, desc: 'InChIKey'
|
|
132
|
+
method_option :smiles, type: :string, desc: 'SMILES'
|
|
133
|
+
method_option :source, type: :string, default: 'pubchem', desc: 'Resolver source'
|
|
134
|
+
method_option :refresh, type: :boolean, default: false, desc: 'Bypass the cache'
|
|
130
135
|
def cite
|
|
131
136
|
convention, value = %i[cas name cid inchikey smiles]
|
|
132
137
|
.filter_map { |k| [k, options[k.to_s]] if options[k.to_s] }
|
|
133
138
|
.first
|
|
134
|
-
raise AsciiChem::Error,
|
|
139
|
+
raise AsciiChem::Error, 'give one of --cas/--name/--cid/--inchikey/--smiles' unless value
|
|
135
140
|
|
|
136
|
-
convention = { cas:
|
|
137
|
-
inchikey:
|
|
141
|
+
convention = { cas: 'cas', name: 'name', cid: 'pubchem-cid',
|
|
142
|
+
inchikey: 'inchikey', smiles: 'smiles' }.fetch(convention)
|
|
138
143
|
substance = AsciiChem::Resolver[options[:source]].new.resolve(
|
|
139
144
|
value: value, convention: convention, refresh: options[:refresh]
|
|
140
145
|
)
|
|
@@ -146,37 +151,43 @@ module AsciiChem
|
|
|
146
151
|
exit 4
|
|
147
152
|
end
|
|
148
153
|
|
|
149
|
-
desc
|
|
150
|
-
method_option :input, aliases:
|
|
154
|
+
desc 'validate -i INPUT', 'Offline identifier validation'
|
|
155
|
+
method_option :input, aliases: '-i', type: :string, required: true
|
|
151
156
|
def validate
|
|
152
157
|
formula = AsciiChem.parse(options[:input])
|
|
153
158
|
annotations = formula.nodes.grep(AsciiChem::Model::Molecule).flat_map(&:identifiers)
|
|
154
159
|
if annotations.empty?
|
|
155
|
-
puts
|
|
160
|
+
puts 'no identifier annotations found'
|
|
156
161
|
return
|
|
157
162
|
end
|
|
158
163
|
annotations.each do |identifier|
|
|
159
164
|
known = AsciiChem::Identifiers.known?(identifier.convention)
|
|
160
165
|
valid = known && AsciiChem::Identifiers.valid?(identifier.convention, identifier.value)
|
|
161
|
-
status = known
|
|
162
|
-
|
|
166
|
+
status = if known
|
|
167
|
+
valid ? 'ok' : 'INVALID'
|
|
168
|
+
else
|
|
169
|
+
'unknown convention'
|
|
170
|
+
end
|
|
171
|
+
puts format('%-12s %-40s %s', identifier.convention, identifier.value, status)
|
|
172
|
+
end
|
|
173
|
+
exit 1 if annotations.any? do |i|
|
|
174
|
+
AsciiChem::Identifiers.known?(i.convention) &&
|
|
175
|
+
!AsciiChem::Identifiers.valid?(i.convention, i.value)
|
|
163
176
|
end
|
|
164
|
-
exit 1 if annotations.any? { |i| AsciiChem::Identifiers.known?(i.convention) &&
|
|
165
|
-
!AsciiChem::Identifiers.valid?(i.convention, i.value) }
|
|
166
177
|
rescue AsciiChem::ParseError => e
|
|
167
178
|
warn "Parse error: #{e.message}"
|
|
168
179
|
exit 1
|
|
169
180
|
end
|
|
170
181
|
|
|
171
|
-
desc
|
|
172
|
-
method_option :input, aliases:
|
|
173
|
-
|
|
174
|
-
method_option :file, aliases:
|
|
175
|
-
|
|
176
|
-
method_option :from, type: :string, default:
|
|
177
|
-
|
|
182
|
+
desc 'identity -i INPUT', 'Derive InChI/InChIKey locally from the structure (offline; requires an InChI engine)'
|
|
183
|
+
method_option :input, aliases: '-i', type: :string,
|
|
184
|
+
desc: "Source text (or '-' for stdin)"
|
|
185
|
+
method_option :file, aliases: '-f', type: :string,
|
|
186
|
+
desc: 'Read source from a file'
|
|
187
|
+
method_option :from, type: :string, default: 'asciichem',
|
|
188
|
+
desc: 'Input grammar: asciichem|smiles|molfile'
|
|
178
189
|
method_option :engine_bin, type: :string,
|
|
179
|
-
desc:
|
|
190
|
+
desc: 'Path to the inchi-1 binary (overrides the configured engine)'
|
|
180
191
|
def identity
|
|
181
192
|
formula = ingest(read_source, options[:from])
|
|
182
193
|
molecules = molecules_in(formula)
|
|
@@ -195,9 +206,18 @@ module AsciiChem
|
|
|
195
206
|
|
|
196
207
|
private
|
|
197
208
|
|
|
209
|
+
# --engine feeds the Engine selector (0.29.0+); :parsanol is an
|
|
210
|
+
# opt-in soft dependency, :parslet stays the default.
|
|
211
|
+
def select_engine(name)
|
|
212
|
+
return if name.to_s == 'parslet' && AsciiChem::Engine.current == AsciiChem::Engine::ParsletEngine
|
|
213
|
+
|
|
214
|
+
require 'asciichem/engine'
|
|
215
|
+
AsciiChem::Engine.use(name.to_s)
|
|
216
|
+
end
|
|
217
|
+
|
|
198
218
|
def read_source
|
|
199
219
|
return File.read(options[:file]) if options[:file]
|
|
200
|
-
return $stdin.read if options[:input] ==
|
|
220
|
+
return $stdin.read if options[:input] == '-'
|
|
201
221
|
|
|
202
222
|
options[:input]
|
|
203
223
|
end
|
|
@@ -207,9 +227,9 @@ module AsciiChem
|
|
|
207
227
|
# format works regardless of the input language.
|
|
208
228
|
def ingest(source, from)
|
|
209
229
|
case from.to_s
|
|
210
|
-
when
|
|
211
|
-
when
|
|
212
|
-
when
|
|
230
|
+
when 'asciichem' then AsciiChem.parse(source)
|
|
231
|
+
when 'smiles' then AsciiChem.parse_smiles(source)
|
|
232
|
+
when 'molfile' then molfile_formula(source)
|
|
213
233
|
else
|
|
214
234
|
raise AsciiChem::ParseError, "unknown --from grammar: #{from}"
|
|
215
235
|
end
|
|
@@ -224,7 +244,7 @@ module AsciiChem
|
|
|
224
244
|
|
|
225
245
|
def render(formula, format)
|
|
226
246
|
return formula.to_cml if format.to_sym == :cml
|
|
227
|
-
return formula.to_model_json if format.to_sym == :
|
|
247
|
+
return formula.to_model_json if format.to_sym == :'model-json'
|
|
228
248
|
return formula.to_smiles if format.to_sym == :smiles
|
|
229
249
|
return formula.nodes.first.to_molfile if format.to_sym == :molfile
|
|
230
250
|
|
|
@@ -246,14 +266,14 @@ module AsciiChem
|
|
|
246
266
|
|
|
247
267
|
def output_lint(diagnostics, format)
|
|
248
268
|
case format.to_s
|
|
249
|
-
when
|
|
269
|
+
when 'json' then output_lint_json(diagnostics)
|
|
250
270
|
else
|
|
251
271
|
diagnostics.each { |d| puts d }
|
|
252
272
|
end
|
|
253
273
|
end
|
|
254
274
|
|
|
255
275
|
def output_lint_json(diagnostics)
|
|
256
|
-
require
|
|
276
|
+
require 'json'
|
|
257
277
|
payload = diagnostics.map do |d|
|
|
258
278
|
{
|
|
259
279
|
severity: d.severity.to_s,
|