asciichem 0.29.0 → 0.29.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 106d03bf9d192f66b4f9ade534ed84f8dcc5cbf0ceeb07d45b633c353fc92aaa
4
- data.tar.gz: 36c047aa3f26f73f94cc698eb905a939025feb9fc34fb86b67163cf3a93f7fbe
3
+ metadata.gz: f372a8539906c658a399b015255e574de5feb73d1b5e7183958901681bfbce34
4
+ data.tar.gz: 7269dc4ad6e39e7c5ea6726b54a8ffe1508c039bdf9ced5c0a74e23efdc9b00f
5
5
  SHA512:
6
- metadata.gz: b9215fa19d5dc3a65d82cfbe5cdfd4f837bd136cbb18cb0241654e7307a314c42088a870a76c3393858bd285e8e5834a95de136849c6f4d1ee4d2469fe616731
7
- data.tar.gz: f7654b2d8257c634e7be6b514011b5875c8060886204c43dd9293b982c8cd241773fb92958d8772d0e28ec07ce50ad82753aba085469e0c3edbc03446201c3bc
6
+ metadata.gz: 6f936bc0e95907d589c0095d34e4e65c554dcbdccd276aab7186db8f1218cbb179652b082f088f1b70a76391e830db53de2259ea8d333b64dd85ba367b16c610
7
+ data.tar.gz: 63d979830e4a8ab9aca450006a9740e83428e26073323af0ed4f8ddeed4be6f7aecb8bcfe1e5173f7a6049f92ff8b0a8e0a2b374fbd13630ae00b0e9e7a0db60
@@ -34,6 +34,15 @@ jobs:
34
34
  with:
35
35
  ruby-version: "3.4"
36
36
  bundler-cache: true
37
+ # CI never pushes to git (read-only contents). `rake release`
38
+ # attempts `git push origin main` after publishing unless the
39
+ # version tag already exists locally — bundler then prints
40
+ # "Tag vX has already been created" and skips its git stage
41
+ # entirely (this is how 0.29.0/0.29.1 released). Pre-create the
42
+ # tag so the gem push is the only remote operation; tags on the
43
+ # remote remain the maintainer's.
44
+ - name: Pre-create the release tag (skips rake's git stage)
45
+ run: git tag "v${{ inputs.version }}"
37
46
  # Builds and pushes using the GitHub OIDC identity — no API keys.
38
47
  - uses: rubygems/release-gem@v1
39
48
  - name: Summary
data/CHANGELOG.md CHANGED
@@ -3,6 +3,26 @@
3
3
  All notable changes to AsciiChem are documented here.
4
4
  This project follows [Semantic Versioning](https://semver.org/).
5
5
 
6
+ ## [0.29.2] - 2026-09-17
7
+
8
+ ### Changed
9
+ - `relaton-bib` constraint widened to `>= 0.1, < 3`: asciichem now
10
+ co-resolves with current metanorma gems (metanorma-standoc and
11
+ friends require relaton-bib 2). The citation track speaks both
12
+ major lines through a single `Citation::RelatonApi` seam —
13
+ relaton-bib 1 (`RelatonBib`) and relaton-bib 2
14
+ (`Relaton::Bib` typed models) both build and serialize the
15
+ dataset-type bibitems; profile data is version-independent.
16
+ - nil resolver links no longer emit an empty `<uri>` element.
17
+
18
+ ## [0.29.1] - 2026-09-16
19
+
20
+ ### Added
21
+ - CLI `convert --engine parslet|parsanol` selects the parsing engine
22
+ per invocation (parslet stays the default); absent parsanol gem
23
+ exits 6 with install guidance. The ASCIICHEM_ENGINE env var keeps
24
+ working for programmatic selection.
25
+
6
26
  ## [0.29.0] - 2026-09-15
7
27
 
8
28
  ### Added
data/asciichem.gemspec CHANGED
@@ -39,7 +39,7 @@ Gem::Specification.new do |spec|
39
39
  spec.add_dependency "mml", "~> 2.3"
40
40
  spec.add_dependency "nokogiri", "~> 1.16"
41
41
  spec.add_dependency "parslet", "~> 2.0"
42
- spec.add_dependency "relaton-bib", ">= 0.1", "< 2"
42
+ spec.add_dependency "relaton-bib", ">= 0.1", "< 3"
43
43
  spec.add_dependency "plurimath", "~> 0.8"
44
44
  spec.add_dependency "thor", "~> 1.3"
45
45
 
data/benchmarks/README.md CHANGED
@@ -104,9 +104,67 @@ fork-per-case gate so Rust aborts are reported, not fatal):
104
104
  but rejected under parslet (optimizer Str/Re run-merging semantics;
105
105
  no corpus case covers these spellings today).
106
106
 
107
- **Verdict: one upstream one-liner from adoption evaluation.** With
108
- `@next_id += 1` fixed, the entire corpus passes under native at
109
- 3x parslet speed — at that point the decision is whether to make the
110
- engine switchable (opt-in, soft dependency) in the gem.
111
-
107
+ ### Re-check 5 (2026-09-15, parsanol 1.3.17)
108
+
109
+ The ffi-gem cdylib tier (Rust engine on every runtime) changes
110
+ nothing for us: gate still **221/221**, 4.40 ms/i on the recheck
111
+ workload. The opt-in engine shipped in asciichem 0.29.0 is
112
+ unaffected; `@next_id` remains unfixed upstream but no longer fires
113
+ on corpus inputs. The re-check-3 verdict below is superseded — the
114
+ engine IS switchable and shipped (TODO.impl 64).
115
+
116
+ **Verdict: superseded — the engine IS switchable and shipped**
117
+ (asciichem 0.29.0, TODO.impl 64).
118
+
119
+ ### Re-check 6 (2026-09-16, parsanol 1.3.18)
120
+
121
+ Gate still **221/221** through the shipped engine, but the "one
122
+ decode path" rework **regressed compat-layer throughput ~60%** for
123
+ this grammar: 7.2 ms/batch (138 i/s) vs 4.3-4.6 ms on 1.3.15/16,
124
+ with the parslet control stable across sessions (11-14 ms
125
+ throughout). The `H2`/`_2O` acceptance divergence also persists.
126
+ Reported upstream (parsanol-ruby#25, fourth comment). We stay on
127
+ the shipped engine; users pinning parsanol for speed should prefer
128
+ 1.3.16/1.3.17 until the regression is addressed.
129
+
130
+ ### Re-check 7 (2026-09-16, parsanol 1.3.20)
131
+
132
+ Three upstream issues closed since 1.3.18:
133
+
134
+ - **#38** (`Dynamic.register` `@next_id` collision panicking the
135
+ Rust core) — fixed; no panic during full-corpus run.
136
+ - **#37** ("one decode path" throughput regression) — fixed as a
137
+ side effect of the optimizer acceptance fix in #39; throughput on
138
+ this grammar is back to and ahead of 1.3.15/16 levels.
139
+ - **#39** (optimizer Str/Re run-merging changed sequence-boundary
140
+ acceptance) — root-caused to Re-run regex-source concatenation
141
+ (proven unsafe: `"a|"+"b"` → `"a|b"` accepts `"a"`); Re runs now
142
+ stay unmerged, Str-run merging stays. Spec-level decision
143
+ recorded: the optimizer may never alter acceptance.
144
+
145
+ Validation against 1.3.20:
146
+
147
+ - Gate **221/221** through the shipped `ParsanolEngine` (fork-per-case,
148
+ no Rust aborts).
149
+ - Head-to-head vs parslet, same Ruby process (3 runs, ±3-15%):
150
+ parsanol **2.6x faster** (4.65–5.19 ms/batch vs 12.18–13.65 ms for
151
+ parslet). Up from the 1.7x under 1.3.18 — the #37 regression is
152
+ gone.
153
+ - Direct `H2` / `_2O` / `Ca2+` / `H22` / `O2` probe across both
154
+ parslet and parsanol (native and ruby backends) shows **identical
155
+ parse outcomes**. The earlier "divergence" framing in re-checks
156
+ 3-6 was a misreading: AsciiChem's `hydrogen_atom` grammar rule
157
+ intentionally permits bare-digit subscripts after `H` ("lets users
158
+ write `H2O` instead of `H_2O`" — grammar_rules.rb:228-231) and
159
+ `isotope_marker` accepts both `^digits` and `_digits`, so `_2O`
160
+ parses as the isotope of `O` and round-trips as `^2O`. The
161
+ parsanol optimizer bug in #39 was real and is fixed, but the
162
+ AsciiChem repro was a misleading example — both engines agree on
163
+ these inputs because they share the same grammar rules.
164
+
165
+ **Verdict: shipped engine fully validated.** 2.6x speedup, 100%
166
+ corpus gate, all four reported upstream issues now resolved or
167
+ non-blocking (#36 bare repeated sibling captures remains open but
168
+ is worked around in `ParsanolEngine` via single `.as(...)` capture
169
+ wrapping).
112
170
 
@@ -1,25 +1,30 @@
1
1
  # frozen_string_literal: true
2
2
 
3
- # Parsanol re-check (parsanol-ruby 1.3.15, issue-25 thread): runs the
4
- # UNMODIFIED AsciiChem grammar over the Parsanol engine via the
5
- # Parslet compat shim, then (1) gates on the shared corpus, (2) gates
6
- # on the issue-25 EOF repro, (3) measures against the parslet path.
3
+ # Parsanol gate + benchmark against the SHIPPED opt-in engine
4
+ # (asciichem 0.29.0+): AsciiChem::Engine.use(:parsanol) runs the same
5
+ # GrammarRules/TransformRules over Parsanol's Rust-backed compat
6
+ # layer. (1) gates on the shared corpus, forked per case so a Rust
7
+ # panic aborts the child — reported, not fatal; (2) gates on the
8
+ # issue-25 EOF repro; (3) measures against the parslet path
9
+ # (benchmarks/engines.rb).
7
10
  #
8
- # Mode: native by default (PARSOLAN_MODE=native — default routing);
9
- # PARSANOL_MODE=ruby forces the pure-Ruby engine. As of 1.3.15 the
10
- # native backend parses 169/170 corpus accepts + all 51 rejects at
11
- # ~2.3x parslet speed; the single fatal case is embedded-math input,
12
- # which panics the Rust core on the known @next_id collision
13
- # (parsanol-ruby#25) — run that one in :ruby until upstream fixes it.
11
+ # Requires the parsanol gem (add to the Gemfile, or point -I at a
12
+ # local checkout):
14
13
  #
15
- # Run from asciichem-ruby/ with the shim dir prepended:
16
- # ruby -I /tmp/parsanol_spike -I ../parsanol/parsanol-ruby/lib benchmarks/parsanol_recheck.rb
14
+ # bundle exec ruby -I ../parsanol/parsanol-ruby/lib benchmarks/parsanol_recheck.rb
15
+ #
16
+ # PARSANOL_MODE=ruby forces Parsanol's pure-Ruby backend (no Rust
17
+ # core) for comparison.
17
18
  require "benchmark/ips"
18
19
  require "asciichem"
20
+ require "asciichem/engine/parsanol_engine"
19
21
  require "json"
20
22
 
21
- puts "parsanol #{Parsanol::VERSION} | parslet-compat Parser=#{Parsanol::Parslet::Parser}"
22
- puts "AsciiChem::Grammar superclass: #{AsciiChem::Grammar.superclass}"
23
+ AsciiChem::Engine.use(:parsanol)
24
+
25
+ engine = AsciiChem::Engine.current
26
+ puts "parsanol #{Parsanol::VERSION} | engine: #{engine}"
27
+ puts "grammar superclass: #{engine.grammar.superclass}"
23
28
  puts "mode: #{ENV.fetch("PARSANOL_MODE", "native")}"
24
29
 
25
30
  if ENV.fetch("PARSANOL_MODE", "native") == "ruby"
@@ -1,6 +1,14 @@
1
1
  # frozen_string_literal: true
2
2
 
3
- require "relaton_bib"
3
+ # relaton-bib 2 renamed the entry file (relaton_bib -> relaton/bib)
4
+ # and reworked the namespace (RelatonBib -> Relaton::Bib). The
5
+ # gemspec admits both major lines, so load whichever is resolved and
6
+ # speak to it through RelatonApi below.
7
+ begin
8
+ require 'relaton/bib'
9
+ rescue LoadError
10
+ require 'relaton_bib'
11
+ end
4
12
 
5
13
  module AsciiChem
6
14
  # Citation track (TODO.v2 08; TODO.impl 44): a bibitem is a function
@@ -16,25 +24,25 @@ module AsciiChem
16
24
  Profile = Struct.new(:publisher, :link_for, :identifier_for, keyword_init: true)
17
25
 
18
26
  PROFILES = {
19
- "pubchem" => Profile.new(
20
- publisher: "PubChem, U.S. National Library of Medicine",
21
- link_for: ->(substance) do
22
- cid = substance.identifier_value("pubchem-cid")
27
+ 'pubchem' => Profile.new(
28
+ publisher: 'PubChem, U.S. National Library of Medicine',
29
+ link_for: lambda do |substance|
30
+ cid = substance.identifier_value('pubchem-cid')
23
31
  "https://pubchem.ncbi.nlm.nih.gov/compound/#{cid}" if cid
24
32
  end,
25
- identifier_for: ->(substance) do
26
- cid = substance.identifier_value("pubchem-cid")
33
+ identifier_for: lambda do |substance|
34
+ cid = substance.identifier_value('pubchem-cid')
27
35
  "PubChem CID #{cid}" if cid
28
36
  end
29
37
  ),
30
- "common_chemistry" => Profile.new(
31
- publisher: "CAS Common Chemistry",
32
- link_for: ->(substance) do
33
- cas = substance.identifier_value("cas")
38
+ 'common_chemistry' => Profile.new(
39
+ publisher: 'CAS Common Chemistry',
40
+ link_for: lambda do |substance|
41
+ cas = substance.identifier_value('cas')
34
42
  "https://commonchemistry.cas.org/detail?cas_rn=#{cas}" if cas
35
43
  end,
36
- identifier_for: ->(substance) do
37
- cas = substance.identifier_value("cas")
44
+ identifier_for: lambda do |substance|
45
+ cas = substance.identifier_value('cas')
38
46
  "CAS RN #{cas}" if cas
39
47
  end
40
48
  )
@@ -42,8 +50,8 @@ module AsciiChem
42
50
 
43
51
  DEFAULT_PROFILE = Profile.new(
44
52
  publisher: nil,
45
- link_for: ->(_substance) { nil },
46
- identifier_for: ->(substance) do
53
+ link_for: ->(_substance) {},
54
+ identifier_for: lambda do |substance|
47
55
  key = substance.identifiers.first
48
56
  "#{key.convention}: #{key.value}" if key
49
57
  end
@@ -56,34 +64,20 @@ module AsciiChem
56
64
  # carries no provenance (hand-built, not resolved).
57
65
  def bibitem(substance)
58
66
  provenance = substance.provenance
59
- unless provenance&.source
60
- raise Error, "substance has no provenance - resolve it first (AsciiChem::Resolver)"
61
- end
67
+ raise Error, 'substance has no provenance - resolve it first (AsciiChem::Resolver)' unless provenance&.source
62
68
 
63
69
  profile = PROFILES.fetch(provenance.source, DEFAULT_PROFILE)
64
- RelatonBib::BibliographicItem.new(
65
- type: "dataset",
66
- title: [{ type: "main",
67
- content: "#{title_base(substance)} - #{profile.publisher || provenance.source} substance record" }],
68
- docid: [RelatonBib::DocumentIdentifier.new(
69
- id: profile.identifier_for.call(substance) || "#{provenance.source} substance",
70
- type: provenance.source)],
71
- contributor: [{ entity: RelatonBib::Organization.new(name: profile.publisher || provenance.source),
72
- role: [{ type: "publisher" }] }],
73
- date: [{ type: "accessed", on: accessed_on(provenance) }],
74
- link: [{ type: "src", content: profile.link_for.call(substance) }].compact,
75
- keyword: substance.identifiers.map { |i| "#{i.convention}=#{i.value}" }
76
- )
70
+ RelatonApi.dataset_bibitem(fields(substance, profile, provenance))
77
71
  end
78
72
 
79
73
  # Convenience: bibitem XML (what a document pipeline embeds).
80
74
  def to_xml(substance)
81
- bibitem(substance).to_xml
75
+ RelatonApi.to_xml(bibitem(substance))
82
76
  end
83
77
 
84
78
  # The cite syntax (TODO.impl 45): a molecule annotated
85
79
  # `@cite("pubchem")` (a property annotation — the grammar needs
86
- # no extension) declares *which source to cite it from*. This
80
+ # no extension) declares *which source to cite it from. This
87
81
  # resolves the molecule's identifiers and emits one bibitem per
88
82
  # cited source. Returns [[source, bibitem]] pairs; empty when the
89
83
  # molecule has no @cite annotations.
@@ -97,13 +91,14 @@ module AsciiChem
97
91
  convention, value = lookup_key(molecule)
98
92
  unless value
99
93
  raise Error,
100
- "molecule carries no resolvable identifier for citation " \
101
- "(annotate @cas/@inchikey/@smiles or @name)"
94
+ 'molecule carries no resolvable identifier for citation ' \
95
+ '(annotate @cas/@inchikey/@smiles or @name)'
102
96
  end
103
97
 
104
98
  sources.filter_map do |source|
105
99
  substance = AsciiChem::Resolver[source].new.resolve(
106
- value: value, convention: convention, cache: cache, fetch: fetch)
100
+ value: value, convention: convention, cache: cache, fetch: fetch
101
+ )
107
102
  next unless substance
108
103
 
109
104
  [source, bibitem(substance)]
@@ -112,11 +107,27 @@ module AsciiChem
112
107
 
113
108
  private
114
109
 
110
+ # The version-independent field payload: one hash describing the
111
+ # citation, translated to Relaton objects by RelatonApi.
112
+ def fields(substance, profile, provenance)
113
+ publisher = profile.publisher || provenance.source
114
+ {
115
+ type: 'dataset',
116
+ title: "#{title_base(substance)} - #{publisher} substance record",
117
+ docid: { id: profile.identifier_for.call(substance) || "#{provenance.source} substance",
118
+ type: provenance.source },
119
+ publisher: publisher,
120
+ accessed_on: accessed_on(provenance),
121
+ link: profile.link_for.call(substance),
122
+ keywords: substance.identifiers.map { |i| "#{i.convention}=#{i.value}" }
123
+ }
124
+ end
125
+
115
126
  # The property annotation whose title is "cite": values are the
116
127
  # source names to cite from.
117
128
  def citation_sources(molecule)
118
129
  molecule.properties
119
- .select { |p| p.title == "cite" && p.value }
130
+ .select { |p| p.title == 'cite' && p.value }
120
131
  .map(&:value)
121
132
  end
122
133
 
@@ -127,23 +138,96 @@ module AsciiChem
127
138
  return [identifier.convention, identifier.value] if identifier
128
139
 
129
140
  name = molecule.names.first
130
- return ["name", name.content] if name
141
+ return ['name', name.content] if name
131
142
 
132
143
  nil
133
144
  end
134
145
 
135
- private
136
-
137
146
  def title_base(substance)
138
- substance.preferred_name || substance.identifier_value("cas") ||
139
- substance.identifier_value("inchikey") || "Substance"
147
+ substance.preferred_name || substance.identifier_value('cas') ||
148
+ substance.identifier_value('inchikey') || 'Substance'
140
149
  end
141
150
 
142
151
  def accessed_on(provenance)
143
152
  return provenance.retrieved_at[0, 10] if provenance.retrieved_at
144
153
 
145
- Time.now.utc.strftime("%Y-%m-%d")
154
+ Time.now.utc.strftime('%Y-%m-%d')
155
+ end
156
+ end
157
+
158
+ # The relaton-bib version seam. Both major lines accept the same
159
+ # field hash (see Citation#fields) and serialize through their own
160
+ # API; the rest of the citation track stays version-agnostic.
161
+ # Adding a future major = one more module here (OCP).
162
+ module RelatonApi
163
+ module_function
164
+
165
+ def dataset_bibitem(fields)
166
+ (defined?(::Relaton::Bib) ? V2 : V1).build(fields)
167
+ end
168
+
169
+ def to_xml(item)
170
+ item.to_xml
171
+ end
172
+
173
+ # relaton-bib 1: RelatonBib::* with hash-argument constructors.
174
+ module V1
175
+ module_function
176
+
177
+ def build(fields)
178
+ RelatonBib::BibliographicItem.new(
179
+ type: fields[:type],
180
+ title: [{ type: 'main', content: fields[:title] }],
181
+ docid: [RelatonBib::DocumentIdentifier.new(id: fields[:docid][:id],
182
+ type: fields[:docid][:type])],
183
+ contributor: [{ entity: RelatonBib::Organization.new(name: fields[:publisher]),
184
+ role: [{ type: 'publisher' }] }],
185
+ date: [{ type: 'accessed', on: fields[:accessed_on] }],
186
+ link: fields[:link] ? [{ type: 'src', content: fields[:link] }] : [],
187
+ keyword: fields[:keywords]
188
+ )
189
+ end
190
+ end
191
+
192
+ # relaton-bib 2: Relaton::Bib::* typed models (lutaml-model).
193
+ # Date's XML <on> element maps to the Ruby `at` attribute;
194
+ # keywords carry their text in a nested vocab LocalizedString;
195
+ # links are source Uri entries serializing to <uri type="src">.
196
+ module V2
197
+ module_function
198
+
199
+ def build(fields)
200
+ Relaton::Bib::ItemData.new(
201
+ type: fields[:type],
202
+ title: [Relaton::Bib::Title.new(type: 'main', content: fields[:title])],
203
+ docidentifier: [docidentifier(fields[:docid])],
204
+ contributor: [contributor(fields[:publisher])],
205
+ date: [Relaton::Bib::Date.new(type: 'accessed', at: fields[:accessed_on])],
206
+ source: fields[:link] ? [Relaton::Bib::Uri.new(type: 'src', content: fields[:link])] : [],
207
+ keyword: fields[:keywords].map { |text| keyword(text) }
208
+ )
209
+ end
210
+
211
+ def docidentifier(docid)
212
+ Relaton::Bib::Docidentifier.new(type: docid[:type], content: docid[:id])
213
+ end
214
+
215
+ def contributor(publisher)
216
+ Relaton::Bib::Contributor.new(
217
+ organization: Relaton::Bib::Organization.new(
218
+ name: [Relaton::Bib::TypedLocalizedString.new(content: publisher)]
219
+ ),
220
+ role: [Relaton::Bib::Contributor::Role.new(type: 'publisher')]
221
+ )
222
+ end
223
+
224
+ def keyword(text)
225
+ Relaton::Bib::Keyword.new(
226
+ vocab: Relaton::Bib::LocalizedString.new(content: text)
227
+ )
228
+ end
146
229
  end
147
230
  end
231
+ private_constant :RelatonApi
148
232
  end
149
233
  end
data/lib/asciichem/cli.rb CHANGED
@@ -1,6 +1,6 @@
1
1
  # frozen_string_literal: true
2
2
 
3
- require "thor"
3
+ require 'thor'
4
4
 
5
5
  module AsciiChem
6
6
  # Thor-based command line interface. Invoked via the `asciichem`
@@ -8,25 +8,30 @@ module AsciiChem
8
8
  class Cli < Thor
9
9
  # Use lowercase 'asciichem' as the program name in help output
10
10
  # and command banners, matching the executable name.
11
- package_name "asciichem"
11
+ package_name 'asciichem'
12
12
 
13
- desc "convert -i INPUT -t FORMAT", "Convert INPUT to FORMAT (mathml|text|html|latex|svg|structural-svg|model-json|cml|smiles|molfile)"
14
- method_option :input, aliases: "-i", type: :string,
15
- desc: "Source text (or '-' for stdin)"
16
- method_option :file, aliases: "-f", type: :string,
17
- desc: "Read source from a file"
18
- method_option :from, type: :string, default: "asciichem",
19
- desc: "Input grammar: asciichem|smiles|molfile"
20
- method_option :format, aliases: "-t", type: :string, default: "mathml",
21
- desc: "Output format"
13
+ desc 'convert -i INPUT -t FORMAT',
14
+ 'Convert INPUT to FORMAT (mathml|text|html|latex|svg|structural-svg|model-json|cml|smiles|molfile)'
15
+ method_option :input, aliases: '-i', type: :string,
16
+ desc: "Source text (or '-' for stdin)"
17
+ method_option :file, aliases: '-f', type: :string,
18
+ desc: 'Read source from a file'
19
+ method_option :from, type: :string, default: 'asciichem',
20
+ desc: 'Input grammar: asciichem|smiles|molfile'
21
+ method_option :format, aliases: '-t', type: :string, default: 'mathml',
22
+ desc: 'Output format'
23
+ method_option :engine, type: :string, default: 'parslet',
24
+ desc: 'Parsing engine: parslet (default) | parsanol (opt-in, needs the parsanol gem)'
22
25
  def convert
23
- unless options["input"] || options["file"]
24
- raise AsciiChem::ParseError, "provide -i INPUT or -f FILE"
25
- end
26
+ raise AsciiChem::ParseError, 'provide -i INPUT or -f FILE' unless options['input'] || options['file']
26
27
 
28
+ select_engine(options[:engine])
27
29
  source = read_source
28
30
  formula = ingest(source, options[:from])
29
31
  puts render(formula, options[:format])
32
+ rescue AsciiChem::Engine::Error => e
33
+ warn "Engine error: #{e.message}"
34
+ exit 6
30
35
  rescue AsciiChem::ParseError => e
31
36
  warn "Parse error: #{e.message}"
32
37
  exit 1
@@ -35,9 +40,9 @@ module AsciiChem
35
40
  exit 2
36
41
  end
37
42
 
38
- desc "parse-cml -i INPUT", "Parse CML XML and emit AsciiChem text"
39
- method_option :input, aliases: "-i", type: :string, required: true,
40
- desc: "CML XML source"
43
+ desc 'parse-cml -i INPUT', 'Parse CML XML and emit AsciiChem text'
44
+ method_option :input, aliases: '-i', type: :string, required: true,
45
+ desc: 'CML XML source'
41
46
  def parse_cml
42
47
  formula = AsciiChem::Cml.parse(options[:input])
43
48
  puts formula.to_text
@@ -46,8 +51,8 @@ module AsciiChem
46
51
  exit 1
47
52
  end
48
53
 
49
- desc "roundtrip -i INPUT", "Parse and re-emit; exit non-zero if not equal"
50
- method_option :input, aliases: "-i", type: :string, required: true
54
+ desc 'roundtrip -i INPUT', 'Parse and re-emit; exit non-zero if not equal'
55
+ method_option :input, aliases: '-i', type: :string, required: true
51
56
  def roundtrip
52
57
  original = options[:input]
53
58
  rendered = AsciiChem.parse(original).to_text
@@ -60,11 +65,11 @@ module AsciiChem
60
65
  end
61
66
  end
62
67
 
63
- desc "lint -i INPUT", "Run chemistry checks; exit 1 on error, 0 if clean"
64
- method_option :input, aliases: "-i", type: :string, required: true,
65
- desc: "AsciiChem source text"
66
- method_option :format, aliases: "-f", type: :string, default: "text",
67
- desc: "Output format: text or json"
68
+ desc 'lint -i INPUT', 'Run chemistry checks; exit 1 on error, 0 if clean'
69
+ method_option :input, aliases: '-i', type: :string, required: true,
70
+ desc: 'AsciiChem source text'
71
+ method_option :format, aliases: '-f', type: :string, default: 'text',
72
+ desc: 'Output format: text or json'
68
73
  def lint
69
74
  formula = AsciiChem.parse(options[:input])
70
75
  diagnostics = AsciiChem::Linter.run(formula)
@@ -76,7 +81,7 @@ module AsciiChem
76
81
  end
77
82
 
78
83
  map %w[--version -v] => :version
79
- desc "version", "Print the asciichem version"
84
+ desc 'version', 'Print the asciichem version'
80
85
  def version
81
86
  puts "asciichem #{AsciiChem::VERSION}"
82
87
  end
@@ -86,32 +91,32 @@ module AsciiChem
86
91
  "asciichem #{command.usage}"
87
92
  end
88
93
 
89
- desc "resolve --cas X | --name X | ...", "Resolve a substance from a source (network; cached)"
90
- method_option :cas, type: :string, desc: "CAS registry number"
91
- method_option :name, type: :string, desc: "Substance name"
92
- method_option :cid, type: :string, desc: "PubChem CID"
93
- method_option :inchikey, type: :string, desc: "InChIKey"
94
- method_option :smiles, type: :string, desc: "SMILES"
95
- method_option :source, type: :string, default: "pubchem", desc: "Resolver source"
96
- method_option :refresh, type: :boolean, default: false, desc: "Bypass the cache"
97
- method_option :format, aliases: "-t", type: :string, default: "model-json",
98
- desc: "Output: model-json | text | smiles"
94
+ desc 'resolve --cas X | --name X | ...', 'Resolve a substance from a source (network; cached)'
95
+ method_option :cas, type: :string, desc: 'CAS registry number'
96
+ method_option :name, type: :string, desc: 'Substance name'
97
+ method_option :cid, type: :string, desc: 'PubChem CID'
98
+ method_option :inchikey, type: :string, desc: 'InChIKey'
99
+ method_option :smiles, type: :string, desc: 'SMILES'
100
+ method_option :source, type: :string, default: 'pubchem', desc: 'Resolver source'
101
+ method_option :refresh, type: :boolean, default: false, desc: 'Bypass the cache'
102
+ method_option :format, aliases: '-t', type: :string, default: 'model-json',
103
+ desc: 'Output: model-json | text | smiles'
99
104
  def resolve
100
105
  convention, value = %i[cas name cid inchikey smiles]
101
106
  .filter_map { |k| [k, options[k.to_s]] if options[k.to_s] }
102
107
  .first
103
- raise AsciiChem::Error, "give one of --cas/--name/--cid/--inchikey/--smiles" unless value
108
+ raise AsciiChem::Error, 'give one of --cas/--name/--cid/--inchikey/--smiles' unless value
104
109
 
105
- convention = { cas: "cas", name: "name", cid: "pubchem-cid",
106
- inchikey: "inchikey", smiles: "smiles" }.fetch(convention)
110
+ convention = { cas: 'cas', name: 'name', cid: 'pubchem-cid',
111
+ inchikey: 'inchikey', smiles: 'smiles' }.fetch(convention)
107
112
  substance = AsciiChem::Resolver[options[:source]].new.resolve(
108
113
  value: value, convention: convention, refresh: options[:refresh]
109
114
  )
110
115
  raise AsciiChem::Error, "#{options[:source]} does not know #{value.inspect}" unless substance
111
116
 
112
117
  puts case options[:format].to_s
113
- when "text" then substance.preferred_name.to_s
114
- when "smiles" then substance.identifier_value("canonical-smiles").to_s
118
+ when 'text' then substance.preferred_name.to_s
119
+ when 'smiles' then substance.identifier_value('canonical-smiles').to_s
115
120
  else substance.to_model_json
116
121
  end
117
122
  rescue AsciiChem::Error => e
@@ -119,22 +124,22 @@ module AsciiChem
119
124
  exit 3
120
125
  end
121
126
 
122
- desc "cite --cas X | --name X | ...", "Resolve a substance and emit a dataset-type Relaton bibitem (XML)"
123
- method_option :cas, type: :string, desc: "CAS registry number"
124
- method_option :name, type: :string, desc: "Substance name"
125
- method_option :cid, type: :string, desc: "PubChem CID"
126
- method_option :inchikey, type: :string, desc: "InChIKey"
127
- method_option :smiles, type: :string, desc: "SMILES"
128
- method_option :source, type: :string, default: "pubchem", desc: "Resolver source"
129
- method_option :refresh, type: :boolean, default: false, desc: "Bypass the cache"
127
+ desc 'cite --cas X | --name X | ...', 'Resolve a substance and emit a dataset-type Relaton bibitem (XML)'
128
+ method_option :cas, type: :string, desc: 'CAS registry number'
129
+ method_option :name, type: :string, desc: 'Substance name'
130
+ method_option :cid, type: :string, desc: 'PubChem CID'
131
+ method_option :inchikey, type: :string, desc: 'InChIKey'
132
+ method_option :smiles, type: :string, desc: 'SMILES'
133
+ method_option :source, type: :string, default: 'pubchem', desc: 'Resolver source'
134
+ method_option :refresh, type: :boolean, default: false, desc: 'Bypass the cache'
130
135
  def cite
131
136
  convention, value = %i[cas name cid inchikey smiles]
132
137
  .filter_map { |k| [k, options[k.to_s]] if options[k.to_s] }
133
138
  .first
134
- raise AsciiChem::Error, "give one of --cas/--name/--cid/--inchikey/--smiles" unless value
139
+ raise AsciiChem::Error, 'give one of --cas/--name/--cid/--inchikey/--smiles' unless value
135
140
 
136
- convention = { cas: "cas", name: "name", cid: "pubchem-cid",
137
- inchikey: "inchikey", smiles: "smiles" }.fetch(convention)
141
+ convention = { cas: 'cas', name: 'name', cid: 'pubchem-cid',
142
+ inchikey: 'inchikey', smiles: 'smiles' }.fetch(convention)
138
143
  substance = AsciiChem::Resolver[options[:source]].new.resolve(
139
144
  value: value, convention: convention, refresh: options[:refresh]
140
145
  )
@@ -146,37 +151,43 @@ module AsciiChem
146
151
  exit 4
147
152
  end
148
153
 
149
- desc "validate -i INPUT", "Offline identifier validation"
150
- method_option :input, aliases: "-i", type: :string, required: true
154
+ desc 'validate -i INPUT', 'Offline identifier validation'
155
+ method_option :input, aliases: '-i', type: :string, required: true
151
156
  def validate
152
157
  formula = AsciiChem.parse(options[:input])
153
158
  annotations = formula.nodes.grep(AsciiChem::Model::Molecule).flat_map(&:identifiers)
154
159
  if annotations.empty?
155
- puts "no identifier annotations found"
160
+ puts 'no identifier annotations found'
156
161
  return
157
162
  end
158
163
  annotations.each do |identifier|
159
164
  known = AsciiChem::Identifiers.known?(identifier.convention)
160
165
  valid = known && AsciiChem::Identifiers.valid?(identifier.convention, identifier.value)
161
- status = known ? (valid ? "ok" : "INVALID") : "unknown convention"
162
- puts format("%-12s %-40s %s", identifier.convention, identifier.value, status)
166
+ status = if known
167
+ valid ? 'ok' : 'INVALID'
168
+ else
169
+ 'unknown convention'
170
+ end
171
+ puts format('%-12s %-40s %s', identifier.convention, identifier.value, status)
172
+ end
173
+ exit 1 if annotations.any? do |i|
174
+ AsciiChem::Identifiers.known?(i.convention) &&
175
+ !AsciiChem::Identifiers.valid?(i.convention, i.value)
163
176
  end
164
- exit 1 if annotations.any? { |i| AsciiChem::Identifiers.known?(i.convention) &&
165
- !AsciiChem::Identifiers.valid?(i.convention, i.value) }
166
177
  rescue AsciiChem::ParseError => e
167
178
  warn "Parse error: #{e.message}"
168
179
  exit 1
169
180
  end
170
181
 
171
- desc "identity -i INPUT", "Derive InChI/InChIKey locally from the structure (offline; requires an InChI engine)"
172
- method_option :input, aliases: "-i", type: :string,
173
- desc: "Source text (or '-' for stdin)"
174
- method_option :file, aliases: "-f", type: :string,
175
- desc: "Read source from a file"
176
- method_option :from, type: :string, default: "asciichem",
177
- desc: "Input grammar: asciichem|smiles|molfile"
182
+ desc 'identity -i INPUT', 'Derive InChI/InChIKey locally from the structure (offline; requires an InChI engine)'
183
+ method_option :input, aliases: '-i', type: :string,
184
+ desc: "Source text (or '-' for stdin)"
185
+ method_option :file, aliases: '-f', type: :string,
186
+ desc: 'Read source from a file'
187
+ method_option :from, type: :string, default: 'asciichem',
188
+ desc: 'Input grammar: asciichem|smiles|molfile'
178
189
  method_option :engine_bin, type: :string,
179
- desc: "Path to the inchi-1 binary (overrides the configured engine)"
190
+ desc: 'Path to the inchi-1 binary (overrides the configured engine)'
180
191
  def identity
181
192
  formula = ingest(read_source, options[:from])
182
193
  molecules = molecules_in(formula)
@@ -195,9 +206,18 @@ module AsciiChem
195
206
 
196
207
  private
197
208
 
209
+ # --engine feeds the Engine selector (0.29.0+); :parsanol is an
210
+ # opt-in soft dependency, :parslet stays the default.
211
+ def select_engine(name)
212
+ return if name.to_s == 'parslet' && AsciiChem::Engine.current == AsciiChem::Engine::ParsletEngine
213
+
214
+ require 'asciichem/engine'
215
+ AsciiChem::Engine.use(name.to_s)
216
+ end
217
+
198
218
  def read_source
199
219
  return File.read(options[:file]) if options[:file]
200
- return $stdin.read if options[:input] == "-"
220
+ return $stdin.read if options[:input] == '-'
201
221
 
202
222
  options[:input]
203
223
  end
@@ -207,9 +227,9 @@ module AsciiChem
207
227
  # format works regardless of the input language.
208
228
  def ingest(source, from)
209
229
  case from.to_s
210
- when "asciichem" then AsciiChem.parse(source)
211
- when "smiles" then AsciiChem.parse_smiles(source)
212
- when "molfile" then molfile_formula(source)
230
+ when 'asciichem' then AsciiChem.parse(source)
231
+ when 'smiles' then AsciiChem.parse_smiles(source)
232
+ when 'molfile' then molfile_formula(source)
213
233
  else
214
234
  raise AsciiChem::ParseError, "unknown --from grammar: #{from}"
215
235
  end
@@ -224,7 +244,7 @@ module AsciiChem
224
244
 
225
245
  def render(formula, format)
226
246
  return formula.to_cml if format.to_sym == :cml
227
- return formula.to_model_json if format.to_sym == :"model-json"
247
+ return formula.to_model_json if format.to_sym == :'model-json'
228
248
  return formula.to_smiles if format.to_sym == :smiles
229
249
  return formula.nodes.first.to_molfile if format.to_sym == :molfile
230
250
 
@@ -246,14 +266,14 @@ module AsciiChem
246
266
 
247
267
  def output_lint(diagnostics, format)
248
268
  case format.to_s
249
- when "json" then output_lint_json(diagnostics)
269
+ when 'json' then output_lint_json(diagnostics)
250
270
  else
251
271
  diagnostics.each { |d| puts d }
252
272
  end
253
273
  end
254
274
 
255
275
  def output_lint_json(diagnostics)
256
- require "json"
276
+ require 'json'
257
277
  payload = diagnostics.map do |d|
258
278
  {
259
279
  severity: d.severity.to_s,
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module AsciiChem
4
- VERSION = "0.29.0"
4
+ VERSION = "0.29.2"
5
5
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: asciichem
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.29.0
4
+ version: 0.29.2
5
5
  platform: ruby
6
6
  authors:
7
7
  - Ribose Inc.
@@ -108,7 +108,7 @@ dependencies:
108
108
  version: '0.1'
109
109
  - - "<"
110
110
  - !ruby/object:Gem::Version
111
- version: '2'
111
+ version: '3'
112
112
  type: :runtime
113
113
  prerelease: false
114
114
  version_requirements: !ruby/object:Gem::Requirement
@@ -118,7 +118,7 @@ dependencies:
118
118
  version: '0.1'
119
119
  - - "<"
120
120
  - !ruby/object:Gem::Version
121
- version: '2'
121
+ version: '3'
122
122
  - !ruby/object:Gem::Dependency
123
123
  name: plurimath
124
124
  requirement: !ruby/object:Gem::Requirement