asciichem 0.29.0 → 0.29.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.github/workflows/release.yml +9 -0
- data/CHANGELOG.md +20 -0
- data/asciichem.gemspec +1 -1
- data/benchmarks/README.md +63 -5
- data/benchmarks/parsanol_recheck.rb +19 -14
- data/lib/asciichem/citation.rb +127 -43
- data/lib/asciichem/cli.rb +93 -73
- data/lib/asciichem/version.rb +1 -1
- metadata +3 -3
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: f372a8539906c658a399b015255e574de5feb73d1b5e7183958901681bfbce34
|
|
4
|
+
data.tar.gz: 7269dc4ad6e39e7c5ea6726b54a8ffe1508c039bdf9ced5c0a74e23efdc9b00f
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 6f936bc0e95907d589c0095d34e4e65c554dcbdccd276aab7186db8f1218cbb179652b082f088f1b70a76391e830db53de2259ea8d333b64dd85ba367b16c610
|
|
7
|
+
data.tar.gz: 63d979830e4a8ab9aca450006a9740e83428e26073323af0ed4f8ddeed4be6f7aecb8bcfe1e5173f7a6049f92ff8b0a8e0a2b374fbd13630ae00b0e9e7a0db60
|
|
@@ -34,6 +34,15 @@ jobs:
|
|
|
34
34
|
with:
|
|
35
35
|
ruby-version: "3.4"
|
|
36
36
|
bundler-cache: true
|
|
37
|
+
# CI never pushes to git (read-only contents). `rake release`
|
|
38
|
+
# attempts `git push origin main` after publishing unless the
|
|
39
|
+
# version tag already exists locally — bundler then prints
|
|
40
|
+
# "Tag vX has already been created" and skips its git stage
|
|
41
|
+
# entirely (this is how 0.29.0/0.29.1 released). Pre-create the
|
|
42
|
+
# tag so the gem push is the only remote operation; tags on the
|
|
43
|
+
# remote remain the maintainer's.
|
|
44
|
+
- name: Pre-create the release tag (skips rake's git stage)
|
|
45
|
+
run: git tag "v${{ inputs.version }}"
|
|
37
46
|
# Builds and pushes using the GitHub OIDC identity — no API keys.
|
|
38
47
|
- uses: rubygems/release-gem@v1
|
|
39
48
|
- name: Summary
|
data/CHANGELOG.md
CHANGED
|
@@ -3,6 +3,26 @@
|
|
|
3
3
|
All notable changes to AsciiChem are documented here.
|
|
4
4
|
This project follows [Semantic Versioning](https://semver.org/).
|
|
5
5
|
|
|
6
|
+
## [0.29.2] - 2026-09-17
|
|
7
|
+
|
|
8
|
+
### Changed
|
|
9
|
+
- `relaton-bib` constraint widened to `>= 0.1, < 3`: asciichem now
|
|
10
|
+
co-resolves with current metanorma gems (metanorma-standoc and
|
|
11
|
+
friends require relaton-bib 2). The citation track speaks both
|
|
12
|
+
major lines through a single `Citation::RelatonApi` seam —
|
|
13
|
+
relaton-bib 1 (`RelatonBib`) and relaton-bib 2
|
|
14
|
+
(`Relaton::Bib` typed models) both build and serialize the
|
|
15
|
+
dataset-type bibitems; profile data is version-independent.
|
|
16
|
+
- nil resolver links no longer emit an empty `<uri>` element.
|
|
17
|
+
|
|
18
|
+
## [0.29.1] - 2026-09-16
|
|
19
|
+
|
|
20
|
+
### Added
|
|
21
|
+
- CLI `convert --engine parslet|parsanol` selects the parsing engine
|
|
22
|
+
per invocation (parslet stays the default); absent parsanol gem
|
|
23
|
+
exits 6 with install guidance. The ASCIICHEM_ENGINE env var keeps
|
|
24
|
+
working for programmatic selection.
|
|
25
|
+
|
|
6
26
|
## [0.29.0] - 2026-09-15
|
|
7
27
|
|
|
8
28
|
### Added
|
data/asciichem.gemspec
CHANGED
|
@@ -39,7 +39,7 @@ Gem::Specification.new do |spec|
|
|
|
39
39
|
spec.add_dependency "mml", "~> 2.3"
|
|
40
40
|
spec.add_dependency "nokogiri", "~> 1.16"
|
|
41
41
|
spec.add_dependency "parslet", "~> 2.0"
|
|
42
|
-
spec.add_dependency "relaton-bib", ">= 0.1", "<
|
|
42
|
+
spec.add_dependency "relaton-bib", ">= 0.1", "< 3"
|
|
43
43
|
spec.add_dependency "plurimath", "~> 0.8"
|
|
44
44
|
spec.add_dependency "thor", "~> 1.3"
|
|
45
45
|
|
data/benchmarks/README.md
CHANGED
|
@@ -104,9 +104,67 @@ fork-per-case gate so Rust aborts are reported, not fatal):
|
|
|
104
104
|
but rejected under parslet (optimizer Str/Re run-merging semantics;
|
|
105
105
|
no corpus case covers these spellings today).
|
|
106
106
|
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
107
|
+
### Re-check 5 (2026-09-15, parsanol 1.3.17)
|
|
108
|
+
|
|
109
|
+
The ffi-gem cdylib tier (Rust engine on every runtime) changes
|
|
110
|
+
nothing for us: gate still **221/221**, 4.40 ms/i on the recheck
|
|
111
|
+
workload. The opt-in engine shipped in asciichem 0.29.0 is
|
|
112
|
+
unaffected; `@next_id` remains unfixed upstream but no longer fires
|
|
113
|
+
on corpus inputs. The re-check-3 verdict below is superseded — the
|
|
114
|
+
engine IS switchable and shipped (TODO.impl 64).
|
|
115
|
+
|
|
116
|
+
**Verdict: superseded — the engine IS switchable and shipped**
|
|
117
|
+
(asciichem 0.29.0, TODO.impl 64).
|
|
118
|
+
|
|
119
|
+
### Re-check 6 (2026-09-16, parsanol 1.3.18)
|
|
120
|
+
|
|
121
|
+
Gate still **221/221** through the shipped engine, but the "one
|
|
122
|
+
decode path" rework **regressed compat-layer throughput ~60%** for
|
|
123
|
+
this grammar: 7.2 ms/batch (138 i/s) vs 4.3-4.6 ms on 1.3.15/16,
|
|
124
|
+
with the parslet control stable across sessions (11-14 ms
|
|
125
|
+
throughout). The `H2`/`_2O` acceptance divergence also persists.
|
|
126
|
+
Reported upstream (parsanol-ruby#25, fourth comment). We stay on
|
|
127
|
+
the shipped engine; users pinning parsanol for speed should prefer
|
|
128
|
+
1.3.16/1.3.17 until the regression is addressed.
|
|
129
|
+
|
|
130
|
+
### Re-check 7 (2026-09-16, parsanol 1.3.20)
|
|
131
|
+
|
|
132
|
+
Three upstream issues closed since 1.3.18:
|
|
133
|
+
|
|
134
|
+
- **#38** (`Dynamic.register` `@next_id` collision panicking the
|
|
135
|
+
Rust core) — fixed; no panic during full-corpus run.
|
|
136
|
+
- **#37** ("one decode path" throughput regression) — fixed as a
|
|
137
|
+
side effect of the optimizer acceptance fix in #39; throughput on
|
|
138
|
+
this grammar is back to and ahead of 1.3.15/16 levels.
|
|
139
|
+
- **#39** (optimizer Str/Re run-merging changed sequence-boundary
|
|
140
|
+
acceptance) — root-caused to Re-run regex-source concatenation
|
|
141
|
+
(proven unsafe: `"a|"+"b"` → `"a|b"` accepts `"a"`); Re runs now
|
|
142
|
+
stay unmerged, Str-run merging stays. Spec-level decision
|
|
143
|
+
recorded: the optimizer may never alter acceptance.
|
|
144
|
+
|
|
145
|
+
Validation against 1.3.20:
|
|
146
|
+
|
|
147
|
+
- Gate **221/221** through the shipped `ParsanolEngine` (fork-per-case,
|
|
148
|
+
no Rust aborts).
|
|
149
|
+
- Head-to-head vs parslet, same Ruby process (3 runs, ±3-15%):
|
|
150
|
+
parsanol **2.6x faster** (4.65–5.19 ms/batch vs 12.18–13.65 ms for
|
|
151
|
+
parslet). Up from the 1.7x under 1.3.18 — the #37 regression is
|
|
152
|
+
gone.
|
|
153
|
+
- Direct `H2` / `_2O` / `Ca2+` / `H22` / `O2` probe across both
|
|
154
|
+
parslet and parsanol (native and ruby backends) shows **identical
|
|
155
|
+
parse outcomes**. The earlier "divergence" framing in re-checks
|
|
156
|
+
3-6 was a misreading: AsciiChem's `hydrogen_atom` grammar rule
|
|
157
|
+
intentionally permits bare-digit subscripts after `H` ("lets users
|
|
158
|
+
write `H2O` instead of `H_2O`" — grammar_rules.rb:228-231) and
|
|
159
|
+
`isotope_marker` accepts both `^digits` and `_digits`, so `_2O`
|
|
160
|
+
parses as the isotope of `O` and round-trips as `^2O`. The
|
|
161
|
+
parsanol optimizer bug in #39 was real and is fixed, but the
|
|
162
|
+
AsciiChem repro was a misleading example — both engines agree on
|
|
163
|
+
these inputs because they share the same grammar rules.
|
|
164
|
+
|
|
165
|
+
**Verdict: shipped engine fully validated.** 2.6x speedup, 100%
|
|
166
|
+
corpus gate, all four reported upstream issues now resolved or
|
|
167
|
+
non-blocking (#36 bare repeated sibling captures remains open but
|
|
168
|
+
is worked around in `ParsanolEngine` via single `.as(...)` capture
|
|
169
|
+
wrapping).
|
|
112
170
|
|
|
@@ -1,25 +1,30 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
|
-
# Parsanol
|
|
4
|
-
#
|
|
5
|
-
#
|
|
6
|
-
# on the
|
|
3
|
+
# Parsanol gate + benchmark against the SHIPPED opt-in engine
|
|
4
|
+
# (asciichem 0.29.0+): AsciiChem::Engine.use(:parsanol) runs the same
|
|
5
|
+
# GrammarRules/TransformRules over Parsanol's Rust-backed compat
|
|
6
|
+
# layer. (1) gates on the shared corpus, forked per case so a Rust
|
|
7
|
+
# panic aborts the child — reported, not fatal; (2) gates on the
|
|
8
|
+
# issue-25 EOF repro; (3) measures against the parslet path
|
|
9
|
+
# (benchmarks/engines.rb).
|
|
7
10
|
#
|
|
8
|
-
#
|
|
9
|
-
#
|
|
10
|
-
# native backend parses 169/170 corpus accepts + all 51 rejects at
|
|
11
|
-
# ~2.3x parslet speed; the single fatal case is embedded-math input,
|
|
12
|
-
# which panics the Rust core on the known @next_id collision
|
|
13
|
-
# (parsanol-ruby#25) — run that one in :ruby until upstream fixes it.
|
|
11
|
+
# Requires the parsanol gem (add to the Gemfile, or point -I at a
|
|
12
|
+
# local checkout):
|
|
14
13
|
#
|
|
15
|
-
#
|
|
16
|
-
#
|
|
14
|
+
# bundle exec ruby -I ../parsanol/parsanol-ruby/lib benchmarks/parsanol_recheck.rb
|
|
15
|
+
#
|
|
16
|
+
# PARSANOL_MODE=ruby forces Parsanol's pure-Ruby backend (no Rust
|
|
17
|
+
# core) for comparison.
|
|
17
18
|
require "benchmark/ips"
|
|
18
19
|
require "asciichem"
|
|
20
|
+
require "asciichem/engine/parsanol_engine"
|
|
19
21
|
require "json"
|
|
20
22
|
|
|
21
|
-
|
|
22
|
-
|
|
23
|
+
AsciiChem::Engine.use(:parsanol)
|
|
24
|
+
|
|
25
|
+
engine = AsciiChem::Engine.current
|
|
26
|
+
puts "parsanol #{Parsanol::VERSION} | engine: #{engine}"
|
|
27
|
+
puts "grammar superclass: #{engine.grammar.superclass}"
|
|
23
28
|
puts "mode: #{ENV.fetch("PARSANOL_MODE", "native")}"
|
|
24
29
|
|
|
25
30
|
if ENV.fetch("PARSANOL_MODE", "native") == "ruby"
|
data/lib/asciichem/citation.rb
CHANGED
|
@@ -1,6 +1,14 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
# relaton-bib 2 renamed the entry file (relaton_bib -> relaton/bib)
|
|
4
|
+
# and reworked the namespace (RelatonBib -> Relaton::Bib). The
|
|
5
|
+
# gemspec admits both major lines, so load whichever is resolved and
|
|
6
|
+
# speak to it through RelatonApi below.
|
|
7
|
+
begin
|
|
8
|
+
require 'relaton/bib'
|
|
9
|
+
rescue LoadError
|
|
10
|
+
require 'relaton_bib'
|
|
11
|
+
end
|
|
4
12
|
|
|
5
13
|
module AsciiChem
|
|
6
14
|
# Citation track (TODO.v2 08; TODO.impl 44): a bibitem is a function
|
|
@@ -16,25 +24,25 @@ module AsciiChem
|
|
|
16
24
|
Profile = Struct.new(:publisher, :link_for, :identifier_for, keyword_init: true)
|
|
17
25
|
|
|
18
26
|
PROFILES = {
|
|
19
|
-
|
|
20
|
-
publisher:
|
|
21
|
-
link_for:
|
|
22
|
-
cid = substance.identifier_value(
|
|
27
|
+
'pubchem' => Profile.new(
|
|
28
|
+
publisher: 'PubChem, U.S. National Library of Medicine',
|
|
29
|
+
link_for: lambda do |substance|
|
|
30
|
+
cid = substance.identifier_value('pubchem-cid')
|
|
23
31
|
"https://pubchem.ncbi.nlm.nih.gov/compound/#{cid}" if cid
|
|
24
32
|
end,
|
|
25
|
-
identifier_for:
|
|
26
|
-
cid = substance.identifier_value(
|
|
33
|
+
identifier_for: lambda do |substance|
|
|
34
|
+
cid = substance.identifier_value('pubchem-cid')
|
|
27
35
|
"PubChem CID #{cid}" if cid
|
|
28
36
|
end
|
|
29
37
|
),
|
|
30
|
-
|
|
31
|
-
publisher:
|
|
32
|
-
link_for:
|
|
33
|
-
cas = substance.identifier_value(
|
|
38
|
+
'common_chemistry' => Profile.new(
|
|
39
|
+
publisher: 'CAS Common Chemistry',
|
|
40
|
+
link_for: lambda do |substance|
|
|
41
|
+
cas = substance.identifier_value('cas')
|
|
34
42
|
"https://commonchemistry.cas.org/detail?cas_rn=#{cas}" if cas
|
|
35
43
|
end,
|
|
36
|
-
identifier_for:
|
|
37
|
-
cas = substance.identifier_value(
|
|
44
|
+
identifier_for: lambda do |substance|
|
|
45
|
+
cas = substance.identifier_value('cas')
|
|
38
46
|
"CAS RN #{cas}" if cas
|
|
39
47
|
end
|
|
40
48
|
)
|
|
@@ -42,8 +50,8 @@ module AsciiChem
|
|
|
42
50
|
|
|
43
51
|
DEFAULT_PROFILE = Profile.new(
|
|
44
52
|
publisher: nil,
|
|
45
|
-
link_for: ->(_substance) {
|
|
46
|
-
identifier_for:
|
|
53
|
+
link_for: ->(_substance) {},
|
|
54
|
+
identifier_for: lambda do |substance|
|
|
47
55
|
key = substance.identifiers.first
|
|
48
56
|
"#{key.convention}: #{key.value}" if key
|
|
49
57
|
end
|
|
@@ -56,34 +64,20 @@ module AsciiChem
|
|
|
56
64
|
# carries no provenance (hand-built, not resolved).
|
|
57
65
|
def bibitem(substance)
|
|
58
66
|
provenance = substance.provenance
|
|
59
|
-
unless provenance&.source
|
|
60
|
-
raise Error, "substance has no provenance - resolve it first (AsciiChem::Resolver)"
|
|
61
|
-
end
|
|
67
|
+
raise Error, 'substance has no provenance - resolve it first (AsciiChem::Resolver)' unless provenance&.source
|
|
62
68
|
|
|
63
69
|
profile = PROFILES.fetch(provenance.source, DEFAULT_PROFILE)
|
|
64
|
-
|
|
65
|
-
type: "dataset",
|
|
66
|
-
title: [{ type: "main",
|
|
67
|
-
content: "#{title_base(substance)} - #{profile.publisher || provenance.source} substance record" }],
|
|
68
|
-
docid: [RelatonBib::DocumentIdentifier.new(
|
|
69
|
-
id: profile.identifier_for.call(substance) || "#{provenance.source} substance",
|
|
70
|
-
type: provenance.source)],
|
|
71
|
-
contributor: [{ entity: RelatonBib::Organization.new(name: profile.publisher || provenance.source),
|
|
72
|
-
role: [{ type: "publisher" }] }],
|
|
73
|
-
date: [{ type: "accessed", on: accessed_on(provenance) }],
|
|
74
|
-
link: [{ type: "src", content: profile.link_for.call(substance) }].compact,
|
|
75
|
-
keyword: substance.identifiers.map { |i| "#{i.convention}=#{i.value}" }
|
|
76
|
-
)
|
|
70
|
+
RelatonApi.dataset_bibitem(fields(substance, profile, provenance))
|
|
77
71
|
end
|
|
78
72
|
|
|
79
73
|
# Convenience: bibitem XML (what a document pipeline embeds).
|
|
80
74
|
def to_xml(substance)
|
|
81
|
-
bibitem(substance)
|
|
75
|
+
RelatonApi.to_xml(bibitem(substance))
|
|
82
76
|
end
|
|
83
77
|
|
|
84
78
|
# The cite syntax (TODO.impl 45): a molecule annotated
|
|
85
79
|
# `@cite("pubchem")` (a property annotation — the grammar needs
|
|
86
|
-
# no extension) declares *which source to cite it from
|
|
80
|
+
# no extension) declares *which source to cite it from. This
|
|
87
81
|
# resolves the molecule's identifiers and emits one bibitem per
|
|
88
82
|
# cited source. Returns [[source, bibitem]] pairs; empty when the
|
|
89
83
|
# molecule has no @cite annotations.
|
|
@@ -97,13 +91,14 @@ module AsciiChem
|
|
|
97
91
|
convention, value = lookup_key(molecule)
|
|
98
92
|
unless value
|
|
99
93
|
raise Error,
|
|
100
|
-
|
|
101
|
-
|
|
94
|
+
'molecule carries no resolvable identifier for citation ' \
|
|
95
|
+
'(annotate @cas/@inchikey/@smiles or @name)'
|
|
102
96
|
end
|
|
103
97
|
|
|
104
98
|
sources.filter_map do |source|
|
|
105
99
|
substance = AsciiChem::Resolver[source].new.resolve(
|
|
106
|
-
value: value, convention: convention, cache: cache, fetch: fetch
|
|
100
|
+
value: value, convention: convention, cache: cache, fetch: fetch
|
|
101
|
+
)
|
|
107
102
|
next unless substance
|
|
108
103
|
|
|
109
104
|
[source, bibitem(substance)]
|
|
@@ -112,11 +107,27 @@ module AsciiChem
|
|
|
112
107
|
|
|
113
108
|
private
|
|
114
109
|
|
|
110
|
+
# The version-independent field payload: one hash describing the
|
|
111
|
+
# citation, translated to Relaton objects by RelatonApi.
|
|
112
|
+
def fields(substance, profile, provenance)
|
|
113
|
+
publisher = profile.publisher || provenance.source
|
|
114
|
+
{
|
|
115
|
+
type: 'dataset',
|
|
116
|
+
title: "#{title_base(substance)} - #{publisher} substance record",
|
|
117
|
+
docid: { id: profile.identifier_for.call(substance) || "#{provenance.source} substance",
|
|
118
|
+
type: provenance.source },
|
|
119
|
+
publisher: publisher,
|
|
120
|
+
accessed_on: accessed_on(provenance),
|
|
121
|
+
link: profile.link_for.call(substance),
|
|
122
|
+
keywords: substance.identifiers.map { |i| "#{i.convention}=#{i.value}" }
|
|
123
|
+
}
|
|
124
|
+
end
|
|
125
|
+
|
|
115
126
|
# The property annotation whose title is "cite": values are the
|
|
116
127
|
# source names to cite from.
|
|
117
128
|
def citation_sources(molecule)
|
|
118
129
|
molecule.properties
|
|
119
|
-
.select { |p| p.title ==
|
|
130
|
+
.select { |p| p.title == 'cite' && p.value }
|
|
120
131
|
.map(&:value)
|
|
121
132
|
end
|
|
122
133
|
|
|
@@ -127,23 +138,96 @@ module AsciiChem
|
|
|
127
138
|
return [identifier.convention, identifier.value] if identifier
|
|
128
139
|
|
|
129
140
|
name = molecule.names.first
|
|
130
|
-
return [
|
|
141
|
+
return ['name', name.content] if name
|
|
131
142
|
|
|
132
143
|
nil
|
|
133
144
|
end
|
|
134
145
|
|
|
135
|
-
private
|
|
136
|
-
|
|
137
146
|
def title_base(substance)
|
|
138
|
-
substance.preferred_name || substance.identifier_value(
|
|
139
|
-
substance.identifier_value(
|
|
147
|
+
substance.preferred_name || substance.identifier_value('cas') ||
|
|
148
|
+
substance.identifier_value('inchikey') || 'Substance'
|
|
140
149
|
end
|
|
141
150
|
|
|
142
151
|
def accessed_on(provenance)
|
|
143
152
|
return provenance.retrieved_at[0, 10] if provenance.retrieved_at
|
|
144
153
|
|
|
145
|
-
Time.now.utc.strftime(
|
|
154
|
+
Time.now.utc.strftime('%Y-%m-%d')
|
|
155
|
+
end
|
|
156
|
+
end
|
|
157
|
+
|
|
158
|
+
# The relaton-bib version seam. Both major lines accept the same
|
|
159
|
+
# field hash (see Citation#fields) and serialize through their own
|
|
160
|
+
# API; the rest of the citation track stays version-agnostic.
|
|
161
|
+
# Adding a future major = one more module here (OCP).
|
|
162
|
+
module RelatonApi
|
|
163
|
+
module_function
|
|
164
|
+
|
|
165
|
+
def dataset_bibitem(fields)
|
|
166
|
+
(defined?(::Relaton::Bib) ? V2 : V1).build(fields)
|
|
167
|
+
end
|
|
168
|
+
|
|
169
|
+
def to_xml(item)
|
|
170
|
+
item.to_xml
|
|
171
|
+
end
|
|
172
|
+
|
|
173
|
+
# relaton-bib 1: RelatonBib::* with hash-argument constructors.
|
|
174
|
+
module V1
|
|
175
|
+
module_function
|
|
176
|
+
|
|
177
|
+
def build(fields)
|
|
178
|
+
RelatonBib::BibliographicItem.new(
|
|
179
|
+
type: fields[:type],
|
|
180
|
+
title: [{ type: 'main', content: fields[:title] }],
|
|
181
|
+
docid: [RelatonBib::DocumentIdentifier.new(id: fields[:docid][:id],
|
|
182
|
+
type: fields[:docid][:type])],
|
|
183
|
+
contributor: [{ entity: RelatonBib::Organization.new(name: fields[:publisher]),
|
|
184
|
+
role: [{ type: 'publisher' }] }],
|
|
185
|
+
date: [{ type: 'accessed', on: fields[:accessed_on] }],
|
|
186
|
+
link: fields[:link] ? [{ type: 'src', content: fields[:link] }] : [],
|
|
187
|
+
keyword: fields[:keywords]
|
|
188
|
+
)
|
|
189
|
+
end
|
|
190
|
+
end
|
|
191
|
+
|
|
192
|
+
# relaton-bib 2: Relaton::Bib::* typed models (lutaml-model).
|
|
193
|
+
# Date's XML <on> element maps to the Ruby `at` attribute;
|
|
194
|
+
# keywords carry their text in a nested vocab LocalizedString;
|
|
195
|
+
# links are source Uri entries serializing to <uri type="src">.
|
|
196
|
+
module V2
|
|
197
|
+
module_function
|
|
198
|
+
|
|
199
|
+
def build(fields)
|
|
200
|
+
Relaton::Bib::ItemData.new(
|
|
201
|
+
type: fields[:type],
|
|
202
|
+
title: [Relaton::Bib::Title.new(type: 'main', content: fields[:title])],
|
|
203
|
+
docidentifier: [docidentifier(fields[:docid])],
|
|
204
|
+
contributor: [contributor(fields[:publisher])],
|
|
205
|
+
date: [Relaton::Bib::Date.new(type: 'accessed', at: fields[:accessed_on])],
|
|
206
|
+
source: fields[:link] ? [Relaton::Bib::Uri.new(type: 'src', content: fields[:link])] : [],
|
|
207
|
+
keyword: fields[:keywords].map { |text| keyword(text) }
|
|
208
|
+
)
|
|
209
|
+
end
|
|
210
|
+
|
|
211
|
+
def docidentifier(docid)
|
|
212
|
+
Relaton::Bib::Docidentifier.new(type: docid[:type], content: docid[:id])
|
|
213
|
+
end
|
|
214
|
+
|
|
215
|
+
def contributor(publisher)
|
|
216
|
+
Relaton::Bib::Contributor.new(
|
|
217
|
+
organization: Relaton::Bib::Organization.new(
|
|
218
|
+
name: [Relaton::Bib::TypedLocalizedString.new(content: publisher)]
|
|
219
|
+
),
|
|
220
|
+
role: [Relaton::Bib::Contributor::Role.new(type: 'publisher')]
|
|
221
|
+
)
|
|
222
|
+
end
|
|
223
|
+
|
|
224
|
+
def keyword(text)
|
|
225
|
+
Relaton::Bib::Keyword.new(
|
|
226
|
+
vocab: Relaton::Bib::LocalizedString.new(content: text)
|
|
227
|
+
)
|
|
228
|
+
end
|
|
146
229
|
end
|
|
147
230
|
end
|
|
231
|
+
private_constant :RelatonApi
|
|
148
232
|
end
|
|
149
233
|
end
|
data/lib/asciichem/cli.rb
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
|
-
require
|
|
3
|
+
require 'thor'
|
|
4
4
|
|
|
5
5
|
module AsciiChem
|
|
6
6
|
# Thor-based command line interface. Invoked via the `asciichem`
|
|
@@ -8,25 +8,30 @@ module AsciiChem
|
|
|
8
8
|
class Cli < Thor
|
|
9
9
|
# Use lowercase 'asciichem' as the program name in help output
|
|
10
10
|
# and command banners, matching the executable name.
|
|
11
|
-
package_name
|
|
11
|
+
package_name 'asciichem'
|
|
12
12
|
|
|
13
|
-
desc
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
13
|
+
desc 'convert -i INPUT -t FORMAT',
|
|
14
|
+
'Convert INPUT to FORMAT (mathml|text|html|latex|svg|structural-svg|model-json|cml|smiles|molfile)'
|
|
15
|
+
method_option :input, aliases: '-i', type: :string,
|
|
16
|
+
desc: "Source text (or '-' for stdin)"
|
|
17
|
+
method_option :file, aliases: '-f', type: :string,
|
|
18
|
+
desc: 'Read source from a file'
|
|
19
|
+
method_option :from, type: :string, default: 'asciichem',
|
|
20
|
+
desc: 'Input grammar: asciichem|smiles|molfile'
|
|
21
|
+
method_option :format, aliases: '-t', type: :string, default: 'mathml',
|
|
22
|
+
desc: 'Output format'
|
|
23
|
+
method_option :engine, type: :string, default: 'parslet',
|
|
24
|
+
desc: 'Parsing engine: parslet (default) | parsanol (opt-in, needs the parsanol gem)'
|
|
22
25
|
def convert
|
|
23
|
-
unless options[
|
|
24
|
-
raise AsciiChem::ParseError, "provide -i INPUT or -f FILE"
|
|
25
|
-
end
|
|
26
|
+
raise AsciiChem::ParseError, 'provide -i INPUT or -f FILE' unless options['input'] || options['file']
|
|
26
27
|
|
|
28
|
+
select_engine(options[:engine])
|
|
27
29
|
source = read_source
|
|
28
30
|
formula = ingest(source, options[:from])
|
|
29
31
|
puts render(formula, options[:format])
|
|
32
|
+
rescue AsciiChem::Engine::Error => e
|
|
33
|
+
warn "Engine error: #{e.message}"
|
|
34
|
+
exit 6
|
|
30
35
|
rescue AsciiChem::ParseError => e
|
|
31
36
|
warn "Parse error: #{e.message}"
|
|
32
37
|
exit 1
|
|
@@ -35,9 +40,9 @@ module AsciiChem
|
|
|
35
40
|
exit 2
|
|
36
41
|
end
|
|
37
42
|
|
|
38
|
-
desc
|
|
39
|
-
method_option :input, aliases:
|
|
40
|
-
|
|
43
|
+
desc 'parse-cml -i INPUT', 'Parse CML XML and emit AsciiChem text'
|
|
44
|
+
method_option :input, aliases: '-i', type: :string, required: true,
|
|
45
|
+
desc: 'CML XML source'
|
|
41
46
|
def parse_cml
|
|
42
47
|
formula = AsciiChem::Cml.parse(options[:input])
|
|
43
48
|
puts formula.to_text
|
|
@@ -46,8 +51,8 @@ module AsciiChem
|
|
|
46
51
|
exit 1
|
|
47
52
|
end
|
|
48
53
|
|
|
49
|
-
desc
|
|
50
|
-
method_option :input, aliases:
|
|
54
|
+
desc 'roundtrip -i INPUT', 'Parse and re-emit; exit non-zero if not equal'
|
|
55
|
+
method_option :input, aliases: '-i', type: :string, required: true
|
|
51
56
|
def roundtrip
|
|
52
57
|
original = options[:input]
|
|
53
58
|
rendered = AsciiChem.parse(original).to_text
|
|
@@ -60,11 +65,11 @@ module AsciiChem
|
|
|
60
65
|
end
|
|
61
66
|
end
|
|
62
67
|
|
|
63
|
-
desc
|
|
64
|
-
method_option :input, aliases:
|
|
65
|
-
|
|
66
|
-
method_option :format, aliases:
|
|
67
|
-
|
|
68
|
+
desc 'lint -i INPUT', 'Run chemistry checks; exit 1 on error, 0 if clean'
|
|
69
|
+
method_option :input, aliases: '-i', type: :string, required: true,
|
|
70
|
+
desc: 'AsciiChem source text'
|
|
71
|
+
method_option :format, aliases: '-f', type: :string, default: 'text',
|
|
72
|
+
desc: 'Output format: text or json'
|
|
68
73
|
def lint
|
|
69
74
|
formula = AsciiChem.parse(options[:input])
|
|
70
75
|
diagnostics = AsciiChem::Linter.run(formula)
|
|
@@ -76,7 +81,7 @@ module AsciiChem
|
|
|
76
81
|
end
|
|
77
82
|
|
|
78
83
|
map %w[--version -v] => :version
|
|
79
|
-
desc
|
|
84
|
+
desc 'version', 'Print the asciichem version'
|
|
80
85
|
def version
|
|
81
86
|
puts "asciichem #{AsciiChem::VERSION}"
|
|
82
87
|
end
|
|
@@ -86,32 +91,32 @@ module AsciiChem
|
|
|
86
91
|
"asciichem #{command.usage}"
|
|
87
92
|
end
|
|
88
93
|
|
|
89
|
-
desc
|
|
90
|
-
method_option :cas, type: :string, desc:
|
|
91
|
-
method_option :name, type: :string, desc:
|
|
92
|
-
method_option :cid, type: :string, desc:
|
|
93
|
-
method_option :inchikey, type: :string, desc:
|
|
94
|
-
method_option :smiles, type: :string, desc:
|
|
95
|
-
method_option :source, type: :string, default:
|
|
96
|
-
method_option :refresh, type: :boolean, default: false, desc:
|
|
97
|
-
method_option :format, aliases:
|
|
98
|
-
|
|
94
|
+
desc 'resolve --cas X | --name X | ...', 'Resolve a substance from a source (network; cached)'
|
|
95
|
+
method_option :cas, type: :string, desc: 'CAS registry number'
|
|
96
|
+
method_option :name, type: :string, desc: 'Substance name'
|
|
97
|
+
method_option :cid, type: :string, desc: 'PubChem CID'
|
|
98
|
+
method_option :inchikey, type: :string, desc: 'InChIKey'
|
|
99
|
+
method_option :smiles, type: :string, desc: 'SMILES'
|
|
100
|
+
method_option :source, type: :string, default: 'pubchem', desc: 'Resolver source'
|
|
101
|
+
method_option :refresh, type: :boolean, default: false, desc: 'Bypass the cache'
|
|
102
|
+
method_option :format, aliases: '-t', type: :string, default: 'model-json',
|
|
103
|
+
desc: 'Output: model-json | text | smiles'
|
|
99
104
|
def resolve
|
|
100
105
|
convention, value = %i[cas name cid inchikey smiles]
|
|
101
106
|
.filter_map { |k| [k, options[k.to_s]] if options[k.to_s] }
|
|
102
107
|
.first
|
|
103
|
-
raise AsciiChem::Error,
|
|
108
|
+
raise AsciiChem::Error, 'give one of --cas/--name/--cid/--inchikey/--smiles' unless value
|
|
104
109
|
|
|
105
|
-
convention = { cas:
|
|
106
|
-
inchikey:
|
|
110
|
+
convention = { cas: 'cas', name: 'name', cid: 'pubchem-cid',
|
|
111
|
+
inchikey: 'inchikey', smiles: 'smiles' }.fetch(convention)
|
|
107
112
|
substance = AsciiChem::Resolver[options[:source]].new.resolve(
|
|
108
113
|
value: value, convention: convention, refresh: options[:refresh]
|
|
109
114
|
)
|
|
110
115
|
raise AsciiChem::Error, "#{options[:source]} does not know #{value.inspect}" unless substance
|
|
111
116
|
|
|
112
117
|
puts case options[:format].to_s
|
|
113
|
-
when
|
|
114
|
-
when
|
|
118
|
+
when 'text' then substance.preferred_name.to_s
|
|
119
|
+
when 'smiles' then substance.identifier_value('canonical-smiles').to_s
|
|
115
120
|
else substance.to_model_json
|
|
116
121
|
end
|
|
117
122
|
rescue AsciiChem::Error => e
|
|
@@ -119,22 +124,22 @@ module AsciiChem
|
|
|
119
124
|
exit 3
|
|
120
125
|
end
|
|
121
126
|
|
|
122
|
-
desc
|
|
123
|
-
method_option :cas, type: :string, desc:
|
|
124
|
-
method_option :name, type: :string, desc:
|
|
125
|
-
method_option :cid, type: :string, desc:
|
|
126
|
-
method_option :inchikey, type: :string, desc:
|
|
127
|
-
method_option :smiles, type: :string, desc:
|
|
128
|
-
method_option :source, type: :string, default:
|
|
129
|
-
method_option :refresh, type: :boolean, default: false, desc:
|
|
127
|
+
desc 'cite --cas X | --name X | ...', 'Resolve a substance and emit a dataset-type Relaton bibitem (XML)'
|
|
128
|
+
method_option :cas, type: :string, desc: 'CAS registry number'
|
|
129
|
+
method_option :name, type: :string, desc: 'Substance name'
|
|
130
|
+
method_option :cid, type: :string, desc: 'PubChem CID'
|
|
131
|
+
method_option :inchikey, type: :string, desc: 'InChIKey'
|
|
132
|
+
method_option :smiles, type: :string, desc: 'SMILES'
|
|
133
|
+
method_option :source, type: :string, default: 'pubchem', desc: 'Resolver source'
|
|
134
|
+
method_option :refresh, type: :boolean, default: false, desc: 'Bypass the cache'
|
|
130
135
|
def cite
|
|
131
136
|
convention, value = %i[cas name cid inchikey smiles]
|
|
132
137
|
.filter_map { |k| [k, options[k.to_s]] if options[k.to_s] }
|
|
133
138
|
.first
|
|
134
|
-
raise AsciiChem::Error,
|
|
139
|
+
raise AsciiChem::Error, 'give one of --cas/--name/--cid/--inchikey/--smiles' unless value
|
|
135
140
|
|
|
136
|
-
convention = { cas:
|
|
137
|
-
inchikey:
|
|
141
|
+
convention = { cas: 'cas', name: 'name', cid: 'pubchem-cid',
|
|
142
|
+
inchikey: 'inchikey', smiles: 'smiles' }.fetch(convention)
|
|
138
143
|
substance = AsciiChem::Resolver[options[:source]].new.resolve(
|
|
139
144
|
value: value, convention: convention, refresh: options[:refresh]
|
|
140
145
|
)
|
|
@@ -146,37 +151,43 @@ module AsciiChem
|
|
|
146
151
|
exit 4
|
|
147
152
|
end
|
|
148
153
|
|
|
149
|
-
desc
|
|
150
|
-
method_option :input, aliases:
|
|
154
|
+
desc 'validate -i INPUT', 'Offline identifier validation'
|
|
155
|
+
method_option :input, aliases: '-i', type: :string, required: true
|
|
151
156
|
def validate
|
|
152
157
|
formula = AsciiChem.parse(options[:input])
|
|
153
158
|
annotations = formula.nodes.grep(AsciiChem::Model::Molecule).flat_map(&:identifiers)
|
|
154
159
|
if annotations.empty?
|
|
155
|
-
puts
|
|
160
|
+
puts 'no identifier annotations found'
|
|
156
161
|
return
|
|
157
162
|
end
|
|
158
163
|
annotations.each do |identifier|
|
|
159
164
|
known = AsciiChem::Identifiers.known?(identifier.convention)
|
|
160
165
|
valid = known && AsciiChem::Identifiers.valid?(identifier.convention, identifier.value)
|
|
161
|
-
status = known
|
|
162
|
-
|
|
166
|
+
status = if known
|
|
167
|
+
valid ? 'ok' : 'INVALID'
|
|
168
|
+
else
|
|
169
|
+
'unknown convention'
|
|
170
|
+
end
|
|
171
|
+
puts format('%-12s %-40s %s', identifier.convention, identifier.value, status)
|
|
172
|
+
end
|
|
173
|
+
exit 1 if annotations.any? do |i|
|
|
174
|
+
AsciiChem::Identifiers.known?(i.convention) &&
|
|
175
|
+
!AsciiChem::Identifiers.valid?(i.convention, i.value)
|
|
163
176
|
end
|
|
164
|
-
exit 1 if annotations.any? { |i| AsciiChem::Identifiers.known?(i.convention) &&
|
|
165
|
-
!AsciiChem::Identifiers.valid?(i.convention, i.value) }
|
|
166
177
|
rescue AsciiChem::ParseError => e
|
|
167
178
|
warn "Parse error: #{e.message}"
|
|
168
179
|
exit 1
|
|
169
180
|
end
|
|
170
181
|
|
|
171
|
-
desc
|
|
172
|
-
method_option :input, aliases:
|
|
173
|
-
|
|
174
|
-
method_option :file, aliases:
|
|
175
|
-
|
|
176
|
-
method_option :from, type: :string, default:
|
|
177
|
-
|
|
182
|
+
desc 'identity -i INPUT', 'Derive InChI/InChIKey locally from the structure (offline; requires an InChI engine)'
|
|
183
|
+
method_option :input, aliases: '-i', type: :string,
|
|
184
|
+
desc: "Source text (or '-' for stdin)"
|
|
185
|
+
method_option :file, aliases: '-f', type: :string,
|
|
186
|
+
desc: 'Read source from a file'
|
|
187
|
+
method_option :from, type: :string, default: 'asciichem',
|
|
188
|
+
desc: 'Input grammar: asciichem|smiles|molfile'
|
|
178
189
|
method_option :engine_bin, type: :string,
|
|
179
|
-
desc:
|
|
190
|
+
desc: 'Path to the inchi-1 binary (overrides the configured engine)'
|
|
180
191
|
def identity
|
|
181
192
|
formula = ingest(read_source, options[:from])
|
|
182
193
|
molecules = molecules_in(formula)
|
|
@@ -195,9 +206,18 @@ module AsciiChem
|
|
|
195
206
|
|
|
196
207
|
private
|
|
197
208
|
|
|
209
|
+
# --engine feeds the Engine selector (0.29.0+); :parsanol is an
|
|
210
|
+
# opt-in soft dependency, :parslet stays the default.
|
|
211
|
+
def select_engine(name)
|
|
212
|
+
return if name.to_s == 'parslet' && AsciiChem::Engine.current == AsciiChem::Engine::ParsletEngine
|
|
213
|
+
|
|
214
|
+
require 'asciichem/engine'
|
|
215
|
+
AsciiChem::Engine.use(name.to_s)
|
|
216
|
+
end
|
|
217
|
+
|
|
198
218
|
def read_source
|
|
199
219
|
return File.read(options[:file]) if options[:file]
|
|
200
|
-
return $stdin.read if options[:input] ==
|
|
220
|
+
return $stdin.read if options[:input] == '-'
|
|
201
221
|
|
|
202
222
|
options[:input]
|
|
203
223
|
end
|
|
@@ -207,9 +227,9 @@ module AsciiChem
|
|
|
207
227
|
# format works regardless of the input language.
|
|
208
228
|
def ingest(source, from)
|
|
209
229
|
case from.to_s
|
|
210
|
-
when
|
|
211
|
-
when
|
|
212
|
-
when
|
|
230
|
+
when 'asciichem' then AsciiChem.parse(source)
|
|
231
|
+
when 'smiles' then AsciiChem.parse_smiles(source)
|
|
232
|
+
when 'molfile' then molfile_formula(source)
|
|
213
233
|
else
|
|
214
234
|
raise AsciiChem::ParseError, "unknown --from grammar: #{from}"
|
|
215
235
|
end
|
|
@@ -224,7 +244,7 @@ module AsciiChem
|
|
|
224
244
|
|
|
225
245
|
def render(formula, format)
|
|
226
246
|
return formula.to_cml if format.to_sym == :cml
|
|
227
|
-
return formula.to_model_json if format.to_sym == :
|
|
247
|
+
return formula.to_model_json if format.to_sym == :'model-json'
|
|
228
248
|
return formula.to_smiles if format.to_sym == :smiles
|
|
229
249
|
return formula.nodes.first.to_molfile if format.to_sym == :molfile
|
|
230
250
|
|
|
@@ -246,14 +266,14 @@ module AsciiChem
|
|
|
246
266
|
|
|
247
267
|
def output_lint(diagnostics, format)
|
|
248
268
|
case format.to_s
|
|
249
|
-
when
|
|
269
|
+
when 'json' then output_lint_json(diagnostics)
|
|
250
270
|
else
|
|
251
271
|
diagnostics.each { |d| puts d }
|
|
252
272
|
end
|
|
253
273
|
end
|
|
254
274
|
|
|
255
275
|
def output_lint_json(diagnostics)
|
|
256
|
-
require
|
|
276
|
+
require 'json'
|
|
257
277
|
payload = diagnostics.map do |d|
|
|
258
278
|
{
|
|
259
279
|
severity: d.severity.to_s,
|
data/lib/asciichem/version.rb
CHANGED
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: asciichem
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.29.
|
|
4
|
+
version: 0.29.2
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Ribose Inc.
|
|
@@ -108,7 +108,7 @@ dependencies:
|
|
|
108
108
|
version: '0.1'
|
|
109
109
|
- - "<"
|
|
110
110
|
- !ruby/object:Gem::Version
|
|
111
|
-
version: '
|
|
111
|
+
version: '3'
|
|
112
112
|
type: :runtime
|
|
113
113
|
prerelease: false
|
|
114
114
|
version_requirements: !ruby/object:Gem::Requirement
|
|
@@ -118,7 +118,7 @@ dependencies:
|
|
|
118
118
|
version: '0.1'
|
|
119
119
|
- - "<"
|
|
120
120
|
- !ruby/object:Gem::Version
|
|
121
|
-
version: '
|
|
121
|
+
version: '3'
|
|
122
122
|
- !ruby/object:Gem::Dependency
|
|
123
123
|
name: plurimath
|
|
124
124
|
requirement: !ruby/object:Gem::Requirement
|