relaton-cli 3.0.0.pre.alpha.1 → 3.0.0.pre.alpha.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.gitignore +4 -0
- data/CLAUDE.md +403 -0
- data/Gemfile +5 -0
- data/Rakefile +15 -0
- data/docs/README.adoc +177 -1
- data/frontend/dist/404.html +18 -0
- data/frontend/dist/app.iife.js +1 -0
- data/frontend/dist/style.css +3 -0
- data/lib/relaton/cli/command.rb +116 -0
- data/lib/relaton/cli/frontend_assets.rb +58 -0
- data/lib/relaton/cli/index_item_normalizer.rb +380 -0
- data/lib/relaton/cli/index_site_generator.rb +860 -0
- data/lib/relaton/cli/subcommand_collection.rb +3 -0
- data/lib/relaton-cli.rb +5 -0
- data/relaton-cli.gemspec +13 -3
- data/templates/index/page.liquid +20 -0
- metadata +27 -9
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 40a33fcc5e9cb123ea61e9d73105524c896a8e79601431980b52589edb6123fd
|
|
4
|
+
data.tar.gz: 72cee8178b5ea4ba8a628f77ed04a9efa05f7d1890d6c0028a72d18eeb8939b6
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 046773f651afaf06a182b5217c8054b32bc2d34f4d4aedc68283741791347fc41b29603cad1d45fae318c0a961a7bc72cd3242b93785241c3cb56ead3436d0aa
|
|
7
|
+
data.tar.gz: 4fe86cad3ad78c8f317c26686f20bcb75780cac66ecac56cd0547a31b0aa3084a8e33d6fb129d9f9f1ddb47708bc27ed8962792addd900ca2eac6b194bd499f3
|
data/.gitignore
CHANGED
data/CLAUDE.md
CHANGED
|
@@ -49,6 +49,32 @@ The executable `exe/relaton` calls `Relaton::Cli.start(ARGV)` which routes to `R
|
|
|
49
49
|
|
|
50
50
|
Current fetch options that use this pattern: `--no-cache`, `--all-parts`, `--keep-year`, `--publication-date-before`, `--publication-date-after`.
|
|
51
51
|
|
|
52
|
+
### Malformed-identifier handling (CLI is the friendly layer)
|
|
53
|
+
|
|
54
|
+
The `relaton` library **raises** `Pubid::Errors::ParseError` when a reference
|
|
55
|
+
can't be parsed — it does not swallow it, so API consumers of `Relaton::Db`
|
|
56
|
+
handle it themselves. Every flavor raises that class, including ITU, whose own
|
|
57
|
+
Parslet grammar (`Relaton::Itu::Pubid`) re-raises its failure as one. The
|
|
58
|
+
**CLI** is the single place that turns it into a user-facing message:
|
|
59
|
+
`fetch_document` rescues the marker module `Pubid::Errors::Error` — which also
|
|
60
|
+
covers `Pubid::Errors::InvalidInputError` (input too long) — (alongside
|
|
61
|
+
`Relaton::RequestError`) and returns `"<code>" is not a recognized
|
|
62
|
+
standards identifier`, and `SubcommandCollection#fetch` — which bypasses
|
|
63
|
+
`fetch_document` — rescues it too, emitting the same message via `Util.error`.
|
|
64
|
+
The paths that actually reach pubid's parser are `relaton fetch` (→
|
|
65
|
+
`Db#fetch`/`#fetch_std` → `processor.get`) and `relaton collection fetch` (→
|
|
66
|
+
`Db#fetch`). `relaton db fetch` is a **cache-only** lookup (`fetch_db: true`
|
|
67
|
+
short-circuits `check_bibliocache` to a cache-key read via `std_id`, which is
|
|
68
|
+
plain string manipulation, never `processor.get`), so it never parses and can't
|
|
69
|
+
raise a parse error — a malformed id there simply misses the cache and
|
|
70
|
+
prints "No matching bibliographic entry found". `command.rb` and
|
|
71
|
+
`subcommand_collection.rb` therefore `require "pubid"` at load time so the
|
|
72
|
+
constant resolves even when an unrelated exception (e.g. an `ArgumentError` from
|
|
73
|
+
`parse_date_option`) is the one propagating through the rescue chain:
|
|
74
|
+
`require "relaton"` loads pubid lazily, so it may not be loaded yet. `Pubid::Errors`
|
|
75
|
+
first shipped in pubid `2.0.0.pre.alpha.11`, which is why `relaton.gemspec`
|
|
76
|
+
pins `~> 2.0.0.pre.alpha.11`.
|
|
77
|
+
|
|
52
78
|
### Core Data Classes
|
|
53
79
|
|
|
54
80
|
- `lib/relaton/bibdata.rb` — `Relaton::Bibdata` wraps `RelatonBib::BibliographicItem`, adding URL type handling and serialization to XML/YAML/Hash. Uses `method_missing` to delegate to the underlying bibitem.
|
|
@@ -62,6 +88,383 @@ Current fetch options that use this pattern: `--no-cache`, `--all-parts`, `--kee
|
|
|
62
88
|
- `lib/relaton/cli/yaml_convertor.rb` — YAML → XML conversion (includes processor detection via doctype)
|
|
63
89
|
- `lib/relaton/cli/xml_to_html_renderer.rb` — Renders XML/YAML to HTML using Liquid templates from `templates/`
|
|
64
90
|
|
|
91
|
+
### Data index site generator (`relaton index`)
|
|
92
|
+
|
|
93
|
+
`relaton index [DATA-DIR]` builds the modern, browsable GitHub Pages index for
|
|
94
|
+
`relaton-data-<flavor>` repos (e.g. https://relaton.github.io/relaton-data-bipm/),
|
|
95
|
+
replacing the shared Jekyll theme (`relaton/jekyll-theme-relaton-data-index`) +
|
|
96
|
+
`relaton/support` `data-deploy.yml` build. Pieces:
|
|
97
|
+
|
|
98
|
+
**One delivery mode, sharded.** The shell carries **no document data** — only
|
|
99
|
+
branding, the inlined bundle, and five `data-*` scalars describing the shard
|
|
100
|
+
layout — so page weight is a function of what the reader opens, not of corpus
|
|
101
|
+
size. Everything else is fetched:
|
|
102
|
+
|
|
103
|
+
```
|
|
104
|
+
_site/index.html ~110 KB regardless of corpus size
|
|
105
|
+
_site/search-0000.json … summary records {r,c,t,s,d,u,l}, `--shard-size` each
|
|
106
|
+
_site/detail-0000.json … the rich fields, `--detail-shard-size` each
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
Two scope decisions are baked in here; both were deliberate, and re-litigating
|
|
110
|
+
either without reading this will lead you back to a 150 MB page:
|
|
111
|
+
|
|
112
|
+
1. **Crawler-indexability was dropped as a requirement.** The old `embedded`/`dom`
|
|
113
|
+
modes rendered one `.document` row per document precisely so crawlers could
|
|
114
|
+
read the corpus; that is what made the page O(corpus). relaton/relaton#48
|
|
115
|
+
listed it as an acceptance criterion, and it was given up knowingly — these
|
|
116
|
+
are machine-data repos whose consumers are relaton fetchers, and the Pages
|
|
117
|
+
site is a "what data is available" browser reached from the repo. Restoring it
|
|
118
|
+
means paginated `page/N/index.html` shells, not un-sharding the data.
|
|
119
|
+
2. **Therefore the three `--mode`s collapsed to one.** `dom` existed *only* to be
|
|
120
|
+
crawlable, and as a data channel `data-*` attributes are strictly worse than
|
|
121
|
+
JSON (bigger, and lossy — 7 fields, no detail). `embedded`'s remaining
|
|
122
|
+
advantage was `file://` self-containment, which a fetched sidecar had already
|
|
123
|
+
given up. `static-json` was just "sharded, with one shard". `--mode` is
|
|
124
|
+
**removed**, not deprecated; see the note at the end of this section.
|
|
125
|
+
|
|
126
|
+
- `lib/relaton/cli/index_site_generator.rb` — scans `DATA-DIR/**/*.{yaml,yml}`
|
|
127
|
+
(default `./data`, skips `index-v*.yaml`) **plus an auto-detected sibling
|
|
128
|
+
`static/` folder** (`STATIC_DIRNAME`, a sibling of `DATA-DIR` = `repo_root/static`;
|
|
129
|
+
`--no-static`/`static: false` opts out), normalizes each doc, and writes the
|
|
130
|
+
shell plus the shards into `--output` (default `_site`).
|
|
131
|
+
Both source dirs go through the same `SKIP_BASENAMES`/`document?` filters. De-dup
|
|
132
|
+
is **cross-folder only**: `data/` is scanned first, so a `static/` doc whose
|
|
133
|
+
normalized id already appeared in `data/` is skipped with a warning (data wins);
|
|
134
|
+
duplicates *within* a single folder are left as-is (preserving the pre-static
|
|
135
|
+
data-scan behavior), and a blank/empty normalized id never de-dups (so docid-less
|
|
136
|
+
title-only docs aren't collapsed). Static docs get correct `static/<path>`
|
|
137
|
+
raw-YAML links for free because `relative_path` is taken from `repo_root` (the
|
|
138
|
+
parent of both `data/` and `static/`). `static/` holds manually-curated bib
|
|
139
|
+
records the crawler can't fetch (ISO/IEC Directives, JCGM/GUM guides, NIST
|
|
140
|
+
research-library metadata) that are already part of the corpus (in `index-vN.yaml`).
|
|
141
|
+
**`each_document` is the streaming spine**: it yields normalized items instead
|
|
142
|
+
of accumulating them, and `ShardWriter` buffers at most one shard of each
|
|
143
|
+
family, so for the **search/detail** shard families peak RSS is O(shard size) —
|
|
144
|
+
measured at ~38 MB above baseline for a 20,000-document, 79 MB corpus. Nothing
|
|
145
|
+
may reintroduce a `collect_documents` that materializes the corpus for those.
|
|
146
|
+
The spec that guards this observes the property rather than asserting a method
|
|
147
|
+
is never called (that passes just as happily once the method is renamed): it
|
|
148
|
+
pins that a completed shard is on disk while the corpus is still being read.
|
|
149
|
+
**`MachineIndex` is the deliberate exception to the O(shard size) bound.**
|
|
150
|
+
Shard assignment needs the corpus size, so it is not knowable until the pass
|
|
151
|
+
ends — the class therefore retains one
|
|
152
|
+
`Row` per document. That is O(corpus), and fine: ~0.44 KB/row, i.e. ~76 MB at
|
|
153
|
+
177k rows. What it must NOT do is retain the pubid object, which costs
|
|
154
|
+
3.14 KB/row — 543 MB on the same corpus, past the ~250 MB where the process
|
|
155
|
+
was observed to hang. `Row` holds only `rendered`/`file`/`key`/`id_hash`, all
|
|
156
|
+
derived inside `add`; anything else needed from pubid must likewise be
|
|
157
|
+
computed there, because by write time the identifier is gone.
|
|
158
|
+
|
|
159
|
+
**The machine-index wire contract.** `relaton index` also publishes a
|
|
160
|
+
machine-consumable index — `index-vN.yaml`/`.zip`, `index/manifest.json` and
|
|
161
|
+
`index/shard-NNNNN.json`. It is specified in `docs/data-repository-format.adoc`
|
|
162
|
+
in the `relaton` repo; that file is the contract, this is why it is shaped that
|
|
163
|
+
way. Four decisions, all settled deliberately — don't re-derive them:
|
|
164
|
+
|
|
165
|
+
- **The name comes from the flavor's own `INDEXFILE`, never from the data.**
|
|
166
|
+
`index-vN` is *per-flavor structure versioning*, not a global generation
|
|
167
|
+
counter: 25 flavors read `index-v2`, IHO reads `index-v3`, JCGM reads
|
|
168
|
+
`index-v1` — and JCGM's v1 is structured, so the number is a name, not a
|
|
169
|
+
claim about the rows. An earlier revision derived it from a parse ratio
|
|
170
|
+
(`structured? ? "v3" : "v1"`), which both re-labelled v2 as v3 and let a crawl
|
|
171
|
+
that shifted the ratio by a few percent *rename the published file*, breaking
|
|
172
|
+
every pinned consumer with nothing but a 404. `--index-name` overrides, for a
|
|
173
|
+
corpus that is not a relaton flavor. Resolution is `Relaton::<Flavor>::INDEXFILE`
|
|
174
|
+
via `FLAVOR_NAMESPACES`, whose only entry is 3GPP (`Pubid::Tgpp` against
|
|
175
|
+
`Relaton::ThreeGpp`); every other flavor capitalizes identically in both gems.
|
|
176
|
+
- **The shard key is `crc32(pubid.root.number.to_s) % N`, with no fallback.**
|
|
177
|
+
That expression is the one `Relaton::Index` sorts and bsearches on (`Type#
|
|
178
|
+
candidates_by_number`, `FileIO#deserialize_pubid`), and the identity is the
|
|
179
|
+
whole point: a client computing the key from its parsed query must land in the
|
|
180
|
+
bucket the rows are in. A rendered-id fallback for an empty root number
|
|
181
|
+
(added once to avoid a hot shard 0) breaks that — the client looks in shard 0,
|
|
182
|
+
the row is elsewhere, and the miss is indistinguishable from not-found because
|
|
183
|
+
empty shards are omitted. An identifier with no root number keys on `""` and
|
|
184
|
+
goes in shard 0, the same degeneracy the gem's bsearch already has.
|
|
185
|
+
- **Every published row is structured, and the shards and the monolith agree.**
|
|
186
|
+
`row_record` and `write_monolith` both carry the pubid hash for every row.
|
|
187
|
+
They cannot differ: `FileIO#deserialize_id` calls `from_hash` on every row and
|
|
188
|
+
rejects the *whole* index on the first one it cannot deserialize, so one
|
|
189
|
+
plain-string row in the monolith would poison the file. That is also why
|
|
190
|
+
`--pubid-flavor` is **required** unless `--no-machine-index` — a machine index
|
|
191
|
+
with no parser could only be the string shape.
|
|
192
|
+
- **An id the parser rejects is taken from the repo's committed index.** The
|
|
193
|
+
generator derives ids by parsing each document's *rendered* docid, while the
|
|
194
|
+
committed `index-vN.yaml` was built by the flavor's `DataFetcher` from source
|
|
195
|
+
metadata; five shipping corpora disagree on a handful of rows (ieee 69,
|
|
196
|
+
itu-r 47, iec 42, itu 3, nist 3). `committed_index` reads that file once and
|
|
197
|
+
`add` uses its hash **verbatim**, so the site's row is byte-identical to the
|
|
198
|
+
repo's. Only a row neither path resolves is dropped, counted in
|
|
199
|
+
`skipped_count`, and warned about.
|
|
200
|
+
|
|
201
|
+
**The monolith's hand-rolled YAML writer must preserve type, not just text.**
|
|
202
|
+
`write_monolith` avoids Psych (40 min on 177k rows), so `yaml_nested`/`yaml_scalar`
|
|
203
|
+
own the whole round trip. Two shapes broke it, both invisible to specs that only
|
|
204
|
+
build `{_type, number}` hashes: an **Array** fell through to `yaml_scalar`, whose
|
|
205
|
+
`["IEC"].to_s` starts with `[` and was quoted into the String `"[\"IEC\"]"` —
|
|
206
|
+
every ISO/IEC copublished id carries `copublishers`, and `from_hash` cannot cast
|
|
207
|
+
it, so `FileIO` rejects the whole index; and a real **Boolean** was stringified
|
|
208
|
+
then quoted (`d_prefix: "true"`, 31 of CIE's 1139 rows), so the published row no
|
|
209
|
+
longer matched the repo's. Arrays are now written as a JSON flow sequence (valid
|
|
210
|
+
YAML, unambiguous escaping); booleans and integers go out bare. Verify a writer
|
|
211
|
+
change against a real corpus, comparing the site's monolith row-by-row to the
|
|
212
|
+
repo's committed index — that is how both were found.
|
|
213
|
+
|
|
214
|
+
**`purge_stale!` reads the previous manifest.** `STALE_GLOBS` only knows
|
|
215
|
+
`index-v*`, but `--index-name` accepts any single file name (`INDEX_NAME`; a path
|
|
216
|
+
segment or a leading dot is refused), so the previous build's `index/manifest.json`
|
|
217
|
+
is the only record of which monolith to delete.
|
|
218
|
+
|
|
219
|
+
**`Util.warn` is not the relaton logger — it is `Kernel#warn`.**
|
|
220
|
+
`Relaton::Cli::Util` gets `info`/`error` from `Relaton::Bib::Util#method_missing`,
|
|
221
|
+
but `warn` already exists as a private `Kernel` instance method, so the explicit
|
|
222
|
+
receiver makes Ruby route it to `method_missing` *after* a private-method
|
|
223
|
+
NoMethodError. The message does reach the logger, but `expect(Relaton::Cli::Util)
|
|
224
|
+
.to receive(:warn)` never intercepts it. Mock `Relaton.logger_pool` instead:
|
|
225
|
+
`expect(Relaton.logger_pool).to receive(:warn).with(/…/, "relaton-cli")`.
|
|
226
|
+
**`purge_stale!` is load-bearing, not hygiene**: with no manifest, a
|
|
227
|
+
`search-0034.json` left by a previous larger corpus is invisible (the shard
|
|
228
|
+
count on the mount node says 34), so it would sit in the deployed site forever.
|
|
229
|
+
**Per-flavor branding** beyond `--title` is
|
|
230
|
+
`--description` and `--favicon` (both `nil` by default — the per-flavor values the
|
|
231
|
+
deleted Jekyll `_config.yml`s carried, now supplied by the caller workflow). The
|
|
232
|
+
description becomes `<meta name="description">` and `data-description` on the
|
|
233
|
+
mount node; the favicon becomes `<link rel="icon">`
|
|
234
|
+
with its href passed through **verbatim** and a `type` sniffed from the extension
|
|
235
|
+
via `FAVICON_TYPES` (unknown extension → no `type`, the browser sniffs). The
|
|
236
|
+
generator copies **no assets** (only the shell + the shards), so a relative
|
|
237
|
+
favicon href is the caller's job to place in the deployed site — normally it's an
|
|
238
|
+
absolute URL. Both tags are emitted by `{%- if … %}` branches in `page.liquid`, so
|
|
239
|
+
a build with neither option is byte-identical to one that never had them —
|
|
240
|
+
`spec/index_fixtures/golden/index.html` pins the whole shell, so regenerate it
|
|
241
|
+
deliberately and read the diff when you touch that template.
|
|
242
|
+
`title`/`description`/`favicon` all go
|
|
243
|
+
through `presence`, which maps a blank string to `nil`: the deploy workflow
|
|
244
|
+
forwards unset inputs as `--favicon ""`, and an empty `href` would make the
|
|
245
|
+
browser re-fetch the page as its own icon.
|
|
246
|
+
- `lib/relaton/cli/index_item_normalizer.rb` — doc-hash → `{id,title,doctype,
|
|
247
|
+
stage,date,link,yaml}` (the always-present summary core) **plus optional
|
|
248
|
+
detail-page fields** (`abstract, edition, languages, keywords, publisher,
|
|
249
|
+
contributors, docids, dates, relations`) added by `details`. Reads the doc's own rendered
|
|
250
|
+
fields (primary `docidentifier`, main/en `title`, `date[].at|on` preferring
|
|
251
|
+
`published`, `ext.doctype`), so **no pubid reconstruction** is needed. Detail
|
|
252
|
+
fields are dropped when empty (via `reject`), so a summary-only doc keeps the
|
|
253
|
+
compact 7-key shape. The generator then splits that one record in two:
|
|
254
|
+
`compact_record` projects the seven summary keys (renamed `r,c,t,s,d,u,l`) into
|
|
255
|
+
the search shards, and `detail_record` takes **everything else**, tagged with
|
|
256
|
+
`r`. The split is defined by one constant, `COMPACT_KEYS` — a key that isn't in
|
|
257
|
+
it is a detail field by definition, so a new normalizer field lands in the
|
|
258
|
+
detail shards automatically and never silently bloats the summary. The frontend
|
|
259
|
+
`types.ts` `IndexDocument` / `DetailRecord` interfaces mirror both halves.
|
|
260
|
+
|
|
261
|
+
**`relations` carries only `{type, id}`, and the id must render exactly like a
|
|
262
|
+
document's own.** The whole point of the field is that the detail page can match
|
|
263
|
+
the target against the rest of the corpus and turn it into a `?doc=` link, so
|
|
264
|
+
`relation_id` shares the primary-DocID logic that computes each document's `id`
|
|
265
|
+
(a relation's nested `bibitem` carries the same `docidentifier`/`primary`
|
|
266
|
+
shape). Don't "improve" it into a separate id renderer — the two would drift and
|
|
267
|
+
every link would silently degrade to plain text. It deliberately shares
|
|
268
|
+
`primary_docidentifier` and **not** `docidentifier` itself: the latter bottoms
|
|
269
|
+
out at the item's `id`, which on a relation bibitem is an internal XML anchor
|
|
270
|
+
(`CENISO/TS21003-7-2008/A1-2010` — CEN records carry exactly this). An anchor can
|
|
271
|
+
never match a corpus document, so the fallback order is docidentifier →
|
|
272
|
+
`formattedref` → `docnumber` → anchor; putting the anchor any earlier strands a
|
|
273
|
+
target that *is* present as plain text under a mangled label. No title/date is
|
|
274
|
+
carried: measured on
|
|
275
|
+
`relaton-data-iso` (79,402 docs) that is 83,074 relations, of which **99.9%
|
|
276
|
+
resolve in-dataset**; the remaining 48 are genuinely-absent `ISO/DIS …` drafts.
|
|
277
|
+
Two guards are load-bearing on real data — **self-references are dropped**
|
|
278
|
+
(BIPM's Metrologia articles relate to themselves once per carrier) and
|
|
279
|
+
identical `{type, id}` pairs are **de-duplicated**.
|
|
280
|
+
|
|
281
|
+
**Detail lookup is positional, verified by id.** A detail shard is a *dense*
|
|
282
|
+
array: slot `i` is either `null` (that document has no detail fields) or
|
|
283
|
+
`{"r": "<id>", …}`. A `null` slot is still written, because the frontend finds a
|
|
284
|
+
document's detail by arithmetic — `shard = ⌊i / detailShardSize⌋`,
|
|
285
|
+
`slot = i % detailShardSize`, where `i` is its position in the corpus — and
|
|
286
|
+
dropping empties would shift everything after it. Nothing stores that mapping,
|
|
287
|
+
which is why `loadDetail` checks the entry's `r` against the document id and
|
|
288
|
+
falls back to scanning the shard on mismatch: showing another document's
|
|
289
|
+
abstract would be worse than showing none. Keep the two families written from
|
|
290
|
+
the same single pass over `each_document`, or the arithmetic stops holding.
|
|
291
|
+
|
|
292
|
+
**Corpus position is carried, never derived from the array index.** The shard
|
|
293
|
+
loader stamps `pos` onto each document from its shard number
|
|
294
|
+
(`i * shardSize + j`), and `App.vue` uses `doc.pos` — falling back to `indexOf`
|
|
295
|
+
only for a document list supplied whole. This is not redundancy. A summary
|
|
296
|
+
shard that fails to load leaves a *hole*: every later document then sits at an
|
|
297
|
+
array index `shardSize` lower than its true position, which at the defaults
|
|
298
|
+
(5000 / 500) addresses a detail shard **ten files early**. The `r` check would
|
|
299
|
+
reject the wrong record, but its fallback scans only the shard it was sent to,
|
|
300
|
+
so it can never recover a cross-shard drift — the failure would be silent
|
|
301
|
+
summary-only detail panels for the entire rest of the corpus. Don't "simplify"
|
|
302
|
+
`doc.pos` back to `indexOf`.
|
|
303
|
+
- `lib/relaton/cli/frontend_assets.rb` — reads the compiled bundle from
|
|
304
|
+
`frontend/dist/` (one relative path works for both gem and git checkout);
|
|
305
|
+
raises `BuildMissingError` with a "run `rake build_frontend`" message if absent.
|
|
306
|
+
`dist_dir` is overridable so specs point at a fixture dist (no Node build).
|
|
307
|
+
- `templates/index/page.liquid` — the whole output HTML: head/branding, the
|
|
308
|
+
inlined IIFE + CSS, and the mount node carrying the five shard scalars
|
|
309
|
+
(`data-total`, `data-shards`, `data-shard-size`, `data-detail-shards`,
|
|
310
|
+
`data-detail-shard-size`). Those attributes are the **only** description of the
|
|
311
|
+
index's shape — there is deliberately no manifest file, which saves a
|
|
312
|
+
round-trip before the first shard can be requested and removes a failure mode
|
|
313
|
+
where the manifest 404s but the shards are fine. Shard file names are derived by
|
|
314
|
+
convention on both sides (`search-%04d.json` / `detail-%04d.json`) and resolve
|
|
315
|
+
relative to the page, which sits beside them. Rendered via a Liquid
|
|
316
|
+
`Environment` (not the deprecated global `Template.file_system=`).
|
|
317
|
+
(`_document.liquid` is gone with the crawler DOM.)
|
|
318
|
+
- `--flavor <f>` / `--repo <org/name>` shallow-clone the data repo to a tempdir
|
|
319
|
+
and read its `data/`; the checked-out branch is auto-detected (bipm's is `v2`)
|
|
320
|
+
and used for the raw-YAML `--base-url`. Default (no flavor/repo) reads the local
|
|
321
|
+
folder — the in-Action path where the repo is already checked out.
|
|
322
|
+
|
|
323
|
+
**Frontend (`frontend/`).** A Vite 8 **library-mode** project
|
|
324
|
+
(`@vitejs/plugin-vue` + `@tailwindcss/vite` + Vue 3) that builds ONE
|
|
325
|
+
`frontend/dist/app.iife.js` + `style.css` (the lutaml-xsd embedding pattern,
|
|
326
|
+
modernized). `src/` is committed; `dist/` + `node_modules/` are gitignored and
|
|
327
|
+
built at package time. The island (`src/App.vue`) does search / doctype+stage
|
|
328
|
+
facets / sort / list-grid / dark-mode / copy-DocID / client pagination.
|
|
329
|
+
|
|
330
|
+
**Hydration is one synchronous path.** `src/lib/hydrate.ts` `resolveIndex(el)`
|
|
331
|
+
just reads the mount node — branding plus `readShardInfo` — and returns an empty
|
|
332
|
+
document list; there is nothing to await, so the app paints immediately and the
|
|
333
|
+
first shard request starts at boot. `src/lib/shards.ts` then owns all fetching:
|
|
334
|
+
|
|
335
|
+
- `loadSearchShards(info, onBatch)` — a bounded worker pool (4) over the summary
|
|
336
|
+
shards. Two properties are load-bearing. **Order**: records are appended in
|
|
337
|
+
shard order no matter what order responses arrive in, because corpus position is
|
|
338
|
+
what the detail lookup is keyed to — a shard that resolves early waits for its
|
|
339
|
+
predecessors. **Throttling**: `App.vue` re-filters and re-sorts the entire array
|
|
340
|
+
on every change, so emitting per shard would re-sort a growing 166k-element
|
|
341
|
+
array dozens of times; emissions are coalesced into one per `flushMs` window,
|
|
342
|
+
with a forced final emit. A shard that 404s becomes `[]` rather than aborting
|
|
343
|
+
the load, and `res.ok` is checked because a 404 on Pages returns an HTML error
|
|
344
|
+
page that would otherwise throw inside `res.json()`.
|
|
345
|
+
- `createDetailLoader(info)` — fetch-on-open with a per-shard cache, so paging
|
|
346
|
+
through neighbouring documents costs one request rather than one per document;
|
|
347
|
+
concurrent opens in the same shard share a single in-flight promise.
|
|
348
|
+
|
|
349
|
+
`src/app.ts` holds `reactive()` `IndexData` and `Hydration` objects and mutates
|
|
350
|
+
them as shards land — `createApp(App, props)` props are **static**, so reassigning
|
|
351
|
+
a prop would never reach the component. The shard load is kicked off via
|
|
352
|
+
`requestIdleCallback` so it stays off the first paint.
|
|
353
|
+
The site `description` comes from `data-description` on the mount node; `App.vue`
|
|
354
|
+
renders it as the header subtitle, falling back to the default "Please use the
|
|
355
|
+
provided Relaton DocID…" sentence.
|
|
356
|
+
|
|
357
|
+
**The loading gates are the subtle part.** While shards are arriving the corpus is
|
|
358
|
+
knowingly incomplete, so anything that reads "absent" as "does not exist" has to
|
|
359
|
+
wait for `hydration.loading` to clear:
|
|
360
|
+
- an unresolved `?doc=` must **not** bounce to the list (that watcher previously
|
|
361
|
+
treated not-found as removed-document; on a large index every deep link would
|
|
362
|
+
break) — instead it calls `requestAll()` to hurry the load, and only drops the
|
|
363
|
+
param once loading finishes;
|
|
364
|
+
- `?page=N` must not be clamped against a `pageCount` computed from a partial
|
|
365
|
+
corpus — `clampPage()` is deferred to a `watch(loading)`, and `onPopState`
|
|
366
|
+
skips clamping while loading;
|
|
367
|
+
- the count line reads `hydration.total` (known up front from `data-total`) so it
|
|
368
|
+
says "12 of 166658", not a denominator that grows.
|
|
369
|
+
The `hydration` and `loadDetail` props are **optional**: mounted without them the
|
|
370
|
+
component behaves exactly as it did before, which is what keeps the pre-existing
|
|
371
|
+
tests meaningful.
|
|
372
|
+
Pagination is **viewport-adaptive**: `computePageSize()`
|
|
373
|
+
sizes a page from the available height and grid/list row density (recomputed on
|
|
374
|
+
resize — debounced — and on view-mode change, clamped `10..100`) rather than a
|
|
375
|
+
fixed 100 rows, and the Prev/Next/page controls live in a shared
|
|
376
|
+
`src/components/Pagination.vue` rendered **both above and below** the list. The
|
|
377
|
+
URL-synced state is **two params**, both in `src/lib/url.ts` (History API,
|
|
378
|
+
preserving each other + any other params, synced back by the `popstate`
|
|
379
|
+
listener): the current page `?page=N` (page 1 omits it —
|
|
380
|
+
`readPageFromUrl`/`writePageToUrl`) and the open **detail document** `?doc=<id>`
|
|
381
|
+
(`readDocFromUrl`/`writeDocToUrl`) so a reload/share/bookmark restores it. User
|
|
382
|
+
paging and opening/closing a detail use `history.pushState` (Back/Forward walk
|
|
383
|
+
pages and in/out of the detail), while restore-clamp and filter-reset use
|
|
384
|
+
`replaceState`. Search/facets/sort stay in-memory (view/theme still use the
|
|
385
|
+
`localStorage` `relaton-index-*` keys).
|
|
386
|
+
|
|
387
|
+
**Detail page.** Clicking a `DocumentRow` (its id/title/chevron — real
|
|
388
|
+
`?doc=<id>` anchors, click-intercepted to SPA-navigate) opens
|
|
389
|
+
`src/components/DocumentDetail.vue` in place of the list (`App.vue` swaps on the
|
|
390
|
+
`selectedDoc` computed; an unknown `?doc=` falls back to the list). It renders the
|
|
391
|
+
enriched fields (abstract, metadata grid, keywords, contributors, all identifiers,
|
|
392
|
+
**relations**, all dates) and the landing + raw-YAML links, each block shown only
|
|
393
|
+
when present so a summary-only doc degrades gracefully.
|
|
394
|
+
|
|
395
|
+
**Relations link within the dataset.** A relation whose target DocID is a document
|
|
396
|
+
in this corpus renders as a real `?doc=<id>` anchor (same `docHref` helper and
|
|
397
|
+
click-intercept as `DocumentRow` — `lib/url.ts` owns that one builder, don't
|
|
398
|
+
re-inline `URLSearchParams`); a target that isn't stays visible as plain text
|
|
399
|
+
rather than linking to a page that would bounce back to the list. Resolution is
|
|
400
|
+
**frontend-side against the loaded corpus**, not baked in at build time — the
|
|
401
|
+
generator is a single streaming pass (`each_document`) that never holds the corpus,
|
|
402
|
+
so it cannot know the full id set when it writes a detail shard. The consequence is
|
|
403
|
+
the usual loading gate: while summary shards are still arriving a real target can
|
|
404
|
+
look absent, so `DocumentDetail` takes a **`hasDoc(id)` predicate** rather than a
|
|
405
|
+
built `Set` (a `computed` Set over 166k ids stays lazy behind the function, and only
|
|
406
|
+
materializes for a document that actually has relations), and `App.vue` calls
|
|
407
|
+
`hydration.requestAll()` when an opened document's detail carries relations. Links
|
|
408
|
+
then upgrade from plain text to anchors as shards land. Pure
|
|
409
|
+
filter/sort logic is `src/lib/filter.ts`. UI icons
|
|
410
|
+
are inline Heroicons-outline SVGs via the shared `src/components/Icon.vue`
|
|
411
|
+
(`<Icon name="…" />`, a name→path map — `fill=none stroke=currentColor` so they
|
|
412
|
+
inherit text colour + dark mode), matching the CalConnect standards site — **not**
|
|
413
|
+
Unicode emoji; add a new glyph to `Icon.vue`'s `PATHS` rather than inlining an
|
|
414
|
+
emoji. Tests: `npm run typecheck` + `vitest`
|
|
415
|
+
(filter/hydrate/shards/url/App/Icon/DocumentDetail via happy-dom).
|
|
416
|
+
Anything touching the shard loaders needs a `fetch` stub — `src/test/fetch.ts`
|
|
417
|
+
(`stubFetch` + `deferred`) is the shared one; `deferred` is how the
|
|
418
|
+
out-of-order-response ordering test forces shard 1 to resolve before shard 0.
|
|
419
|
+
|
|
420
|
+
**`npm run dev`.** `src/dev/main.ts` shards `sample-data.json` in memory, stubs
|
|
421
|
+
`fetch` to serve it, and sets the mount-node scalars — i.e. the dev harness goes
|
|
422
|
+
through the *real* hydration path rather than a shortcut that only exists there.
|
|
423
|
+
Its shard sizes are deliberately tiny (3 / 2) so progressive loading is visible.
|
|
424
|
+
|
|
425
|
+
**Build wiring.** `rake build_frontend` (`npm ci`/`install` + `npm run build`) is
|
|
426
|
+
a prerequisite of `build`/`release`; the root `rake build_all` runs it before
|
|
427
|
+
`gem build`. The gemspec ships the gitignored `frontend/dist/*` explicitly
|
|
428
|
+
(`git ls-files` can't see it) and excludes the tracked `frontend/` sources.
|
|
429
|
+
**Release note:** the release pipeline must have Node available or the gem ships
|
|
430
|
+
without the bundle. The reusable Pages-deploy workflow (install relaton-cli →
|
|
431
|
+
`relaton index` → deploy-pages) belongs in `relaton/support` — data repos call it
|
|
432
|
+
via `uses: relaton/support/.github/workflows/data-deploy.yml@main` — not in this
|
|
433
|
+
repo (a reusable workflow only runs from the calling repo's `.github/workflows/`,
|
|
434
|
+
and an in-tree copy would also get packaged into the gem).
|
|
435
|
+
|
|
436
|
+
**`--mode` removal and the caller contract.** `--mode` is **gone**, not
|
|
437
|
+
deprecated. `relaton/support`'s `data-deploy.yml` passed it at two call sites,
|
|
438
|
+
and every `relaton-data-*` caller runs `source: git` — building relaton-cli from
|
|
439
|
+
`relaton/relaton` **main**, not a released gem — so there is no release gate:
|
|
440
|
+
a stale caller breaks on its next deploy run, not at some later `gem push`.
|
|
441
|
+
`/work/HANDOFFS/relaton__support__data-deploy-drop-mode-input.md` asks support to
|
|
442
|
+
drop the input, and **must land first**.
|
|
443
|
+
|
|
444
|
+
That is also why `Command.exit_on_failure?` now returns `true`. Thor's legacy
|
|
445
|
+
default is to print a usage error and exit **0**, so a caller passing the removed
|
|
446
|
+
flag would write no site at all and still go green — publishing an empty index
|
|
447
|
+
across ~29 unattended Pages deploys with nothing red. An acceptance spec pins the
|
|
448
|
+
loud failure. Don't remove `exit_on_failure?` to quiet a test.
|
|
449
|
+
|
|
450
|
+
**Don't pre-compress the shards, and don't add a single archive.** GitHub Pages
|
|
451
|
+
already gzips them in transit (measured: a 32.9 MB `search.json` transfers in
|
|
452
|
+
4.8 MB, and the old 151 MB 3gpp page in 6.4 MB — the hand-off's headline byte
|
|
453
|
+
counts were `curl` *without* `--compressed`, i.e. uncompressed sizes). Storing
|
|
454
|
+
`.json.gz` would be served as `application/gzip` with no `Content-Encoding`, so
|
|
455
|
+
the browser would not inflate it transparently: you would add a decompressor to
|
|
456
|
+
the bundle to reach bytes Fastly already gives you, and forfeit brotli. A single
|
|
457
|
+
zip (the `index-vN.zip` pattern the data repos use for their *machine* index) is
|
|
458
|
+
worse still here — nothing renders until all of it arrives, which is exactly the
|
|
459
|
+
O(corpus) first paint this design removes. The real remaining cost was never
|
|
460
|
+
transfer; it is parse/DOM/heap, which compression does nothing for.
|
|
461
|
+
|
|
462
|
+
**Known ceiling.** Sharding fixes *transfer and parse*, not *compute*. 166k
|
|
463
|
+
documents is still ~40–60 MB of JS heap, and `applyFilters` runs a full
|
|
464
|
+
`Intl.Collator` sort over the whole array on every keystroke. The real fix is a
|
|
465
|
+
prebuilt inverted index or server-side search; that is a separate design, not
|
|
466
|
+
something to bolt onto the shard loader.
|
|
467
|
+
|
|
65
468
|
### File Operations
|
|
66
469
|
|
|
67
470
|
`lib/relaton/cli/relaton_file.rb` — Static methods for extract (pull bibdata from Metanorma XML), concatenate (combine files into a collection), and split (break a collection into individual files).
|
data/Gemfile
CHANGED
|
@@ -10,6 +10,11 @@ gem "relaton", path: "../.."
|
|
|
10
10
|
# add openssl as an explicit gem to fix SSL verification issues in GHA.
|
|
11
11
|
# gem "openssl"
|
|
12
12
|
|
|
13
|
+
# `canon` comes in through `lutaml-model`. Keep it below 0.3.52, as the root
|
|
14
|
+
# Gemfile does: from that version it installs `yeptris`, which replaces
|
|
15
|
+
# `Object#to_yaml` and raises `NameError` on Windows. See the root Gemfile.
|
|
16
|
+
gem "canon", "< 0.3.52"
|
|
17
|
+
|
|
13
18
|
gem "byebug", "~> 11.0"
|
|
14
19
|
gem "debug"
|
|
15
20
|
gem "equivalent-xml", "~> 0.6"
|
data/Rakefile
CHANGED
|
@@ -3,4 +3,19 @@ require "rspec/core/rake_task"
|
|
|
3
3
|
|
|
4
4
|
RSpec::Core::RakeTask.new(:spec)
|
|
5
5
|
|
|
6
|
+
FRONTEND_DIR = File.expand_path("frontend", __dir__)
|
|
7
|
+
|
|
8
|
+
desc "Build the JS/CSS frontend (Vite → frontend/dist/app.iife.js + style.css)"
|
|
9
|
+
task :build_frontend do
|
|
10
|
+
Dir.chdir(FRONTEND_DIR) do
|
|
11
|
+
# `npm ci` when a lockfile is present (reproducible), else `npm install`.
|
|
12
|
+
sh(File.exist?("package-lock.json") ? "npm ci" : "npm install")
|
|
13
|
+
sh "npm run build"
|
|
14
|
+
end
|
|
15
|
+
end
|
|
16
|
+
|
|
17
|
+
# The gem must ship the compiled frontend, so build it before packaging/release.
|
|
18
|
+
task build: :build_frontend
|
|
19
|
+
task release: :build_frontend
|
|
20
|
+
|
|
6
21
|
task default: :spec
|