@konneal/engine 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +29 -0
- package/README.md +13 -0
- package/dist/admin.d.ts +26 -0
- package/dist/ai.d.ts +6 -0
- package/dist/anchors.d.ts +6 -0
- package/dist/answercache.d.ts +22 -0
- package/dist/ask.d.ts +5 -0
- package/dist/auth.d.ts +12 -0
- package/dist/bubble.d.ts +14 -0
- package/dist/chunk-LLWPT2XV.js +49 -0
- package/dist/chunk-MB74PTRM.js +114 -0
- package/dist/chunk-WOGQM7DJ.js +197 -0
- package/dist/chunk-WWNCWKKC.js +42 -0
- package/dist/completion.d.ts +5 -0
- package/dist/config.d.ts +154 -0
- package/dist/config.js +37 -0
- package/dist/context.d.ts +115 -0
- package/dist/conversations.d.ts +5 -0
- package/dist/drafts.d.ts +129 -0
- package/dist/env.d.ts +57 -0
- package/dist/faithfulness.d.ts +5 -0
- package/dist/grader.d.ts +3 -0
- package/dist/graph.d.ts +13 -0
- package/dist/hybrid.d.ts +7 -0
- package/dist/index.d.ts +9 -0
- package/dist/index.js +5373 -0
- package/dist/internal_gateway.d.ts +14 -0
- package/dist/lexical.d.ts +7 -0
- package/dist/livedata.d.ts +77 -0
- package/dist/memories.d.ts +10 -0
- package/dist/modelplane.d.ts +61 -0
- package/dist/oidc.d.ts +73 -0
- package/dist/pipeline.d.ts +57 -0
- package/dist/profile.d.ts +2 -0
- package/dist/profile.gen.d.ts +70 -0
- package/dist/profile.js +8 -0
- package/dist/projects.d.ts +8 -0
- package/dist/prompts/conversational.md +8 -0
- package/dist/prompts/enrichment.md +3 -0
- package/dist/prompts/faithfulness.md +1 -0
- package/dist/prompts/grader.md +5 -0
- package/dist/prompts/listwise.md +3 -0
- package/dist/prompts/precision.md +1 -0
- package/dist/prompts/reflect.md +1 -0
- package/dist/prompts/relevancy.md +1 -0
- package/dist/prompts/research.md +10 -0
- package/dist/prompts/section-summary.md +5 -0
- package/dist/prompts/summarize.md +1 -0
- package/dist/prompts/system.md +18 -0
- package/dist/prompts/understanding.md +17 -0
- package/dist/quota.d.ts +13 -0
- package/dist/reflect.d.ts +5 -0
- package/dist/refs.d.ts +40 -0
- package/dist/refusal.d.ts +9 -0
- package/dist/refusal.js +9 -0
- package/dist/requestScope.d.ts +26 -0
- package/dist/requestScope.js +10 -0
- package/dist/research.d.ts +8 -0
- package/dist/search.d.ts +4 -0
- package/dist/selfquery.d.ts +7 -0
- package/dist/session.d.ts +1 -0
- package/dist/share.d.ts +2 -0
- package/dist/structural.d.ts +27 -0
- package/dist/tablecontext.d.ts +11 -0
- package/dist/understand.d.ts +11 -0
- package/dist/understandContract.d.ts +29 -0
- package/dist/verdict.d.ts +24 -0
- package/docs/API.md +451 -0
- package/docs/ARCHITECTURE.md +302 -0
- package/docs/AUDIT-2026-08-24.md +71 -0
- package/docs/CONTRIBUTOR-AUDIT-2026-08-25.md +147 -0
- package/docs/INGEST-ARCHITECTURE.md +158 -0
- package/docs/MCP.md +92 -0
- package/docs/METANORMA-AI-SERIALIZATION.md +247 -0
- package/docs/MKO-EXPORT-PIPELINE.md +147 -0
- package/docs/REDESIGN-NORMATIVE-RAG-ETSI.md +485 -0
- package/docs/RESEARCH-SOTA-2026.md +243 -0
- package/docs/ROADMAP-SOTA.md +130 -0
- package/docs/SOTA-STAGE-SPECS.md +509 -0
- package/docs/annealment/F1-verdict.md +27 -0
- package/docs/annealment/F10-notes.md +19 -0
- package/docs/annealment/F11-composition.md +17 -0
- package/docs/annealment/F12-passport.md +17 -0
- package/docs/annealment/F2-counterfactual.md +20 -0
- package/docs/annealment/F3-absence.md +21 -0
- package/docs/annealment/F4-instance.md +18 -0
- package/docs/annealment/F5-workflow.md +21 -0
- package/docs/annealment/F6-impact.md +21 -0
- package/docs/annealment/F7-editions.md +18 -0
- package/docs/annealment/F8-selfverify.md +19 -0
- package/docs/annealment/F9-projection-qa.md +17 -0
- package/docs/annealment/L0-locate.md +19 -0
- package/docs/annealment/L1-extract.md +18 -0
- package/docs/annealment/L2-nomenclature.md +22 -0
- package/docs/annealment/L3-geometry.md +23 -0
- package/docs/annealment/L4-composition.md +21 -0
- package/docs/annealment/L5-cross-standard.md +20 -0
- package/docs/annealment/L6-diachrony.md +21 -0
- package/docs/annealment/L7-perception.md +20 -0
- package/docs/annealment/L8-computation.md +22 -0
- package/docs/annealment/L9-instance-process.md +23 -0
- package/docs/annealment/README.md +10 -0
- package/docs/guidelines-metanorma-ai-programme.md +279 -0
- package/docs/identity-onboarding-rag.md +65 -0
- package/docs/identity-service.md +219 -0
- package/docs/knowledge-annealment.md +273 -0
- package/docs/konneal-extraction-plan.md +481 -0
- package/docs/metanorma-for-ai.md +270 -0
- package/docs/mirror-plan.md +36 -0
- package/docs/multi-sdo-architecture.md +191 -0
- package/docs/paper-annealment-comparison.md +259 -0
- package/docs/paper-assets/architecture.svg +94 -0
- package/docs/paper-assets/contract-v2.svg +94 -0
- package/docs/paper-assets/mko-ingest.svg +91 -0
- package/docs/paper-oiml-bulletin.md +402 -0
- package/docs/paper-oiml-bulletin.mdx +419 -0
- package/docs/product-branding-options.md +172 -0
- package/docs/projects-design.md +88 -0
- package/docs/sota-mechanisms.md +184 -0
- package/docs/spec-api.md +77 -0
- package/docs/spec-pipeline.md +126 -0
- package/docs/vector-adapter.md +88 -0
- package/package.json +70 -0
- package/profile/corpora.yaml +5 -0
- package/profile/datasets.yaml +14 -0
- package/profile/prompts.yaml +5 -0
- package/profile/publisher.yaml +17 -0
- package/profile/retrieval.yaml +1 -0
- package/profile/sources.yaml +5 -0
- package/profile/ui.yaml +7 -0
- package/scripts/gen_profile.mjs +33 -0
- package/workers/shared/ai.ts +21 -0
- package/workers/shared/auth.ts +16 -0
- package/workers/shared/chunk.ts +108 -0
- package/workers/shared/oidc.ts +312 -0
- package/workers/shared/router.ts +45 -0
- package/workers/shared/session.ts +104 -0
- package/workers/worker_internal/src/index.ts +157 -0
- package/workers/worker_internal/tsconfig.json +15 -0
- package/workers/worker_internal/wrangler.toml +32 -0
- package/workers/worker_mcp/src/index.ts +175 -0
- package/workers/worker_mcp/tsconfig.json +13 -0
- package/workers/worker_mcp/wrangler.toml +18 -0
- package/workers/worker_public/migrations/0002_conversations.sql +22 -0
- package/workers/worker_public/migrations/0003_shared_conversations.sql +9 -0
- package/workers/worker_public/migrations/0004_graph.sql +16 -0
- package/workers/worker_public/migrations/0005_documents.sql +19 -0
- package/workers/worker_public/migrations/0006_conversation_entities.sql +11 -0
- package/workers/worker_public/migrations/0007_chunks_fts.sql +43 -0
- package/workers/worker_public/migrations/0008_unit_payloads.sql +15 -0
- package/workers/worker_public/migrations/0009_chunks_unit.sql +8 -0
- package/workers/worker_public/migrations/0009_message_context.sql +7 -0
- package/workers/worker_public/migrations/0010_model_nodes.sql +33 -0
- package/workers/worker_public/migrations/0011_schema_union.sql +38 -0
- package/workers/worker_public/migrations/0012_memories.sql +15 -0
- package/workers/worker_public/migrations/0013_projects.sql +21 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.d.ts +16306 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.ts +16261 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.d.ts +16373 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.ts +16328 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.d.ts +16382 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.ts +16337 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.d.ts +16383 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.ts +16338 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.d.ts +16403 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.ts +16358 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.d.ts +16408 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.ts +16363 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.d.ts +16414 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.ts +16369 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.d.ts +16414 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.ts +16369 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/README.md +135 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/entrypoints.svg +53 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.d.ts +17095 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.ts +17050 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/index.d.ts +16306 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/index.ts +16261 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.d.ts +16447 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.ts +16402 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.d.ts +16306 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.ts +16261 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/package.json +11 -0
- package/workers/worker_public/package.json +13 -0
- package/workers/worker_public/prompts/conversational.md +8 -0
- package/workers/worker_public/prompts/enrichment.md +3 -0
- package/workers/worker_public/prompts/faithfulness.md +1 -0
- package/workers/worker_public/prompts/grader.md +5 -0
- package/workers/worker_public/prompts/listwise.md +3 -0
- package/workers/worker_public/prompts/precision.md +1 -0
- package/workers/worker_public/prompts/reflect.md +1 -0
- package/workers/worker_public/prompts/relevancy.md +1 -0
- package/workers/worker_public/prompts/research.md +10 -0
- package/workers/worker_public/prompts/section-summary.md +5 -0
- package/workers/worker_public/prompts/summarize.md +1 -0
- package/workers/worker_public/prompts/system.md +18 -0
- package/workers/worker_public/prompts/understanding.md +17 -0
- package/workers/worker_public/public/app.js +166 -0
- package/workers/worker_public/public/index.html +48 -0
- package/workers/worker_public/public/style.css +147 -0
- package/workers/worker_public/schema.sql +248 -0
- package/workers/worker_public/src/admin.ts +358 -0
- package/workers/worker_public/src/ai.ts +71 -0
- package/workers/worker_public/src/anchors.ts +41 -0
- package/workers/worker_public/src/answercache.ts +72 -0
- package/workers/worker_public/src/ask.ts +1094 -0
- package/workers/worker_public/src/auth.ts +252 -0
- package/workers/worker_public/src/bubble.ts +111 -0
- package/workers/worker_public/src/completion.ts +75 -0
- package/workers/worker_public/src/config.ts +238 -0
- package/workers/worker_public/src/context.ts +238 -0
- package/workers/worker_public/src/conversations.ts +162 -0
- package/workers/worker_public/src/drafts.ts +497 -0
- package/workers/worker_public/src/env.ts +90 -0
- package/workers/worker_public/src/faithfulness.ts +63 -0
- package/workers/worker_public/src/grader.ts +89 -0
- package/workers/worker_public/src/graph.ts +63 -0
- package/workers/worker_public/src/hybrid.ts +77 -0
- package/workers/worker_public/src/index.ts +441 -0
- package/workers/worker_public/src/internal_gateway.ts +41 -0
- package/workers/worker_public/src/lexical.ts +86 -0
- package/workers/worker_public/src/lib/hit.ts +4 -0
- package/workers/worker_public/src/lib/http.ts +83 -0
- package/workers/worker_public/src/lib/router.ts +4 -0
- package/workers/worker_public/src/livedata.ts +334 -0
- package/workers/worker_public/src/memories.ts +81 -0
- package/workers/worker_public/src/modelplane.ts +213 -0
- package/workers/worker_public/src/oidc.ts +333 -0
- package/workers/worker_public/src/pipeline.ts +377 -0
- package/workers/worker_public/src/ports/blobs.ts +7 -0
- package/workers/worker_public/src/ports/cloudflare/adapters.ts +177 -0
- package/workers/worker_public/src/ports/kv.ts +8 -0
- package/workers/worker_public/src/ports/model.ts +28 -0
- package/workers/worker_public/src/ports/runtime.ts +13 -0
- package/workers/worker_public/src/ports/store.ts +20 -0
- package/workers/worker_public/src/ports/vector.ts +26 -0
- package/workers/worker_public/src/profile.gen.ts +101 -0
- package/workers/worker_public/src/profile.ts +16 -0
- package/workers/worker_public/src/projects.ts +108 -0
- package/workers/worker_public/src/prompts.d.ts +6 -0
- package/workers/worker_public/src/quota.ts +54 -0
- package/workers/worker_public/src/reflect.ts +67 -0
- package/workers/worker_public/src/refs.ts +107 -0
- package/workers/worker_public/src/refusal.ts +65 -0
- package/workers/worker_public/src/requestScope.ts +71 -0
- package/workers/worker_public/src/research.ts +126 -0
- package/workers/worker_public/src/search.ts +56 -0
- package/workers/worker_public/src/selfquery.ts +25 -0
- package/workers/worker_public/src/session.ts +4 -0
- package/workers/worker_public/src/share.ts +53 -0
- package/workers/worker_public/src/stages/conceptGraph.ts +68 -0
- package/workers/worker_public/src/stages/conceptSteer.ts +39 -0
- package/workers/worker_public/src/stages/corpusScope.ts +25 -0
- package/workers/worker_public/src/stages/dedup.ts +10 -0
- package/workers/worker_public/src/stages/dense.ts +73 -0
- package/workers/worker_public/src/stages/diversity.ts +33 -0
- package/workers/worker_public/src/stages/editionCover.ts +63 -0
- package/workers/worker_public/src/stages/editionSteer.ts +88 -0
- package/workers/worker_public/src/stages/familyBoost.ts +22 -0
- package/workers/worker_public/src/stages/federate.ts +22 -0
- package/workers/worker_public/src/stages/glossary.ts +65 -0
- package/workers/worker_public/src/stages/graphLane.ts +31 -0
- package/workers/worker_public/src/stages/hyde.ts +29 -0
- package/workers/worker_public/src/stages/index.ts +69 -0
- package/workers/worker_public/src/stages/lexicalUnion.ts +21 -0
- package/workers/worker_public/src/stages/multiQuery.ts +57 -0
- package/workers/worker_public/src/stages/overviewDemote.ts +14 -0
- package/workers/worker_public/src/stages/poolOpen.ts +10 -0
- package/workers/worker_public/src/stages/propagate.ts +15 -0
- package/workers/worker_public/src/stages/rerank.ts +47 -0
- package/workers/worker_public/src/stages/seal.ts +16 -0
- package/workers/worker_public/src/stages/sectionDescent.ts +61 -0
- package/workers/worker_public/src/stages/stdRefNudge.ts +35 -0
- package/workers/worker_public/src/stages/subQuery.ts +42 -0
- package/workers/worker_public/src/stages/termNudge.ts +24 -0
- package/workers/worker_public/src/stages/typedPin.ts +131 -0
- package/workers/worker_public/src/stages/types.ts +112 -0
- package/workers/worker_public/src/stages/windowFloor.ts +23 -0
- package/workers/worker_public/src/structural.ts +171 -0
- package/workers/worker_public/src/tablecontext.ts +41 -0
- package/workers/worker_public/src/understand.ts +72 -0
- package/workers/worker_public/src/understandContract.ts +67 -0
- package/workers/worker_public/src/verdict.ts +255 -0
- package/workers/worker_public/tsconfig.json +18 -0
- package/workers/worker_public/wrangler.toml +104 -0
|
@@ -0,0 +1,402 @@
|
|
|
1
|
+
*Draft article for the OIML Bulletin — 2026-08-30. Style: Bulletin
|
|
2
|
+
technical article (MS Word single-column on submission; this is the
|
|
3
|
+
authoring source). Numbers are production-measured unless noted.*
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## Abstract
|
|
8
|
+
|
|
9
|
+
The OIML publishes close to 900 documents — Recommendations, Basic
|
|
10
|
+
publications, Guides, Documents and Expert reports — that together form
|
|
11
|
+
the reference library of legal metrology. Finding what they say has
|
|
12
|
+
required knowing the corpus before querying it. We describe a
|
|
13
|
+
question-answering service, **OIML SMART AI** (ai.oimlsmart.org), that
|
|
14
|
+
answers natural-language questions from the indexed publications with
|
|
15
|
+
clause-level citations, verbatim quote anchors for normative values, and
|
|
16
|
+
typed renderings of the tables and equations the answers depend on — and
|
|
17
|
+
it reads figures and user-supplied photographs directly, from pixels. The
|
|
18
|
+
service measures every component it ships: a golden question set,
|
|
19
|
+
retrieval metrics in the tradition of the information-retrieval
|
|
20
|
+
literature, and judged faithfulness. We report the measured design
|
|
21
|
+
decisions, the corpus-structure findings that shaped them, and the open
|
|
22
|
+
path: producer-native ingestion of the Metanorma document model so that
|
|
23
|
+
any Metanorma-authored corpus can be indexed the same way.
|
|
24
|
+
|
|
25
|
+
---
|
|
26
|
+
|
|
27
|
+
## 1. Introduction
|
|
28
|
+
|
|
29
|
+
Legal metrology's knowledge lives in a large, cross-referential,
|
|
30
|
+
multi-edition corpus. A typical working question — *"what is the maximum
|
|
31
|
+
permissible error for a class III nonautomatic weighing instrument?"* —
|
|
32
|
+
requires knowing which Recommendation applies (OIML R 76-1), which
|
|
33
|
+
edition is current, which clause holds the table, and how to read the
|
|
34
|
+
table's rows against the instrument's verification scale intervals.
|
|
35
|
+
Every one of those steps assumes corpus knowledge the questioner may not
|
|
36
|
+
have.
|
|
37
|
+
|
|
38
|
+
Question answering over such corpora has been transformed by
|
|
39
|
+
retrieval-augmented generation (RAG): a language model is given passages
|
|
40
|
+
retrieved from a trusted corpus and may use only those. The engineering
|
|
41
|
+
challenge is no longer whether a model can write fluent prose — it can —
|
|
42
|
+
but whether the *right passages* are found, whether the answer stays
|
|
43
|
+
*faithful* to them, and whether the reader can *verify* the result. Our
|
|
44
|
+
service is built around those three properties.
|
|
45
|
+
|
|
46
|
+
This article describes the architecture as deployed, the measurements
|
|
47
|
+
that shaped it, and the corpus findings — including errors in the
|
|
48
|
+
bibliographic record itself — that surfaced because the system keeps
|
|
49
|
+
score. It closes with the producer-native ingestion path (Metanorma
|
|
50
|
+
Knowledge Objects) and what it makes possible for other standards bodies.
|
|
51
|
+
|
|
52
|
+
## 2. What the service does
|
|
53
|
+
|
|
54
|
+
A user asks a question in any language. The service:
|
|
55
|
+
|
|
56
|
+
1. **Understands the question** with a small language model — language,
|
|
57
|
+
named publication, edition, defined terms, complexity. There are no
|
|
58
|
+
keyword rules; the same model makes these judgements for every
|
|
59
|
+
question.
|
|
60
|
+
2. **Retrieves in parallel** — three lanes leave the question at the
|
|
61
|
+
same instant: the vector index is queried with the question's
|
|
62
|
+
embedding while understanding is still running (the dense lane's
|
|
63
|
+
results are reused unchanged when understanding adds nothing), a
|
|
64
|
+
full-corpus keyword index catches exact identifiers, part numbers
|
|
65
|
+
and defined terms that vector similarity alone misses, and
|
|
66
|
+
understanding's refinements — a named publication, an edition —
|
|
67
|
+
arrive as metadata filters on the dense query.
|
|
68
|
+
3. **Ranks** the fused candidates with a cross-encoder; complex
|
|
69
|
+
questions get a second, stronger re-ranking pass.
|
|
70
|
+
4. **Answers from the retrieved passages only**, under a contract that
|
|
71
|
+
requires inline citations (`[OIML R 60-1:2021 §4.4.2]`) and verbatim
|
|
72
|
+
quote anchors for normative values.
|
|
73
|
+
5. **Verifies** the answer after generation: a deterministic check that
|
|
74
|
+
every quoted phrase exists in a cited passage; a separate judge that
|
|
75
|
+
scores faithfulness; a correction-and-retry when either fails. An
|
|
76
|
+
answer that cannot be verified is never written to the answer cache.
|
|
77
|
+
|
|
78
|
+

|
|
79
|
+
|
|
80
|
+
*Figure 1 — A question's path through the service. Three retrieval lanes run concurrently; every answer passes mechanical verification before it is served or cached. The edition registry derives publication status from supersession links, never from a status field.*
|
|
81
|
+
|
|
82
|
+
|
|
83
|
+
When the question names a publication, answers are steered to the
|
|
84
|
+
current edition by a registry derived from the bibliographic record's
|
|
85
|
+
supersession links (§5). When the answer depends on a table or
|
|
86
|
+
equation, the interface renders the *producer's* typed object — the
|
|
87
|
+
actual rows and columns — never the model's re-typing of it (§6).
|
|
88
|
+
|
|
89
|
+
If the corpus does not contain the answer, the system says so in one
|
|
90
|
+
canonical sentence and redirects; refusals are never cached, because a
|
|
91
|
+
refusal is a property of the moment, not of the question.
|
|
92
|
+
|
|
93
|
+
## 3. Corpus and structure
|
|
94
|
+
|
|
95
|
+
The indexed corpus comprises the English editions of the OIML
|
|
96
|
+
publications (~900 documents, ~9.3 million words). Documents are chunked
|
|
97
|
+
along clause boundaries — never arbitrary token windows — and each chunk
|
|
98
|
+
carries provenance: publication identifier, edition, language, clause
|
|
99
|
+
anchor, and bibliographic status.
|
|
100
|
+
|
|
101
|
+
Three corpus-level findings shaped the design:
|
|
102
|
+
|
|
103
|
+
**The bibliographic status field disagrees with its own record.** In 58
|
|
104
|
+
of 224 publication families, the machine-readable status field claims a
|
|
105
|
+
document is current while the same record's supersession links say
|
|
106
|
+
otherwise. The service therefore *derives* status from the links and
|
|
107
|
+
ignores the field. More consequentially, 36 of 224 families have no
|
|
108
|
+
recorded successor link at all, so the current edition cannot be derived
|
|
109
|
+
for them; the registry surfaces these gaps rather than guessing. Both
|
|
110
|
+
numbers are the worklist for upstream bibliographic corrections — the
|
|
111
|
+
kind that serving surfaces only because it keeps score.
|
|
112
|
+
|
|
113
|
+
**Identifiers are identity, not display.** A cross-referenced audit of
|
|
114
|
+
the dirty corpus found documents whose machine-readable identifier was
|
|
115
|
+
a placeholder ("OIML D 0:0000", "OIML D X") or polluted with language
|
|
116
|
+
markers — around 200 records — making them unreachable by
|
|
117
|
+
doc-scoped retrieval. Identifiers are now derived from slug identity
|
|
118
|
+
with placeholders never winning, and the index is reconciled against
|
|
119
|
+
the canonical chunk set: the first census measured 49,553 live vectors
|
|
120
|
+
against 31,512 canonical — 18,041 strays from every prior re-chunking,
|
|
121
|
+
deleted in place (upserts never delete). A retrieval service that never
|
|
122
|
+
enumerates its own index serves ghosts.
|
|
123
|
+
|
|
124
|
+
**Tables are where the normative values live.** Maximum permissible
|
|
125
|
+
errors, accuracy-class limits, verification interval bounds — the values
|
|
126
|
+
practitioners ask for — are tabular. Flattening tables into prose loses
|
|
127
|
+
the row/column geometry that makes them answerable. The index now treats
|
|
128
|
+
tables as atomic typed objects (§6).
|
|
129
|
+
|
|
130
|
+
**Structure is a retrieval signal.** Standards are hierarchically
|
|
131
|
+
organized with extensive cross-references; the system indexes the
|
|
132
|
+
section hierarchy and the citation graph as a queryable graph (7,128
|
|
133
|
+
nodes, 6,486 edges) alongside the text, so a defined term can be
|
|
134
|
+
resolved to the documents that define it even when the question's
|
|
135
|
+
vocabulary does not match the corpus's.
|
|
136
|
+
|
|
137
|
+
## 4. Retrieval: what the measurements showed
|
|
138
|
+
|
|
139
|
+
We evaluate retrieval with a golden question set (now 29 questions
|
|
140
|
+
including eight table-value questions with witness rows) using
|
|
141
|
+
recall@5, average precision@5, and mean reciprocal rank@5 — the protocol
|
|
142
|
+
of the standards-retrieval literature. Single runs proved noisy (±5
|
|
143
|
+
percentage points); all reported numbers are means over three runs on a
|
|
144
|
+
quiet account.
|
|
145
|
+
|
|
146
|
+
The single largest measured improvement came from making keyword
|
|
147
|
+
retrieval a **full-corpus first stage** rather than a re-scoring of the
|
|
148
|
+
vector results: recall@5 rose from 86% to 95%. The reason is structural
|
|
149
|
+
— standards vocabulary is exact ("n_LC", "R 60-3", "creep"), and dense
|
|
150
|
+
embeddings alone miss exact identifiers that a lexical scan of the whole
|
|
151
|
+
corpus recovers. Fusing both rankings (reciprocal rank fusion) lifted
|
|
152
|
+
precision and MRR without costing recall.
|
|
153
|
+
|
|
154
|
+
The second improvement came from **contextual enrichment**: a one-time
|
|
155
|
+
pass that prepends a short model-written preamble to every chunk ("this
|
|
156
|
+
clause defines the accuracy-class limits for load cells") before
|
|
157
|
+
embedding and indexing. This is the contextual-retrieval recipe
|
|
158
|
+
validated industry-wide; our corpus measurement confirmed it, and the
|
|
159
|
+
one-time cost (~$0.0017 per chunk) is amortized over every future
|
|
160
|
+
answer.
|
|
161
|
+
|
|
162
|
+
With the typed-table lane live (§6), retrieval lanes running
|
|
163
|
+
concurrently with query understanding, edition steering active (§5) and
|
|
164
|
+
multimodal generation on (§6), the current baseline is:
|
|
165
|
+
|
|
166
|
+
| Metric | Value |
|
|
167
|
+
|---|---|
|
|
168
|
+
| Recall@5 (mean of 3) | **94.3%** (range 93.1–96.6) |
|
|
169
|
+
| Average precision@5 | 0.82 |
|
|
170
|
+
| MRR@5 | 0.87 |
|
|
171
|
+
| End-to-end golden cases | 14/14 |
|
|
172
|
+
|
|
173
|
+
These numbers are on our own golden set, not on the benchmarks of the
|
|
174
|
+
works we build on — the honest comparison is per-technique, on the same
|
|
175
|
+
corpus, before and after. So read: the full-corpus lexical stage is the
|
|
176
|
+
ETSI study's recipe (ref 1) — adopting it lifted recall@5 here from 86%
|
|
177
|
+
to 95%; the contextual preambles are Anthropic's contextual retrieval
|
|
178
|
+
(ref 3) — confirmed on this corpus at ~$0.0017 per chunk one-time; the
|
|
179
|
+
clause-boundary, structure-preserving chunking follows the same
|
|
180
|
+
structure-aware line as refs 2 and 7, which our typed-unit pin extends
|
|
181
|
+
from "chunk better" to "guarantee the answering object a slot." Two
|
|
182
|
+
elements of the deployed system have no counterpart in the cited work:
|
|
183
|
+
the symbolic-reference contract, under which table data never passes
|
|
184
|
+
through the model at all (refs 2's error-reduction approach still
|
|
185
|
+
re-generates tables; we removed the corruption channel), and the
|
|
186
|
+
mechanical post-generation verification of every answer (quote anchors,
|
|
187
|
+
unit references, retyped-table detection) with a faithfulness judge that
|
|
188
|
+
sees the passages actually used.
|
|
189
|
+
|
|
190
|
+
Two structural additions complete the retrieval story. First, the
|
|
191
|
+
corpus is a tree — every chunk carries its clause anchor — and the
|
|
192
|
+
serving path uses the tree: a hit's score blends its neighbouring
|
|
193
|
+
clauses' scores, evidence is presented to the answer model in document
|
|
194
|
+
reading order, and a ranked section summary resolves to its quotable
|
|
195
|
+
child clauses (adapted from FABLE/BEAR, ref 7, at none of its index-time
|
|
196
|
+
cost, because Metanorma documents arrive as trees rather than needing
|
|
197
|
+
one inferred). Second, everyday words rarely match defined terms — the
|
|
198
|
+
one gap that failed for every corpus representation equally — so a
|
|
199
|
+
terminology index of the corpus's defined concepts links a question's
|
|
200
|
+
phrasing to candidate terms, and the answer model adjudicates and leads
|
|
201
|
+
with the corpus's own term, quoting its definition ("my output keeps
|
|
202
|
+
drifting" resolves to *span stability*, with the verbatim definition and
|
|
203
|
+
its defining clause cited).
|
|
204
|
+
|
|
205
|
+
## 5. Editions and trust
|
|
206
|
+
|
|
207
|
+
Which edition applies is as consequential as what it says. The service's
|
|
208
|
+
registry of publication families and editions is derived from
|
|
209
|
+
supersession links, not from status fields (§3). Citations carry the
|
|
210
|
+
edition and status of every source; superseded sources are marked.
|
|
211
|
+
Answers to questions that name a publication are steered to the current
|
|
212
|
+
edition; answers that must cite a superseded edition (the 2021 passages
|
|
213
|
+
do not reproduce a table the 2017 edition carries) say so explicitly.
|
|
214
|
+
|
|
215
|
+
Steering is measured, not assumed, and two mechanisms earned their place
|
|
216
|
+
by fixing observed failures. Superseded editions match archaic phrasing
|
|
217
|
+
strongly — a complex verification question once cited R 76-1:1992/1988
|
|
218
|
+
while the current 2006 edition was crowded out of the passage window —
|
|
219
|
+
so ranking now demotes older-edition chunks of a publication whenever a
|
|
220
|
+
newer edition of the same publication is present in the evidence
|
|
221
|
+
(cross-publication recency is deliberately untouched: a 1992 document
|
|
222
|
+
that is still current must not lose to unrelated 2024 ones). And because
|
|
223
|
+
query understanding can pin an edition for the wrong family, an edition
|
|
224
|
+
pin that the corpus itself fails to corroborate is dropped to the
|
|
225
|
+
document-level filter. After both fixes the probe cites R 76-1:2006 in
|
|
226
|
+
force, with one legitimate 1988 tail citation for a clause only that
|
|
227
|
+
edition carries; the window's measured average precision (0.85) and MRR
|
|
228
|
+
(0.88) are the best the service has recorded.
|
|
229
|
+
|
|
230
|
+
## 6. Tables, equations, and figures as typed objects
|
|
231
|
+
|
|
232
|
+
The most consequential rendering decision: the language model never
|
|
233
|
+
re-types normative data. When an answer depends on a table, the model
|
|
234
|
+
emits a symbolic reference to the table's unit (`[[u:table-1]]`); the
|
|
235
|
+
server validates the reference against the passages actually used and
|
|
236
|
+
resolves it to the producer's typed payload — columns, rows, units. The
|
|
237
|
+
interface renders the exact object. Small verbatim values in prose are
|
|
238
|
+
permitted and mechanically verified against the source. The payload
|
|
239
|
+
never passes through the model's output, so it cannot be corrupted by
|
|
240
|
+
generation.
|
|
241
|
+
|
|
242
|
+
The same principle extends from rendering to judgment: the model never
|
|
243
|
+
judges what a machine can compute. Where the corpus's machine-readable
|
|
244
|
+
models carry checkable objects — a constraint with an OCL expression, a
|
|
245
|
+
requirement with a threshold limit — a question naming one gets it
|
|
246
|
+
EXECUTED: the question's values bind to the rule's own parameters, the
|
|
247
|
+
check evaluates deterministically, and the verdict (pass, or the
|
|
248
|
+
standard's own violation word, with the arithmetic shown, or void naming
|
|
249
|
+
exactly what is missing) is attached to the answer as computed data.
|
|
250
|
+
Two sibling capabilities complete the set: provable absence (a topic
|
|
251
|
+
that a standard does not constrain returns an enumeration certificate
|
|
252
|
+
over its entire model plane — N nodes enumerated, zero matches — rather
|
|
253
|
+
than a refusal) and answer verification (any answer checkable against
|
|
254
|
+
the corpus: verbatim-quote containment, object-reference resolution, a
|
|
255
|
+
judged faithfulness score labeled as judged).
|
|
256
|
+
|
|
257
|
+

|
|
258
|
+
|
|
259
|
+
*Figure 2 — The answer contract. The model writes prose with inline citations and symbolic references; three mechanical checks run after generation; the server resolves surviving references to the producer's payloads and the client renders them.*
|
|
260
|
+
|
|
261
|
+
|
|
262
|
+
Figures are described by a vision-capable model that reads the actual
|
|
263
|
+
image pixels; equations are carried in their authoring-native formats
|
|
264
|
+
(AsciiMath and MathML) with a plain-language description. Both arrive as typed objects with the same validation.
|
|
265
|
+
|
|
266
|
+
Figures also flow the other way: when the retrieved passages contain
|
|
267
|
+
figure units, their actual images are attached to the answer model's
|
|
268
|
+
input, so descriptions and reasoning come from the drawing itself — the
|
|
269
|
+
model reads labels that exist only in the pixels (the pixel-label
|
|
270
|
+
probe — "what are the labeled example cases in the R 60-2 design-shapes
|
|
271
|
+
figure?" — is answered A, B and C from the drawing, stable across
|
|
272
|
+
repeated runs). Two measured invariants keep that true: assets must be
|
|
273
|
+
readable by vision pipelines (vector-sourced rasters that draw black on
|
|
274
|
+
transparent alpha flatten to a solid black rectangle inside a vision
|
|
275
|
+
model, whatever a browser shows — a corpus sweep detects and repairs
|
|
276
|
+
them), and images ride their own message in the generation call (long
|
|
277
|
+
passage text and image parts in one message triggers provider errors
|
|
278
|
+
that scale with payload). Users can likewise attach a photograph (an
|
|
279
|
+
instrument nameplate, a scale dial, a schematic) to their question; the
|
|
280
|
+
text still drives retrieval, and the image gives the model the visual
|
|
281
|
+
context, under the same citation contract.
|
|
282
|
+
|
|
283
|
+
## 7. Models, cost, and sovereignty
|
|
284
|
+
|
|
285
|
+
All models are open-weight and served on a single cloud platform
|
|
286
|
+
(Cloudflare Workers AI); no corpus data is sent to proprietary model
|
|
287
|
+
providers. The answer model (GLM-5.3 Flash, natively multimodal) serves
|
|
288
|
+
all tiers at approximately $0.001 per answer. The full one-time corpus
|
|
289
|
+
preparation — contextual enrichment of 42,000 chunks — cost
|
|
290
|
+
approximately $75. Daily serving at current traffic is under $5 per
|
|
291
|
+
month including infrastructure. The model policy is deliberately
|
|
292
|
+
cost-first on the serving path (every question pays it) and quality-first
|
|
293
|
+
on one-time work (enrichment, captioning, evaluation), where quality
|
|
294
|
+
persists into every future answer.
|
|
295
|
+
|
|
296
|
+
One lesson generalizes: read the model card before wiring a model. Three
|
|
297
|
+
separate live failures traced to defaults we never set — a
|
|
298
|
+
reasoning-effort parameter that silently defaults to maximum (starving
|
|
299
|
+
small output budgets), sampling defaults that let a thinking model loop
|
|
300
|
+
(repetition consuming the token budget that carried the structured
|
|
301
|
+
output), and a degraded non-thinking mode we were unknowingly paying
|
|
302
|
+
for. Every call site now states its reasoning mode, its sampling, and a
|
|
303
|
+
budget the reasoning cannot starve; the discipline costs nothing and
|
|
304
|
+
removed an entire class of silent failure.
|
|
305
|
+
|
|
306
|
+
## 8. The producer-native path: Metanorma Knowledge Objects
|
|
307
|
+
|
|
308
|
+
OIML publications are authored in Metanorma, a model-driven document
|
|
309
|
+
system: the source of truth is a typed document model, not any rendered
|
|
310
|
+
PDF or HTML. The service's newest ingestion path consumes that model
|
|
311
|
+
directly — one bundle per document containing typed units (clauses,
|
|
312
|
+
tables, terms, equations, figures, requirements), a section graph, a
|
|
313
|
+
native Glossarist glossary, and Relaton bibliographic objects — and
|
|
314
|
+
replaces HTML scraping end to end for the cleanly-authored portion of
|
|
315
|
+
the corpus. The format, Metanorma Knowledge Objects (MKO), is specified
|
|
316
|
+
as Metanorma note 116 with the consumer contract (symbolic unit
|
|
317
|
+
references and typed excerpts) contributed from this work.
|
|
318
|
+
|
|
319
|
+

|
|
320
|
+
|
|
321
|
+
*Figure 3 — Producer-native ingestion. Each authored document exports as one MKO bundle; chunking, enrichment and indexing are derived from the bundle, and re-ingest is incremental by content hash.*
|
|
322
|
+
|
|
323
|
+
|
|
324
|
+
This path matters beyond OIML: any standards body whose publications are
|
|
325
|
+
authored in Metanorma gets structured, table-aware, graph-connected
|
|
326
|
+
question answering over its corpus without scraping renderings. The
|
|
327
|
+
guidelines for building such a service from scratch accompany this
|
|
328
|
+
article.
|
|
329
|
+
|
|
330
|
+
## 9. Evaluation as a discipline
|
|
331
|
+
|
|
332
|
+
The service's answers are evaluated on three axes, continuously:
|
|
333
|
+
|
|
334
|
+
- **Retrieval** (recall, precision, MRR against the golden set with
|
|
335
|
+
witness citations)
|
|
336
|
+
- **Faithfulness** (an independent judge scores whether each claim is
|
|
337
|
+
supported by the retrieved passages — with the important lesson that
|
|
338
|
+
the judge must see the passages the answer was actually built from,
|
|
339
|
+
not a fresh retrieval)
|
|
340
|
+
- **User feedback** (thumbs up/down on every answer, logged to the same
|
|
341
|
+
evaluation loop)
|
|
342
|
+
|
|
343
|
+
Every change ships through the same gate — literally one command:
|
|
344
|
+
golden suite ×3 and the capability battery ×6 against the live service,
|
|
345
|
+
where any failed run fails the command. The suites also run in CI and
|
|
346
|
+
the service's cache versions flush on every retrieval change so no
|
|
347
|
+
answer is served from a superseded index. A humility note from the
|
|
348
|
+
measurement machinery itself: the capability battery's runner once
|
|
349
|
+
omitted to set a non-zero exit code on failure, so a failing battery
|
|
350
|
+
reported green through the gate — the same class of silent failure the
|
|
351
|
+
cache-version discipline exists to prevent. Exit codes are part of the
|
|
352
|
+
measurement contract now. The gate rejects as often as it accepts.
|
|
353
|
+
A candidate change that unioned additional retrieval candidates into a
|
|
354
|
+
rewritten query's pool — a plausible-looking "more evidence" improvement
|
|
355
|
+
— dropped recall@5 from 94.3% to 89.7%: topically close but wrong
|
|
356
|
+
documents outscored the right ones under the cross-encoder, and the
|
|
357
|
+
per-case diff named the three failing questions. The fix (additive
|
|
358
|
+
candidates may only replace an identical query, never dilute a rewritten
|
|
359
|
+
one) is now a comment in the code. A second candidate — grading
|
|
360
|
+
document-scoped retrieval more aggressively — was measured, found to buy
|
|
361
|
+
nothing, and reverted the same day. Additive is not free; the suite, not
|
|
362
|
+
the author, decides.
|
|
363
|
+
|
|
364
|
+
## 10. Conclusions and outlook
|
|
365
|
+
|
|
366
|
+
A grounded, citation-linked, typed-rendering question-answering service
|
|
367
|
+
over the OIML corpus is measurable, economical, and — with the
|
|
368
|
+
producer-native ingestion path — increasingly maintained by the
|
|
369
|
+
documents' own structure rather than by scraping their renderings. The
|
|
370
|
+
open work is equally concrete: upstream bibliographic corrections for
|
|
371
|
+
the 36 families without derivable current editions; unit-level language
|
|
372
|
+
tagging so bilingual annexes inside English editions stop masquerading
|
|
373
|
+
as main text; a publisher-specific identifier flavor so citation joins
|
|
374
|
+
become exact; collection-level bundles for cross-document reasoning; and
|
|
375
|
+
interlingual unit alignment so the same clause can be answered in every
|
|
376
|
+
OIML language. The requirements behind several of these belong to the
|
|
377
|
+
authoring system itself, and we have filed them with the Metanorma
|
|
378
|
+
document model team as the producer side of an AI-native publishing
|
|
379
|
+
stack.
|
|
380
|
+
|
|
381
|
+
The service is live at ai.oimlsmart.org. Try it, and tell us when it is
|
|
382
|
+
wrong — the feedback button is the fastest path into the evaluation
|
|
383
|
+
loop that everything above runs on.
|
|
384
|
+
|
|
385
|
+
---
|
|
386
|
+
|
|
387
|
+
### References
|
|
388
|
+
|
|
389
|
+
1. Al Masoud, A., Arazzi, M., Germani, S., Nocera, A. *Exploring
|
|
390
|
+
Structural Complexity in Normative RAG with Graph-based approaches: A
|
|
391
|
+
case study on the ETSI Standards.* arXiv:2604.09868 (2026).
|
|
392
|
+
2. Guttal, P., et al. *Structure-Aware Chunking for Tabular Data in
|
|
393
|
+
Retrieval-Augmented Generation.* arXiv:2605.00318 (2026).
|
|
394
|
+
3. Anthropic. *Contextual Retrieval.* Engineering blog (2024).
|
|
395
|
+
4. Günther, M., et al. *Late Chunking: Contextual Chunk Embeddings Using
|
|
396
|
+
Long-Context Embedding Models.* arXiv:2409.04701 (2024).
|
|
397
|
+
5. Lewis, P., et al. *Retrieval-Augmented Generation for
|
|
398
|
+
Knowledge-Intensive NLP Tasks.* NeurIPS (2020).
|
|
399
|
+
6. Metanorma. *MN 116: Metanorma Knowledge Objects (MKO) machine
|
|
400
|
+
serialization format.* Metanorma documentation (2026).
|
|
401
|
+
7. Xu, L., et al. *Equipping Retrieval-Augmented LLMs with Document
|
|
402
|
+
Structure Awareness.* arXiv:2510.04293 (2025).
|