@konneal/engine 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (285) hide show
  1. package/LICENSE +29 -0
  2. package/README.md +13 -0
  3. package/dist/admin.d.ts +26 -0
  4. package/dist/ai.d.ts +6 -0
  5. package/dist/anchors.d.ts +6 -0
  6. package/dist/answercache.d.ts +22 -0
  7. package/dist/ask.d.ts +5 -0
  8. package/dist/auth.d.ts +12 -0
  9. package/dist/bubble.d.ts +14 -0
  10. package/dist/chunk-LLWPT2XV.js +49 -0
  11. package/dist/chunk-MB74PTRM.js +114 -0
  12. package/dist/chunk-WOGQM7DJ.js +197 -0
  13. package/dist/chunk-WWNCWKKC.js +42 -0
  14. package/dist/completion.d.ts +5 -0
  15. package/dist/config.d.ts +154 -0
  16. package/dist/config.js +37 -0
  17. package/dist/context.d.ts +115 -0
  18. package/dist/conversations.d.ts +5 -0
  19. package/dist/drafts.d.ts +129 -0
  20. package/dist/env.d.ts +57 -0
  21. package/dist/faithfulness.d.ts +5 -0
  22. package/dist/grader.d.ts +3 -0
  23. package/dist/graph.d.ts +13 -0
  24. package/dist/hybrid.d.ts +7 -0
  25. package/dist/index.d.ts +9 -0
  26. package/dist/index.js +5373 -0
  27. package/dist/internal_gateway.d.ts +14 -0
  28. package/dist/lexical.d.ts +7 -0
  29. package/dist/livedata.d.ts +77 -0
  30. package/dist/memories.d.ts +10 -0
  31. package/dist/modelplane.d.ts +61 -0
  32. package/dist/oidc.d.ts +73 -0
  33. package/dist/pipeline.d.ts +57 -0
  34. package/dist/profile.d.ts +2 -0
  35. package/dist/profile.gen.d.ts +70 -0
  36. package/dist/profile.js +8 -0
  37. package/dist/projects.d.ts +8 -0
  38. package/dist/prompts/conversational.md +8 -0
  39. package/dist/prompts/enrichment.md +3 -0
  40. package/dist/prompts/faithfulness.md +1 -0
  41. package/dist/prompts/grader.md +5 -0
  42. package/dist/prompts/listwise.md +3 -0
  43. package/dist/prompts/precision.md +1 -0
  44. package/dist/prompts/reflect.md +1 -0
  45. package/dist/prompts/relevancy.md +1 -0
  46. package/dist/prompts/research.md +10 -0
  47. package/dist/prompts/section-summary.md +5 -0
  48. package/dist/prompts/summarize.md +1 -0
  49. package/dist/prompts/system.md +18 -0
  50. package/dist/prompts/understanding.md +17 -0
  51. package/dist/quota.d.ts +13 -0
  52. package/dist/reflect.d.ts +5 -0
  53. package/dist/refs.d.ts +40 -0
  54. package/dist/refusal.d.ts +9 -0
  55. package/dist/refusal.js +9 -0
  56. package/dist/requestScope.d.ts +26 -0
  57. package/dist/requestScope.js +10 -0
  58. package/dist/research.d.ts +8 -0
  59. package/dist/search.d.ts +4 -0
  60. package/dist/selfquery.d.ts +7 -0
  61. package/dist/session.d.ts +1 -0
  62. package/dist/share.d.ts +2 -0
  63. package/dist/structural.d.ts +27 -0
  64. package/dist/tablecontext.d.ts +11 -0
  65. package/dist/understand.d.ts +11 -0
  66. package/dist/understandContract.d.ts +29 -0
  67. package/dist/verdict.d.ts +24 -0
  68. package/docs/API.md +451 -0
  69. package/docs/ARCHITECTURE.md +302 -0
  70. package/docs/AUDIT-2026-08-24.md +71 -0
  71. package/docs/CONTRIBUTOR-AUDIT-2026-08-25.md +147 -0
  72. package/docs/INGEST-ARCHITECTURE.md +158 -0
  73. package/docs/MCP.md +92 -0
  74. package/docs/METANORMA-AI-SERIALIZATION.md +247 -0
  75. package/docs/MKO-EXPORT-PIPELINE.md +147 -0
  76. package/docs/REDESIGN-NORMATIVE-RAG-ETSI.md +485 -0
  77. package/docs/RESEARCH-SOTA-2026.md +243 -0
  78. package/docs/ROADMAP-SOTA.md +130 -0
  79. package/docs/SOTA-STAGE-SPECS.md +509 -0
  80. package/docs/annealment/F1-verdict.md +27 -0
  81. package/docs/annealment/F10-notes.md +19 -0
  82. package/docs/annealment/F11-composition.md +17 -0
  83. package/docs/annealment/F12-passport.md +17 -0
  84. package/docs/annealment/F2-counterfactual.md +20 -0
  85. package/docs/annealment/F3-absence.md +21 -0
  86. package/docs/annealment/F4-instance.md +18 -0
  87. package/docs/annealment/F5-workflow.md +21 -0
  88. package/docs/annealment/F6-impact.md +21 -0
  89. package/docs/annealment/F7-editions.md +18 -0
  90. package/docs/annealment/F8-selfverify.md +19 -0
  91. package/docs/annealment/F9-projection-qa.md +17 -0
  92. package/docs/annealment/L0-locate.md +19 -0
  93. package/docs/annealment/L1-extract.md +18 -0
  94. package/docs/annealment/L2-nomenclature.md +22 -0
  95. package/docs/annealment/L3-geometry.md +23 -0
  96. package/docs/annealment/L4-composition.md +21 -0
  97. package/docs/annealment/L5-cross-standard.md +20 -0
  98. package/docs/annealment/L6-diachrony.md +21 -0
  99. package/docs/annealment/L7-perception.md +20 -0
  100. package/docs/annealment/L8-computation.md +22 -0
  101. package/docs/annealment/L9-instance-process.md +23 -0
  102. package/docs/annealment/README.md +10 -0
  103. package/docs/guidelines-metanorma-ai-programme.md +279 -0
  104. package/docs/identity-onboarding-rag.md +65 -0
  105. package/docs/identity-service.md +219 -0
  106. package/docs/knowledge-annealment.md +273 -0
  107. package/docs/konneal-extraction-plan.md +481 -0
  108. package/docs/metanorma-for-ai.md +270 -0
  109. package/docs/mirror-plan.md +36 -0
  110. package/docs/multi-sdo-architecture.md +191 -0
  111. package/docs/paper-annealment-comparison.md +259 -0
  112. package/docs/paper-assets/architecture.svg +94 -0
  113. package/docs/paper-assets/contract-v2.svg +94 -0
  114. package/docs/paper-assets/mko-ingest.svg +91 -0
  115. package/docs/paper-oiml-bulletin.md +402 -0
  116. package/docs/paper-oiml-bulletin.mdx +419 -0
  117. package/docs/product-branding-options.md +172 -0
  118. package/docs/projects-design.md +88 -0
  119. package/docs/sota-mechanisms.md +184 -0
  120. package/docs/spec-api.md +77 -0
  121. package/docs/spec-pipeline.md +126 -0
  122. package/docs/vector-adapter.md +88 -0
  123. package/package.json +70 -0
  124. package/profile/corpora.yaml +5 -0
  125. package/profile/datasets.yaml +14 -0
  126. package/profile/prompts.yaml +5 -0
  127. package/profile/publisher.yaml +17 -0
  128. package/profile/retrieval.yaml +1 -0
  129. package/profile/sources.yaml +5 -0
  130. package/profile/ui.yaml +7 -0
  131. package/scripts/gen_profile.mjs +33 -0
  132. package/workers/shared/ai.ts +21 -0
  133. package/workers/shared/auth.ts +16 -0
  134. package/workers/shared/chunk.ts +108 -0
  135. package/workers/shared/oidc.ts +312 -0
  136. package/workers/shared/router.ts +45 -0
  137. package/workers/shared/session.ts +104 -0
  138. package/workers/worker_internal/src/index.ts +157 -0
  139. package/workers/worker_internal/tsconfig.json +15 -0
  140. package/workers/worker_internal/wrangler.toml +32 -0
  141. package/workers/worker_mcp/src/index.ts +175 -0
  142. package/workers/worker_mcp/tsconfig.json +13 -0
  143. package/workers/worker_mcp/wrangler.toml +18 -0
  144. package/workers/worker_public/migrations/0002_conversations.sql +22 -0
  145. package/workers/worker_public/migrations/0003_shared_conversations.sql +9 -0
  146. package/workers/worker_public/migrations/0004_graph.sql +16 -0
  147. package/workers/worker_public/migrations/0005_documents.sql +19 -0
  148. package/workers/worker_public/migrations/0006_conversation_entities.sql +11 -0
  149. package/workers/worker_public/migrations/0007_chunks_fts.sql +43 -0
  150. package/workers/worker_public/migrations/0008_unit_payloads.sql +15 -0
  151. package/workers/worker_public/migrations/0009_chunks_unit.sql +8 -0
  152. package/workers/worker_public/migrations/0009_message_context.sql +7 -0
  153. package/workers/worker_public/migrations/0010_model_nodes.sql +33 -0
  154. package/workers/worker_public/migrations/0011_schema_union.sql +38 -0
  155. package/workers/worker_public/migrations/0012_memories.sql +15 -0
  156. package/workers/worker_public/migrations/0013_projects.sql +21 -0
  157. package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.d.ts +16306 -0
  158. package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.ts +16261 -0
  159. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.d.ts +16373 -0
  160. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.ts +16328 -0
  161. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.d.ts +16382 -0
  162. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.ts +16337 -0
  163. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.d.ts +16383 -0
  164. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.ts +16338 -0
  165. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.d.ts +16403 -0
  166. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.ts +16358 -0
  167. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.d.ts +16408 -0
  168. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.ts +16363 -0
  169. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.d.ts +16414 -0
  170. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.ts +16369 -0
  171. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.d.ts +16414 -0
  172. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.ts +16369 -0
  173. package/workers/worker_public/node_modules/@cloudflare/workers-types/README.md +135 -0
  174. package/workers/worker_public/node_modules/@cloudflare/workers-types/entrypoints.svg +53 -0
  175. package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.d.ts +17095 -0
  176. package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.ts +17050 -0
  177. package/workers/worker_public/node_modules/@cloudflare/workers-types/index.d.ts +16306 -0
  178. package/workers/worker_public/node_modules/@cloudflare/workers-types/index.ts +16261 -0
  179. package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.d.ts +16447 -0
  180. package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.ts +16402 -0
  181. package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.d.ts +16306 -0
  182. package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.ts +16261 -0
  183. package/workers/worker_public/node_modules/@cloudflare/workers-types/package.json +11 -0
  184. package/workers/worker_public/package.json +13 -0
  185. package/workers/worker_public/prompts/conversational.md +8 -0
  186. package/workers/worker_public/prompts/enrichment.md +3 -0
  187. package/workers/worker_public/prompts/faithfulness.md +1 -0
  188. package/workers/worker_public/prompts/grader.md +5 -0
  189. package/workers/worker_public/prompts/listwise.md +3 -0
  190. package/workers/worker_public/prompts/precision.md +1 -0
  191. package/workers/worker_public/prompts/reflect.md +1 -0
  192. package/workers/worker_public/prompts/relevancy.md +1 -0
  193. package/workers/worker_public/prompts/research.md +10 -0
  194. package/workers/worker_public/prompts/section-summary.md +5 -0
  195. package/workers/worker_public/prompts/summarize.md +1 -0
  196. package/workers/worker_public/prompts/system.md +18 -0
  197. package/workers/worker_public/prompts/understanding.md +17 -0
  198. package/workers/worker_public/public/app.js +166 -0
  199. package/workers/worker_public/public/index.html +48 -0
  200. package/workers/worker_public/public/style.css +147 -0
  201. package/workers/worker_public/schema.sql +248 -0
  202. package/workers/worker_public/src/admin.ts +358 -0
  203. package/workers/worker_public/src/ai.ts +71 -0
  204. package/workers/worker_public/src/anchors.ts +41 -0
  205. package/workers/worker_public/src/answercache.ts +72 -0
  206. package/workers/worker_public/src/ask.ts +1094 -0
  207. package/workers/worker_public/src/auth.ts +252 -0
  208. package/workers/worker_public/src/bubble.ts +111 -0
  209. package/workers/worker_public/src/completion.ts +75 -0
  210. package/workers/worker_public/src/config.ts +238 -0
  211. package/workers/worker_public/src/context.ts +238 -0
  212. package/workers/worker_public/src/conversations.ts +162 -0
  213. package/workers/worker_public/src/drafts.ts +497 -0
  214. package/workers/worker_public/src/env.ts +90 -0
  215. package/workers/worker_public/src/faithfulness.ts +63 -0
  216. package/workers/worker_public/src/grader.ts +89 -0
  217. package/workers/worker_public/src/graph.ts +63 -0
  218. package/workers/worker_public/src/hybrid.ts +77 -0
  219. package/workers/worker_public/src/index.ts +441 -0
  220. package/workers/worker_public/src/internal_gateway.ts +41 -0
  221. package/workers/worker_public/src/lexical.ts +86 -0
  222. package/workers/worker_public/src/lib/hit.ts +4 -0
  223. package/workers/worker_public/src/lib/http.ts +83 -0
  224. package/workers/worker_public/src/lib/router.ts +4 -0
  225. package/workers/worker_public/src/livedata.ts +334 -0
  226. package/workers/worker_public/src/memories.ts +81 -0
  227. package/workers/worker_public/src/modelplane.ts +213 -0
  228. package/workers/worker_public/src/oidc.ts +333 -0
  229. package/workers/worker_public/src/pipeline.ts +377 -0
  230. package/workers/worker_public/src/ports/blobs.ts +7 -0
  231. package/workers/worker_public/src/ports/cloudflare/adapters.ts +177 -0
  232. package/workers/worker_public/src/ports/kv.ts +8 -0
  233. package/workers/worker_public/src/ports/model.ts +28 -0
  234. package/workers/worker_public/src/ports/runtime.ts +13 -0
  235. package/workers/worker_public/src/ports/store.ts +20 -0
  236. package/workers/worker_public/src/ports/vector.ts +26 -0
  237. package/workers/worker_public/src/profile.gen.ts +101 -0
  238. package/workers/worker_public/src/profile.ts +16 -0
  239. package/workers/worker_public/src/projects.ts +108 -0
  240. package/workers/worker_public/src/prompts.d.ts +6 -0
  241. package/workers/worker_public/src/quota.ts +54 -0
  242. package/workers/worker_public/src/reflect.ts +67 -0
  243. package/workers/worker_public/src/refs.ts +107 -0
  244. package/workers/worker_public/src/refusal.ts +65 -0
  245. package/workers/worker_public/src/requestScope.ts +71 -0
  246. package/workers/worker_public/src/research.ts +126 -0
  247. package/workers/worker_public/src/search.ts +56 -0
  248. package/workers/worker_public/src/selfquery.ts +25 -0
  249. package/workers/worker_public/src/session.ts +4 -0
  250. package/workers/worker_public/src/share.ts +53 -0
  251. package/workers/worker_public/src/stages/conceptGraph.ts +68 -0
  252. package/workers/worker_public/src/stages/conceptSteer.ts +39 -0
  253. package/workers/worker_public/src/stages/corpusScope.ts +25 -0
  254. package/workers/worker_public/src/stages/dedup.ts +10 -0
  255. package/workers/worker_public/src/stages/dense.ts +73 -0
  256. package/workers/worker_public/src/stages/diversity.ts +33 -0
  257. package/workers/worker_public/src/stages/editionCover.ts +63 -0
  258. package/workers/worker_public/src/stages/editionSteer.ts +88 -0
  259. package/workers/worker_public/src/stages/familyBoost.ts +22 -0
  260. package/workers/worker_public/src/stages/federate.ts +22 -0
  261. package/workers/worker_public/src/stages/glossary.ts +65 -0
  262. package/workers/worker_public/src/stages/graphLane.ts +31 -0
  263. package/workers/worker_public/src/stages/hyde.ts +29 -0
  264. package/workers/worker_public/src/stages/index.ts +69 -0
  265. package/workers/worker_public/src/stages/lexicalUnion.ts +21 -0
  266. package/workers/worker_public/src/stages/multiQuery.ts +57 -0
  267. package/workers/worker_public/src/stages/overviewDemote.ts +14 -0
  268. package/workers/worker_public/src/stages/poolOpen.ts +10 -0
  269. package/workers/worker_public/src/stages/propagate.ts +15 -0
  270. package/workers/worker_public/src/stages/rerank.ts +47 -0
  271. package/workers/worker_public/src/stages/seal.ts +16 -0
  272. package/workers/worker_public/src/stages/sectionDescent.ts +61 -0
  273. package/workers/worker_public/src/stages/stdRefNudge.ts +35 -0
  274. package/workers/worker_public/src/stages/subQuery.ts +42 -0
  275. package/workers/worker_public/src/stages/termNudge.ts +24 -0
  276. package/workers/worker_public/src/stages/typedPin.ts +131 -0
  277. package/workers/worker_public/src/stages/types.ts +112 -0
  278. package/workers/worker_public/src/stages/windowFloor.ts +23 -0
  279. package/workers/worker_public/src/structural.ts +171 -0
  280. package/workers/worker_public/src/tablecontext.ts +41 -0
  281. package/workers/worker_public/src/understand.ts +72 -0
  282. package/workers/worker_public/src/understandContract.ts +67 -0
  283. package/workers/worker_public/src/verdict.ts +255 -0
  284. package/workers/worker_public/tsconfig.json +18 -0
  285. package/workers/worker_public/wrangler.toml +104 -0
@@ -0,0 +1,402 @@
1
+ *Draft article for the OIML Bulletin — 2026-08-30. Style: Bulletin
2
+ technical article (MS Word single-column on submission; this is the
3
+ authoring source). Numbers are production-measured unless noted.*
4
+
5
+ ---
6
+
7
+ ## Abstract
8
+
9
+ The OIML publishes close to 900 documents — Recommendations, Basic
10
+ publications, Guides, Documents and Expert reports — that together form
11
+ the reference library of legal metrology. Finding what they say has
12
+ required knowing the corpus before querying it. We describe a
13
+ question-answering service, **OIML SMART AI** (ai.oimlsmart.org), that
14
+ answers natural-language questions from the indexed publications with
15
+ clause-level citations, verbatim quote anchors for normative values, and
16
+ typed renderings of the tables and equations the answers depend on — and
17
+ it reads figures and user-supplied photographs directly, from pixels. The
18
+ service measures every component it ships: a golden question set,
19
+ retrieval metrics in the tradition of the information-retrieval
20
+ literature, and judged faithfulness. We report the measured design
21
+ decisions, the corpus-structure findings that shaped them, and the open
22
+ path: producer-native ingestion of the Metanorma document model so that
23
+ any Metanorma-authored corpus can be indexed the same way.
24
+
25
+ ---
26
+
27
+ ## 1. Introduction
28
+
29
+ Legal metrology's knowledge lives in a large, cross-referential,
30
+ multi-edition corpus. A typical working question — *"what is the maximum
31
+ permissible error for a class III nonautomatic weighing instrument?"* —
32
+ requires knowing which Recommendation applies (OIML R 76-1), which
33
+ edition is current, which clause holds the table, and how to read the
34
+ table's rows against the instrument's verification scale intervals.
35
+ Every one of those steps assumes corpus knowledge the questioner may not
36
+ have.
37
+
38
+ Question answering over such corpora has been transformed by
39
+ retrieval-augmented generation (RAG): a language model is given passages
40
+ retrieved from a trusted corpus and may use only those. The engineering
41
+ challenge is no longer whether a model can write fluent prose — it can —
42
+ but whether the *right passages* are found, whether the answer stays
43
+ *faithful* to them, and whether the reader can *verify* the result. Our
44
+ service is built around those three properties.
45
+
46
+ This article describes the architecture as deployed, the measurements
47
+ that shaped it, and the corpus findings — including errors in the
48
+ bibliographic record itself — that surfaced because the system keeps
49
+ score. It closes with the producer-native ingestion path (Metanorma
50
+ Knowledge Objects) and what it makes possible for other standards bodies.
51
+
52
+ ## 2. What the service does
53
+
54
+ A user asks a question in any language. The service:
55
+
56
+ 1. **Understands the question** with a small language model — language,
57
+ named publication, edition, defined terms, complexity. There are no
58
+ keyword rules; the same model makes these judgements for every
59
+ question.
60
+ 2. **Retrieves in parallel** — three lanes leave the question at the
61
+ same instant: the vector index is queried with the question's
62
+ embedding while understanding is still running (the dense lane's
63
+ results are reused unchanged when understanding adds nothing), a
64
+ full-corpus keyword index catches exact identifiers, part numbers
65
+ and defined terms that vector similarity alone misses, and
66
+ understanding's refinements — a named publication, an edition —
67
+ arrive as metadata filters on the dense query.
68
+ 3. **Ranks** the fused candidates with a cross-encoder; complex
69
+ questions get a second, stronger re-ranking pass.
70
+ 4. **Answers from the retrieved passages only**, under a contract that
71
+ requires inline citations (`[OIML R 60-1:2021 §4.4.2]`) and verbatim
72
+ quote anchors for normative values.
73
+ 5. **Verifies** the answer after generation: a deterministic check that
74
+ every quoted phrase exists in a cited passage; a separate judge that
75
+ scores faithfulness; a correction-and-retry when either fails. An
76
+ answer that cannot be verified is never written to the answer cache.
77
+
78
+ ![Figure 1 — A question's path through the service. Three retrieval lanes run concurrently; every answer passes mechanical verification before it is served or cached. The edition registry derives publication status from supersession links, never from a status field.](paper-assets/architecture.svg)
79
+
80
+ *Figure 1 — A question's path through the service. Three retrieval lanes run concurrently; every answer passes mechanical verification before it is served or cached. The edition registry derives publication status from supersession links, never from a status field.*
81
+
82
+
83
+ When the question names a publication, answers are steered to the
84
+ current edition by a registry derived from the bibliographic record's
85
+ supersession links (§5). When the answer depends on a table or
86
+ equation, the interface renders the *producer's* typed object — the
87
+ actual rows and columns — never the model's re-typing of it (§6).
88
+
89
+ If the corpus does not contain the answer, the system says so in one
90
+ canonical sentence and redirects; refusals are never cached, because a
91
+ refusal is a property of the moment, not of the question.
92
+
93
+ ## 3. Corpus and structure
94
+
95
+ The indexed corpus comprises the English editions of the OIML
96
+ publications (~900 documents, ~9.3 million words). Documents are chunked
97
+ along clause boundaries — never arbitrary token windows — and each chunk
98
+ carries provenance: publication identifier, edition, language, clause
99
+ anchor, and bibliographic status.
100
+
101
+ Three corpus-level findings shaped the design:
102
+
103
+ **The bibliographic status field disagrees with its own record.** In 58
104
+ of 224 publication families, the machine-readable status field claims a
105
+ document is current while the same record's supersession links say
106
+ otherwise. The service therefore *derives* status from the links and
107
+ ignores the field. More consequentially, 36 of 224 families have no
108
+ recorded successor link at all, so the current edition cannot be derived
109
+ for them; the registry surfaces these gaps rather than guessing. Both
110
+ numbers are the worklist for upstream bibliographic corrections — the
111
+ kind that serving surfaces only because it keeps score.
112
+
113
+ **Identifiers are identity, not display.** A cross-referenced audit of
114
+ the dirty corpus found documents whose machine-readable identifier was
115
+ a placeholder ("OIML D 0:0000", "OIML D X") or polluted with language
116
+ markers — around 200 records — making them unreachable by
117
+ doc-scoped retrieval. Identifiers are now derived from slug identity
118
+ with placeholders never winning, and the index is reconciled against
119
+ the canonical chunk set: the first census measured 49,553 live vectors
120
+ against 31,512 canonical — 18,041 strays from every prior re-chunking,
121
+ deleted in place (upserts never delete). A retrieval service that never
122
+ enumerates its own index serves ghosts.
123
+
124
+ **Tables are where the normative values live.** Maximum permissible
125
+ errors, accuracy-class limits, verification interval bounds — the values
126
+ practitioners ask for — are tabular. Flattening tables into prose loses
127
+ the row/column geometry that makes them answerable. The index now treats
128
+ tables as atomic typed objects (§6).
129
+
130
+ **Structure is a retrieval signal.** Standards are hierarchically
131
+ organized with extensive cross-references; the system indexes the
132
+ section hierarchy and the citation graph as a queryable graph (7,128
133
+ nodes, 6,486 edges) alongside the text, so a defined term can be
134
+ resolved to the documents that define it even when the question's
135
+ vocabulary does not match the corpus's.
136
+
137
+ ## 4. Retrieval: what the measurements showed
138
+
139
+ We evaluate retrieval with a golden question set (now 29 questions
140
+ including eight table-value questions with witness rows) using
141
+ recall@5, average precision@5, and mean reciprocal rank@5 — the protocol
142
+ of the standards-retrieval literature. Single runs proved noisy (±5
143
+ percentage points); all reported numbers are means over three runs on a
144
+ quiet account.
145
+
146
+ The single largest measured improvement came from making keyword
147
+ retrieval a **full-corpus first stage** rather than a re-scoring of the
148
+ vector results: recall@5 rose from 86% to 95%. The reason is structural
149
+ — standards vocabulary is exact ("n_LC", "R 60-3", "creep"), and dense
150
+ embeddings alone miss exact identifiers that a lexical scan of the whole
151
+ corpus recovers. Fusing both rankings (reciprocal rank fusion) lifted
152
+ precision and MRR without costing recall.
153
+
154
+ The second improvement came from **contextual enrichment**: a one-time
155
+ pass that prepends a short model-written preamble to every chunk ("this
156
+ clause defines the accuracy-class limits for load cells") before
157
+ embedding and indexing. This is the contextual-retrieval recipe
158
+ validated industry-wide; our corpus measurement confirmed it, and the
159
+ one-time cost (~$0.0017 per chunk) is amortized over every future
160
+ answer.
161
+
162
+ With the typed-table lane live (§6), retrieval lanes running
163
+ concurrently with query understanding, edition steering active (§5) and
164
+ multimodal generation on (§6), the current baseline is:
165
+
166
+ | Metric | Value |
167
+ |---|---|
168
+ | Recall@5 (mean of 3) | **94.3%** (range 93.1–96.6) |
169
+ | Average precision@5 | 0.82 |
170
+ | MRR@5 | 0.87 |
171
+ | End-to-end golden cases | 14/14 |
172
+
173
+ These numbers are on our own golden set, not on the benchmarks of the
174
+ works we build on — the honest comparison is per-technique, on the same
175
+ corpus, before and after. So read: the full-corpus lexical stage is the
176
+ ETSI study's recipe (ref 1) — adopting it lifted recall@5 here from 86%
177
+ to 95%; the contextual preambles are Anthropic's contextual retrieval
178
+ (ref 3) — confirmed on this corpus at ~$0.0017 per chunk one-time; the
179
+ clause-boundary, structure-preserving chunking follows the same
180
+ structure-aware line as refs 2 and 7, which our typed-unit pin extends
181
+ from "chunk better" to "guarantee the answering object a slot." Two
182
+ elements of the deployed system have no counterpart in the cited work:
183
+ the symbolic-reference contract, under which table data never passes
184
+ through the model at all (refs 2's error-reduction approach still
185
+ re-generates tables; we removed the corruption channel), and the
186
+ mechanical post-generation verification of every answer (quote anchors,
187
+ unit references, retyped-table detection) with a faithfulness judge that
188
+ sees the passages actually used.
189
+
190
+ Two structural additions complete the retrieval story. First, the
191
+ corpus is a tree — every chunk carries its clause anchor — and the
192
+ serving path uses the tree: a hit's score blends its neighbouring
193
+ clauses' scores, evidence is presented to the answer model in document
194
+ reading order, and a ranked section summary resolves to its quotable
195
+ child clauses (adapted from FABLE/BEAR, ref 7, at none of its index-time
196
+ cost, because Metanorma documents arrive as trees rather than needing
197
+ one inferred). Second, everyday words rarely match defined terms — the
198
+ one gap that failed for every corpus representation equally — so a
199
+ terminology index of the corpus's defined concepts links a question's
200
+ phrasing to candidate terms, and the answer model adjudicates and leads
201
+ with the corpus's own term, quoting its definition ("my output keeps
202
+ drifting" resolves to *span stability*, with the verbatim definition and
203
+ its defining clause cited).
204
+
205
+ ## 5. Editions and trust
206
+
207
+ Which edition applies is as consequential as what it says. The service's
208
+ registry of publication families and editions is derived from
209
+ supersession links, not from status fields (§3). Citations carry the
210
+ edition and status of every source; superseded sources are marked.
211
+ Answers to questions that name a publication are steered to the current
212
+ edition; answers that must cite a superseded edition (the 2021 passages
213
+ do not reproduce a table the 2017 edition carries) say so explicitly.
214
+
215
+ Steering is measured, not assumed, and two mechanisms earned their place
216
+ by fixing observed failures. Superseded editions match archaic phrasing
217
+ strongly — a complex verification question once cited R 76-1:1992/1988
218
+ while the current 2006 edition was crowded out of the passage window —
219
+ so ranking now demotes older-edition chunks of a publication whenever a
220
+ newer edition of the same publication is present in the evidence
221
+ (cross-publication recency is deliberately untouched: a 1992 document
222
+ that is still current must not lose to unrelated 2024 ones). And because
223
+ query understanding can pin an edition for the wrong family, an edition
224
+ pin that the corpus itself fails to corroborate is dropped to the
225
+ document-level filter. After both fixes the probe cites R 76-1:2006 in
226
+ force, with one legitimate 1988 tail citation for a clause only that
227
+ edition carries; the window's measured average precision (0.85) and MRR
228
+ (0.88) are the best the service has recorded.
229
+
230
+ ## 6. Tables, equations, and figures as typed objects
231
+
232
+ The most consequential rendering decision: the language model never
233
+ re-types normative data. When an answer depends on a table, the model
234
+ emits a symbolic reference to the table's unit (`[[u:table-1]]`); the
235
+ server validates the reference against the passages actually used and
236
+ resolves it to the producer's typed payload — columns, rows, units. The
237
+ interface renders the exact object. Small verbatim values in prose are
238
+ permitted and mechanically verified against the source. The payload
239
+ never passes through the model's output, so it cannot be corrupted by
240
+ generation.
241
+
242
+ The same principle extends from rendering to judgment: the model never
243
+ judges what a machine can compute. Where the corpus's machine-readable
244
+ models carry checkable objects — a constraint with an OCL expression, a
245
+ requirement with a threshold limit — a question naming one gets it
246
+ EXECUTED: the question's values bind to the rule's own parameters, the
247
+ check evaluates deterministically, and the verdict (pass, or the
248
+ standard's own violation word, with the arithmetic shown, or void naming
249
+ exactly what is missing) is attached to the answer as computed data.
250
+ Two sibling capabilities complete the set: provable absence (a topic
251
+ that a standard does not constrain returns an enumeration certificate
252
+ over its entire model plane — N nodes enumerated, zero matches — rather
253
+ than a refusal) and answer verification (any answer checkable against
254
+ the corpus: verbatim-quote containment, object-reference resolution, a
255
+ judged faithfulness score labeled as judged).
256
+
257
+ ![Figure 2 — The answer contract. The model writes prose with inline citations and symbolic references; three mechanical checks run after generation; the server resolves surviving references to the producer's payloads and the client renders them.](paper-assets/contract-v2.svg)
258
+
259
+ *Figure 2 — The answer contract. The model writes prose with inline citations and symbolic references; three mechanical checks run after generation; the server resolves surviving references to the producer's payloads and the client renders them.*
260
+
261
+
262
+ Figures are described by a vision-capable model that reads the actual
263
+ image pixels; equations are carried in their authoring-native formats
264
+ (AsciiMath and MathML) with a plain-language description. Both arrive as typed objects with the same validation.
265
+
266
+ Figures also flow the other way: when the retrieved passages contain
267
+ figure units, their actual images are attached to the answer model's
268
+ input, so descriptions and reasoning come from the drawing itself — the
269
+ model reads labels that exist only in the pixels (the pixel-label
270
+ probe — "what are the labeled example cases in the R 60-2 design-shapes
271
+ figure?" — is answered A, B and C from the drawing, stable across
272
+ repeated runs). Two measured invariants keep that true: assets must be
273
+ readable by vision pipelines (vector-sourced rasters that draw black on
274
+ transparent alpha flatten to a solid black rectangle inside a vision
275
+ model, whatever a browser shows — a corpus sweep detects and repairs
276
+ them), and images ride their own message in the generation call (long
277
+ passage text and image parts in one message triggers provider errors
278
+ that scale with payload). Users can likewise attach a photograph (an
279
+ instrument nameplate, a scale dial, a schematic) to their question; the
280
+ text still drives retrieval, and the image gives the model the visual
281
+ context, under the same citation contract.
282
+
283
+ ## 7. Models, cost, and sovereignty
284
+
285
+ All models are open-weight and served on a single cloud platform
286
+ (Cloudflare Workers AI); no corpus data is sent to proprietary model
287
+ providers. The answer model (GLM-5.3 Flash, natively multimodal) serves
288
+ all tiers at approximately $0.001 per answer. The full one-time corpus
289
+ preparation — contextual enrichment of 42,000 chunks — cost
290
+ approximately $75. Daily serving at current traffic is under $5 per
291
+ month including infrastructure. The model policy is deliberately
292
+ cost-first on the serving path (every question pays it) and quality-first
293
+ on one-time work (enrichment, captioning, evaluation), where quality
294
+ persists into every future answer.
295
+
296
+ One lesson generalizes: read the model card before wiring a model. Three
297
+ separate live failures traced to defaults we never set — a
298
+ reasoning-effort parameter that silently defaults to maximum (starving
299
+ small output budgets), sampling defaults that let a thinking model loop
300
+ (repetition consuming the token budget that carried the structured
301
+ output), and a degraded non-thinking mode we were unknowingly paying
302
+ for. Every call site now states its reasoning mode, its sampling, and a
303
+ budget the reasoning cannot starve; the discipline costs nothing and
304
+ removed an entire class of silent failure.
305
+
306
+ ## 8. The producer-native path: Metanorma Knowledge Objects
307
+
308
+ OIML publications are authored in Metanorma, a model-driven document
309
+ system: the source of truth is a typed document model, not any rendered
310
+ PDF or HTML. The service's newest ingestion path consumes that model
311
+ directly — one bundle per document containing typed units (clauses,
312
+ tables, terms, equations, figures, requirements), a section graph, a
313
+ native Glossarist glossary, and Relaton bibliographic objects — and
314
+ replaces HTML scraping end to end for the cleanly-authored portion of
315
+ the corpus. The format, Metanorma Knowledge Objects (MKO), is specified
316
+ as Metanorma note 116 with the consumer contract (symbolic unit
317
+ references and typed excerpts) contributed from this work.
318
+
319
+ ![Figure 3 — Producer-native ingestion. Each authored document exports as one MKO bundle; chunking, enrichment and indexing are derived from the bundle, and re-ingest is incremental by content hash.](paper-assets/mko-ingest.svg)
320
+
321
+ *Figure 3 — Producer-native ingestion. Each authored document exports as one MKO bundle; chunking, enrichment and indexing are derived from the bundle, and re-ingest is incremental by content hash.*
322
+
323
+
324
+ This path matters beyond OIML: any standards body whose publications are
325
+ authored in Metanorma gets structured, table-aware, graph-connected
326
+ question answering over its corpus without scraping renderings. The
327
+ guidelines for building such a service from scratch accompany this
328
+ article.
329
+
330
+ ## 9. Evaluation as a discipline
331
+
332
+ The service's answers are evaluated on three axes, continuously:
333
+
334
+ - **Retrieval** (recall, precision, MRR against the golden set with
335
+ witness citations)
336
+ - **Faithfulness** (an independent judge scores whether each claim is
337
+ supported by the retrieved passages — with the important lesson that
338
+ the judge must see the passages the answer was actually built from,
339
+ not a fresh retrieval)
340
+ - **User feedback** (thumbs up/down on every answer, logged to the same
341
+ evaluation loop)
342
+
343
+ Every change ships through the same gate — literally one command:
344
+ golden suite ×3 and the capability battery ×6 against the live service,
345
+ where any failed run fails the command. The suites also run in CI and
346
+ the service's cache versions flush on every retrieval change so no
347
+ answer is served from a superseded index. A humility note from the
348
+ measurement machinery itself: the capability battery's runner once
349
+ omitted to set a non-zero exit code on failure, so a failing battery
350
+ reported green through the gate — the same class of silent failure the
351
+ cache-version discipline exists to prevent. Exit codes are part of the
352
+ measurement contract now. The gate rejects as often as it accepts.
353
+ A candidate change that unioned additional retrieval candidates into a
354
+ rewritten query's pool — a plausible-looking "more evidence" improvement
355
+ — dropped recall@5 from 94.3% to 89.7%: topically close but wrong
356
+ documents outscored the right ones under the cross-encoder, and the
357
+ per-case diff named the three failing questions. The fix (additive
358
+ candidates may only replace an identical query, never dilute a rewritten
359
+ one) is now a comment in the code. A second candidate — grading
360
+ document-scoped retrieval more aggressively — was measured, found to buy
361
+ nothing, and reverted the same day. Additive is not free; the suite, not
362
+ the author, decides.
363
+
364
+ ## 10. Conclusions and outlook
365
+
366
+ A grounded, citation-linked, typed-rendering question-answering service
367
+ over the OIML corpus is measurable, economical, and — with the
368
+ producer-native ingestion path — increasingly maintained by the
369
+ documents' own structure rather than by scraping their renderings. The
370
+ open work is equally concrete: upstream bibliographic corrections for
371
+ the 36 families without derivable current editions; unit-level language
372
+ tagging so bilingual annexes inside English editions stop masquerading
373
+ as main text; a publisher-specific identifier flavor so citation joins
374
+ become exact; collection-level bundles for cross-document reasoning; and
375
+ interlingual unit alignment so the same clause can be answered in every
376
+ OIML language. The requirements behind several of these belong to the
377
+ authoring system itself, and we have filed them with the Metanorma
378
+ document model team as the producer side of an AI-native publishing
379
+ stack.
380
+
381
+ The service is live at ai.oimlsmart.org. Try it, and tell us when it is
382
+ wrong — the feedback button is the fastest path into the evaluation
383
+ loop that everything above runs on.
384
+
385
+ ---
386
+
387
+ ### References
388
+
389
+ 1. Al Masoud, A., Arazzi, M., Germani, S., Nocera, A. *Exploring
390
+ Structural Complexity in Normative RAG with Graph-based approaches: A
391
+ case study on the ETSI Standards.* arXiv:2604.09868 (2026).
392
+ 2. Guttal, P., et al. *Structure-Aware Chunking for Tabular Data in
393
+ Retrieval-Augmented Generation.* arXiv:2605.00318 (2026).
394
+ 3. Anthropic. *Contextual Retrieval.* Engineering blog (2024).
395
+ 4. Günther, M., et al. *Late Chunking: Contextual Chunk Embeddings Using
396
+ Long-Context Embedding Models.* arXiv:2409.04701 (2024).
397
+ 5. Lewis, P., et al. *Retrieval-Augmented Generation for
398
+ Knowledge-Intensive NLP Tasks.* NeurIPS (2020).
399
+ 6. Metanorma. *MN 116: Metanorma Knowledge Objects (MKO) machine
400
+ serialization format.* Metanorma documentation (2026).
401
+ 7. Xu, L., et al. *Equipping Retrieval-Augmented LLMs with Document
402
+ Structure Awareness.* arXiv:2510.04293 (2025).