@konneal/engine 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (285) hide show
  1. package/LICENSE +29 -0
  2. package/README.md +13 -0
  3. package/dist/admin.d.ts +26 -0
  4. package/dist/ai.d.ts +6 -0
  5. package/dist/anchors.d.ts +6 -0
  6. package/dist/answercache.d.ts +22 -0
  7. package/dist/ask.d.ts +5 -0
  8. package/dist/auth.d.ts +12 -0
  9. package/dist/bubble.d.ts +14 -0
  10. package/dist/chunk-LLWPT2XV.js +49 -0
  11. package/dist/chunk-MB74PTRM.js +114 -0
  12. package/dist/chunk-WOGQM7DJ.js +197 -0
  13. package/dist/chunk-WWNCWKKC.js +42 -0
  14. package/dist/completion.d.ts +5 -0
  15. package/dist/config.d.ts +154 -0
  16. package/dist/config.js +37 -0
  17. package/dist/context.d.ts +115 -0
  18. package/dist/conversations.d.ts +5 -0
  19. package/dist/drafts.d.ts +129 -0
  20. package/dist/env.d.ts +57 -0
  21. package/dist/faithfulness.d.ts +5 -0
  22. package/dist/grader.d.ts +3 -0
  23. package/dist/graph.d.ts +13 -0
  24. package/dist/hybrid.d.ts +7 -0
  25. package/dist/index.d.ts +9 -0
  26. package/dist/index.js +5373 -0
  27. package/dist/internal_gateway.d.ts +14 -0
  28. package/dist/lexical.d.ts +7 -0
  29. package/dist/livedata.d.ts +77 -0
  30. package/dist/memories.d.ts +10 -0
  31. package/dist/modelplane.d.ts +61 -0
  32. package/dist/oidc.d.ts +73 -0
  33. package/dist/pipeline.d.ts +57 -0
  34. package/dist/profile.d.ts +2 -0
  35. package/dist/profile.gen.d.ts +70 -0
  36. package/dist/profile.js +8 -0
  37. package/dist/projects.d.ts +8 -0
  38. package/dist/prompts/conversational.md +8 -0
  39. package/dist/prompts/enrichment.md +3 -0
  40. package/dist/prompts/faithfulness.md +1 -0
  41. package/dist/prompts/grader.md +5 -0
  42. package/dist/prompts/listwise.md +3 -0
  43. package/dist/prompts/precision.md +1 -0
  44. package/dist/prompts/reflect.md +1 -0
  45. package/dist/prompts/relevancy.md +1 -0
  46. package/dist/prompts/research.md +10 -0
  47. package/dist/prompts/section-summary.md +5 -0
  48. package/dist/prompts/summarize.md +1 -0
  49. package/dist/prompts/system.md +18 -0
  50. package/dist/prompts/understanding.md +17 -0
  51. package/dist/quota.d.ts +13 -0
  52. package/dist/reflect.d.ts +5 -0
  53. package/dist/refs.d.ts +40 -0
  54. package/dist/refusal.d.ts +9 -0
  55. package/dist/refusal.js +9 -0
  56. package/dist/requestScope.d.ts +26 -0
  57. package/dist/requestScope.js +10 -0
  58. package/dist/research.d.ts +8 -0
  59. package/dist/search.d.ts +4 -0
  60. package/dist/selfquery.d.ts +7 -0
  61. package/dist/session.d.ts +1 -0
  62. package/dist/share.d.ts +2 -0
  63. package/dist/structural.d.ts +27 -0
  64. package/dist/tablecontext.d.ts +11 -0
  65. package/dist/understand.d.ts +11 -0
  66. package/dist/understandContract.d.ts +29 -0
  67. package/dist/verdict.d.ts +24 -0
  68. package/docs/API.md +451 -0
  69. package/docs/ARCHITECTURE.md +302 -0
  70. package/docs/AUDIT-2026-08-24.md +71 -0
  71. package/docs/CONTRIBUTOR-AUDIT-2026-08-25.md +147 -0
  72. package/docs/INGEST-ARCHITECTURE.md +158 -0
  73. package/docs/MCP.md +92 -0
  74. package/docs/METANORMA-AI-SERIALIZATION.md +247 -0
  75. package/docs/MKO-EXPORT-PIPELINE.md +147 -0
  76. package/docs/REDESIGN-NORMATIVE-RAG-ETSI.md +485 -0
  77. package/docs/RESEARCH-SOTA-2026.md +243 -0
  78. package/docs/ROADMAP-SOTA.md +130 -0
  79. package/docs/SOTA-STAGE-SPECS.md +509 -0
  80. package/docs/annealment/F1-verdict.md +27 -0
  81. package/docs/annealment/F10-notes.md +19 -0
  82. package/docs/annealment/F11-composition.md +17 -0
  83. package/docs/annealment/F12-passport.md +17 -0
  84. package/docs/annealment/F2-counterfactual.md +20 -0
  85. package/docs/annealment/F3-absence.md +21 -0
  86. package/docs/annealment/F4-instance.md +18 -0
  87. package/docs/annealment/F5-workflow.md +21 -0
  88. package/docs/annealment/F6-impact.md +21 -0
  89. package/docs/annealment/F7-editions.md +18 -0
  90. package/docs/annealment/F8-selfverify.md +19 -0
  91. package/docs/annealment/F9-projection-qa.md +17 -0
  92. package/docs/annealment/L0-locate.md +19 -0
  93. package/docs/annealment/L1-extract.md +18 -0
  94. package/docs/annealment/L2-nomenclature.md +22 -0
  95. package/docs/annealment/L3-geometry.md +23 -0
  96. package/docs/annealment/L4-composition.md +21 -0
  97. package/docs/annealment/L5-cross-standard.md +20 -0
  98. package/docs/annealment/L6-diachrony.md +21 -0
  99. package/docs/annealment/L7-perception.md +20 -0
  100. package/docs/annealment/L8-computation.md +22 -0
  101. package/docs/annealment/L9-instance-process.md +23 -0
  102. package/docs/annealment/README.md +10 -0
  103. package/docs/guidelines-metanorma-ai-programme.md +279 -0
  104. package/docs/identity-onboarding-rag.md +65 -0
  105. package/docs/identity-service.md +219 -0
  106. package/docs/knowledge-annealment.md +273 -0
  107. package/docs/konneal-extraction-plan.md +481 -0
  108. package/docs/metanorma-for-ai.md +270 -0
  109. package/docs/mirror-plan.md +36 -0
  110. package/docs/multi-sdo-architecture.md +191 -0
  111. package/docs/paper-annealment-comparison.md +259 -0
  112. package/docs/paper-assets/architecture.svg +94 -0
  113. package/docs/paper-assets/contract-v2.svg +94 -0
  114. package/docs/paper-assets/mko-ingest.svg +91 -0
  115. package/docs/paper-oiml-bulletin.md +402 -0
  116. package/docs/paper-oiml-bulletin.mdx +419 -0
  117. package/docs/product-branding-options.md +172 -0
  118. package/docs/projects-design.md +88 -0
  119. package/docs/sota-mechanisms.md +184 -0
  120. package/docs/spec-api.md +77 -0
  121. package/docs/spec-pipeline.md +126 -0
  122. package/docs/vector-adapter.md +88 -0
  123. package/package.json +70 -0
  124. package/profile/corpora.yaml +5 -0
  125. package/profile/datasets.yaml +14 -0
  126. package/profile/prompts.yaml +5 -0
  127. package/profile/publisher.yaml +17 -0
  128. package/profile/retrieval.yaml +1 -0
  129. package/profile/sources.yaml +5 -0
  130. package/profile/ui.yaml +7 -0
  131. package/scripts/gen_profile.mjs +33 -0
  132. package/workers/shared/ai.ts +21 -0
  133. package/workers/shared/auth.ts +16 -0
  134. package/workers/shared/chunk.ts +108 -0
  135. package/workers/shared/oidc.ts +312 -0
  136. package/workers/shared/router.ts +45 -0
  137. package/workers/shared/session.ts +104 -0
  138. package/workers/worker_internal/src/index.ts +157 -0
  139. package/workers/worker_internal/tsconfig.json +15 -0
  140. package/workers/worker_internal/wrangler.toml +32 -0
  141. package/workers/worker_mcp/src/index.ts +175 -0
  142. package/workers/worker_mcp/tsconfig.json +13 -0
  143. package/workers/worker_mcp/wrangler.toml +18 -0
  144. package/workers/worker_public/migrations/0002_conversations.sql +22 -0
  145. package/workers/worker_public/migrations/0003_shared_conversations.sql +9 -0
  146. package/workers/worker_public/migrations/0004_graph.sql +16 -0
  147. package/workers/worker_public/migrations/0005_documents.sql +19 -0
  148. package/workers/worker_public/migrations/0006_conversation_entities.sql +11 -0
  149. package/workers/worker_public/migrations/0007_chunks_fts.sql +43 -0
  150. package/workers/worker_public/migrations/0008_unit_payloads.sql +15 -0
  151. package/workers/worker_public/migrations/0009_chunks_unit.sql +8 -0
  152. package/workers/worker_public/migrations/0009_message_context.sql +7 -0
  153. package/workers/worker_public/migrations/0010_model_nodes.sql +33 -0
  154. package/workers/worker_public/migrations/0011_schema_union.sql +38 -0
  155. package/workers/worker_public/migrations/0012_memories.sql +15 -0
  156. package/workers/worker_public/migrations/0013_projects.sql +21 -0
  157. package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.d.ts +16306 -0
  158. package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.ts +16261 -0
  159. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.d.ts +16373 -0
  160. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.ts +16328 -0
  161. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.d.ts +16382 -0
  162. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.ts +16337 -0
  163. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.d.ts +16383 -0
  164. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.ts +16338 -0
  165. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.d.ts +16403 -0
  166. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.ts +16358 -0
  167. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.d.ts +16408 -0
  168. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.ts +16363 -0
  169. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.d.ts +16414 -0
  170. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.ts +16369 -0
  171. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.d.ts +16414 -0
  172. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.ts +16369 -0
  173. package/workers/worker_public/node_modules/@cloudflare/workers-types/README.md +135 -0
  174. package/workers/worker_public/node_modules/@cloudflare/workers-types/entrypoints.svg +53 -0
  175. package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.d.ts +17095 -0
  176. package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.ts +17050 -0
  177. package/workers/worker_public/node_modules/@cloudflare/workers-types/index.d.ts +16306 -0
  178. package/workers/worker_public/node_modules/@cloudflare/workers-types/index.ts +16261 -0
  179. package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.d.ts +16447 -0
  180. package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.ts +16402 -0
  181. package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.d.ts +16306 -0
  182. package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.ts +16261 -0
  183. package/workers/worker_public/node_modules/@cloudflare/workers-types/package.json +11 -0
  184. package/workers/worker_public/package.json +13 -0
  185. package/workers/worker_public/prompts/conversational.md +8 -0
  186. package/workers/worker_public/prompts/enrichment.md +3 -0
  187. package/workers/worker_public/prompts/faithfulness.md +1 -0
  188. package/workers/worker_public/prompts/grader.md +5 -0
  189. package/workers/worker_public/prompts/listwise.md +3 -0
  190. package/workers/worker_public/prompts/precision.md +1 -0
  191. package/workers/worker_public/prompts/reflect.md +1 -0
  192. package/workers/worker_public/prompts/relevancy.md +1 -0
  193. package/workers/worker_public/prompts/research.md +10 -0
  194. package/workers/worker_public/prompts/section-summary.md +5 -0
  195. package/workers/worker_public/prompts/summarize.md +1 -0
  196. package/workers/worker_public/prompts/system.md +18 -0
  197. package/workers/worker_public/prompts/understanding.md +17 -0
  198. package/workers/worker_public/public/app.js +166 -0
  199. package/workers/worker_public/public/index.html +48 -0
  200. package/workers/worker_public/public/style.css +147 -0
  201. package/workers/worker_public/schema.sql +248 -0
  202. package/workers/worker_public/src/admin.ts +358 -0
  203. package/workers/worker_public/src/ai.ts +71 -0
  204. package/workers/worker_public/src/anchors.ts +41 -0
  205. package/workers/worker_public/src/answercache.ts +72 -0
  206. package/workers/worker_public/src/ask.ts +1094 -0
  207. package/workers/worker_public/src/auth.ts +252 -0
  208. package/workers/worker_public/src/bubble.ts +111 -0
  209. package/workers/worker_public/src/completion.ts +75 -0
  210. package/workers/worker_public/src/config.ts +238 -0
  211. package/workers/worker_public/src/context.ts +238 -0
  212. package/workers/worker_public/src/conversations.ts +162 -0
  213. package/workers/worker_public/src/drafts.ts +497 -0
  214. package/workers/worker_public/src/env.ts +90 -0
  215. package/workers/worker_public/src/faithfulness.ts +63 -0
  216. package/workers/worker_public/src/grader.ts +89 -0
  217. package/workers/worker_public/src/graph.ts +63 -0
  218. package/workers/worker_public/src/hybrid.ts +77 -0
  219. package/workers/worker_public/src/index.ts +441 -0
  220. package/workers/worker_public/src/internal_gateway.ts +41 -0
  221. package/workers/worker_public/src/lexical.ts +86 -0
  222. package/workers/worker_public/src/lib/hit.ts +4 -0
  223. package/workers/worker_public/src/lib/http.ts +83 -0
  224. package/workers/worker_public/src/lib/router.ts +4 -0
  225. package/workers/worker_public/src/livedata.ts +334 -0
  226. package/workers/worker_public/src/memories.ts +81 -0
  227. package/workers/worker_public/src/modelplane.ts +213 -0
  228. package/workers/worker_public/src/oidc.ts +333 -0
  229. package/workers/worker_public/src/pipeline.ts +377 -0
  230. package/workers/worker_public/src/ports/blobs.ts +7 -0
  231. package/workers/worker_public/src/ports/cloudflare/adapters.ts +177 -0
  232. package/workers/worker_public/src/ports/kv.ts +8 -0
  233. package/workers/worker_public/src/ports/model.ts +28 -0
  234. package/workers/worker_public/src/ports/runtime.ts +13 -0
  235. package/workers/worker_public/src/ports/store.ts +20 -0
  236. package/workers/worker_public/src/ports/vector.ts +26 -0
  237. package/workers/worker_public/src/profile.gen.ts +101 -0
  238. package/workers/worker_public/src/profile.ts +16 -0
  239. package/workers/worker_public/src/projects.ts +108 -0
  240. package/workers/worker_public/src/prompts.d.ts +6 -0
  241. package/workers/worker_public/src/quota.ts +54 -0
  242. package/workers/worker_public/src/reflect.ts +67 -0
  243. package/workers/worker_public/src/refs.ts +107 -0
  244. package/workers/worker_public/src/refusal.ts +65 -0
  245. package/workers/worker_public/src/requestScope.ts +71 -0
  246. package/workers/worker_public/src/research.ts +126 -0
  247. package/workers/worker_public/src/search.ts +56 -0
  248. package/workers/worker_public/src/selfquery.ts +25 -0
  249. package/workers/worker_public/src/session.ts +4 -0
  250. package/workers/worker_public/src/share.ts +53 -0
  251. package/workers/worker_public/src/stages/conceptGraph.ts +68 -0
  252. package/workers/worker_public/src/stages/conceptSteer.ts +39 -0
  253. package/workers/worker_public/src/stages/corpusScope.ts +25 -0
  254. package/workers/worker_public/src/stages/dedup.ts +10 -0
  255. package/workers/worker_public/src/stages/dense.ts +73 -0
  256. package/workers/worker_public/src/stages/diversity.ts +33 -0
  257. package/workers/worker_public/src/stages/editionCover.ts +63 -0
  258. package/workers/worker_public/src/stages/editionSteer.ts +88 -0
  259. package/workers/worker_public/src/stages/familyBoost.ts +22 -0
  260. package/workers/worker_public/src/stages/federate.ts +22 -0
  261. package/workers/worker_public/src/stages/glossary.ts +65 -0
  262. package/workers/worker_public/src/stages/graphLane.ts +31 -0
  263. package/workers/worker_public/src/stages/hyde.ts +29 -0
  264. package/workers/worker_public/src/stages/index.ts +69 -0
  265. package/workers/worker_public/src/stages/lexicalUnion.ts +21 -0
  266. package/workers/worker_public/src/stages/multiQuery.ts +57 -0
  267. package/workers/worker_public/src/stages/overviewDemote.ts +14 -0
  268. package/workers/worker_public/src/stages/poolOpen.ts +10 -0
  269. package/workers/worker_public/src/stages/propagate.ts +15 -0
  270. package/workers/worker_public/src/stages/rerank.ts +47 -0
  271. package/workers/worker_public/src/stages/seal.ts +16 -0
  272. package/workers/worker_public/src/stages/sectionDescent.ts +61 -0
  273. package/workers/worker_public/src/stages/stdRefNudge.ts +35 -0
  274. package/workers/worker_public/src/stages/subQuery.ts +42 -0
  275. package/workers/worker_public/src/stages/termNudge.ts +24 -0
  276. package/workers/worker_public/src/stages/typedPin.ts +131 -0
  277. package/workers/worker_public/src/stages/types.ts +112 -0
  278. package/workers/worker_public/src/stages/windowFloor.ts +23 -0
  279. package/workers/worker_public/src/structural.ts +171 -0
  280. package/workers/worker_public/src/tablecontext.ts +41 -0
  281. package/workers/worker_public/src/understand.ts +72 -0
  282. package/workers/worker_public/src/understandContract.ts +67 -0
  283. package/workers/worker_public/src/verdict.ts +255 -0
  284. package/workers/worker_public/tsconfig.json +18 -0
  285. package/workers/worker_public/wrangler.toml +104 -0
@@ -0,0 +1,270 @@
1
+ # Metanorma for AI — a feature proposal from the RAG trenches
2
+
3
+ **From:** the OIML RAG service (ai.oimlsmart.org — 875 Metanorma documents,
4
+ ~42k indexed clause chunks, live citation-grounded Q&A).
5
+ **Claim:** every item below maps to a concrete cost we paid or a capability
6
+ we could not ship. This is not speculative — it is the wish list of a
7
+ production RAG consumer of Metanorma output.
8
+
9
+ ## Evidence base (2025–26 primary literature)
10
+
11
+ Each proposal below is now grounded in peer-reviewed/preprint research on
12
+ RAG over standards and structured documents, not only in our production
13
+ experience:
14
+
15
+ - **[ETSI]** Al Masoud, Arazzi, Germani, Nocera. *Exploring Structural
16
+ Complexity in Normative RAG with Graph-based approaches: A case study on
17
+ the ETSI Standards.* arXiv:2604.09868 (2026). The only empirical RAG
18
+ study on industrial standards (ETSI EN 301 489-X; 800+ Q&A). Findings
19
+ used below: hierarchical section structure ↑precision/↑MRR;
20
+ structure-preserving chunking is the "overall best compromise";
21
+ tables are exempted from chunking (atomic); parthood (P) and citation
22
+ (C) edges are the core information model; neighbor-expansion re-ranking
23
+ *failed*; embedding *smoothing* is the lightweight recall lever. Their
24
+ §II explicitly points at machine-readable standards (IEC Smart
25
+ Standards) as the intended end state — this document is the Metanorma
26
+ version of that end state.
27
+ - **[STC]** Guttal et al. *Structure-Aware Chunking for Tabular Data in
28
+ RAG.* arXiv:2605.00318 (2026). Row-level key-value units, structural
29
+ boundaries: hybrid MRR 0.36→0.59, BM25-only Recall@1 0.37→0.75,
30
+ chunk count −40–56%. Quantifies P1-tables.
31
+ - **[RDR2]** Xu et al. *Equipping Retrieval-Augmented LLMs with Document
32
+ Structure Awareness.* arXiv:2510.04293 (2025). Document structure trees
33
+ as first-class retrieval input; flattened chunks are the named failure
34
+ mode. Grounds P0-manifest.
35
+ - **[SF-RAG]** Yu et al. *SF-RAG: Structure-Fidelity RAG for Academic
36
+ QA.* arXiv:2602.13647 (2026). Native hierarchy as a low-entropy
37
+ retrieval prior; flattening "destroys the native hierarchical
38
+ structure". Grounds P0-manifest + P2 stable ids.
39
+ - **[SPIRE]** *Structure-Preserving Interpretable Retrieval of Evidence.*
40
+ arXiv:2604.20849 (2026). Linearization "obscures section structure,
41
+ lists, and tables"; wants citation-ready subdocuments. Grounds P0/P1.
42
+ - **[MAHA]** Rashmi & Upadhya. *Modality-Aware Hybrid retrieval
43
+ Architecture.* arXiv:2510.14592 (2025). Tables→HTML-structured,
44
+ equations→LaTeX + text description, modality-aware knowledge graph.
45
+ Grounds P1-tables/equations.
46
+ - **[ANTHROPIC]** Anthropic. *Contextual Retrieval* (2024): contextual
47
+ preambles cut top-20 retrieval failures 35%→67% (with BM25+rerank).
48
+ Grounds the manifest's role: preamble generation needs typed units,
49
+ not scraped prose.
50
+ - **[LATE]** Günther et al. *Late Chunking.* arXiv:2409.04701 (2024);
51
+ **[CMP]** Merola & Singh. *Reconstructing Context.*
52
+ arXiv:2504.19754 (2025): contextual retrieval preserves coherence
53
+ better; late chunking is cheaper. Both need long-context/typed units —
54
+ neither works on GUID-anchored scraped HTML.
55
+
56
+ ## The core problem
57
+
58
+ RAG consumers today scrape the **compiled HTML** — the least machine-shaped
59
+ artifact Metanorma emits. We parse heading text to recover clause numbers
60
+ (because element ids are GUIDs in OCR-derived docs), regex out boilerplate,
61
+ flatten tables into pipe-joined rows, skip equations entirely, and join
62
+ document status from an external bibliography. Every one of those workarounds
63
+ is a proposal in disguise. This is not just our experience: it is the exact
64
+ failure mode the structure-first literature names — flattened chunks lose
65
+ the hierarchy, tables, and cross-references that make a standard a standard
66
+ ([ETSI] §I; [RDR2]; [SF-RAG]; [SPIRE]).
67
+
68
+ ## P0 — the unit manifest (sidecar alongside every render)
69
+
70
+ Emit `document.rag.json` — an array of *addressable semantic units* the
71
+ renderer already understands while building the HTML:
72
+
73
+ ```json
74
+ {
75
+ "id": "oiml:r60-1:2021#3.9",
76
+ "type": "term", // clause | term | table | figure | equation | example | note | bibliography | annex
77
+ "number": "3.9",
78
+ "title": "load cell",
79
+ "text": "measuring transducer which …",
80
+ "status": "in-force",
81
+ "parent": "oiml:r60-1:2021#3",
82
+ "lang": "en",
83
+ "spans": { "source": "sections/03-terms.adoc:104-112" }
84
+ }
85
+ ```
86
+
87
+ What this replaces, item for item:
88
+
89
+ | Today (our pipeline) | With the manifest |
90
+ |---|---|
91
+ | BeautifulSoup over `document.html` | read JSON |
92
+ | clause numbers re-derived from heading text | `number` + canonical `id` |
93
+ | GUID anchors dropped by regex | stable canonical ids |
94
+ | boilerplate filtered by regex (© OIML, rue Turgot…) | boilerplate simply absent |
95
+ | OCR-escaped HTML tags stripped from text | clean source text |
96
+ | `type` unknown (tables/definitions found by shape) | `type` drives retrieval specialization (exact term lookup; table-value lookups) |
97
+ | no audit trail | `spans` point back to source |
98
+
99
+ The HTML remains for humans; the manifest is the machine truth. Embedding
100
+ one inside the HTML (`<script type="application/json" id="rag-manifest">`)
101
+ also works and keeps a single artifact.
102
+
103
+ **Evidence:** [ETSI] builds exactly this model (InfoUnits with title+body
104
+ in a parthood/citation graph) by *recovering* it from PDFs — ToC parsing,
105
+ section-code prefix geometry, reference-resolution heuristics — and shows
106
+ structure preservation improves precision and MRR. The manifest emits what
107
+ they recover, for free, from the model the renderer already holds.
108
+ [RDR2]'s structure trees and [SF-RAG]'s structure-fidelity index consume
109
+ the same shape. [SPIRE] shows the payoff is citation-ready evidence
110
+ subdocuments — which is what our answers cite.
111
+
112
+ ## P0 — self-contained identity and status block
113
+
114
+ JSON front matter in every output:
115
+
116
+ ```json
117
+ {
118
+ "docidentifier": { "series": "OIML R", "number": "60", "part": "1", "year": "2021" },
119
+ "edition": "2021",
120
+ "status": "in-force", // in-force | superseded | withdrawn | joint
121
+ "successor": "oiml:r60-1:2021", // from relaton relation hasSuccessor
122
+ "relaton_ref": "r60-1_2021",
123
+ "published": "2021-04", "withdrawn": null,
124
+ "part_of": "oiml:r60", "translation_of": "oiml:r60-1:2021(E)"
125
+ }
126
+ ```
127
+
128
+ Today we join status and successor from `relaton-data-oiml` at ingest
129
+ (5,707 YAML records; ~37% of editions are superseded or withdrawn). The
130
+ join works but is external state that can diverge — the render should
131
+ carry the bibliographic truth it was built from. Identity parsing is our
132
+ single largest source of dirty-corpus defects (identifiers without the
133
+ series letter, edition fields holding part numbers, slug/identifier
134
+ mismatches); a parsed, structured docid eliminates the whole class.
135
+
136
+ ## P1 — tables as data, not prose
137
+
138
+ Normative values live in tables (MPE tables, accuracy-class limits). We
139
+ flatten them to `caption + " | "-joined rows` and hope the embedding
140
+ model copes. Instead, emit per table: caption, column headers **with
141
+ quantity and unit** (Metanorma already knows `[q]`uantities in many
142
+ flavors), rows as typed values, and the enclosing clause context. A JSON
143
+ or CSV serialization per table turns "what is the MPE for class III at
144
+ 500 g?" from a fuzzy vector match into an exact lookup the serving layer
145
+ can execute and cite (`OIML R 76:2006, Table 3, row 4`).
146
+
147
+ **Evidence:** [ETSI] §II-A exempts tabular sections from chunking
148
+ entirely — atomic units, never split. [STC] quantifies the payoff of
149
+ row-level key-value structure: MRR +66% hybrid, Recall@1 +106% BM25-only,
150
+ chunk count −40–56%. [MAHA] parses tables into HTML structure and
151
+ equations into LaTeX as first-class modalities. Our own G1 prototype
152
+ (extracting 10,388 tables from adoc `|===` sources) reproduces the shape
153
+ by hand — e.g. R 76-1 §3.5's MPE table with rowspan-merged class columns
154
+ and clean `ClassⅢ = 0≤m≤500` rows — precisely what the manifest should
155
+ emit natively.
156
+
157
+ ## P1 — the glossary layer as first-class output
158
+
159
+ Term entries (`term:: [preferred,definition]`) should also emit a
160
+ structured per-document glossary: term, grammar info, definition,
161
+ non-preferred terms, symbols, source citation. "What is a load cell?"
162
+ then becomes an exact-match against the glossary before any vector
163
+ search — the single highest-value retrieval specialization for standards
164
+ Q&A, and it aligns with the Glossarist model OIML already uses.
165
+
166
+ ## P1 — equations that survive retrieval
167
+
168
+ `stem:[...]` (AsciiMath) is invisible to retrieval today; normative
169
+ formulas (MPE = 0.5·e …) are effectively unanswerable. Emit MathML plus a
170
+ generated plain-language fallback ("maximum permissible error equals half
171
+ the verification interval value") in the unit manifest. The fallback is
172
+ what gets indexed; the MathML is what gets rendered.
173
+
174
+ **Evidence:** [MAHA] treats equations as a distinct modality (LaTeX +
175
+ textual description) and shows retrieval gains from modality-aware
176
+ indexing; our METANORMA-AI-SERIALIZATION conformance rule already requires
177
+ ≥2 of {asciimath, latex, described} for exactly this reason.
178
+
179
+ ## P2 — stable unit ids across revisions
180
+
181
+ Content-hash-derived ids (or editor-stable anchors) so an editorial
182
+ change re-indexes only changed units. Today any text change re-embeds the
183
+ whole document; across a 900-document corpus with weekly revisions this
184
+ is the difference between incremental and full re-index cost.
185
+
186
+ **Evidence:** [SF-RAG] shows hierarchy-stable indexing is what makes
187
+ structure-fidelity cheap to maintain; [ETSI]'s InfoUnit ids are only as
188
+ stable as their recovered section codes — model-native anchors (what
189
+ Metanorma has) are strictly better. Stability is also prerequisite for
190
+ [ANTHROPIC] contextual preambles to be cacheable across revisions (our
191
+ enrichment cache is keyed by content-hashed unit ids).
192
+
193
+ ## P2 — interlinear translation alignment
194
+
195
+ Same unit id, `lang` attribute — the renderer knows which French clause
196
+ corresponds to which English clause. That unlocks parallel corpora,
197
+ cross-language retrieval with source-language provenance, and
198
+ English-first indexes that can still cite the user's language edition.
199
+ (We currently ingest English-only because unaligned multilingual chunks
200
+ degrade retrieval; alignment would let us reverse that decision.)
201
+
202
+ ## P2 — requirement metadata
203
+
204
+ Requirements are the product of legal metrology documents. Tag units
205
+ with obligation (`shall`/`should`/`may`) and subject where the markup
206
+ knows them → "list every requirement for load cell marking" becomes a
207
+ filter, not a prayer.
208
+
209
+ ## P2 — a render profile for ingestion
210
+
211
+ `metanorma compile --profile rag`: boilerplate excluded (or typed as
212
+ such), notes/examples toggleable, source spans on, manifest emitted.
213
+ Consumers stop shipping regexes that guess at these boundaries.
214
+
215
+ ## Anti-goals (keep the contract clean)
216
+
217
+ - No embeddings, no model output, no opinions inside Metanorma
218
+ artifacts — deterministic renders only. AI-layer choices stay with
219
+ consumers.
220
+ - Metanorma should not own chunking **policy**. Emit semantic units;
221
+ consumers compose units into chunks for their model. What Metanorma
222
+ owns is **addressability** (stable ids, types, clean text, spans).
223
+ ([ETSI]'s own taxonomy supports this split: their "Structured +
224
+ Chunks" winner is a *consumer-side* composition over
225
+ producer-emitted structure — the producer's job is the structure.)
226
+
227
+ ## Sequencing
228
+
229
+ 1. Manifest + identity/status block (P0) — unblocks every RAG consumer,
230
+ mostly exporter work in existing render code paths.
231
+ 2. Tables + glossary (P1) — the two biggest answer-quality wins.
232
+ 3. Equations, stable ids, alignment, requirement flags, profiles (P2).
233
+
234
+ ## Beyond RAG — who else needs this
235
+
236
+ The unit manifest is not a chatbot feature; it is **addressability for
237
+ documents**, which unlocks a family of consumers:
238
+
239
+ - **Agentic conformity assessment.** Certification bodies and, soon,
240
+ AI agents that draft type-evaluation reports: map measured test results
241
+ to requirement unit-ids, auto-generate test plans from a Recommendation's
242
+ requirements, and produce audit trails (result → requirement → source
243
+ span). OIML R-series test-report annexes are already structured tables —
244
+ a machine layer turns report preparation from copywork into assembly.
245
+ - **Regulatory transposition and comparison.** National bodies transpose
246
+ OIML Recommendations into national law. Requirement-level units +
247
+ superseded→successor chains make "what changed between the 2000 and
248
+ 2017 edition" and "which national clauses diverge" diffable questions
249
+ instead of expert reading marathons.
250
+ - **Knowledge graphs / ontologies.** Terms × definitions × symbols ×
251
+ requirements × documents, emitted per render, compose directly into the
252
+ semantic layer legal metrology is already building — the glossary layer
253
+ is the same graph in miniature.
254
+ - **LLM training and evaluation corpora.** Standards are high-quality,
255
+ normatively precise domain text. Unit manifests make them citable
256
+ training/eval material (grounded QA benchmarks for legal metrology,
257
+ like legal benchmarks did for law) — with per-unit provenance.
258
+ - **Cross-publisher standards search.** Canonical unit ids (series,
259
+ number, part, year, clause) federate search across OIML/ISO/IEC
260
+ ecosystems instead of each publisher's PDF silo.
261
+ - **Translation QA.** Interlinear alignment turns "is the term set
262
+ consistent across the E/F/A editions" into a diff.
263
+ - **Accessibility and plain language.** The semantic layer (structured
264
+ definitions, requirement subjects) is the substrate for plain-language
265
+ summaries and assistive reading — the same data, different renderer.
266
+
267
+ The common thread: today all of these consumers re-derive semantics from
268
+ prose. The renderer already holds these semantics in memory while it
269
+ builds the HTML. Emitting them is cheaper than every consumer
270
+ re-guessing them.
@@ -0,0 +1,36 @@
1
+ # metanorma-mirror integration — phase 1: inventory and design
2
+
3
+ > The goal: answers that open the original document at the clause — a
4
+ > citation becomes a door, not a label. This phase inventories what
5
+ > exists and designs the integration; implementation is a follow-up.
6
+
7
+ ## 1. What exists today
8
+
9
+ | Piece | State |
10
+ |---|---|
11
+ | Clean-corpus sources | `document.xml` (Metanorma XML) per document — the rendering input mirror-js consumes |
12
+ | Dirty-corpus sources | Metanorma authoring trees (`metanorma/sections/*.adoc`) — renderable, lower fidelity |
13
+ | mirror-js renderer | lives in the smart repo (`browser/src/metanorma-mirror-js.d.ts` + runtime) — client-side XML→DOM rendering |
14
+ | R2 public bucket | `rag-public-assets` (unit assets only today) — no document renderings |
15
+ | Citation cards | docidentifier + edition + clause + snippet; no in-document link |
16
+ | Architecture note | CLAUDE.md already plans "rendered HTML in R2 for clause-anchored deep links (OIML docs only)" |
17
+
18
+ ## 2. The design (three layers, each valuable alone)
19
+
20
+ 1. **Rendered document per publication (build-time).** Render each clean-corpus `document.xml` to a single self-contained HTML (mirror-js headless, or `metanorma build` HTML) with stable clause anchors (`id="cls-4.2.1"`); upload to R2 under `docs/<doc_id>.html`; never the ISO/IEC internal corpus (copyright — public OIML only, same rule as every R2 public object).
21
+ 2. **Citation deep-links (one-line change once (1) exists).** The citation card's clause becomes `<a href="/docs/<doc_id>.html#<anchor>" target="_blank">` — the worker already emits doc_id + clause_anchor in every citation.
22
+ 3. **In-context pane (the payoff).** An embedded, scroll-to-clause view inside the answer: clicking a citation opens a slide-over rendering the document AT the anchor with the clause highlighted and its neighbors visible. Load strategy: fetch the R2 HTML and let the browser scroll (`#anchor` + a highlight script), no iframe sandboxing needed if the HTML is same-origin and static.
23
+
24
+ ## 3. Effort and risks
25
+
26
+ - (1) is a build-script + storage pass over ~36 clean documents; dirty-corpus rendering (880 docs) is a later wave and depends on metanorma build succeeding per tree.
27
+ - Clause-anchor stability: the anchors are the corpus's own (`cls-x.y.z`) — the same keys the retrieval index cites, so no mapping layer.
28
+ - Payload size: single-file HTML per doc with inlined CSS (~100–500 KB); R2 + CDN caching makes this free at read.
29
+ - The figure units interplay: the rendered document embeds its own figures — the pane and the typed-unit blocks will coexist (block = the answer's object; pane = the document's context).
30
+
31
+ ## 4. Follow-up board
32
+
33
+ - Render + upload the clean corpus (script + R2 layout).
34
+ - Citation deep-links in the site.
35
+ - The slide-over pane (site component).
36
+ - Dirty-corpus rendering pass.
@@ -0,0 +1,191 @@
1
+ # Multi-SDO architecture — one engine, many publishers
2
+
3
+ > Status: design (2026-09-12). The question this document answers: how does
4
+ > this RAG / Metanorma / Primmel pipeline serve other standards developers,
5
+ > with clean encapsulation, so that each SDO manages its own content and its
6
+ > own deployment? The short answer: the engine becomes publisher-agnostic by
7
+ > construction, every publisher-specific fact moves into a declarative
8
+ > **publisher profile** owned by that SDO, and the engine consumes profiles
9
+ > the same way it already consumes prompts and datasets — as data.
10
+
11
+ ## 1. The principle
12
+
13
+ The system already learned this principle the hard way, several times:
14
+
15
+ - Corpus behavior travels with the dataset declaration, not with pipeline
16
+ code (the `DATASETS` notes injected into prompts).
17
+ - Prompts are data files, never inline strings.
18
+ - The vector adapter is the single door where a **registry**, not an `if`
19
+ chain, decides which corpora may enter which index.
20
+ - The model plane consumes each publisher's own package repository at a
21
+ pinned reference, with a freshness gate that detects content drift.
22
+
23
+ The multi-SDO design extends the same principle to its conclusion: nothing
24
+ in the engine may know which publisher it serves. Every fact that is
25
+ currently true because "this is OIML" becomes a field in a profile.
26
+
27
+ ## 2. The honest audit — what is publisher-specific today
28
+
29
+ | Surface | Current state | Where it must live |
30
+ |---|---|---|
31
+ | Dataset declarations | `DATASETS` in `config.ts` (ids, labels, permission codes, prompt notes) | profile: `datasets` |
32
+ | Corpus registry | `PRODUCTION_CORPORA` / lane registries in `vector_adapter.py` | profile: `corpora` (registered names + target gates) |
33
+ | System prompt voice | `prompts/system.md` mentions OIML publications and the corpus shape | profile: prompt template variables |
34
+ | Identifier grammar | `OIML R/D/B/G/E` patterns in slugs, the PubID parser, mirror upload | profile: an identifier codec (per-SDO PubID module) |
35
+ | Terminology source | 13 Glossarist datasets under the OIML vocab repo | profile: vocabulary sources |
36
+ | Bibliography | `relaton-data-oiml` (5,707 records) | profile: relaton source (relaton is already one dataset per SDO across the ecosystem — direct leverage) |
37
+ | Document models | `primmel-packages` (oiml-r60, oiml-cs, …) | per-SDO packages repo at a pinned ref (the pattern already exists) |
38
+ | Rendered documents | mirror upload from the clean corpus; public-OIML-only rule | profile: rendering sources + the copyright policy flag per corpus |
39
+ | Access policy | two physically separate indexes; `ai-preview` estate permission | profile: audiences and their index bindings; permission codes map to the SDO's identity provider |
40
+ | Evaluation sets | golden 38 + annealment 18 (OIML content) | profile: per-SDO eval suites; the harness is shared |
41
+ | Branding and nav | OIML branding, the site shell | profile: theme, logo, labels, domain |
42
+ | Identity | RAG is an OIDC relying party of id.oimlsmart.org | profile: the OP per deployment; the role→permission mapping per estate |
43
+
44
+ The engine-side list — the stages, the retrieval mechanics, the answer
45
+ contract, the verdict engine, the gate harness, the quota and cache
46
+ machinery — is publisher-agnostic today and must stay that way. Where a
47
+ stage reads OIML-shaped assumptions (for example, edition steering's
48
+ family semantics), the assumption is metadata-driven already and travels
49
+ with the corpus.
50
+
51
+ ## 3. The publisher profile
52
+
53
+ One declarative unit, owned by the SDO, versioned with their content:
54
+
55
+ ```
56
+ publisher/
57
+ profile.yaml # id, names, branding, domains, identifier codec id
58
+ datasets.yaml # dataset declarations + permission codes + notes
59
+ corpora.yaml # corpus registry entries + audience bindings
60
+ prompts/ # template variables + overrides (rare)
61
+ evals/ # golden + annealment suites over THEIR corpus
62
+ sources.yaml # pinned refs: metanorma repos, relaton, glossarist,
63
+ # primmel packages repo, renderings bucket
64
+ identity.yaml # the estate OP, role → permission mapping
65
+ theme/ # logo, colors, nav labels
66
+ ```
67
+
68
+ The engine consumes exactly this at build and deploy time. A deployment
69
+ instance is then `engine@version + profile + content repos`, and the
70
+ profile repository is the only thing an SDO edits in the normal course
71
+ of business. Content repositories (Metanorma sources, relaton data,
72
+ Glossarist datasets, Primmel packages) stay under the SDO's own control
73
+ on their own infrastructure; the pipeline reads them at pinned
74
+ references and never writes to them — the same contract this system
75
+ already keeps with its upstreams.
76
+
77
+ ## 4. Deployment topology
78
+
79
+ Each SDO runs its own deployment, in its own Cloudflare account, under
80
+ its own domain:
81
+
82
+ - **Structural isolation per publisher.** One SDO's corpus, indexes,
83
+ caches and keys physically cannot reach another's, because they are
84
+ separate accounts with no shared bindings. This is the same
85
+ two-index isolation discipline applied one level up.
86
+ - **No cross-publisher federation by default.** If an SDO legitimately
87
+ holds another's content (as this deployment holds ISO/IEC material
88
+ internally), that relationship is declared in the SDO's own profile
89
+ with an explicit audience and copyright policy, and it inherits the
90
+ same structural separation between public and internal indexes.
91
+ - **One engine release train.** The engine is versioned as a package
92
+ (workers + stages + site shell + ingest CLI + gate harness).
93
+ Publisher deployments pin a version and upgrade deliberately, the way
94
+ they would adopt any dependency. Breaking changes carry migration
95
+ notes for profiles.
96
+
97
+ ## 5. What already generalizes (the leverage list)
98
+
99
+ The ecosystem was multi-SDO before this system was:
100
+
101
+ - **Metanorma** renders documents for many SDO flavors; the rendering
102
+ input and clause-anchor structure are common.
103
+ - **Relaton** bibliographies are one dataset per SDO; the registry
104
+ logic (families, editions, supersession-derived currency) is generic.
105
+ - **Glossarist** terminology datasets carry the same shape for any SDO;
106
+ the vocabulary lane is dataset-agnostic.
107
+ - **Primmel** packages are per-publisher repositories by design; the
108
+ model plane and verdict engine consume them through the projection,
109
+ not through OIML names.
110
+ - **The site shell** is already a package (`@oimlsmart/site-shell`)
111
+ consumed by this site — the same boundary serves other themes.
112
+ - **The evaluation harness** grades declarative expectations; only the
113
+ cases are content.
114
+
115
+ ## 6. Content governance per SDO
116
+
117
+ - **Single source of truth.** Each SDO's authoritative content lives in
118
+ their own repositories. The pipeline derives chunks, indexes, models
119
+ and renderings from them, and the derivation is reproducible from
120
+ pinned refs.
121
+ - **The freshness gate.** The model plane already fails its gate when a
122
+ package's source hash moves (CI re-indexes on content change). The
123
+ same mechanism watches every profile-declared source, so an SDO's
124
+ content update flows to their deployment through their own CI, with
125
+ their own gates.
126
+ - **Their own evaluation bar.** The promotion gate (golden ×N +
127
+ annealment ×M against their deployment) runs on the SDO's cases. An
128
+ SDO cannot ship a regression against their own corpus silently, and
129
+ the engine cannot ship one against any profile in a reference matrix
130
+ (engine CI runs the reference profiles' suites against fixture
131
+ corpora).
132
+
133
+ > Migration status: **step 1 shipped** (2026-09-12) — `profile/` exists,
134
+ > datasets and corpora registries are consumed from it by both the
135
+ > TypeScript and Python sides, with a drift test. See
136
+ > TODO.impl/70-profile-extraction.md for the board.
137
+
138
+ ## 7. Migration path (incremental, no rewrite)
139
+
140
+ 1. **Profile extraction, no behavior change.** *(done)* Move the OIML-specific
141
+ tables (`DATASETS`, corpus registries, prompt voice variables) into
142
+ `profile/` in this repository, consumed at build time. The OIML
143
+ deployment becomes the first profile; every test stays green because
144
+ nothing observable changes.
145
+ 2. **Identifier codec.** Generalize the slug/PubID surface into a
146
+ per-profile codec module (this repo already depends on a PubID
147
+ library for OIML). The mirror upload and the citation deep-link path
148
+ consume the codec.
149
+ 3. **Prompt templates.** Interpolate publisher variables in
150
+ `prompts/system.md`; overrides live in the profile and are rare.
151
+ 4. **Eval suites per profile.** Move the golden and annealment cases
152
+ under `profile/evals/`; the harness reads them from the declared
153
+ path. This repo's suites become the OIML profile's suites.
154
+ 5. **Engine packaging.** Extract the workers, stages, site shell, CLI
155
+ and harness into the engine package; keep this repository as
156
+ `publisher-oiml` — a profile plus deployment configuration — which
157
+ becomes the reference implementation and the template other SDOs
158
+ copy.
159
+ 6. **Reference matrix.** Engine CI grows a second fixture profile (a
160
+ small public corpus from another flavor) to prove no OIML assumption
161
+ leaks into the engine.
162
+
163
+ Each step ships independently and keeps the production gates green.
164
+
165
+ ## 8. Open questions (named, not hidden)
166
+
167
+ - **Versioning discipline.** How quickly do SDOs want to track engine
168
+ releases — pinned with scheduled upgrades, or floating with gates?
169
+ The gate machinery supports both; the governance choice is theirs.
170
+ - **Shared evaluation of the engine itself.** Content quality is
171
+ per-SDO, but retrieval mechanics regressions should be caught once,
172
+ centrally, on the reference matrix rather than re-measured by every
173
+ SDO.
174
+ - **Localization policy.** The index is English-only today by explicit
175
+ decision; other SDOs may decide differently, and the language gate
176
+ (`ingest/langid.py`) is already profile-shaped (a declared language
177
+ set plus exclusions).
178
+ - **Commercial and licensing terms.** Who may run the engine, under
179
+ what license, and what support looks like — an organizational
180
+ question that precedes any technical packaging.
181
+
182
+ ## 9. What this buys each SDO
183
+
184
+ A publisher with Metanorma-authored documents gets, for the cost of a
185
+ profile and their existing content: hybrid retrieval with clause-level
186
+ citations, contextual enrichment, typed tables and formulas, terminology
187
+ binding, edition-aware answers over their bibliography, conformance
188
+ checking by execution over their Primmel models, citation deep links
189
+ into their own renderings, and a promotion gate that measures the
190
+ result on their own questions — on infrastructure they control, with no
191
+ dependency on any other publisher's deployment.