@konneal/engine 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (285) hide show
  1. package/LICENSE +29 -0
  2. package/README.md +13 -0
  3. package/dist/admin.d.ts +26 -0
  4. package/dist/ai.d.ts +6 -0
  5. package/dist/anchors.d.ts +6 -0
  6. package/dist/answercache.d.ts +22 -0
  7. package/dist/ask.d.ts +5 -0
  8. package/dist/auth.d.ts +12 -0
  9. package/dist/bubble.d.ts +14 -0
  10. package/dist/chunk-LLWPT2XV.js +49 -0
  11. package/dist/chunk-MB74PTRM.js +114 -0
  12. package/dist/chunk-WOGQM7DJ.js +197 -0
  13. package/dist/chunk-WWNCWKKC.js +42 -0
  14. package/dist/completion.d.ts +5 -0
  15. package/dist/config.d.ts +154 -0
  16. package/dist/config.js +37 -0
  17. package/dist/context.d.ts +115 -0
  18. package/dist/conversations.d.ts +5 -0
  19. package/dist/drafts.d.ts +129 -0
  20. package/dist/env.d.ts +57 -0
  21. package/dist/faithfulness.d.ts +5 -0
  22. package/dist/grader.d.ts +3 -0
  23. package/dist/graph.d.ts +13 -0
  24. package/dist/hybrid.d.ts +7 -0
  25. package/dist/index.d.ts +9 -0
  26. package/dist/index.js +5373 -0
  27. package/dist/internal_gateway.d.ts +14 -0
  28. package/dist/lexical.d.ts +7 -0
  29. package/dist/livedata.d.ts +77 -0
  30. package/dist/memories.d.ts +10 -0
  31. package/dist/modelplane.d.ts +61 -0
  32. package/dist/oidc.d.ts +73 -0
  33. package/dist/pipeline.d.ts +57 -0
  34. package/dist/profile.d.ts +2 -0
  35. package/dist/profile.gen.d.ts +70 -0
  36. package/dist/profile.js +8 -0
  37. package/dist/projects.d.ts +8 -0
  38. package/dist/prompts/conversational.md +8 -0
  39. package/dist/prompts/enrichment.md +3 -0
  40. package/dist/prompts/faithfulness.md +1 -0
  41. package/dist/prompts/grader.md +5 -0
  42. package/dist/prompts/listwise.md +3 -0
  43. package/dist/prompts/precision.md +1 -0
  44. package/dist/prompts/reflect.md +1 -0
  45. package/dist/prompts/relevancy.md +1 -0
  46. package/dist/prompts/research.md +10 -0
  47. package/dist/prompts/section-summary.md +5 -0
  48. package/dist/prompts/summarize.md +1 -0
  49. package/dist/prompts/system.md +18 -0
  50. package/dist/prompts/understanding.md +17 -0
  51. package/dist/quota.d.ts +13 -0
  52. package/dist/reflect.d.ts +5 -0
  53. package/dist/refs.d.ts +40 -0
  54. package/dist/refusal.d.ts +9 -0
  55. package/dist/refusal.js +9 -0
  56. package/dist/requestScope.d.ts +26 -0
  57. package/dist/requestScope.js +10 -0
  58. package/dist/research.d.ts +8 -0
  59. package/dist/search.d.ts +4 -0
  60. package/dist/selfquery.d.ts +7 -0
  61. package/dist/session.d.ts +1 -0
  62. package/dist/share.d.ts +2 -0
  63. package/dist/structural.d.ts +27 -0
  64. package/dist/tablecontext.d.ts +11 -0
  65. package/dist/understand.d.ts +11 -0
  66. package/dist/understandContract.d.ts +29 -0
  67. package/dist/verdict.d.ts +24 -0
  68. package/docs/API.md +451 -0
  69. package/docs/ARCHITECTURE.md +302 -0
  70. package/docs/AUDIT-2026-08-24.md +71 -0
  71. package/docs/CONTRIBUTOR-AUDIT-2026-08-25.md +147 -0
  72. package/docs/INGEST-ARCHITECTURE.md +158 -0
  73. package/docs/MCP.md +92 -0
  74. package/docs/METANORMA-AI-SERIALIZATION.md +247 -0
  75. package/docs/MKO-EXPORT-PIPELINE.md +147 -0
  76. package/docs/REDESIGN-NORMATIVE-RAG-ETSI.md +485 -0
  77. package/docs/RESEARCH-SOTA-2026.md +243 -0
  78. package/docs/ROADMAP-SOTA.md +130 -0
  79. package/docs/SOTA-STAGE-SPECS.md +509 -0
  80. package/docs/annealment/F1-verdict.md +27 -0
  81. package/docs/annealment/F10-notes.md +19 -0
  82. package/docs/annealment/F11-composition.md +17 -0
  83. package/docs/annealment/F12-passport.md +17 -0
  84. package/docs/annealment/F2-counterfactual.md +20 -0
  85. package/docs/annealment/F3-absence.md +21 -0
  86. package/docs/annealment/F4-instance.md +18 -0
  87. package/docs/annealment/F5-workflow.md +21 -0
  88. package/docs/annealment/F6-impact.md +21 -0
  89. package/docs/annealment/F7-editions.md +18 -0
  90. package/docs/annealment/F8-selfverify.md +19 -0
  91. package/docs/annealment/F9-projection-qa.md +17 -0
  92. package/docs/annealment/L0-locate.md +19 -0
  93. package/docs/annealment/L1-extract.md +18 -0
  94. package/docs/annealment/L2-nomenclature.md +22 -0
  95. package/docs/annealment/L3-geometry.md +23 -0
  96. package/docs/annealment/L4-composition.md +21 -0
  97. package/docs/annealment/L5-cross-standard.md +20 -0
  98. package/docs/annealment/L6-diachrony.md +21 -0
  99. package/docs/annealment/L7-perception.md +20 -0
  100. package/docs/annealment/L8-computation.md +22 -0
  101. package/docs/annealment/L9-instance-process.md +23 -0
  102. package/docs/annealment/README.md +10 -0
  103. package/docs/guidelines-metanorma-ai-programme.md +279 -0
  104. package/docs/identity-onboarding-rag.md +65 -0
  105. package/docs/identity-service.md +219 -0
  106. package/docs/knowledge-annealment.md +273 -0
  107. package/docs/konneal-extraction-plan.md +481 -0
  108. package/docs/metanorma-for-ai.md +270 -0
  109. package/docs/mirror-plan.md +36 -0
  110. package/docs/multi-sdo-architecture.md +191 -0
  111. package/docs/paper-annealment-comparison.md +259 -0
  112. package/docs/paper-assets/architecture.svg +94 -0
  113. package/docs/paper-assets/contract-v2.svg +94 -0
  114. package/docs/paper-assets/mko-ingest.svg +91 -0
  115. package/docs/paper-oiml-bulletin.md +402 -0
  116. package/docs/paper-oiml-bulletin.mdx +419 -0
  117. package/docs/product-branding-options.md +172 -0
  118. package/docs/projects-design.md +88 -0
  119. package/docs/sota-mechanisms.md +184 -0
  120. package/docs/spec-api.md +77 -0
  121. package/docs/spec-pipeline.md +126 -0
  122. package/docs/vector-adapter.md +88 -0
  123. package/package.json +70 -0
  124. package/profile/corpora.yaml +5 -0
  125. package/profile/datasets.yaml +14 -0
  126. package/profile/prompts.yaml +5 -0
  127. package/profile/publisher.yaml +17 -0
  128. package/profile/retrieval.yaml +1 -0
  129. package/profile/sources.yaml +5 -0
  130. package/profile/ui.yaml +7 -0
  131. package/scripts/gen_profile.mjs +33 -0
  132. package/workers/shared/ai.ts +21 -0
  133. package/workers/shared/auth.ts +16 -0
  134. package/workers/shared/chunk.ts +108 -0
  135. package/workers/shared/oidc.ts +312 -0
  136. package/workers/shared/router.ts +45 -0
  137. package/workers/shared/session.ts +104 -0
  138. package/workers/worker_internal/src/index.ts +157 -0
  139. package/workers/worker_internal/tsconfig.json +15 -0
  140. package/workers/worker_internal/wrangler.toml +32 -0
  141. package/workers/worker_mcp/src/index.ts +175 -0
  142. package/workers/worker_mcp/tsconfig.json +13 -0
  143. package/workers/worker_mcp/wrangler.toml +18 -0
  144. package/workers/worker_public/migrations/0002_conversations.sql +22 -0
  145. package/workers/worker_public/migrations/0003_shared_conversations.sql +9 -0
  146. package/workers/worker_public/migrations/0004_graph.sql +16 -0
  147. package/workers/worker_public/migrations/0005_documents.sql +19 -0
  148. package/workers/worker_public/migrations/0006_conversation_entities.sql +11 -0
  149. package/workers/worker_public/migrations/0007_chunks_fts.sql +43 -0
  150. package/workers/worker_public/migrations/0008_unit_payloads.sql +15 -0
  151. package/workers/worker_public/migrations/0009_chunks_unit.sql +8 -0
  152. package/workers/worker_public/migrations/0009_message_context.sql +7 -0
  153. package/workers/worker_public/migrations/0010_model_nodes.sql +33 -0
  154. package/workers/worker_public/migrations/0011_schema_union.sql +38 -0
  155. package/workers/worker_public/migrations/0012_memories.sql +15 -0
  156. package/workers/worker_public/migrations/0013_projects.sql +21 -0
  157. package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.d.ts +16306 -0
  158. package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.ts +16261 -0
  159. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.d.ts +16373 -0
  160. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.ts +16328 -0
  161. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.d.ts +16382 -0
  162. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.ts +16337 -0
  163. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.d.ts +16383 -0
  164. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.ts +16338 -0
  165. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.d.ts +16403 -0
  166. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.ts +16358 -0
  167. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.d.ts +16408 -0
  168. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.ts +16363 -0
  169. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.d.ts +16414 -0
  170. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.ts +16369 -0
  171. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.d.ts +16414 -0
  172. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.ts +16369 -0
  173. package/workers/worker_public/node_modules/@cloudflare/workers-types/README.md +135 -0
  174. package/workers/worker_public/node_modules/@cloudflare/workers-types/entrypoints.svg +53 -0
  175. package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.d.ts +17095 -0
  176. package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.ts +17050 -0
  177. package/workers/worker_public/node_modules/@cloudflare/workers-types/index.d.ts +16306 -0
  178. package/workers/worker_public/node_modules/@cloudflare/workers-types/index.ts +16261 -0
  179. package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.d.ts +16447 -0
  180. package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.ts +16402 -0
  181. package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.d.ts +16306 -0
  182. package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.ts +16261 -0
  183. package/workers/worker_public/node_modules/@cloudflare/workers-types/package.json +11 -0
  184. package/workers/worker_public/package.json +13 -0
  185. package/workers/worker_public/prompts/conversational.md +8 -0
  186. package/workers/worker_public/prompts/enrichment.md +3 -0
  187. package/workers/worker_public/prompts/faithfulness.md +1 -0
  188. package/workers/worker_public/prompts/grader.md +5 -0
  189. package/workers/worker_public/prompts/listwise.md +3 -0
  190. package/workers/worker_public/prompts/precision.md +1 -0
  191. package/workers/worker_public/prompts/reflect.md +1 -0
  192. package/workers/worker_public/prompts/relevancy.md +1 -0
  193. package/workers/worker_public/prompts/research.md +10 -0
  194. package/workers/worker_public/prompts/section-summary.md +5 -0
  195. package/workers/worker_public/prompts/summarize.md +1 -0
  196. package/workers/worker_public/prompts/system.md +18 -0
  197. package/workers/worker_public/prompts/understanding.md +17 -0
  198. package/workers/worker_public/public/app.js +166 -0
  199. package/workers/worker_public/public/index.html +48 -0
  200. package/workers/worker_public/public/style.css +147 -0
  201. package/workers/worker_public/schema.sql +248 -0
  202. package/workers/worker_public/src/admin.ts +358 -0
  203. package/workers/worker_public/src/ai.ts +71 -0
  204. package/workers/worker_public/src/anchors.ts +41 -0
  205. package/workers/worker_public/src/answercache.ts +72 -0
  206. package/workers/worker_public/src/ask.ts +1094 -0
  207. package/workers/worker_public/src/auth.ts +252 -0
  208. package/workers/worker_public/src/bubble.ts +111 -0
  209. package/workers/worker_public/src/completion.ts +75 -0
  210. package/workers/worker_public/src/config.ts +238 -0
  211. package/workers/worker_public/src/context.ts +238 -0
  212. package/workers/worker_public/src/conversations.ts +162 -0
  213. package/workers/worker_public/src/drafts.ts +497 -0
  214. package/workers/worker_public/src/env.ts +90 -0
  215. package/workers/worker_public/src/faithfulness.ts +63 -0
  216. package/workers/worker_public/src/grader.ts +89 -0
  217. package/workers/worker_public/src/graph.ts +63 -0
  218. package/workers/worker_public/src/hybrid.ts +77 -0
  219. package/workers/worker_public/src/index.ts +441 -0
  220. package/workers/worker_public/src/internal_gateway.ts +41 -0
  221. package/workers/worker_public/src/lexical.ts +86 -0
  222. package/workers/worker_public/src/lib/hit.ts +4 -0
  223. package/workers/worker_public/src/lib/http.ts +83 -0
  224. package/workers/worker_public/src/lib/router.ts +4 -0
  225. package/workers/worker_public/src/livedata.ts +334 -0
  226. package/workers/worker_public/src/memories.ts +81 -0
  227. package/workers/worker_public/src/modelplane.ts +213 -0
  228. package/workers/worker_public/src/oidc.ts +333 -0
  229. package/workers/worker_public/src/pipeline.ts +377 -0
  230. package/workers/worker_public/src/ports/blobs.ts +7 -0
  231. package/workers/worker_public/src/ports/cloudflare/adapters.ts +177 -0
  232. package/workers/worker_public/src/ports/kv.ts +8 -0
  233. package/workers/worker_public/src/ports/model.ts +28 -0
  234. package/workers/worker_public/src/ports/runtime.ts +13 -0
  235. package/workers/worker_public/src/ports/store.ts +20 -0
  236. package/workers/worker_public/src/ports/vector.ts +26 -0
  237. package/workers/worker_public/src/profile.gen.ts +101 -0
  238. package/workers/worker_public/src/profile.ts +16 -0
  239. package/workers/worker_public/src/projects.ts +108 -0
  240. package/workers/worker_public/src/prompts.d.ts +6 -0
  241. package/workers/worker_public/src/quota.ts +54 -0
  242. package/workers/worker_public/src/reflect.ts +67 -0
  243. package/workers/worker_public/src/refs.ts +107 -0
  244. package/workers/worker_public/src/refusal.ts +65 -0
  245. package/workers/worker_public/src/requestScope.ts +71 -0
  246. package/workers/worker_public/src/research.ts +126 -0
  247. package/workers/worker_public/src/search.ts +56 -0
  248. package/workers/worker_public/src/selfquery.ts +25 -0
  249. package/workers/worker_public/src/session.ts +4 -0
  250. package/workers/worker_public/src/share.ts +53 -0
  251. package/workers/worker_public/src/stages/conceptGraph.ts +68 -0
  252. package/workers/worker_public/src/stages/conceptSteer.ts +39 -0
  253. package/workers/worker_public/src/stages/corpusScope.ts +25 -0
  254. package/workers/worker_public/src/stages/dedup.ts +10 -0
  255. package/workers/worker_public/src/stages/dense.ts +73 -0
  256. package/workers/worker_public/src/stages/diversity.ts +33 -0
  257. package/workers/worker_public/src/stages/editionCover.ts +63 -0
  258. package/workers/worker_public/src/stages/editionSteer.ts +88 -0
  259. package/workers/worker_public/src/stages/familyBoost.ts +22 -0
  260. package/workers/worker_public/src/stages/federate.ts +22 -0
  261. package/workers/worker_public/src/stages/glossary.ts +65 -0
  262. package/workers/worker_public/src/stages/graphLane.ts +31 -0
  263. package/workers/worker_public/src/stages/hyde.ts +29 -0
  264. package/workers/worker_public/src/stages/index.ts +69 -0
  265. package/workers/worker_public/src/stages/lexicalUnion.ts +21 -0
  266. package/workers/worker_public/src/stages/multiQuery.ts +57 -0
  267. package/workers/worker_public/src/stages/overviewDemote.ts +14 -0
  268. package/workers/worker_public/src/stages/poolOpen.ts +10 -0
  269. package/workers/worker_public/src/stages/propagate.ts +15 -0
  270. package/workers/worker_public/src/stages/rerank.ts +47 -0
  271. package/workers/worker_public/src/stages/seal.ts +16 -0
  272. package/workers/worker_public/src/stages/sectionDescent.ts +61 -0
  273. package/workers/worker_public/src/stages/stdRefNudge.ts +35 -0
  274. package/workers/worker_public/src/stages/subQuery.ts +42 -0
  275. package/workers/worker_public/src/stages/termNudge.ts +24 -0
  276. package/workers/worker_public/src/stages/typedPin.ts +131 -0
  277. package/workers/worker_public/src/stages/types.ts +112 -0
  278. package/workers/worker_public/src/stages/windowFloor.ts +23 -0
  279. package/workers/worker_public/src/structural.ts +171 -0
  280. package/workers/worker_public/src/tablecontext.ts +41 -0
  281. package/workers/worker_public/src/understand.ts +72 -0
  282. package/workers/worker_public/src/understandContract.ts +67 -0
  283. package/workers/worker_public/src/verdict.ts +255 -0
  284. package/workers/worker_public/tsconfig.json +18 -0
  285. package/workers/worker_public/wrangler.toml +104 -0
@@ -0,0 +1,21 @@
1
+ # F5 — Certification workflow statefulness
2
+
3
+ **Powering objects:** `evaluation/processes.yaml` (layer-composed
4
+ process tree; `validate_provision` binds the campaign to requirement
5
+ URNs like /req/metrological/mpe), `evaluation/gateways.yaml`,
6
+ `execution/` (test-report forms, checklist), entities/performance-test-
7
+ evaluations, evaluation/sample-selection-rules.
8
+
9
+ **The win:** the assistant knows WHERE an evaluation stands and WHAT
10
+ GATES WHAT — not as prose but as traversable state:
11
+ - *"what must be true before the creep test?"* → the sequence rule
12
+ (MDLO baseline first) + the gateway that checks it;
13
+ - *"which requirements does this test campaign validate?"* → the
14
+ validate_provision URNs, exhaustively;
15
+ - agentic follow-through: checklist items flip state as evidence
16
+ lands; the assistant can say "3 of 62 tests outstanding, blocked on
17
+ the humidity chamber" — a document lane can quote the procedure but
18
+ cannot know the POSITION.
19
+
20
+ **Why documents can't follow:** procedures in prose are instructions;
21
+ here they are a state machine with recorded progress.
@@ -0,0 +1,21 @@
1
+ # F6 — Impact analysis (committee tooling)
2
+
3
+ **Powering objects:** the binding graph — requirements ↔ conformance
4
+ tests ↔ formulas (`formulas-used.yaml` binds MDLO to conversion_factor_f,
5
+ E_L, E_R, C_M), tables (mpe_tiers feeds lookupMPE), aspects
6
+ (requirement-binding-targets), execution forms (E_R's report field).
7
+
8
+ **The win:** DRAFTING becomes queryable. *"If the committee raises the
9
+ class C limit_factor, what changes?"* → traverse the graph: the MPE
10
+ verdicts for class C, the tests whose verdicts derive from lookupMPE,
11
+ the R 60-3 report fields that record them, the instances already
12
+ evaluated. The answer is an IMPACT SET — machine-derived, exhaustive
13
+ (F3's completeness), with every hop's clause anchor.
14
+
15
+ **Why it matters:** this is the first capability whose customer is the
16
+ STANDARDS DEVELOPER, not the reader — the model pays for its own
17
+ maintenance by making revision risk computable.
18
+
19
+ **Why documents can't follow:** the binding "this test's verdict
20
+ consumes this table's factor" exists nowhere in any single document —
21
+ it is model knowledge spanning three parts.
@@ -0,0 +1,18 @@
1
+ # F7 — Semantic edition diffs + temporal jurisdiction
2
+
3
+ **Powering objects:** editions/2017/*.prl — a FULL prior-edition model
4
+ package (not a separate PDF) — plus edition lifecycle (status,
5
+ supersedes, validity.from).
6
+
7
+ **The win, beyond L6:** diffs become SEMANTIC, not textual:
8
+ *"what changed in creep between 2017 and 2021?"* → diff the two MODELS:
9
+ the constraint set (new OCL checks), attribute definitions (added
10
+ fields, moved scopes), test sequences (reordered steps), tier tables
11
+ (changed breakpoints) — each delta with both editions' clause anchors.
12
+ Text diffs of two PDFs produce noise; model diffs produce CHANGE RECORDS.
13
+
14
+ **Temporal jurisdiction:** validity windows answer retro-questions —
15
+ *"an evaluation performed 2019-11: which edition governed it, and is
16
+ its verdict still valid today?"* → 2017 governs; superseded-by chain
17
+ says whether re-evaluation is required. Legal-metrology gold: the
18
+ corpus knows TIME, not just content.
@@ -0,0 +1,19 @@
1
+ # F8 — Self-verification as a query (meta-grounding)
2
+
3
+ **Powering objects:** the model-linker rule set (quantity-coherence,
4
+ requirement-binding-targets) and the allowlist ledger whose history
5
+ reads "burned to empty — every formerly allowlisted site was fixed at
6
+ the data layer": the model is MACHINE-CHECKED consistent.
7
+
8
+ **The win:** for model-grounded claims, execution replaces the judge.
9
+ When the D lane answers "the class C limit factor is 0.35", it can
10
+ VERIFY that claim against the model (the value is read from the typed
11
+ table, not generated) — and the user can ask *"is this corpus
12
+ internally consistent?"* and receive the linker verdict. The
13
+ faithfulness judge (an LLM scoring prose) remains for synthesis; for
14
+ model-grounded facts, the check is arithmetic.
15
+
16
+ **Why it matters:** it closes the trust loop — the same mechanism that
17
+ answers also proves. The strongest possible grounding story for the
18
+ paper: not "the model was cited" but "the model was EXECUTED, and the
19
+ execution is checkable".
@@ -0,0 +1,17 @@
1
+ # F9 — Document-as-projection verification
2
+
3
+ **Powering objects:** documents/{1,2,3,a}/document.presentation.xml —
4
+ the model CARRIES its own rendered document projections per part.
5
+
6
+ **The win:** publishing QA becomes a query. *"Does the published
7
+ R 60-1 Table 4 match the model's mpe_tiers?"* → parse the projection's
8
+ table, compare cell-by-cell against the model's typed table → a
9
+ diff-and-verdict. The direction inverts: instead of the document being
10
+ the source the model was derived from, the MODEL becomes the source and
11
+ the document a VIEW to be verified. Forward: model → document
12
+ GENERATION (the view rendered FROM truth, drift structurally
13
+ impossible).
14
+
15
+ **Why it matters:** every standards body's dirty secret is render drift
16
+ (errata). Projection QA makes errata a computable diff. And it is the
17
+ on-ramp to the fully closed loop: author → model → verified document.
@@ -0,0 +1,19 @@
1
+ # L0 — LOCATE (primitives P1–P2)
2
+
3
+ **Definition:** find WHERE the corpus addresses a topic; return the clause
4
+ anchor. No value, no synthesis — presence and position only.
5
+
6
+ **Why it's the floor:** it isolates retrieval from understanding. A lane
7
+ failing L0 is mis-built (chunking/embedding bug), not "less annealed" —
8
+ the floor validates the harness before the ladder means anything.
9
+
10
+ **Example** — *"Where does R 60 address creep?"*
11
+ Witness: citation anchor ∈ {R 60-1 §3.x (definition), R 60-2 §2.11.x
12
+ (test)} — either is a pass; the anchor must EXIST and be normative.
13
+ Lane expectations: A/B/C/D/E all pass (sanity). Primmel additionally
14
+ binds the term to its `behaviors: [creep]` entry with stimulus/response
15
+ and the `load-cell` entity — same anchor, richer provenance.
16
+
17
+ **Harness notes:** 6 questions spanning definition-anchors, test-anchors
18
+ and annex-anchors; paraphrases use colloquial synonyms ("time drift under
19
+ load") so lexical lanes can't win by string match alone.
@@ -0,0 +1,18 @@
1
+ # L1 — EXTRACT (P1–P2)
2
+
3
+ **Definition:** return a normative value VERBATIM from prose — no
4
+ disambiguation beyond finding the sentence.
5
+
6
+ **Example** — *"Quote the note that governs which p_LC value the creep
7
+ MPE must use."*
8
+ Witness (real, from notes.yaml): "p_LC = 0.7" AND the qualifier
9
+ "regardless of any value declared by the manufacturer"
10
+ (note_creep_plc_always_0_7 — see also F10: this note is an OVERRIDE in
11
+ the model, prose everywhere else).
12
+
13
+ **Example** — *"What validity date does the current R 60 edition carry?"*
14
+ Witness: "2021-01-01" (standard.yaml edition.validity.from — P2
15
+ provenance makes this trivial in D, findable in C via the registry).
16
+
17
+ Lane expectations: all pass; separation is noise-level. L1 exists to
18
+ calibrate the harness's regex discipline before structure matters.
@@ -0,0 +1,22 @@
1
+ # L2 — NOMENCLATURE (P4)
2
+
3
+ **Definition:** resolve the user's vocabulary to the corpus's DEFINED
4
+ terms, return the authoritative definition and its vocabulary anchor.
5
+
6
+ **Sub-probes:**
7
+ - **a) corpus-local** — *"my load cell keeps drifting over months — what
8
+ does R 60 call that?"* → durability; witness: the definition text
9
+ ("ability of a measuring instrument to maintain its performance
10
+ characteristics over a period of use") + anchor R 60-1 §3.1 with VIML
11
+ 5.15 register (terminology.yaml: vocab_ref viml-2022 5.15).
12
+ - **b) cross-register** — *"which VIM/VIML concept does 'durability'
13
+ come from?"* → viml-2022 clause 5.15. Only lanes carrying vocab_ref
14
+ answer the REGISTER part (D natively; C via the glossary lane's
15
+ concept sources).
16
+ - **c) multilingual** — term spellings per language (Primmel
17
+ eng-Latn entries; multilingual spellings land with more languages).
18
+
19
+ Lane expectations: A/B luck-dependent (embedding proximity); C strong at
20
+ (a); D strong a–c. The colloquial gap ("drift"≠"durability") is the
21
+ measurement: without P4 the lanes must bridge vocabulary by embeddings
22
+ alone.
@@ -0,0 +1,23 @@
1
+ # L3 — TABULAR GEOMETRY (P3 + P8)
2
+
3
+ **Definition:** a value whose retrieval requires row×column
4
+ disambiguation; delivery requires the TYPED BLOCK (the producer's
5
+ object, never re-typed prose).
6
+
7
+ **Sub-probes:**
8
+ - **a) cell lookup** — *"minimum n_LC for class B"* → R 60-1 Table 3:
9
+ row B, lower-limit column. Witness: 5000 (family values: A 50000, B
10
+ 5000, C 500, D 100 — verified in the corpus) + block artifact
11
+ `block:"table"`.
12
+ - **b) unit-aware cell** — *"what quantity does `load_min` carry in the
13
+ MPE tier table, and in what unit?"* → tables.yaml column
14
+ `{name: load_min, type: number, unit: v}` → "load in verification
15
+ intervals (v)". Only D answers the UNIT question structurally (P8);
16
+ C's payload has unit strings (metanorma-document#55 GAP-1).
17
+ - **c) derived cell** — *"the MPE limit factor for class C at a load
18
+ between the tier bounds"* → needs the tier BREAKPOINTS + factor: the
19
+ lookup operation (formula lookupMPE) — transitional to L8.
20
+
21
+ Lane expectations: A ≤40% (flattened geometry), B partial (pipes as
22
+ noise), C ≥90% at (a), D ≥90% a–c. **First hard separator; the
23
+ side-by-side demo on the site.**
@@ -0,0 +1,21 @@
1
+ # L4 — INTRA-DOCUMENT COMPOSITION (P2 + P5)
2
+
3
+ **Definition:** an answer that JOINS ≥2 clauses (often + a table) and
4
+ returns both anchors plus the derived relation.
5
+
6
+ **Sub-probes:**
7
+ - **a) intra-document** — *"initial-verification MPE vs in-service
8
+ limits for a class C load cell, and how they scale across the
9
+ range"* → R 60-1 MPE table (tier breakpoints, limit_factor) + the
10
+ applicability clauses. Witness: both anchors + the scaling statement
11
+ expressed in v-units.
12
+ - **b) cross-part** — *"which R 60-2 test validates the R 60-1
13
+ repeatability requirement, and where is its result recorded in
14
+ R 60-3?"* → requirement `/req/metrological/repeatability` (R 60-1) ↔
15
+ test `/conf/metrological-tests/...` (R 60-2) ↔ test-report form
16
+ (R 60-3 §2.1.3 E_R). Witness: the three part-anchors. In D the chain
17
+ is EDGES (formulas_used binds the test to
18
+ `repeatabilityError`); in C it must be co-retrieved.
19
+
20
+ Lane expectations: C/D strong at (a); (b) separates C (partial — needs
21
+ all three parts in the window) from D (edges). A/B partial.
@@ -0,0 +1,20 @@
1
+ # L5 — CROSS-STANDARD LINKING (P5)
2
+
3
+ **Definition:** answers that traverse a reference to ANOTHER standard,
4
+ returning the linked identity and clause.
5
+
6
+ **Sub-probes:**
7
+ - **a) cites-edge** — *"which ISO/IEC test method does R 60-2 invoke for
8
+ humidity?"* → witness: the ISO/IEC document id + clause from the
9
+ cites edges / references registry. ISOLATION: public lanes assert the
10
+ LINK only — ISO/IEC text never enters a public index; content probes
11
+ are member/internal-only, structurally enforced.
12
+ - **b) composition** — *"which CASCO vocabulary governs R 60's
13
+ certification activities, and which aspects of the load cell do its
14
+ requirements bind?"* → `uses: iso-iec-17000` + `iso-iec-17065` + the
15
+ aspects registry (markings, accompanying_document… via
16
+ requirement-binding-targets). Witness: both package ids + ≥1 aspect.
17
+
18
+ Lane expectations: D strong (composition is a first-class relation);
19
+ C partial at (a) (cites edges, no composition); A/B near-zero —
20
+ cross-standard vocabulary rarely co-embeds.
@@ -0,0 +1,21 @@
1
+ # L6 — DIACHRONY / EDITIONS (P6)
2
+
3
+ **Definition:** answers that depend on WHICH edition, or on WHAT
4
+ CHANGED between editions.
5
+
6
+ **Sub-probes:**
7
+ - **a) current-edition selection** — *"which R 60 edition applies to a
8
+ type evaluation started this year?"* → 2021 (lifecycle: status
9
+ current, supersedes 2017). Witness: edition + status.
10
+ - **b) delta extraction** — *"what changed in the creep requirements
11
+ between 2017 and 2021?"* → D computes the diff between EDITION
12
+ PACKAGES (editions/2017/*.prl vs current) at the model level — the
13
+ witness is the delta fact (constraint added/changed, limits moved).
14
+ C can align units by hash but has no semantic diff (metanorma-document
15
+ #53 item 5). A/B near-zero.
16
+ - **c) temporal jurisdiction** — *"which edition governed an evaluation
17
+ performed in 2019, and under what validity rule?"* → 2017 (validity
18
+ windows: 2017 until 2021-01-01). Only D (validity.from on editions).
19
+
20
+ Lane expectations: D strong a–c; C partial (a, steering already ships);
21
+ A/B fail b/c. Deepened by frontier F7.
@@ -0,0 +1,20 @@
1
+ # L7 — MULTIMODAL PERCEPTION (P7)
2
+
3
+ **Definition:** answers whose evidence is in PIXELS, not text.
4
+
5
+ **Sub-probes:**
6
+ - **a) caption/asset** — *"what does the R 60-2 test setup figure
7
+ show?"* → figure unit + stored caption/vision description.
8
+ - **b) pixel-only content** — *"what are the three labeled cases in the
9
+ design-shapes example?"* → labels A/B/C that exist ONLY in the
10
+ drawing; witness: the label values (verified: the deployed system
11
+ reads them from pixels — attachFigureImages).
12
+ - **c) user-image grounding** — a nameplate photo → "which accuracy
13
+ class marking does this carry / which OIML R applies?" (image
14
+ questions, shipped). Witness: classification.
15
+
16
+ Lane expectations: C strong a–c (figure lane live + multimodal
17
+ generation); D partial (references figures in sequences —
18
+ `fig-2/fig-3` — asset binding pending); A/B zero (no units, no assets).
19
+ The most DEMO-VISIBLE separator: same question, lane C shows the image
20
+ and reads it, lane A shows prose soup.
@@ -0,0 +1,22 @@
1
+ # L8 — COMPUTATION (P9 + P8)
2
+
3
+ **Definition:** answers the documents DEFINE but never PRINT — values
4
+ that must be computed. Witnesses are COMPUTED OFFLINE from the model
5
+ (deterministic; no judge, no regex luck).
6
+
7
+ **Sub-probes:**
8
+ - **a) pure calculation** — *"the conversion factor f for these
9
+ indications"* → calculations.yaml conversionFactor (inputs:
10
+ avgIndicationAt75pct, indicationAtDmin; R 60-3 2.1.2.4). Witness: the
11
+ computed number.
12
+ - **b) constraint verdict** — *"is D_max = 0.8·E_max acceptable?"* → OCL
13
+ `dead_load_max_geometry` FAILS: d_max must lie in [0.9·E_max, E_max];
14
+ witness: verdict + the recorded violation_meaning ("the type
15
+ evaluation of this load cell is void") — F1's ladder form.
16
+ - **c) unit coherence** — *"are 3000 kgf and 30 kN the same force?"* →
17
+ the unit REGISTER converts (quantity_kind force; conversion to SI);
18
+ witness: the equality verdict. P8 pure.
19
+
20
+ Lane expectations: **D only, by construction** — A/B/C/E cannot compute
21
+ (they can at best quote a formula; E proves it's not prose quality).
22
+ The ladder's most objective rung: ground truth is arithmetic.
@@ -0,0 +1,23 @@
1
+ # L9 — INSTANCE & PROCESS (P10)
2
+
3
+ **Definition:** answers about PARTICULAR things and ORDERED activities —
4
+ knowledge that exists only where the model records instances and
5
+ sequences.
6
+
7
+ **Sub-probes:**
8
+ - **a) process order & consequence** — *"what is contaminated if creep
9
+ runs before MDLO on the same sample?"* → test_sequences.yaml
10
+ `mdlo-creep-dr`: MDLO is the BASELINE (order 1), creep FOLLOWS (order
11
+ 2), DR reads against creep's D_max; witness: the ordering verdict +
12
+ "contaminates the baseline". The rationale is ENCODED as sequence
13
+ semantics — prose lanes only pass if the corpus happens to print it
14
+ (containment-checked; the printed form exists in R 60-2 — so the
15
+ probe asks the COUNTERFACTUAL form: "may I run them in my own
16
+ order?" → D answers from the model rule).
17
+ - **b) instance-grounded fact** — a sample instrument profile →
18
+ "what is ITS E_R budget?" — the model evaluates at the instance.
19
+ - **c) execution form** — *"where is E_R recorded in the test report?"*
20
+ → execution/test-report.yaml + subforms (R 60-3 2.1.3).
21
+
22
+ Lane expectations: **D only.** The deepest binding: the corpus doesn't
23
+ describe the general — it RECORDED the particular and the ORDER.
@@ -0,0 +1,10 @@
1
+ # The annealment documentation set
2
+
3
+ One file per ladder rung (L0–L9) and per frontier win (F1–F12). The
4
+ instrument spec with the primitive taxonomy lives in
5
+ ../knowledge-annealment.md; each file here ELABORATES its level with
6
+ worked examples, witnesses and lane expectations.
7
+
8
+ Ladder: [L0](L0-locate.md) · [L1](L1-extract.md) · [L2](L2-nomenclature.md) · [L3](L3-geometry.md) · [L4](L4-composition.md) · [L5](L5-cross-standard.md) · [L6](L6-diachrony.md) · [L7](L7-perception.md) · [L8](L8-computation.md) · [L9](L9-instance-process.md)
9
+
10
+ Frontier: [F1](F1-verdict.md) · [F2](F2-counterfactual.md) · [F3](F3-absence.md) · [F4](F4-instance.md) · [F5](F5-workflow.md) · [F6](F6-impact.md) · [F7](F7-editions.md) · [F8](F8-selfverify.md) · [F9](F9-projection-qa.md) · [F10](F10-notes.md) · [F11](F11-composition.md) · [F12](F12-passport.md)
@@ -0,0 +1,279 @@
1
+ # Metanorma AI — Building a Grounded RAG Service from Metanorma Sources
2
+
3
+ *Programme guidelines: how any standards body or Metanorma user builds a
4
+ citation-grounded question-answering service over their own corpus,
5
+ from scratch, using the Metanorma document model as the source of
6
+ truth. Companion to the OIML Bulletin article (2026-08-30). Everything
7
+ below is extracted from a production deployment — the OIML SMART AI
8
+ service — and is reproducible.*
9
+
10
+ ---
11
+
12
+ ## Who this is for
13
+
14
+ You have documents authored in Metanorma (or convertible to it) and you
15
+ want people to *ask questions* and get *verifiable, cited answers* —
16
+ not a search box over PDFs. The path below assumes no prior RAG
17
+ experience; every stage names the decision, the cost, and the
18
+ measurement that proves it works.
19
+
20
+ ## Principles (the short list)
21
+
22
+ 1. **Grounding is the product.** Fluency is free; trust is engineered.
23
+ Every claim cites the exact publication, edition and clause. When
24
+ the corpus doesn't contain the answer, the system says so — one
25
+ canonical refusal sentence, never cached.
26
+ 2. **Facts from the source model, never from renderings.** Metanorma
27
+ documents are typed object models; consuming the model (via MKO)
28
+ beats scraping HTML or PDF on every axis that matters.
29
+ 3. **Input = full fidelity; output = references.** The model *reads*
30
+ tables, equations, figures; it *never re-types* them. It writes a
31
+ symbolic reference; your renderer draws the producer's payload.
32
+ 4. **Measure everything, ship nothing unmeasured.** A golden question
33
+ set with witness citations, retrieval metrics over repeated runs,
34
+ and a faithfulness judge that sees the passages the answer was
35
+ actually built from. CI runs the suites; a change without a number
36
+ is a guess.
37
+ 5. **Cost-first on the hot path, quality-first on one-time work.**
38
+ Every question pays the serving model; enrichment, captioning and
39
+ evaluation are one-time costs whose quality persists.
40
+
41
+ ---
42
+
43
+ ## Stage 0 — Decide your corpus and its identity backbone
44
+
45
+ **What:** the documents you will index, and the canonical identifier
46
+ that makes two spellings of the same publication collapse into one.
47
+
48
+ - Establish a registry: publication families → editions → status
49
+ (current / superseded / withdrawn). **Derive status from
50
+ supersession links, never from a status field** — fields lie; edges
51
+ are structure.
52
+ - Surface the gaps: families with no derivable current edition are your
53
+ upstream bibliographic worklist, not something to hide.
54
+ - Record language and edition metadata on every document *before*
55
+ chunking; it becomes the filter your retrieval uses.
56
+
57
+ **Cost:** $0. **Measurement:** registry completeness (families with a
58
+ derivable current edition / total families).
59
+
60
+ ## Stage 1 — Ingest the Metanorma model (MKO), not renderings
61
+
62
+ **What:** export each document as a Metanorma Knowledge Objects bundle
63
+ (MN 116): typed units (clauses, tables, terms, equations, figures,
64
+ requirements), a section graph (part_of, cites, defines edges), a
65
+ native glossary, and native bibliographic objects.
66
+
67
+ ```
68
+ metanorma compile document.adoc -x mko # (or the Ruby exporter today)
69
+ ```
70
+
71
+ **Key decisions:**
72
+ - Chunks follow **clause boundaries** — the unit a human cites — never
73
+ token windows. Tables, equations and figures are **atomic units**
74
+ (never split; cite their parent clause).
75
+ - Each chunk carries: docidentifier, edition, language, clause anchor,
76
+ status, and its **unit id** (stable content hash) for incremental
77
+ re-ingest.
78
+ - Write a **data-quality report** on every build: counts by type,
79
+ missing anchors, oversized payloads. Serving keeps score; upstream
80
+ fixes follow.
81
+
82
+ **Cost:** $0 (local compute). **Measurement:** unit counts per document
83
+ against a hand-check; zero empty chunks for preface-only documents
84
+ (a real producer bug we hit).
85
+
86
+ ## Stage 2 — Contextual enrichment (the highest-ROI one-time spend)
87
+
88
+ **What:** a one-time pass where a strong model writes a 2–3 sentence
89
+ preamble situating each chunk in its document ("This clause defines the
90
+ accuracy-class limits for load cells in R 60-1:2017 §4.1.2"), then you
91
+ embed **preamble + chunk text** and index that.
92
+
93
+ - Cache preambles by chunk id (KV or object store); unchanged chunks
94
+ are free on re-runs. Invalidate by content hash.
95
+ - Budget: ~260 in / ~210 out tokens per chunk. At our models that is
96
+ ≈$0.0017 per chunk; a 40,000-chunk corpus is ≈$70 one-time.
97
+ - Expect a measurable retrieval lift and, more visibly, the recovery of
98
+ paraphrased questions that share no vocabulary with the corpus.
99
+
100
+ **Cost:** ≈$0.002 per chunk, one-time. **Measurement:** retrieval
101
+ recall@5 before/after on a paraphrase probe set.
102
+
103
+ ## Stage 3 — Hybrid retrieval (dense + full-corpus lexical)
104
+
105
+ **What:** two retrieval lanes run in parallel on every question:
106
+
107
+ 1. **Dense** — the question is embedded (same embedding model as the
108
+ index; this is mandatory) and queried against the vector index.
109
+ 2. **Lexical** — a full-corpus BM25/FTS index over the *enriched* chunk
110
+ text, catching exact identifiers, part numbers and defined terms
111
+ that dense similarity misses.
112
+
113
+ Fuse with reciprocal rank fusion (RRF, k=60). Then a cross-encoder
114
+ reranks the fused candidates; optionally a stronger listwise pass for
115
+ complex questions.
116
+
117
+ **The lesson that cost us a baseline:** lexical retrieval must scan the
118
+ *whole corpus* as a first stage. Re-scoring only what dense already
119
+ found cannot recover what dense missed. On our corpus this single
120
+ change lifted recall@5 from 86% to 95%.
121
+
122
+ **Cost:** embedding ≈$0.00001 per query; lexical ≈free (D1/FTS5).
123
+ **Measurement:** recall@5, AP@5, MRR@5 on the golden set, **mean over
124
+ ≥3 runs** (single runs are noise).
125
+
126
+ ## Stage 4 — Query understanding (LLM-decided, zero rules)
127
+
128
+ **What:** a small model reads the question first: language, named
129
+ publication, edition, defined terms, complexity, whether it's
130
+ conversational. There are no keyword rules and no regex on user input —
131
+ the same model judges every question, in every language.
132
+
133
+ - Its output drives **metadata filters** (doc number, edition) on the
134
+ dense query — the exact-identifier query class embeddings handle
135
+ poorly.
136
+ - A **terminology graph** (defined term → defining documents) resolves
137
+ colloquial phrasing to the corpus's own vocabulary.
138
+ - Run the understanding call **in parallel** with the query embedding;
139
+ they have no dependency when no filter is emitted.
140
+
141
+ **Cost:** ≈$0.0002 per question. **Measurement:** filter precision
142
+ (named publication → correct doc_number).
143
+
144
+ ## Stage 5 — The answer contract
145
+
146
+ **What:** the model generates from the retrieved passages only, under a
147
+ data-file prompt contract:
148
+
149
+ - **Inline citations** on every claim: `[OIML R 60-1:2021 §4.4.2]`
150
+ - **Verbatim quote anchors** for normative values — the quoted phrase
151
+ must exist, character-for-character, in a cited passage
152
+ - **Symbolic unit references** (`[[u:table-1]]`) when presenting a
153
+ whole table/equation/figure — the model *points*, your renderer
154
+ draws the producer's payload
155
+ - **One canonical refusal sentence** for off-corpus questions
156
+
157
+ **Verification (all mechanical, all post-generation):**
158
+
159
+ | Check | Method | On failure |
160
+ |---|---|---|
161
+ | Quote anchors | deterministic text match against used passages | one corrective regeneration; never cache the unverified answer |
162
+ | Unit references | every `[[u:…]]` must exist in the used passages | dropped, never rendered |
163
+ | Table retyping | markdown table in output while a typed unit was available | corrective regeneration with a targeted note |
164
+ | Faithfulness | independent judge scores claims vs the passages used | answer withheld or regenerated |
165
+
166
+ The crucial correctness rule: **the faithfulness judge sees the
167
+ passages the answer was actually built from** (return them in the
168
+ response), not a fresh retrieval — otherwise you are measuring a
169
+ different question's evidence.
170
+
171
+ **Cost:** judge ≈$0.0003 per answer. **Measurement:** faithfulness
172
+ score distribution; corrective-regen rate; zero cached violations.
173
+
174
+ ## Stage 6 — Serving, caching, and cost control
175
+
176
+ - **Exact answer cache** (keyed by query hash + index version) and a
177
+ **semantic cache** (leading-dimension signature, cosine ≥0.97) serve
178
+ near-duplicate questions instantly. Refusals and conversational turns
179
+ are never cached. `fresh=true` bypasses *every* cache (we shipped a
180
+ bug where it missed the semantic one — test this).
181
+ - **Cache versioning:** bump the version on every retrieval or prompt
182
+ change; old answers expire by key.
183
+ - **Model policy:** all tiers can share one good model if it's cheap
184
+ enough (ours: ≈$0.001/answer). Hot-path understanding stays on a
185
+ smaller model; only the *answer* needs the quality.
186
+
187
+ ## Stage 7 — Evaluation as a CI gate
188
+
189
+ Build these suites **before** you need them:
190
+
191
+ 1. **Golden set** — 20–30 questions with expected citations (regex over
192
+ docidentifier) and witness values for table questions. Run on every
193
+ change; report R/AP/MRR means over ≥3 runs.
194
+ 2. **Paraphrase probes** — the same question asked 8 ways must hit the
195
+ same sources.
196
+ 3. **End-to-end behavioural cases** — refusal correctness, language
197
+ fidelity, edition steering, block rendering, auth tiers.
198
+ 4. **Faithfulness battery** — judged scores over the golden set,
199
+ re-run after any generation-model change.
200
+
201
+ Our suites caught: a corpus lane invisible to filtered retrieval (an
202
+ identity-parse bug), a mangled CSS rule that blanked dark mode, a cache
203
+ path that served stale answers to regeneration requests, and a
204
+ diversity cap that structurally evicted typed tables. None were visible
205
+ in code review; all were visible in a suite.
206
+
207
+ ## Stage 8 — The typed-block pipeline (tables, equations, figures)
208
+
209
+ This is the stage that separates a chatbot from a *standards*
210
+ assistant:
211
+
212
+ 1. **Typed payloads in a serving store** (D1/KV table: unit_id →
213
+ payload), written at ingest, validated against the MN 116 schemas.
214
+ 2. **Passage headers declare their units** so the model can reference
215
+ them: `[3] OIML R 60-1:2021 §5.1.2 unit u:table-1 (table) — …`
216
+ 3. **Server-side reference resolution**: validate against used
217
+ passages, fetch the payload, return a `blocks` array alongside the
218
+ prose.
219
+ 4. **Client rendering**: real `<table>` from columns/rows, KaTeX from
220
+ LaTeX, `<img>` from the asset store, term cards from designations +
221
+ definition. Unknown block types render nothing (forward
222
+ compatibility).
223
+ 5. **Figures**: upload image assets to immutable URLs; a vision-capable
224
+ model describes each from pixels (one-time, ≈cents per figure);
225
+ store the description so text-only tiers can explain figures too.
226
+
227
+ ## Stage 9 — Operations
228
+
229
+ - **Incremental re-ingest:** diff bundles by (document, unit_id) +
230
+ content hash; a wording change re-processes only the affected chunks.
231
+ Our corpus-proven case: a synonym change diffed exactly 34 of 3,053
232
+ chunks; a no-change re-ingest reports zero.
233
+ - **Spend ledger:** log per-request model spend to a queryable table;
234
+ review weekly against the model policy.
235
+ - **Upstream loop:** every corpus defect the service surfaces (missing
236
+ successor links, malformed table serializations, preface-only
237
+ documents) goes back as an upstream PR or issue — the corpus gets
238
+ better because serving keeps score.
239
+
240
+ ---
241
+
242
+ ## Reference stack (as deployed)
243
+
244
+ | Component | Choice | Notes |
245
+ |---|---|---|
246
+ | Platform | Cloudflare Workers only | minimal accounts, no proprietary model providers |
247
+ | Answer model | GLM-5.3 Flash (open-weight, multimodal) | ≈$0.001/answer, all tiers |
248
+ | Understanding | Qwen3-30B-A3B | cost-first hot path |
249
+ | Embeddings | Qwen3-Embedding-0.6B | same model both sides (mandatory) |
250
+ | Reranker | bge-reranker-base | ≈$0.00003/query |
251
+ | Vector index | Cloudflare Vectorize | dense lane |
252
+ | Lexical index | D1 + FTS5 (porter) | full-corpus BM25 |
253
+ | Registry + graph | D1 (documents, graph_nodes/edges) | derived status |
254
+ | Caches | KV (answer + semantic) | versioned by index |
255
+ | Typed payloads | D1 (unit_payloads) | MN 116 schemas |
256
+ | Assets | R2 (unit-keyed, immutable) | /assets/u:<id>.<ext> |
257
+ | Ingest | Python (pydantic models) | MKO → chunks → enrich → embed → upsert |
258
+ | Eval | Node (golden, paraphrase, e2e, RAGAS-style) | CI-gated |
259
+
260
+ ## What not to do (each one cost us a measurement)
261
+
262
+ - Don't scrape rendered HTML when the document model exists.
263
+ - Don't let the model re-type tables; it will, and the values will be
264
+ wrong in ways that pass review.
265
+ - Don't trust a status field over the record's own edges.
266
+ - Don't judge faithfulness against a fresh retrieval.
267
+ - Don't report single-run metrics; the noise band is ±5pp.
268
+ - Don't run your heavy eval/enrichment traffic against the same account
269
+ that serves users — you'll throttle yourself and mis-measure
270
+ everything.
271
+ - Don't skip the refusal path in your test suite; it's the contract
272
+ users trust most.
273
+
274
+ ---
275
+
276
+ *Production reference: ai.oimlsmart.org — OIML SMART AI. Format spec:
277
+ Metanorma MN 116 (Metanorma Knowledge Objects). This guide is extracted
278
+ from the system's own documentation set and can be adapted per corpus;
279
+ the measurement discipline transfers unchanged.*