@konneal/engine 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (285) hide show
  1. package/LICENSE +29 -0
  2. package/README.md +13 -0
  3. package/dist/admin.d.ts +26 -0
  4. package/dist/ai.d.ts +6 -0
  5. package/dist/anchors.d.ts +6 -0
  6. package/dist/answercache.d.ts +22 -0
  7. package/dist/ask.d.ts +5 -0
  8. package/dist/auth.d.ts +12 -0
  9. package/dist/bubble.d.ts +14 -0
  10. package/dist/chunk-LLWPT2XV.js +49 -0
  11. package/dist/chunk-MB74PTRM.js +114 -0
  12. package/dist/chunk-WOGQM7DJ.js +197 -0
  13. package/dist/chunk-WWNCWKKC.js +42 -0
  14. package/dist/completion.d.ts +5 -0
  15. package/dist/config.d.ts +154 -0
  16. package/dist/config.js +37 -0
  17. package/dist/context.d.ts +115 -0
  18. package/dist/conversations.d.ts +5 -0
  19. package/dist/drafts.d.ts +129 -0
  20. package/dist/env.d.ts +57 -0
  21. package/dist/faithfulness.d.ts +5 -0
  22. package/dist/grader.d.ts +3 -0
  23. package/dist/graph.d.ts +13 -0
  24. package/dist/hybrid.d.ts +7 -0
  25. package/dist/index.d.ts +9 -0
  26. package/dist/index.js +5373 -0
  27. package/dist/internal_gateway.d.ts +14 -0
  28. package/dist/lexical.d.ts +7 -0
  29. package/dist/livedata.d.ts +77 -0
  30. package/dist/memories.d.ts +10 -0
  31. package/dist/modelplane.d.ts +61 -0
  32. package/dist/oidc.d.ts +73 -0
  33. package/dist/pipeline.d.ts +57 -0
  34. package/dist/profile.d.ts +2 -0
  35. package/dist/profile.gen.d.ts +70 -0
  36. package/dist/profile.js +8 -0
  37. package/dist/projects.d.ts +8 -0
  38. package/dist/prompts/conversational.md +8 -0
  39. package/dist/prompts/enrichment.md +3 -0
  40. package/dist/prompts/faithfulness.md +1 -0
  41. package/dist/prompts/grader.md +5 -0
  42. package/dist/prompts/listwise.md +3 -0
  43. package/dist/prompts/precision.md +1 -0
  44. package/dist/prompts/reflect.md +1 -0
  45. package/dist/prompts/relevancy.md +1 -0
  46. package/dist/prompts/research.md +10 -0
  47. package/dist/prompts/section-summary.md +5 -0
  48. package/dist/prompts/summarize.md +1 -0
  49. package/dist/prompts/system.md +18 -0
  50. package/dist/prompts/understanding.md +17 -0
  51. package/dist/quota.d.ts +13 -0
  52. package/dist/reflect.d.ts +5 -0
  53. package/dist/refs.d.ts +40 -0
  54. package/dist/refusal.d.ts +9 -0
  55. package/dist/refusal.js +9 -0
  56. package/dist/requestScope.d.ts +26 -0
  57. package/dist/requestScope.js +10 -0
  58. package/dist/research.d.ts +8 -0
  59. package/dist/search.d.ts +4 -0
  60. package/dist/selfquery.d.ts +7 -0
  61. package/dist/session.d.ts +1 -0
  62. package/dist/share.d.ts +2 -0
  63. package/dist/structural.d.ts +27 -0
  64. package/dist/tablecontext.d.ts +11 -0
  65. package/dist/understand.d.ts +11 -0
  66. package/dist/understandContract.d.ts +29 -0
  67. package/dist/verdict.d.ts +24 -0
  68. package/docs/API.md +451 -0
  69. package/docs/ARCHITECTURE.md +302 -0
  70. package/docs/AUDIT-2026-08-24.md +71 -0
  71. package/docs/CONTRIBUTOR-AUDIT-2026-08-25.md +147 -0
  72. package/docs/INGEST-ARCHITECTURE.md +158 -0
  73. package/docs/MCP.md +92 -0
  74. package/docs/METANORMA-AI-SERIALIZATION.md +247 -0
  75. package/docs/MKO-EXPORT-PIPELINE.md +147 -0
  76. package/docs/REDESIGN-NORMATIVE-RAG-ETSI.md +485 -0
  77. package/docs/RESEARCH-SOTA-2026.md +243 -0
  78. package/docs/ROADMAP-SOTA.md +130 -0
  79. package/docs/SOTA-STAGE-SPECS.md +509 -0
  80. package/docs/annealment/F1-verdict.md +27 -0
  81. package/docs/annealment/F10-notes.md +19 -0
  82. package/docs/annealment/F11-composition.md +17 -0
  83. package/docs/annealment/F12-passport.md +17 -0
  84. package/docs/annealment/F2-counterfactual.md +20 -0
  85. package/docs/annealment/F3-absence.md +21 -0
  86. package/docs/annealment/F4-instance.md +18 -0
  87. package/docs/annealment/F5-workflow.md +21 -0
  88. package/docs/annealment/F6-impact.md +21 -0
  89. package/docs/annealment/F7-editions.md +18 -0
  90. package/docs/annealment/F8-selfverify.md +19 -0
  91. package/docs/annealment/F9-projection-qa.md +17 -0
  92. package/docs/annealment/L0-locate.md +19 -0
  93. package/docs/annealment/L1-extract.md +18 -0
  94. package/docs/annealment/L2-nomenclature.md +22 -0
  95. package/docs/annealment/L3-geometry.md +23 -0
  96. package/docs/annealment/L4-composition.md +21 -0
  97. package/docs/annealment/L5-cross-standard.md +20 -0
  98. package/docs/annealment/L6-diachrony.md +21 -0
  99. package/docs/annealment/L7-perception.md +20 -0
  100. package/docs/annealment/L8-computation.md +22 -0
  101. package/docs/annealment/L9-instance-process.md +23 -0
  102. package/docs/annealment/README.md +10 -0
  103. package/docs/guidelines-metanorma-ai-programme.md +279 -0
  104. package/docs/identity-onboarding-rag.md +65 -0
  105. package/docs/identity-service.md +219 -0
  106. package/docs/knowledge-annealment.md +273 -0
  107. package/docs/konneal-extraction-plan.md +481 -0
  108. package/docs/metanorma-for-ai.md +270 -0
  109. package/docs/mirror-plan.md +36 -0
  110. package/docs/multi-sdo-architecture.md +191 -0
  111. package/docs/paper-annealment-comparison.md +259 -0
  112. package/docs/paper-assets/architecture.svg +94 -0
  113. package/docs/paper-assets/contract-v2.svg +94 -0
  114. package/docs/paper-assets/mko-ingest.svg +91 -0
  115. package/docs/paper-oiml-bulletin.md +402 -0
  116. package/docs/paper-oiml-bulletin.mdx +419 -0
  117. package/docs/product-branding-options.md +172 -0
  118. package/docs/projects-design.md +88 -0
  119. package/docs/sota-mechanisms.md +184 -0
  120. package/docs/spec-api.md +77 -0
  121. package/docs/spec-pipeline.md +126 -0
  122. package/docs/vector-adapter.md +88 -0
  123. package/package.json +70 -0
  124. package/profile/corpora.yaml +5 -0
  125. package/profile/datasets.yaml +14 -0
  126. package/profile/prompts.yaml +5 -0
  127. package/profile/publisher.yaml +17 -0
  128. package/profile/retrieval.yaml +1 -0
  129. package/profile/sources.yaml +5 -0
  130. package/profile/ui.yaml +7 -0
  131. package/scripts/gen_profile.mjs +33 -0
  132. package/workers/shared/ai.ts +21 -0
  133. package/workers/shared/auth.ts +16 -0
  134. package/workers/shared/chunk.ts +108 -0
  135. package/workers/shared/oidc.ts +312 -0
  136. package/workers/shared/router.ts +45 -0
  137. package/workers/shared/session.ts +104 -0
  138. package/workers/worker_internal/src/index.ts +157 -0
  139. package/workers/worker_internal/tsconfig.json +15 -0
  140. package/workers/worker_internal/wrangler.toml +32 -0
  141. package/workers/worker_mcp/src/index.ts +175 -0
  142. package/workers/worker_mcp/tsconfig.json +13 -0
  143. package/workers/worker_mcp/wrangler.toml +18 -0
  144. package/workers/worker_public/migrations/0002_conversations.sql +22 -0
  145. package/workers/worker_public/migrations/0003_shared_conversations.sql +9 -0
  146. package/workers/worker_public/migrations/0004_graph.sql +16 -0
  147. package/workers/worker_public/migrations/0005_documents.sql +19 -0
  148. package/workers/worker_public/migrations/0006_conversation_entities.sql +11 -0
  149. package/workers/worker_public/migrations/0007_chunks_fts.sql +43 -0
  150. package/workers/worker_public/migrations/0008_unit_payloads.sql +15 -0
  151. package/workers/worker_public/migrations/0009_chunks_unit.sql +8 -0
  152. package/workers/worker_public/migrations/0009_message_context.sql +7 -0
  153. package/workers/worker_public/migrations/0010_model_nodes.sql +33 -0
  154. package/workers/worker_public/migrations/0011_schema_union.sql +38 -0
  155. package/workers/worker_public/migrations/0012_memories.sql +15 -0
  156. package/workers/worker_public/migrations/0013_projects.sql +21 -0
  157. package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.d.ts +16306 -0
  158. package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.ts +16261 -0
  159. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.d.ts +16373 -0
  160. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.ts +16328 -0
  161. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.d.ts +16382 -0
  162. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.ts +16337 -0
  163. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.d.ts +16383 -0
  164. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.ts +16338 -0
  165. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.d.ts +16403 -0
  166. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.ts +16358 -0
  167. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.d.ts +16408 -0
  168. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.ts +16363 -0
  169. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.d.ts +16414 -0
  170. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.ts +16369 -0
  171. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.d.ts +16414 -0
  172. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.ts +16369 -0
  173. package/workers/worker_public/node_modules/@cloudflare/workers-types/README.md +135 -0
  174. package/workers/worker_public/node_modules/@cloudflare/workers-types/entrypoints.svg +53 -0
  175. package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.d.ts +17095 -0
  176. package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.ts +17050 -0
  177. package/workers/worker_public/node_modules/@cloudflare/workers-types/index.d.ts +16306 -0
  178. package/workers/worker_public/node_modules/@cloudflare/workers-types/index.ts +16261 -0
  179. package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.d.ts +16447 -0
  180. package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.ts +16402 -0
  181. package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.d.ts +16306 -0
  182. package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.ts +16261 -0
  183. package/workers/worker_public/node_modules/@cloudflare/workers-types/package.json +11 -0
  184. package/workers/worker_public/package.json +13 -0
  185. package/workers/worker_public/prompts/conversational.md +8 -0
  186. package/workers/worker_public/prompts/enrichment.md +3 -0
  187. package/workers/worker_public/prompts/faithfulness.md +1 -0
  188. package/workers/worker_public/prompts/grader.md +5 -0
  189. package/workers/worker_public/prompts/listwise.md +3 -0
  190. package/workers/worker_public/prompts/precision.md +1 -0
  191. package/workers/worker_public/prompts/reflect.md +1 -0
  192. package/workers/worker_public/prompts/relevancy.md +1 -0
  193. package/workers/worker_public/prompts/research.md +10 -0
  194. package/workers/worker_public/prompts/section-summary.md +5 -0
  195. package/workers/worker_public/prompts/summarize.md +1 -0
  196. package/workers/worker_public/prompts/system.md +18 -0
  197. package/workers/worker_public/prompts/understanding.md +17 -0
  198. package/workers/worker_public/public/app.js +166 -0
  199. package/workers/worker_public/public/index.html +48 -0
  200. package/workers/worker_public/public/style.css +147 -0
  201. package/workers/worker_public/schema.sql +248 -0
  202. package/workers/worker_public/src/admin.ts +358 -0
  203. package/workers/worker_public/src/ai.ts +71 -0
  204. package/workers/worker_public/src/anchors.ts +41 -0
  205. package/workers/worker_public/src/answercache.ts +72 -0
  206. package/workers/worker_public/src/ask.ts +1094 -0
  207. package/workers/worker_public/src/auth.ts +252 -0
  208. package/workers/worker_public/src/bubble.ts +111 -0
  209. package/workers/worker_public/src/completion.ts +75 -0
  210. package/workers/worker_public/src/config.ts +238 -0
  211. package/workers/worker_public/src/context.ts +238 -0
  212. package/workers/worker_public/src/conversations.ts +162 -0
  213. package/workers/worker_public/src/drafts.ts +497 -0
  214. package/workers/worker_public/src/env.ts +90 -0
  215. package/workers/worker_public/src/faithfulness.ts +63 -0
  216. package/workers/worker_public/src/grader.ts +89 -0
  217. package/workers/worker_public/src/graph.ts +63 -0
  218. package/workers/worker_public/src/hybrid.ts +77 -0
  219. package/workers/worker_public/src/index.ts +441 -0
  220. package/workers/worker_public/src/internal_gateway.ts +41 -0
  221. package/workers/worker_public/src/lexical.ts +86 -0
  222. package/workers/worker_public/src/lib/hit.ts +4 -0
  223. package/workers/worker_public/src/lib/http.ts +83 -0
  224. package/workers/worker_public/src/lib/router.ts +4 -0
  225. package/workers/worker_public/src/livedata.ts +334 -0
  226. package/workers/worker_public/src/memories.ts +81 -0
  227. package/workers/worker_public/src/modelplane.ts +213 -0
  228. package/workers/worker_public/src/oidc.ts +333 -0
  229. package/workers/worker_public/src/pipeline.ts +377 -0
  230. package/workers/worker_public/src/ports/blobs.ts +7 -0
  231. package/workers/worker_public/src/ports/cloudflare/adapters.ts +177 -0
  232. package/workers/worker_public/src/ports/kv.ts +8 -0
  233. package/workers/worker_public/src/ports/model.ts +28 -0
  234. package/workers/worker_public/src/ports/runtime.ts +13 -0
  235. package/workers/worker_public/src/ports/store.ts +20 -0
  236. package/workers/worker_public/src/ports/vector.ts +26 -0
  237. package/workers/worker_public/src/profile.gen.ts +101 -0
  238. package/workers/worker_public/src/profile.ts +16 -0
  239. package/workers/worker_public/src/projects.ts +108 -0
  240. package/workers/worker_public/src/prompts.d.ts +6 -0
  241. package/workers/worker_public/src/quota.ts +54 -0
  242. package/workers/worker_public/src/reflect.ts +67 -0
  243. package/workers/worker_public/src/refs.ts +107 -0
  244. package/workers/worker_public/src/refusal.ts +65 -0
  245. package/workers/worker_public/src/requestScope.ts +71 -0
  246. package/workers/worker_public/src/research.ts +126 -0
  247. package/workers/worker_public/src/search.ts +56 -0
  248. package/workers/worker_public/src/selfquery.ts +25 -0
  249. package/workers/worker_public/src/session.ts +4 -0
  250. package/workers/worker_public/src/share.ts +53 -0
  251. package/workers/worker_public/src/stages/conceptGraph.ts +68 -0
  252. package/workers/worker_public/src/stages/conceptSteer.ts +39 -0
  253. package/workers/worker_public/src/stages/corpusScope.ts +25 -0
  254. package/workers/worker_public/src/stages/dedup.ts +10 -0
  255. package/workers/worker_public/src/stages/dense.ts +73 -0
  256. package/workers/worker_public/src/stages/diversity.ts +33 -0
  257. package/workers/worker_public/src/stages/editionCover.ts +63 -0
  258. package/workers/worker_public/src/stages/editionSteer.ts +88 -0
  259. package/workers/worker_public/src/stages/familyBoost.ts +22 -0
  260. package/workers/worker_public/src/stages/federate.ts +22 -0
  261. package/workers/worker_public/src/stages/glossary.ts +65 -0
  262. package/workers/worker_public/src/stages/graphLane.ts +31 -0
  263. package/workers/worker_public/src/stages/hyde.ts +29 -0
  264. package/workers/worker_public/src/stages/index.ts +69 -0
  265. package/workers/worker_public/src/stages/lexicalUnion.ts +21 -0
  266. package/workers/worker_public/src/stages/multiQuery.ts +57 -0
  267. package/workers/worker_public/src/stages/overviewDemote.ts +14 -0
  268. package/workers/worker_public/src/stages/poolOpen.ts +10 -0
  269. package/workers/worker_public/src/stages/propagate.ts +15 -0
  270. package/workers/worker_public/src/stages/rerank.ts +47 -0
  271. package/workers/worker_public/src/stages/seal.ts +16 -0
  272. package/workers/worker_public/src/stages/sectionDescent.ts +61 -0
  273. package/workers/worker_public/src/stages/stdRefNudge.ts +35 -0
  274. package/workers/worker_public/src/stages/subQuery.ts +42 -0
  275. package/workers/worker_public/src/stages/termNudge.ts +24 -0
  276. package/workers/worker_public/src/stages/typedPin.ts +131 -0
  277. package/workers/worker_public/src/stages/types.ts +112 -0
  278. package/workers/worker_public/src/stages/windowFloor.ts +23 -0
  279. package/workers/worker_public/src/structural.ts +171 -0
  280. package/workers/worker_public/src/tablecontext.ts +41 -0
  281. package/workers/worker_public/src/understand.ts +72 -0
  282. package/workers/worker_public/src/understandContract.ts +67 -0
  283. package/workers/worker_public/src/verdict.ts +255 -0
  284. package/workers/worker_public/tsconfig.json +18 -0
  285. package/workers/worker_public/wrangler.toml +104 -0
@@ -0,0 +1,509 @@
1
+ # SOTA Architecture — Detailed Stage-by-Stage Gaps and Target Specifications
2
+
3
+ *Companion to `RESEARCH-SOTA-2026.md` (technique survey) and
4
+ `ROADMAP-SOTA.md` (phasing). This document is the ENGINEERING SPEC: for
5
+ every stage — source documents, ingestion, structuring, index management,
6
+ query processing, retrieval, answer formulation, follow-up handling,
7
+ evaluation — it states what we have, the precise gaps, and the TARGET SHAPE
8
+ (schemas, policies, thresholds, models) with migration steps. 2026-08-26.*
9
+
10
+ ---
11
+
12
+ ## Stage 0 — Source contract: the desired Metanorma document shape
13
+
14
+ ### Current
15
+ We consume two shapes: clean collections (`sources/<id>/document.adoc` +
16
+ `collection.yml` + per-clause `metanorma/sections/*.adoc`, 29 docs) and
17
+ dirty OCR slates (880 docs: adoc + `images/extracted` + manifest). Parsing
18
+ is **HTML-first** (BeautifulSoup over compiled Metanorma HTML): we recover
19
+ clause anchors from heading text, tables as caption+rows, drop GUID
20
+ anchors and OCR artifacts. Identity comes from a precedence ladder
21
+ (readable `:docidentifier:` → title-declared annex volume → slug).
22
+
23
+ ### Gaps
24
+ - Identity fields live in prose-ish adoc attributes with no schema
25
+ enforcement; every consumer re-derives structure from HTML scraping.
26
+ - Tables arrive as display HTML; column semantics (units, header
27
+ hierarchy, row scope) are lost or mangled by OCR.
28
+ - Equations are `stem:[...]` AsciiMath strings embedded in prose — no
29
+ MathML/LaTeX dual form, no description.
30
+ - Terms/bibliography are untyped text (Glossarist + relaton data exist in
31
+ sibling repos but are joined post-hoc, not at source).
32
+ - No per-document quality signals (OCR confidence, completeness) beyond
33
+ shell flags.
34
+
35
+ ### Target source contract — consume the metanorma-document MODEL HUB
36
+
37
+ Metanorma is model-driven, and the toolchain already provides the right
38
+ substrate: **`metanorma-document`** deserializes semantic XML into typed
39
+ lutaml-model classes (`Metanorma::IsoDocument::Root.from_xml`,
40
+ `basic_document` → `standard_document` → `iso_document` hierarchies,
41
+ collections, mirror round-trips) — and, being lutaml-model based, the
42
+ same model serializes natively to **XML, YAML, or JSON**. Requirements,
43
+ permissions, and recommendations are ALREADY first-class block classes
44
+ (`RequirementModel`, `PermissionModel`). The upstream sibling
45
+ **`modspec-ruby`** models normative statements and conformance tests as
46
+ addressable objects (`/req/<class>/<name>` URIs, obligation, inheritance,
47
+ suites) with YAML/JSON round-trips; **Glossarist** and **Relaton**
48
+ exports provide the terminology and citation datasets.
49
+
50
+ The ingest contract is therefore NOT a file format — it is a **projection
51
+ over the model**:
52
+
53
+ 1. Deserialize: semantic XML (or YAML/JSON) → metanorma-document model.
54
+ 2. Project: a small adapter walks typed model nodes and emits the
55
+ canonical node JSON below (the adapter can live upstream as a
56
+ metanorma-document serialization flavor, or in our ingest as a Ruby
57
+ step). ChunkRecordV2 (Stage 2) is filled mechanically from nodes.
58
+ 3. Requirement/conformance nodes additionally project through the ModSpec
59
+ model (identifier, obligation, class, linked conformance tests).
60
+ 4. Glossarist/Relaton dataset exports feed definition chunks and graph
61
+ edges directly (no side-repo joins).
62
+ 5. Adoc model elements are parsed directly when no compiled model exists;
63
+ compiled HTML remains the OCR-slate fallback ONLY.
64
+
65
+ Renderings (HTML/PDF/DOCX) are never ingestion inputs outside the OCR
66
+ fallback. Encoding is irrelevant — the node schema is the contract.
67
+
68
+ For reference, the node projection shape:
69
+
70
+ **1. Semantic XML (primary, clean corpus; ask upstream in #592 to emit it
71
+ as a first-class build artifact):**
72
+ ```xml
73
+ <standard-document>
74
+ <bibdata>
75
+ <docidentifier type="oiml">OIML R 60-1</docidentifier>
76
+ <edition>2017</edition> <language>en</language>
77
+ </bibdata>
78
+ <sections>
79
+ <clause id="cl-4.1.2" obligation="normative">
80
+ <title>Maximum number of verification intervals</title>
81
+ <p>…</p>
82
+ <table id="tbl-4.1.2-1">
83
+ <name>Maximum permissible errors</name>
84
+ <thead><th>Load m (in units of e)</th><th>MPE</th></thead>
85
+ <tbody><tr><td>0 ≤ m ≤ 5·10³</td><td>0.5e</td></tr></tbody>
86
+ </table>
87
+ <formula id="frm-1"><stem type="AsciiMath">n_LC &lt;= ...</stem></formula>
88
+ </clause>
89
+ </sections>
90
+ <terms>
91
+ <term id="term-load-cell"><preferred>load cell</preferred>
92
+ <definition>…</definition></term>
93
+ </terms>
94
+ <bibliography>
95
+ <bibitem id="IEC61000-4-2">…</bibitem>
96
+ </bibliography>
97
+ </standard-document>
98
+ ```
99
+
100
+ **2. Canonical AI-friendly JSON projection (what RAG actually wants;
101
+ derived from the model, upstream or at ingest):**
102
+ ```jsonc
103
+ {
104
+ "doc": { "docidentifier": "OIML R 60-1", "part": "1", "edition": "2017",
105
+ "language": "en", "doctype": "R", "status": "in-force",
106
+ "family": "R-60", "superseded_by": null },
107
+ "nodes": [
108
+ { "type": "clause", "anchor": "4.1.2", "obligation": "normative",
109
+ "breadcrumb": ["4 Testing", "4.1 Classification"], "text": "…" },
110
+ { "type": "table", "anchor": "tbl-4.1.2-1",
111
+ "caption": "Maximum permissible errors",
112
+ "columns": [{ "label": "Load m (in units of e)", "unit": "e" },
113
+ { "label": "MPE" }],
114
+ "rows": [["0 ≤ m ≤ 5·10³", "0.5e"]] },
115
+ { "type": "formula", "anchor": "frm-1", "asciimath": "n_LC <= ...",
116
+ "latex": "n_{LC} \leq …", "described": "limit on verification intervals" },
117
+ { "type": "term", "anchor": "term-load-cell", "concept": "load-cell",
118
+ "definition": "…" },
119
+ { "type": "reference", "anchor": "IEC61000-4-2",
120
+ "cited": "IEC 61000-4-2:2008" }
121
+ ]
122
+ }
123
+ ```
124
+ This projection maps 1:1 onto `ChunkRecordV2` (Stage 2) — ingest becomes a
125
+ mechanical walk of typed model nodes, with zero scraping and zero
126
+ structure guessing. If a specific serialization (e.g. TOML or another
127
+ AI-oriented form) is preferred upstream, the projection schema is the
128
+ contract; the encoding is swappable.
129
+
130
+ **3. Adoc model elements parsed directly** — where semantic XML is not
131
+ available, the adoc source IS model-bearing (`stem:[]`, adoc tables,
132
+ terms sections); parse the model, not a rendering.
133
+
134
+ **4. HTML only as OCR-slate fallback** — the dirty corpus's compiled HTML
135
+ remains the last-resort input for slates whose adoc is too degraded;
136
+ everything else moves off renderings entirely.
137
+
138
+ **Migration:** clean corpus first (compile → semantic XML → JSON
139
+ projection → existing chunker consumes nodes); dirty corpus stays on the
140
+ HTML-first parser until per-slate re-OCR/re-compile upgrades it. Upstream
141
+ ask (#592): ship the semantic XML (and ideally the JSON projection) as an
142
+ official output flavor so AI consumers never parse renderings.
143
+
144
+ ---
145
+
146
+ ## Stage 1 — Canonical document model (`DocumentRecord`)
147
+
148
+ ### Current
149
+ No explicit document record — identity is reconstructed per chunk at parse
150
+ time and denormalized into chunk metadata. Relaton join adds
151
+ `status`/`superseded_by` per doc.
152
+
153
+ ### Gaps
154
+ - Identity resolution runs inline in the parser; no persisted, verifiable
155
+ document registry (we found identifier corruption the hard way — R 60-A).
156
+ - No edition/part graph (which parts belong to which family, which edition
157
+ supersedes which) outside relaton YAMLs we don't control.
158
+
159
+ ### Target
160
+ ```ts
161
+ interface DocumentRecord {
162
+ id: string; // "OIML R 60-1:2017:en" (canonical identity)
163
+ docidentifier: string; // "OIML R 60-1"
164
+ family: string; // "R-60" (graph node)
165
+ part: string | null; // "1" | "A" | "annexes"
166
+ edition: string; // "2017"
167
+ language: string; // "en"
168
+ doctype: "R"|"D"|"B"|"G"|"E"|"V";
169
+ status: "in-force"|"superseded"|"withdrawn"|"joint";
170
+ superseded_by: string | null; // canonical id
171
+ title: string;
172
+ source: { repo: string; slug: string; tier: "clean"|"dirty"|"synthetic" };
173
+ quality: { ocr_confidence?: number; shell: boolean; word_count: number };
174
+ clause_tree: Array<{ anchor: string; title: string; path: string[] }>;
175
+ content_hash: string; // re-ingest dedup / enrichment invalidation
176
+ }
177
+ ```
178
+ Persisted to D1 (`documents` table) at ingest; chunk metadata is JOINED
179
+ from it, never hand-assembled. **This is the SSOT for identity** — the
180
+ R 60-A class of bug becomes a one-row fix with a re-chunk of one doc.
181
+
182
+ ---
183
+
184
+ ## Stage 2 — Chunk model v2: typed blocks, not just prose
185
+
186
+ ### Current
187
+ `ChunkRecord = { id, doc_id, chunk_ref, text, metadata }` where metadata
188
+ carries identity + `chunk_text` (context-enriched after Phase 0). All
189
+ blocks — prose, tables, equations, definitions, references — are stringified
190
+ into `text`. Tables are linearized caption+rows; equations stay AsciiMath
191
+ inside prose.
192
+
193
+ ### Gaps (G1, G11)
194
+ - Table values (our most-queried content: MPE tables, accuracy classes)
195
+ are flattened: header/units/row semantics lost; the model reads
196
+ linearized rows and can transcribe wrong cells.
197
+ - Equations are retrievable only via surrounding prose.
198
+ - Definitions and normative references aren't typed, so no dedicated
199
+ retrieval lane or display treatment.
200
+
201
+ ### Target — one chunk schema, typed payloads
202
+ ```ts
203
+ type BlockType = "clause" | "table" | "equation" | "definition" | "reference"
204
+ | "requirement" | "conformance_test" | "family";
205
+
206
+ interface ChunkRecordV2 {
207
+ id: string; // content-hash (stable across re-chunk)
208
+ doc: DocumentRecord["id"]; // FK, joined at upsert
209
+ block: BlockType;
210
+ anchor: string; // "4.1.2" | "t4.1.2-1" | "term:load-cell"
211
+ breadcrumb: string[]; // ["4 Testing", "4.1 Classification"]
212
+ text: string; // DISPLAY text (human-readable)
213
+ embed_input: string; // EMBEDDING input = context + display/serialization
214
+ context?: string; // LLM situating context (Phase 0)
215
+ table?: { // block === "table"
216
+ caption: string;
217
+ columns: Array<{ label: string; unit?: string; scope?: string }>;
218
+ rows: string[][]; // exact cell values, verbatim
219
+ };
220
+ equation?: { asciimath: string; latex: string; described: string };
221
+ term?: { concept: string; definition: string; source_vocab: string };
222
+ requirement?: { // ModSpec projection (metanorma-document
223
+ identifier: string; // RequirementModel → modspec-ruby)
224
+ class: string; // "/req/oiml-r60-1/classification"
225
+ obligation: "requirement" | "recommendation" | "permission";
226
+ statement: string; // the normative statement, verbatim
227
+ inherits: string[]; // parent statement URIs
228
+ };
229
+ conformance_test?: {
230
+ identifier: string; // "/conf/oiml-r60-1/nlc-limit"
231
+ class: string;
232
+ requirement: string; // URI of the tested statement
233
+ method: string; // verification method text
234
+ };
235
+ reference?: { cited: string; relaton_key: string | null };
236
+ quality: { ocr_confidence?: number };
237
+ }
238
+ ```
239
+ - **Derivation:** `ChunkRecordV2` fields come from the Stage 0 model
240
+ projection (semantic XML → JSON), not from HTML scraping — the parser
241
+ walks typed model nodes and fills payloads mechanically.
242
+ - **Embedding input per type:** tables embed as `context + caption +
243
+ header map + row tuples serialized` (structure-preserving, TabRAG
244
+ lesson); equations embed as `described + latex`; definitions embed the
245
+ verbatim definition.
246
+ - **Display text stays human-form** — the UI renders tables as tables.
247
+ - **Generation prompt gets the table JSON** for row-precise answers and
248
+ quote-anchoring (Stage 7).
249
+
250
+ ---
251
+
252
+ ## Stage 3 — Ingestion pipeline
253
+
254
+ ### Current (worked, proven)
255
+ parse (HTML-first, identity ladder, English-only) → family chunks →
256
+ relaton status join → chunks.jsonl → embed (resumable, 25/batch) → upsert
257
+ (idempotent, orphan-deleted) → contextual enrichment (running, in-place
258
+ upsert + KV context cache).
259
+
260
+ ### Gaps
261
+ - Chunking policy is prose-only; no per-block typing (Stage 2).
262
+ - No ingest VERIFICATION stage: we ship counts, not invariants (e.g.,
263
+ "every table chunk has ≥1 row", "every doc has an overview chunk").
264
+ - Re-ingest of one document is manual (full-corpus muscle).
265
+ - Enrichment re-run invalidation is manual (`content_hash` unused).
266
+
267
+ ### Target pipeline (each stage gated by invariants)
268
+ ```
269
+ parse → identity (DocumentRecord, D1) → structure (clause tree, typed blocks)
270
+ → chunk (policy per BlockType) → enrich (quality-first model, KV-cached)
271
+ → embed (per-type embed_input) → graph project (Stage 6 lanes)
272
+ → upsert (public|internal by corpus) → VERIFY → INDEX_VERSION bump
273
+ ```
274
+ **Verify stage (new):** per-doc chunk counts vs clause tree; table chunks
275
+ carry `table` payload; definition chunks resolvable to vocab concepts;
276
+ no orphan ids; embed_input ≤ model max; report to D1 `ingest_runs` table.
277
+ **Partial re-ingest:** `ingest one --doc OIML-R-60-1-2017-en` deletes that
278
+ doc's chunk ids, re-runs stages 3–7 for one document. **Enrichment
279
+ invalidation:** content_hash changes → KV context key evicted, re-enrich
280
+ paid only for changed chunks.
281
+
282
+ ---
283
+
284
+ ## Stage 4 — Index management (Vectorize)
285
+
286
+ ### Current
287
+ `idx_oiml_public_v2` (31k vectors) + `idx_iso_internal` (structural
288
+ isolation). Metadata indexes exist for `doctype, doc_number, edition,
289
+ language`. topK ≤ 50 with `returnMetadata: "all"`. `INDEX_VERSION` gates
290
+ the KV answer cache; deletes/upserts ride `/admin/*` endpoints.
291
+
292
+ ### Gaps (G3, G4)
293
+ - No sparse lane in Vectorize; lexical recall only over already-retrieved
294
+ candidates (in-worker keywordRank) — exact-term recall depends on dense
295
+ top-50 having caught the term.
296
+ - No dimension/quantization control (Vectorize limitation).
297
+ - Index rebuild story is "re-run ingest" (acceptable at 31k, must stay
298
+ scripted and tested).
299
+
300
+ ### Target
301
+ - **In-worker lexical index (the practical sparse lane):** D1 inverted
302
+ index (`term → chunk_ids`) over enriched `embed_input` built during
303
+ ingest; query-side: understanding-extracted key terms → top-200 lexical
304
+ candidates → RRF with dense candidates (today's keywordRank, but
305
+ corpus-wide rather than post-hoc). Cost: one D1 table, one build pass.
306
+ - **Metadata indexes:** add `block` (Stage 2 typing) for
307
+ lane-selective retrieval (`block = table` boosts on value questions).
308
+ - **Keep hard rules written down:** filter-first query pattern, fallback
309
+ unfiltered merge (already implemented); per-field metadata index
310
+ requirement; INDEX_VERSION bump protocol on any content/logic change.
311
+
312
+ ---
313
+
314
+ ## Stage 5 — Chat-side processing: intent, state, caches
315
+
316
+ ### Current
317
+ LLM understanding per ask (JSON: intent, docidentifier/number, edition,
318
+ language, process_intent, term, standalone_query, complexity,
319
+ query_variants, sub_queries, hypothetical_answer), 1200 max_tokens, 7s/4s
320
+ timeouts, warm embedding parallel, conversational route (identity/greeting
321
+ — any language), 16k context budget with history compaction (summarized
322
+ overflow), exact-match KV answer cache.
323
+
324
+ ### Gaps (G5, G6, G7)
325
+ - No cross-turn entity state — each turn re-derives referents from raw
326
+ history text only.
327
+ - Cache is exact-text; near-duplicate queries re-pay the whole pipeline.
328
+ - No follow-up suggestions; clarifying questions permitted but never
329
+ driven by an explicit ambiguity signal.
330
+
331
+ ### Target
332
+ **Understanding v2 output (additive):**
333
+ ```jsonc
334
+ {
335
+ ...existing fields...,
336
+ "entities": [ { "type": "document"|"term"|"edition"|"unit", "value": "R 60", "resolved": "OIML R 60:2017" } ],
337
+ "ambiguity": { "flag": true, "reason": "edition unspecified; R 60 has 2 in-force editions", "clarify": "Which edition — 2000 or 2017?" },
338
+ "follow_ups": [ "What are the accuracy classes?", "How is n_LC limited?" ]
339
+ }
340
+ ```
341
+ - **Conversation entity map (D1 per conversation):** entities upserted per
342
+ turn; understanding consumes the map so "the 2017 one / it / that table"
343
+ resolve O(1). Privacy: conversation-scoped, dies with the conversation.
344
+ - **Semantic cache:** KV `sc:<quantized-query-embedding> → answer+ts`;
345
+ cosine ≥ 0.96 and same corpus tier → serve with "similar question"
346
+ badge; write-through after generation. Eviction via INDEX_VERSION.
347
+ - **Follow-up chips:** rendered from `follow_ups` (API-driven, per Stage-1
348
+ suggestion pattern); logged CTR in telemetry.
349
+ - **Clarification flow:** when `ambiguity.flag` and knowledge-intent, ask
350
+ the generated clarify question INSTEAD of answering (one level deep,
351
+ no loops).
352
+
353
+ ---
354
+
355
+ ## Stage 6 — Retrieval accuracy
356
+
357
+ ### Current
358
+ Dense (query + variants + HyDE + sub-queries) → RRF fusion → filter-first
359
+ + unfiltered merge → overview penalty / family boost → bge cross-encoder
360
+ rerank → in-worker keyword RRF → term/language/edition-recency boosts →
361
+ diversity caps → CRAG grade → corrective re-retrieval (one shot).
362
+
363
+ ### Gaps (G8, G9, G10, G3)
364
+ - No graph lane; relationship queries ("what references R 60?") depend on
365
+ bibliographic chunks happening to surface.
366
+ - No LLM-listwise tier for the final ordering.
367
+ - Correction is single-shot, not a loop.
368
+ - Lexical recall is post-hoc (Stage 4 fixes the index side).
369
+
370
+ ### Target — retrieval as LANES + CASCADE
371
+ ```
372
+ lanes (parallel, budget-capped):
373
+ dense q + variants + HyDE + sub-queries (existing)
374
+ lexical inverted-index top-200 on key terms (Stage 4 build)
375
+ graph entities → D1 graph edges → candidate doc_numbers
376
+ (relaton citation edges + vocab concept relations; PUBLIC
377
+ projection = OIML nodes/edges only, enforced at build)
378
+ table block=table filtered dense search when value-question
379
+ cascade:
380
+ fuse (RRF k=60) → cross-encoder to top-50 → LLM listwise over top-12
381
+ (member/hard queries only, glm-4.7-flash, ~$0.0002/q) → diversity caps
382
+ loop (agentic, Workflows "research mode", spend-capped ≤3 iterations):
383
+ retrieve → sufficiency judge (grader model) → re-retrieve with gap
384
+ terms → answer
385
+ ```
386
+ **Graph projection schema (D1):**
387
+ ```sql
388
+ CREATE TABLE graph_nodes (id TEXT PRIMARY KEY, kind TEXT, label TEXT); -- doc|concept|term
389
+ CREATE TABLE graph_edges (src TEXT, dst TEXT, kind TEXT, meta TEXT);
390
+ -- kinds: cites|supersedes|part_of|defines|related_to|
391
+ -- tested_by|requirement_of (from ModSpec projection)
392
+ -- seeded from relaton (cites/supersedes/part_of) + vocab glossarist
393
+ -- (defines/related_to); PUBLIC projection excludes ISO labels entirely
394
+ ```
395
+
396
+ ---
397
+
398
+ ## Stage 7 — Answer formulation
399
+
400
+ ### Current
401
+ Data-file system prompt; verbatim normative values required; inline
402
+ passage-level citations `[OIML R 60-1:2017 §4.1.2]`; canonical refusal
403
+ sentence + redirect; streaming SSE; CRAG + self-RAG reflection (one shot);
404
+ feedback buttons; deep links to R2 renderings.
405
+
406
+ ### Gaps (G11, G12, G1)
407
+ - Citations are passage-level; no required quote anchor.
408
+ - Tables answered from linearized text (Stage 2 fixes the input).
409
+ - Single candidate; no sample-and-verify.
410
+
411
+ ### Target — quote-anchor protocol (the normative-corpus trust feature)
412
+ System prompt (data file) requires, for every normative claim:
413
+ ```
414
+ …MPE is 0.5e for 0 ≤ m ≤ 5·10³ [3: "the maximum permissible error shall
415
+ not exceed 0.5e"]…
416
+ ```
417
+ i.e. `[passage#: "verbatim source phrase ≤12 words"]`. Verification
418
+ (evals + optional runtime check): the quoted string is a substring of the
419
+ cited passage's text — mechanically checkable, zero LLM cost. Table
420
+ answers render the `table` payload as a real table with the cited cells.
421
+ **Sample-and-verify lane (hard queries):** 2 candidates, faithfulness +
422
+ answer-relevancy judges select; rollout eval-gated.
423
+
424
+ ---
425
+
426
+ ## Stage 8 — Follow-up handling
427
+
428
+ ### Current
429
+ Standalone-query rewriting (understanding), structural ≤8-word fold for
430
+ warm embed, history compaction (summarized overflow), latest-message
431
+ discipline in the prompt.
432
+
433
+ ### Gaps
434
+ Covered by Stage 5 targets (entity map, follow-up chips, clarification
435
+ flow). Add: **follow-up eval probes** — MTRAG-style multi-turn scripts in
436
+ the retrieval eval (turn 2 uses "and for class III?", turn 3 "what about
437
+ the 2000 edition?") so conversation quality is measured, not assumed.
438
+
439
+ ---
440
+
441
+ ## Stage 9 — Evaluation and operations
442
+
443
+ ### Current
444
+ Golden set (19 answer-level), live e2e (13), retrieval eval hit@5 +
445
+ 8 paraphrase probes, faithfulness judge, D1 telemetry + spend ledger,
446
+ feedback buttons.
447
+
448
+ ### Gaps (G13)
449
+ - Missing answer-relevancy + context-precision metrics.
450
+ - No derived dashboards (win rates over time, latency percentiles, cache
451
+ hit rate, refusal rate).
452
+ - No promotion gate automation (golden threshold check as deploy gate).
453
+
454
+ ### Target
455
+ - `tests/eval-suite.mjs`: RAGAS-style battery over golden cases —
456
+ faithfulness ✓, answer relevancy (judge scores answer↔question),
457
+ context precision (judge ranks relevance of each retrieved passage),
458
+ quote-anchor coverage (mechanical). Nightly snapshot to
459
+ `artifacts/eval/`; fail-threshold in CI for the mechanical parts.
460
+ - `/v1/admin/stats` extensions: latency p50/p95, cache-hit rate, refusal
461
+ rate, feedback ratios, eval-latest summary (single JSON).
462
+ - Promotion gate: `npm run eval:gate` exits non-zero if mechanical
463
+ metrics regress — wired as the deploy pre-step for index-affecting
464
+ changes.
465
+
466
+ ---
467
+
468
+ ## Migration order (dependency-aware)
469
+
470
+ 1. Stage 2 chunk typing + Stage 3 verify (index shape changes first)
471
+ 2. Stage 1 document registry (D1) — identity SSOT
472
+ 3. Stage 4 lexical index + `block` metadata index
473
+ 4. Stage 5 understanding v2 (entities/ambiguity/follow-ups) + semantic cache
474
+ 5. Stage 6 graph projection + listwise cascade + research mode
475
+ 6. Stage 7 quote anchors + table rendering + sample-and-verify
476
+ 7. Stage 9 metric battery + dashboards + gate
477
+ Each step ships with its own probes; no step requires the previous step's
478
+ deploy to be simultaneous (schemas are additive).
479
+
480
+
481
+ ---
482
+
483
+ ## Metanorma as semantic publisher — the forward play
484
+
485
+ What the toolchain already supports changes the upstream ask
486
+ (metanorma/metanorma#592) from "please emit AI-friendly output" to
487
+ "please PACKAGE what you already have":
488
+
489
+ 1. **Model serializations are done** — metanorma-document (lutaml-model)
490
+ round-trips semantic XML/YAML/JSON. The missing piece is only a
491
+ curated *node projection* flavor (the schema in this document) so AI
492
+ consumers don't each re-invent their own walk of the model.
493
+ 2. **Requirements/conformance as data** — ModSpec projection makes every
494
+ "shall" statement addressable (`/req/<class>/<name>`), with conformance
495
+ tests linked. For standards bodies this turns compliance questions
496
+ ("what does R 60-1 require? how is it verified?") from prose retrieval
497
+ into OBJECT retrieval — the single biggest structural win available to
498
+ a standards-domain RAG.
499
+ 3. **Glossarist + Relaton exports as first-class ingestion lanes** —
500
+ terminology and citation graphs ship WITH the document instead of
501
+ being joined from downstream repos.
502
+ 4. **The bundle ask:** one document source → human renderings (HTML/PDF)
503
+ AND machine serializations (node projection JSON, ModSpec requirement
504
+ export, Glossarist/Relaton datasets). Metanorma becomes the semantic
505
+ publisher for the AI ecosystem, exactly as it is the rendering
506
+ publisher today.
507
+
508
+ Our side: ingest v2 (roadmap Phase 2) implements the projection adapter
509
+ and the requirement/conformance retrieval lane against this contract.
@@ -0,0 +1,27 @@
1
+ # F1 — Verdict as data (conformance-as-a-service)
2
+
3
+ **Powering objects:** `entities/workflow.yaml` Verdict data class ("the
4
+ canonical verdict chain"), `specification/constraints.yaml` (OCL check +
5
+ violation_meaning + on_violation), `entities/examination-reports.yaml`,
6
+ `specification/conformance/*`.
7
+
8
+ **The win:** conformance questions stop being answered with prose and
9
+ start being answered with a **VERDICT BLOCK** — an answer-contract
10
+ artifact like the table block, resolved from the model by EXECUTION:
11
+
12
+ ```json
13
+ {"type": "verdict", "unit_id": "u:con-dead-load-max-geometry",
14
+ "payload": {"verdict": "fail", "on_violation": "invalid",
15
+ "meaning": "D_max lies outside [0.9·E_max, E_max] — … the type
16
+ evaluation of this load cell is void.",
17
+ "check": "ocl{d_max >= 0.9*e_max and d_max <= e_max}",
18
+ "source": "urn:oiml:pub:r:60-1:2021#clause-3.6"}}
19
+ ```
20
+
21
+ The model's own violation_meaning IS the explanation — the standard
22
+ explains itself in its recorded words. **Why documents can't follow:**
23
+ a verdict requires executing semantics; prose can only be quoted at.
24
+
25
+ **Build:** verdict evaluation in the D lane's serving path (constraints
26
+ interpreter over provided parameters), the verdict block type (contract
27
+ v2 already resolves any typed payload — only the UI card is new).
@@ -0,0 +1,19 @@
1
+ # F10 — Normative notes as overrides
2
+
3
+ **Powering objects:** notes.yaml — first-class NOTE/EXAMPLE constructs
4
+ with ids, e.g. note_creep_plc_always_0_7: "The MPE for creep shall
5
+ always be determined using p_LC = 0.7 regardless of any value declared
6
+ by the manufacturer."
7
+
8
+ **The win:** notes stop being decoration and become RULES. In document
9
+ lanes that sentence is prose a retrieval window may or may not include,
10
+ and a generation may or may not honor under a user's "but my p_LC is
11
+ 0.9" prompt. In the model lane the note is an OVERRIDE the computation
12
+ APPLIES: any F4 instance evaluation with declared p_LC ≠ 0.7 gets
13
+ normalized before lookupMPE runs — the assistant cannot be talked out
14
+ of the rule because the rule executes.
15
+
16
+ **Ladder form (L1 example already uses it):** the note probe extracts
17
+ the sentence; the FRONTIER probe asks the adversarial variant — "my
18
+ datasheet says 0.9, use that" — where D answers "no: p_LC = 0.7
19
+ always" and the computation proves it.
@@ -0,0 +1,17 @@
1
+ # F11 — Composition-aware answers
2
+
3
+ **Powering objects:** `uses: [iso-iec-17000, iso-iec-17065]` package
4
+ composition + layer-overlay semantics (evaluation/processes.yaml is a
5
+ PARTIAL overlay — "shared structure defined once in the core layer;
6
+ REC WINS per layer-composer semantics; identity-keyed union").
7
+
8
+ **The win:** answers that span the composition carry LAYER PROVENANCE:
9
+ *"what does the certification process require?"* → "the CORE layer
10
+ defines the process tree; R 60 overlays validate_provision (its
11
+ metrological URNs) and its testing gateways" — the answer knows which
12
+ stratum each statement comes from. Co-located documents (a shelf of
13
+ PDFs) invite conflation; composed models make provenance structural.
14
+
15
+ **Audit form:** *"where does R 60 deviate from the core process?"* →
16
+ the overlay diff — a question about the COMPOSITION ITSELF, meaningless
17
+ for document corpora.
@@ -0,0 +1,17 @@
1
+ # F12 — The machine passport (r60-to-dpp.prm)
2
+
3
+ **Powering objects:** evaluation/r60-to-dpp.prm — a composed artifact
4
+ that projects the model toward a Digital Product Passport — plus
5
+ certificate-template.yaml and the verdict chain.
6
+
7
+ **The win:** answers become EXPORTS. A conformity answer in the D lane
8
+ can terminate in a structured `.prm`/DPP payload — the clause chain,
9
+ the verdicts, the parameters — that downstream systems (customs,
10
+ procurement, market surveillance) consume as DATA. The Q&A service
11
+ stops being the end of the pipeline and becomes an oracle other
12
+ machines query.
13
+
14
+ **Answer-contract form:** the ultimate block type — alongside table,
15
+ formula, verdict — a `passport` block: machine-readable, signed by
16
+ provenance, rendered in the UI as a downloadable artifact. Prose is for
17
+ humans; the passport is for systems.
@@ -0,0 +1,20 @@
1
+ # F2 — Counterfactual simulation
2
+
3
+ **Powering objects:** constraints (OCL), formulas-as-operations
4
+ (table_lookup with params), calculations (typed IO).
5
+
6
+ **The win:** hypotheticals become answerable: *"what if D_max were
7
+ 0.8·E_max?"* → run the constraint set on the hypothetical profile →
8
+ F1's verdict. *"which accuracy class should I pick so that n_LC = 3 000
9
+ clears the minimum?"* → INVERT lookupMPE's tier table (class C needs
10
+ n_LC ≥ … ; solve). Documents state facts about the world as written;
11
+ the model evaluates worlds as they could be.
12
+
13
+ **Example interaction:**
14
+ > Q: "My E_max is 30 000 v and I want to test to D_max = 26 000 v."
15
+ > A: VERDICT fail — d_max (26 000) < 0.9·E_max (27 000) … evaluation
16
+ > void [R 60-1 §3.6 chain]. The minimum acceptable D_max is 27 000 v.
17
+
18
+ **Why documents can't follow:** nothing to execute; an LLM on prose
19
+ would HALLUCINATE a rule application. Containment rule: the
20
+ counterfactual probe must not be answerable by printed text.
@@ -0,0 +1,21 @@
1
+ # F3 — Exhaustiveness and provable absence
2
+
3
+ **Powering objects:** the CLOSED WORLD — aspects.yaml's audit line
4
+ ("audit of all 60 requirements + 62 tests"), requirements/ and
5
+ conformance/ as enumerated registries, forAll semantics in constraints.
6
+
7
+ **The win — two directions:**
8
+ 1. **Complete listing:** *"list EVERY requirement that binds the
9
+ load cell's markings"* → the aspects registry (markings ← R 60-1
10
+ 6.2.1/6.2.4, examined by /conf/examinations/inscriptions…). The
11
+ answer is EXHAUSTIVE by construction, with the audit count as
12
+ metadata. Retrieval lanes return a top-k subset and cannot promise
13
+ completeness.
14
+ 2. **Provable NO:** *"does R 60 constrain the load cell's packaging
15
+ for shipping?"* → no requirement, no aspect, no test binds
16
+ "packaging" → the answer is NO with the enumeration as proof. A
17
+ document lane must refuse ("not found") — absence of evidence.
18
+ The model lane PROVES absence of the requirement.
19
+
20
+ **Why it matters:** "no" answers are the scariest in conformity
21
+ assessment; provable absence turns them into the safest.
@@ -0,0 +1,18 @@
1
+ # F4 — Instance-grounded, parameterized answers
2
+
3
+ **Powering objects:** attributes with `scope: family|instance` and
4
+ `origin`, entities/load-cell-instance.yaml (serial, type ref),
5
+ sample-data.yaml, the unit register for input normalization.
6
+
7
+ **The win:** the corpus answers FOR YOUR INSTRUMENT. The MPE table
8
+ stops being a table to read and becomes a FUNCTION to evaluate:
9
+
10
+ > Q (member): "Model LC-3000, Max 30 kg, e 5 g, class C — my MPE at
11
+ > 12 kg load?" → lookupMPE(load=2400 v, class C) → ±0.35·v … with the
12
+ > full clause chain and the unit register normalizing whatever units
13
+ > the user typed (kgf, kN, lb) before computing (quantity-kind
14
+ coherence, never string matching).
15
+
16
+ **Tiering:** instance data is member/internal by default; sample data
17
+ is public. **Why documents can't follow:** there are no instances in a
18
+ document — only the general case. This is the personalization rung.