@konneal/engine 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (285) hide show
  1. package/LICENSE +29 -0
  2. package/README.md +13 -0
  3. package/dist/admin.d.ts +26 -0
  4. package/dist/ai.d.ts +6 -0
  5. package/dist/anchors.d.ts +6 -0
  6. package/dist/answercache.d.ts +22 -0
  7. package/dist/ask.d.ts +5 -0
  8. package/dist/auth.d.ts +12 -0
  9. package/dist/bubble.d.ts +14 -0
  10. package/dist/chunk-LLWPT2XV.js +49 -0
  11. package/dist/chunk-MB74PTRM.js +114 -0
  12. package/dist/chunk-WOGQM7DJ.js +197 -0
  13. package/dist/chunk-WWNCWKKC.js +42 -0
  14. package/dist/completion.d.ts +5 -0
  15. package/dist/config.d.ts +154 -0
  16. package/dist/config.js +37 -0
  17. package/dist/context.d.ts +115 -0
  18. package/dist/conversations.d.ts +5 -0
  19. package/dist/drafts.d.ts +129 -0
  20. package/dist/env.d.ts +57 -0
  21. package/dist/faithfulness.d.ts +5 -0
  22. package/dist/grader.d.ts +3 -0
  23. package/dist/graph.d.ts +13 -0
  24. package/dist/hybrid.d.ts +7 -0
  25. package/dist/index.d.ts +9 -0
  26. package/dist/index.js +5373 -0
  27. package/dist/internal_gateway.d.ts +14 -0
  28. package/dist/lexical.d.ts +7 -0
  29. package/dist/livedata.d.ts +77 -0
  30. package/dist/memories.d.ts +10 -0
  31. package/dist/modelplane.d.ts +61 -0
  32. package/dist/oidc.d.ts +73 -0
  33. package/dist/pipeline.d.ts +57 -0
  34. package/dist/profile.d.ts +2 -0
  35. package/dist/profile.gen.d.ts +70 -0
  36. package/dist/profile.js +8 -0
  37. package/dist/projects.d.ts +8 -0
  38. package/dist/prompts/conversational.md +8 -0
  39. package/dist/prompts/enrichment.md +3 -0
  40. package/dist/prompts/faithfulness.md +1 -0
  41. package/dist/prompts/grader.md +5 -0
  42. package/dist/prompts/listwise.md +3 -0
  43. package/dist/prompts/precision.md +1 -0
  44. package/dist/prompts/reflect.md +1 -0
  45. package/dist/prompts/relevancy.md +1 -0
  46. package/dist/prompts/research.md +10 -0
  47. package/dist/prompts/section-summary.md +5 -0
  48. package/dist/prompts/summarize.md +1 -0
  49. package/dist/prompts/system.md +18 -0
  50. package/dist/prompts/understanding.md +17 -0
  51. package/dist/quota.d.ts +13 -0
  52. package/dist/reflect.d.ts +5 -0
  53. package/dist/refs.d.ts +40 -0
  54. package/dist/refusal.d.ts +9 -0
  55. package/dist/refusal.js +9 -0
  56. package/dist/requestScope.d.ts +26 -0
  57. package/dist/requestScope.js +10 -0
  58. package/dist/research.d.ts +8 -0
  59. package/dist/search.d.ts +4 -0
  60. package/dist/selfquery.d.ts +7 -0
  61. package/dist/session.d.ts +1 -0
  62. package/dist/share.d.ts +2 -0
  63. package/dist/structural.d.ts +27 -0
  64. package/dist/tablecontext.d.ts +11 -0
  65. package/dist/understand.d.ts +11 -0
  66. package/dist/understandContract.d.ts +29 -0
  67. package/dist/verdict.d.ts +24 -0
  68. package/docs/API.md +451 -0
  69. package/docs/ARCHITECTURE.md +302 -0
  70. package/docs/AUDIT-2026-08-24.md +71 -0
  71. package/docs/CONTRIBUTOR-AUDIT-2026-08-25.md +147 -0
  72. package/docs/INGEST-ARCHITECTURE.md +158 -0
  73. package/docs/MCP.md +92 -0
  74. package/docs/METANORMA-AI-SERIALIZATION.md +247 -0
  75. package/docs/MKO-EXPORT-PIPELINE.md +147 -0
  76. package/docs/REDESIGN-NORMATIVE-RAG-ETSI.md +485 -0
  77. package/docs/RESEARCH-SOTA-2026.md +243 -0
  78. package/docs/ROADMAP-SOTA.md +130 -0
  79. package/docs/SOTA-STAGE-SPECS.md +509 -0
  80. package/docs/annealment/F1-verdict.md +27 -0
  81. package/docs/annealment/F10-notes.md +19 -0
  82. package/docs/annealment/F11-composition.md +17 -0
  83. package/docs/annealment/F12-passport.md +17 -0
  84. package/docs/annealment/F2-counterfactual.md +20 -0
  85. package/docs/annealment/F3-absence.md +21 -0
  86. package/docs/annealment/F4-instance.md +18 -0
  87. package/docs/annealment/F5-workflow.md +21 -0
  88. package/docs/annealment/F6-impact.md +21 -0
  89. package/docs/annealment/F7-editions.md +18 -0
  90. package/docs/annealment/F8-selfverify.md +19 -0
  91. package/docs/annealment/F9-projection-qa.md +17 -0
  92. package/docs/annealment/L0-locate.md +19 -0
  93. package/docs/annealment/L1-extract.md +18 -0
  94. package/docs/annealment/L2-nomenclature.md +22 -0
  95. package/docs/annealment/L3-geometry.md +23 -0
  96. package/docs/annealment/L4-composition.md +21 -0
  97. package/docs/annealment/L5-cross-standard.md +20 -0
  98. package/docs/annealment/L6-diachrony.md +21 -0
  99. package/docs/annealment/L7-perception.md +20 -0
  100. package/docs/annealment/L8-computation.md +22 -0
  101. package/docs/annealment/L9-instance-process.md +23 -0
  102. package/docs/annealment/README.md +10 -0
  103. package/docs/guidelines-metanorma-ai-programme.md +279 -0
  104. package/docs/identity-onboarding-rag.md +65 -0
  105. package/docs/identity-service.md +219 -0
  106. package/docs/knowledge-annealment.md +273 -0
  107. package/docs/konneal-extraction-plan.md +481 -0
  108. package/docs/metanorma-for-ai.md +270 -0
  109. package/docs/mirror-plan.md +36 -0
  110. package/docs/multi-sdo-architecture.md +191 -0
  111. package/docs/paper-annealment-comparison.md +259 -0
  112. package/docs/paper-assets/architecture.svg +94 -0
  113. package/docs/paper-assets/contract-v2.svg +94 -0
  114. package/docs/paper-assets/mko-ingest.svg +91 -0
  115. package/docs/paper-oiml-bulletin.md +402 -0
  116. package/docs/paper-oiml-bulletin.mdx +419 -0
  117. package/docs/product-branding-options.md +172 -0
  118. package/docs/projects-design.md +88 -0
  119. package/docs/sota-mechanisms.md +184 -0
  120. package/docs/spec-api.md +77 -0
  121. package/docs/spec-pipeline.md +126 -0
  122. package/docs/vector-adapter.md +88 -0
  123. package/package.json +70 -0
  124. package/profile/corpora.yaml +5 -0
  125. package/profile/datasets.yaml +14 -0
  126. package/profile/prompts.yaml +5 -0
  127. package/profile/publisher.yaml +17 -0
  128. package/profile/retrieval.yaml +1 -0
  129. package/profile/sources.yaml +5 -0
  130. package/profile/ui.yaml +7 -0
  131. package/scripts/gen_profile.mjs +33 -0
  132. package/workers/shared/ai.ts +21 -0
  133. package/workers/shared/auth.ts +16 -0
  134. package/workers/shared/chunk.ts +108 -0
  135. package/workers/shared/oidc.ts +312 -0
  136. package/workers/shared/router.ts +45 -0
  137. package/workers/shared/session.ts +104 -0
  138. package/workers/worker_internal/src/index.ts +157 -0
  139. package/workers/worker_internal/tsconfig.json +15 -0
  140. package/workers/worker_internal/wrangler.toml +32 -0
  141. package/workers/worker_mcp/src/index.ts +175 -0
  142. package/workers/worker_mcp/tsconfig.json +13 -0
  143. package/workers/worker_mcp/wrangler.toml +18 -0
  144. package/workers/worker_public/migrations/0002_conversations.sql +22 -0
  145. package/workers/worker_public/migrations/0003_shared_conversations.sql +9 -0
  146. package/workers/worker_public/migrations/0004_graph.sql +16 -0
  147. package/workers/worker_public/migrations/0005_documents.sql +19 -0
  148. package/workers/worker_public/migrations/0006_conversation_entities.sql +11 -0
  149. package/workers/worker_public/migrations/0007_chunks_fts.sql +43 -0
  150. package/workers/worker_public/migrations/0008_unit_payloads.sql +15 -0
  151. package/workers/worker_public/migrations/0009_chunks_unit.sql +8 -0
  152. package/workers/worker_public/migrations/0009_message_context.sql +7 -0
  153. package/workers/worker_public/migrations/0010_model_nodes.sql +33 -0
  154. package/workers/worker_public/migrations/0011_schema_union.sql +38 -0
  155. package/workers/worker_public/migrations/0012_memories.sql +15 -0
  156. package/workers/worker_public/migrations/0013_projects.sql +21 -0
  157. package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.d.ts +16306 -0
  158. package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.ts +16261 -0
  159. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.d.ts +16373 -0
  160. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.ts +16328 -0
  161. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.d.ts +16382 -0
  162. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.ts +16337 -0
  163. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.d.ts +16383 -0
  164. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.ts +16338 -0
  165. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.d.ts +16403 -0
  166. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.ts +16358 -0
  167. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.d.ts +16408 -0
  168. package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.ts +16363 -0
  169. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.d.ts +16414 -0
  170. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.ts +16369 -0
  171. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.d.ts +16414 -0
  172. package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.ts +16369 -0
  173. package/workers/worker_public/node_modules/@cloudflare/workers-types/README.md +135 -0
  174. package/workers/worker_public/node_modules/@cloudflare/workers-types/entrypoints.svg +53 -0
  175. package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.d.ts +17095 -0
  176. package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.ts +17050 -0
  177. package/workers/worker_public/node_modules/@cloudflare/workers-types/index.d.ts +16306 -0
  178. package/workers/worker_public/node_modules/@cloudflare/workers-types/index.ts +16261 -0
  179. package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.d.ts +16447 -0
  180. package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.ts +16402 -0
  181. package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.d.ts +16306 -0
  182. package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.ts +16261 -0
  183. package/workers/worker_public/node_modules/@cloudflare/workers-types/package.json +11 -0
  184. package/workers/worker_public/package.json +13 -0
  185. package/workers/worker_public/prompts/conversational.md +8 -0
  186. package/workers/worker_public/prompts/enrichment.md +3 -0
  187. package/workers/worker_public/prompts/faithfulness.md +1 -0
  188. package/workers/worker_public/prompts/grader.md +5 -0
  189. package/workers/worker_public/prompts/listwise.md +3 -0
  190. package/workers/worker_public/prompts/precision.md +1 -0
  191. package/workers/worker_public/prompts/reflect.md +1 -0
  192. package/workers/worker_public/prompts/relevancy.md +1 -0
  193. package/workers/worker_public/prompts/research.md +10 -0
  194. package/workers/worker_public/prompts/section-summary.md +5 -0
  195. package/workers/worker_public/prompts/summarize.md +1 -0
  196. package/workers/worker_public/prompts/system.md +18 -0
  197. package/workers/worker_public/prompts/understanding.md +17 -0
  198. package/workers/worker_public/public/app.js +166 -0
  199. package/workers/worker_public/public/index.html +48 -0
  200. package/workers/worker_public/public/style.css +147 -0
  201. package/workers/worker_public/schema.sql +248 -0
  202. package/workers/worker_public/src/admin.ts +358 -0
  203. package/workers/worker_public/src/ai.ts +71 -0
  204. package/workers/worker_public/src/anchors.ts +41 -0
  205. package/workers/worker_public/src/answercache.ts +72 -0
  206. package/workers/worker_public/src/ask.ts +1094 -0
  207. package/workers/worker_public/src/auth.ts +252 -0
  208. package/workers/worker_public/src/bubble.ts +111 -0
  209. package/workers/worker_public/src/completion.ts +75 -0
  210. package/workers/worker_public/src/config.ts +238 -0
  211. package/workers/worker_public/src/context.ts +238 -0
  212. package/workers/worker_public/src/conversations.ts +162 -0
  213. package/workers/worker_public/src/drafts.ts +497 -0
  214. package/workers/worker_public/src/env.ts +90 -0
  215. package/workers/worker_public/src/faithfulness.ts +63 -0
  216. package/workers/worker_public/src/grader.ts +89 -0
  217. package/workers/worker_public/src/graph.ts +63 -0
  218. package/workers/worker_public/src/hybrid.ts +77 -0
  219. package/workers/worker_public/src/index.ts +441 -0
  220. package/workers/worker_public/src/internal_gateway.ts +41 -0
  221. package/workers/worker_public/src/lexical.ts +86 -0
  222. package/workers/worker_public/src/lib/hit.ts +4 -0
  223. package/workers/worker_public/src/lib/http.ts +83 -0
  224. package/workers/worker_public/src/lib/router.ts +4 -0
  225. package/workers/worker_public/src/livedata.ts +334 -0
  226. package/workers/worker_public/src/memories.ts +81 -0
  227. package/workers/worker_public/src/modelplane.ts +213 -0
  228. package/workers/worker_public/src/oidc.ts +333 -0
  229. package/workers/worker_public/src/pipeline.ts +377 -0
  230. package/workers/worker_public/src/ports/blobs.ts +7 -0
  231. package/workers/worker_public/src/ports/cloudflare/adapters.ts +177 -0
  232. package/workers/worker_public/src/ports/kv.ts +8 -0
  233. package/workers/worker_public/src/ports/model.ts +28 -0
  234. package/workers/worker_public/src/ports/runtime.ts +13 -0
  235. package/workers/worker_public/src/ports/store.ts +20 -0
  236. package/workers/worker_public/src/ports/vector.ts +26 -0
  237. package/workers/worker_public/src/profile.gen.ts +101 -0
  238. package/workers/worker_public/src/profile.ts +16 -0
  239. package/workers/worker_public/src/projects.ts +108 -0
  240. package/workers/worker_public/src/prompts.d.ts +6 -0
  241. package/workers/worker_public/src/quota.ts +54 -0
  242. package/workers/worker_public/src/reflect.ts +67 -0
  243. package/workers/worker_public/src/refs.ts +107 -0
  244. package/workers/worker_public/src/refusal.ts +65 -0
  245. package/workers/worker_public/src/requestScope.ts +71 -0
  246. package/workers/worker_public/src/research.ts +126 -0
  247. package/workers/worker_public/src/search.ts +56 -0
  248. package/workers/worker_public/src/selfquery.ts +25 -0
  249. package/workers/worker_public/src/session.ts +4 -0
  250. package/workers/worker_public/src/share.ts +53 -0
  251. package/workers/worker_public/src/stages/conceptGraph.ts +68 -0
  252. package/workers/worker_public/src/stages/conceptSteer.ts +39 -0
  253. package/workers/worker_public/src/stages/corpusScope.ts +25 -0
  254. package/workers/worker_public/src/stages/dedup.ts +10 -0
  255. package/workers/worker_public/src/stages/dense.ts +73 -0
  256. package/workers/worker_public/src/stages/diversity.ts +33 -0
  257. package/workers/worker_public/src/stages/editionCover.ts +63 -0
  258. package/workers/worker_public/src/stages/editionSteer.ts +88 -0
  259. package/workers/worker_public/src/stages/familyBoost.ts +22 -0
  260. package/workers/worker_public/src/stages/federate.ts +22 -0
  261. package/workers/worker_public/src/stages/glossary.ts +65 -0
  262. package/workers/worker_public/src/stages/graphLane.ts +31 -0
  263. package/workers/worker_public/src/stages/hyde.ts +29 -0
  264. package/workers/worker_public/src/stages/index.ts +69 -0
  265. package/workers/worker_public/src/stages/lexicalUnion.ts +21 -0
  266. package/workers/worker_public/src/stages/multiQuery.ts +57 -0
  267. package/workers/worker_public/src/stages/overviewDemote.ts +14 -0
  268. package/workers/worker_public/src/stages/poolOpen.ts +10 -0
  269. package/workers/worker_public/src/stages/propagate.ts +15 -0
  270. package/workers/worker_public/src/stages/rerank.ts +47 -0
  271. package/workers/worker_public/src/stages/seal.ts +16 -0
  272. package/workers/worker_public/src/stages/sectionDescent.ts +61 -0
  273. package/workers/worker_public/src/stages/stdRefNudge.ts +35 -0
  274. package/workers/worker_public/src/stages/subQuery.ts +42 -0
  275. package/workers/worker_public/src/stages/termNudge.ts +24 -0
  276. package/workers/worker_public/src/stages/typedPin.ts +131 -0
  277. package/workers/worker_public/src/stages/types.ts +112 -0
  278. package/workers/worker_public/src/stages/windowFloor.ts +23 -0
  279. package/workers/worker_public/src/structural.ts +171 -0
  280. package/workers/worker_public/src/tablecontext.ts +41 -0
  281. package/workers/worker_public/src/understand.ts +72 -0
  282. package/workers/worker_public/src/understandContract.ts +67 -0
  283. package/workers/worker_public/src/verdict.ts +255 -0
  284. package/workers/worker_public/tsconfig.json +18 -0
  285. package/workers/worker_public/wrangler.toml +104 -0
@@ -0,0 +1,419 @@
1
+ ---
2
+ title: "Grounding legal metrology: a retrieval-augmented question-answering service over the OIML publications corpus"
3
+ subtitle: "Draft article for the OIML Bulletin"
4
+ date: 2026-08-30
5
+ figures: paper-assets
6
+ ---
7
+
8
+ export const Figure = ({ id, src, caption, width = 760 }) => (
9
+ <figure id={id} style={{ margin: "24px auto", textAlign: "center" }}>
10
+ <img src={src} alt={caption} width={width} style={{ maxWidth: "100%", height: "auto" }} />
11
+ <figcaption style={{ fontSize: "0.85em", color: "#475569", marginTop: 8 }}>{caption}</figcaption>
12
+ </figure>
13
+ );
14
+
15
+ *Draft article for the OIML Bulletin — 2026-08-30. Style: Bulletin
16
+ technical article (MS Word single-column on submission; this is the
17
+ authoring source). Numbers are production-measured unless noted.*
18
+
19
+ ---
20
+
21
+ ## Abstract
22
+
23
+ The OIML publishes close to 900 documents — Recommendations, Basic
24
+ publications, Guides, Documents and Expert reports — that together form
25
+ the reference library of legal metrology. Finding what they say has
26
+ required knowing the corpus before querying it. We describe a
27
+ question-answering service, **OIML SMART AI** (ai.oimlsmart.org), that
28
+ answers natural-language questions from the indexed publications with
29
+ clause-level citations, verbatim quote anchors for normative values, and
30
+ typed renderings of the tables and equations the answers depend on — and
31
+ it reads figures and user-supplied photographs directly, from pixels. The
32
+ service measures every component it ships: a golden question set,
33
+ retrieval metrics in the tradition of the information-retrieval
34
+ literature, and judged faithfulness. We report the measured design
35
+ decisions, the corpus-structure findings that shaped them, and the open
36
+ path: producer-native ingestion of the Metanorma document model so that
37
+ any Metanorma-authored corpus can be indexed the same way.
38
+
39
+ ---
40
+
41
+ ## 1. Introduction
42
+
43
+ Legal metrology's knowledge lives in a large, cross-referential,
44
+ multi-edition corpus. A typical working question — *"what is the maximum
45
+ permissible error for a class III nonautomatic weighing instrument?"* —
46
+ requires knowing which Recommendation applies (OIML R 76-1), which
47
+ edition is current, which clause holds the table, and how to read the
48
+ table's rows against the instrument's verification scale intervals.
49
+ Every one of those steps assumes corpus knowledge the questioner may not
50
+ have.
51
+
52
+ Question answering over such corpora has been transformed by
53
+ retrieval-augmented generation (RAG): a language model is given passages
54
+ retrieved from a trusted corpus and may use only those. The engineering
55
+ challenge is no longer whether a model can write fluent prose — it can —
56
+ but whether the *right passages* are found, whether the answer stays
57
+ *faithful* to them, and whether the reader can *verify* the result. Our
58
+ service is built around those three properties.
59
+
60
+ This article describes the architecture as deployed, the measurements
61
+ that shaped it, and the corpus findings — including errors in the
62
+ bibliographic record itself — that surfaced because the system keeps
63
+ score. It closes with the producer-native ingestion path (Metanorma
64
+ Knowledge Objects) and what it makes possible for other standards bodies.
65
+
66
+ ## 2. What the service does
67
+
68
+ A user asks a question in any language. The service:
69
+
70
+ 1. **Understands the question** with a small language model — language,
71
+ named publication, edition, defined terms, complexity. There are no
72
+ keyword rules; the same model makes these judgements for every
73
+ question.
74
+ 2. **Retrieves in parallel** — three lanes leave the question at the
75
+ same instant: the vector index is queried with the question's
76
+ embedding while understanding is still running (the dense lane's
77
+ results are reused unchanged when understanding adds nothing), a
78
+ full-corpus keyword index catches exact identifiers, part numbers
79
+ and defined terms that vector similarity alone misses, and
80
+ understanding's refinements — a named publication, an edition —
81
+ arrive as metadata filters on the dense query.
82
+ 3. **Ranks** the fused candidates with a cross-encoder; complex
83
+ questions get a second, stronger re-ranking pass.
84
+ 4. **Answers from the retrieved passages only**, under a contract that
85
+ requires inline citations (`[OIML R 60-1:2021 §4.4.2]`) and verbatim
86
+ quote anchors for normative values.
87
+ 5. **Verifies** the answer after generation: a deterministic check that
88
+ every quoted phrase exists in a cited passage; a separate judge that
89
+ scores faithfulness; a correction-and-retry when either fails. An
90
+ answer that cannot be verified is never written to the answer cache.
91
+
92
+ <Figure
93
+ id="fig-architecture"
94
+ src="paper-assets/architecture.svg"
95
+ caption="Figure 1 — A question's path through the service. Three retrieval lanes run concurrently; every answer passes mechanical verification before it is served or cached. The edition registry derives publication status from supersession links, never from a status field."
96
+ />
97
+
98
+ When the question names a publication, answers are steered to the
99
+ current edition by a registry derived from the bibliographic record's
100
+ supersession links (§5). When the answer depends on a table or
101
+ equation, the interface renders the *producer's* typed object — the
102
+ actual rows and columns — never the model's re-typing of it (§6).
103
+
104
+ If the corpus does not contain the answer, the system says so in one
105
+ canonical sentence and redirects; refusals are never cached, because a
106
+ refusal is a property of the moment, not of the question.
107
+
108
+ ## 3. Corpus and structure
109
+
110
+ The indexed corpus comprises the English editions of the OIML
111
+ publications (~900 documents, ~9.3 million words). Documents are chunked
112
+ along clause boundaries — never arbitrary token windows — and each chunk
113
+ carries provenance: publication identifier, edition, language, clause
114
+ anchor, and bibliographic status.
115
+
116
+ Three corpus-level findings shaped the design:
117
+
118
+ **The bibliographic status field disagrees with its own record.** In 58
119
+ of 224 publication families, the machine-readable status field claims a
120
+ document is current while the same record's supersession links say
121
+ otherwise. The service therefore *derives* status from the links and
122
+ ignores the field. More consequentially, 36 of 224 families have no
123
+ recorded successor link at all, so the current edition cannot be derived
124
+ for them; the registry surfaces these gaps rather than guessing. Both
125
+ numbers are the worklist for upstream bibliographic corrections — the
126
+ kind that serving surfaces only because it keeps score.
127
+
128
+ **Identifiers are identity, not display.** A cross-referenced audit of
129
+ the dirty corpus found documents whose machine-readable identifier was
130
+ a placeholder ("OIML D 0:0000", "OIML D X") or polluted with language
131
+ markers — around 200 records — making them unreachable by
132
+ doc-scoped retrieval. Identifiers are now derived from slug identity
133
+ with placeholders never winning, and the index is reconciled against
134
+ the canonical chunk set: the first census measured 49,553 live vectors
135
+ against 31,512 canonical — 18,041 strays from every prior re-chunking,
136
+ deleted in place (upserts never delete). A retrieval service that never
137
+ enumerates its own index serves ghosts.
138
+
139
+ **Tables are where the normative values live.** Maximum permissible
140
+ errors, accuracy-class limits, verification interval bounds — the values
141
+ practitioners ask for — are tabular. Flattening tables into prose loses
142
+ the row/column geometry that makes them answerable. The index now treats
143
+ tables as atomic typed objects (§6).
144
+
145
+ **Structure is a retrieval signal.** Standards are hierarchically
146
+ organized with extensive cross-references; the system indexes the
147
+ section hierarchy and the citation graph as a queryable graph (7,128
148
+ nodes, 6,486 edges) alongside the text, so a defined term can be
149
+ resolved to the documents that define it even when the question's
150
+ vocabulary does not match the corpus's.
151
+
152
+ ## 4. Retrieval: what the measurements showed
153
+
154
+ We evaluate retrieval with a golden question set (now 29 questions
155
+ including eight table-value questions with witness rows) using
156
+ recall@5, average precision@5, and mean reciprocal rank@5 — the protocol
157
+ of the standards-retrieval literature. Single runs proved noisy (±5
158
+ percentage points); all reported numbers are means over three runs on a
159
+ quiet account.
160
+
161
+ The single largest measured improvement came from making keyword
162
+ retrieval a **full-corpus first stage** rather than a re-scoring of the
163
+ vector results: recall@5 rose from 86% to 95%. The reason is structural
164
+ — standards vocabulary is exact ("n_LC", "R 60-3", "creep"), and dense
165
+ embeddings alone miss exact identifiers that a lexical scan of the whole
166
+ corpus recovers. Fusing both rankings (reciprocal rank fusion) lifted
167
+ precision and MRR without costing recall.
168
+
169
+ The second improvement came from **contextual enrichment**: a one-time
170
+ pass that prepends a short model-written preamble to every chunk ("this
171
+ clause defines the accuracy-class limits for load cells") before
172
+ embedding and indexing. This is the contextual-retrieval recipe
173
+ validated industry-wide; our corpus measurement confirmed it, and the
174
+ one-time cost (~$0.0017 per chunk) is amortized over every future
175
+ answer.
176
+
177
+ With the typed-table lane live (§6), retrieval lanes running
178
+ concurrently with query understanding, edition steering active (§5) and
179
+ multimodal generation on (§6), the current baseline is:
180
+
181
+ | Metric | Value |
182
+ |---|---|
183
+ | Recall@5 (mean of 3) | **94.3%** (range 93.1–96.6) |
184
+ | Average precision@5 | 0.82 |
185
+ | MRR@5 | 0.87 |
186
+ | End-to-end golden cases | 14/14 |
187
+
188
+ These numbers are on our own golden set, not on the benchmarks of the
189
+ works we build on — the honest comparison is per-technique, on the same
190
+ corpus, before and after. So read: the full-corpus lexical stage is the
191
+ ETSI study's recipe (ref 1) — adopting it lifted recall@5 here from 86%
192
+ to 95%; the contextual preambles are Anthropic's contextual retrieval
193
+ (ref 3) — confirmed on this corpus at ~$0.0017 per chunk one-time; the
194
+ clause-boundary, structure-preserving chunking follows the same
195
+ structure-aware line as refs 2 and 7, which our typed-unit pin extends
196
+ from "chunk better" to "guarantee the answering object a slot." Two
197
+ elements of the deployed system have no counterpart in the cited work:
198
+ the symbolic-reference contract, under which table data never passes
199
+ through the model at all (refs 2's error-reduction approach still
200
+ re-generates tables; we removed the corruption channel), and the
201
+ mechanical post-generation verification of every answer (quote anchors,
202
+ unit references, retyped-table detection) with a faithfulness judge that
203
+ sees the passages actually used.
204
+
205
+ Two structural additions complete the retrieval story. First, the
206
+ corpus is a tree — every chunk carries its clause anchor — and the
207
+ serving path uses the tree: a hit's score blends its neighbouring
208
+ clauses' scores, evidence is presented to the answer model in document
209
+ reading order, and a ranked section summary resolves to its quotable
210
+ child clauses (adapted from FABLE/BEAR, ref 7, at none of its index-time
211
+ cost, because Metanorma documents arrive as trees rather than needing
212
+ one inferred). Second, everyday words rarely match defined terms — the
213
+ one gap that failed for every corpus representation equally — so a
214
+ terminology index of the corpus's defined concepts links a question's
215
+ phrasing to candidate terms, and the answer model adjudicates and leads
216
+ with the corpus's own term, quoting its definition ("my output keeps
217
+ drifting" resolves to *span stability*, with the verbatim definition and
218
+ its defining clause cited).
219
+
220
+ ## 5. Editions and trust
221
+
222
+ Which edition applies is as consequential as what it says. The service's
223
+ registry of publication families and editions is derived from
224
+ supersession links, not from status fields (§3). Citations carry the
225
+ edition and status of every source; superseded sources are marked.
226
+ Answers to questions that name a publication are steered to the current
227
+ edition; answers that must cite a superseded edition (the 2021 passages
228
+ do not reproduce a table the 2017 edition carries) say so explicitly.
229
+
230
+ Steering is measured, not assumed, and two mechanisms earned their place
231
+ by fixing observed failures. Superseded editions match archaic phrasing
232
+ strongly — a complex verification question once cited R 76-1:1992/1988
233
+ while the current 2006 edition was crowded out of the passage window —
234
+ so ranking now demotes older-edition chunks of a publication whenever a
235
+ newer edition of the same publication is present in the evidence
236
+ (cross-publication recency is deliberately untouched: a 1992 document
237
+ that is still current must not lose to unrelated 2024 ones). And because
238
+ query understanding can pin an edition for the wrong family, an edition
239
+ pin that the corpus itself fails to corroborate is dropped to the
240
+ document-level filter. After both fixes the probe cites R 76-1:2006 in
241
+ force, with one legitimate 1988 tail citation for a clause only that
242
+ edition carries; the window's measured average precision (0.85) and MRR
243
+ (0.88) are the best the service has recorded.
244
+
245
+ ## 6. Tables, equations, and figures as typed objects
246
+
247
+ The most consequential rendering decision: the language model never
248
+ re-types normative data. When an answer depends on a table, the model
249
+ emits a symbolic reference to the table's unit (`[[u:table-1]]`); the
250
+ server validates the reference against the passages actually used and
251
+ resolves it to the producer's typed payload — columns, rows, units. The
252
+ interface renders the exact object. Small verbatim values in prose are
253
+ permitted and mechanically verified against the source. The payload
254
+ never passes through the model's output, so it cannot be corrupted by
255
+ generation.
256
+
257
+ The same principle extends from rendering to judgment: the model never
258
+ judges what a machine can compute. Where the corpus's machine-readable
259
+ models carry checkable objects — a constraint with an OCL expression, a
260
+ requirement with a threshold limit — a question naming one gets it
261
+ EXECUTED: the question's values bind to the rule's own parameters, the
262
+ check evaluates deterministically, and the verdict (pass, or the
263
+ standard's own violation word, with the arithmetic shown, or void naming
264
+ exactly what is missing) is attached to the answer as computed data.
265
+ Two sibling capabilities complete the set: provable absence (a topic
266
+ that a standard does not constrain returns an enumeration certificate
267
+ over its entire model plane — N nodes enumerated, zero matches — rather
268
+ than a refusal) and answer verification (any answer checkable against
269
+ the corpus: verbatim-quote containment, object-reference resolution, a
270
+ judged faithfulness score labeled as judged).
271
+
272
+ <Figure
273
+ id="fig-contract"
274
+ src="paper-assets/contract-v2.svg"
275
+ caption="Figure 2 — The answer contract. The model writes prose with inline citations and symbolic references; three mechanical checks run after generation; the server resolves surviving references to the producer's payloads and the client renders them."
276
+ />
277
+
278
+ Figures are described by a vision-capable model that reads the actual
279
+ image pixels; equations are carried in their authoring-native formats
280
+ (AsciiMath and MathML) with a plain-language description. Both arrive as typed objects with the same validation.
281
+
282
+ Figures also flow the other way: when the retrieved passages contain
283
+ figure units, their actual images are attached to the answer model's
284
+ input, so descriptions and reasoning come from the drawing itself — the
285
+ model reads labels that exist only in the pixels (the pixel-label
286
+ probe — "what are the labeled example cases in the R 60-2 design-shapes
287
+ figure?" — is answered A, B and C from the drawing, stable across
288
+ repeated runs). Two measured invariants keep that true: assets must be
289
+ readable by vision pipelines (vector-sourced rasters that draw black on
290
+ transparent alpha flatten to a solid black rectangle inside a vision
291
+ model, whatever a browser shows — a corpus sweep detects and repairs
292
+ them), and images ride their own message in the generation call (long
293
+ passage text and image parts in one message triggers provider errors
294
+ that scale with payload). Users can likewise attach a photograph (an
295
+ instrument nameplate, a scale dial, a schematic) to their question; the
296
+ text still drives retrieval, and the image gives the model the visual
297
+ context, under the same citation contract.
298
+
299
+ ## 7. Models, cost, and sovereignty
300
+
301
+ All models are open-weight and served on a single cloud platform
302
+ (Cloudflare Workers AI); no corpus data is sent to proprietary model
303
+ providers. The answer model (GLM-5.3 Flash, natively multimodal) serves
304
+ all tiers at approximately $0.001 per answer. The full one-time corpus
305
+ preparation — contextual enrichment of 42,000 chunks — cost
306
+ approximately $75. Daily serving at current traffic is under $5 per
307
+ month including infrastructure. The model policy is deliberately
308
+ cost-first on the serving path (every question pays it) and quality-first
309
+ on one-time work (enrichment, captioning, evaluation), where quality
310
+ persists into every future answer.
311
+
312
+ One lesson generalizes: read the model card before wiring a model. Three
313
+ separate live failures traced to defaults we never set — a
314
+ reasoning-effort parameter that silently defaults to maximum (starving
315
+ small output budgets), sampling defaults that let a thinking model loop
316
+ (repetition consuming the token budget that carried the structured
317
+ output), and a degraded non-thinking mode we were unknowingly paying
318
+ for. Every call site now states its reasoning mode, its sampling, and a
319
+ budget the reasoning cannot starve; the discipline costs nothing and
320
+ removed an entire class of silent failure.
321
+
322
+ ## 8. The producer-native path: Metanorma Knowledge Objects
323
+
324
+ OIML publications are authored in Metanorma, a model-driven document
325
+ system: the source of truth is a typed document model, not any rendered
326
+ PDF or HTML. The service's newest ingestion path consumes that model
327
+ directly — one bundle per document containing typed units (clauses,
328
+ tables, terms, equations, figures, requirements), a section graph, a
329
+ native Glossarist glossary, and Relaton bibliographic objects — and
330
+ replaces HTML scraping end to end for the cleanly-authored portion of
331
+ the corpus. The format, Metanorma Knowledge Objects (MKO), is specified
332
+ as Metanorma note 116 with the consumer contract (symbolic unit
333
+ references and typed excerpts) contributed from this work.
334
+
335
+ <Figure
336
+ id="fig-mko"
337
+ src="paper-assets/mko-ingest.svg"
338
+ caption="Figure 3 — Producer-native ingestion. Each authored document exports as one MKO bundle; chunking, enrichment and indexing are derived from the bundle, and re-ingest is incremental by content hash."
339
+ />
340
+
341
+ This path matters beyond OIML: any standards body whose publications are
342
+ authored in Metanorma gets structured, table-aware, graph-connected
343
+ question answering over its corpus without scraping renderings. The
344
+ guidelines for building such a service from scratch accompany this
345
+ article.
346
+
347
+ ## 9. Evaluation as a discipline
348
+
349
+ The service's answers are evaluated on three axes, continuously:
350
+
351
+ - **Retrieval** (recall, precision, MRR against the golden set with
352
+ witness citations)
353
+ - **Faithfulness** (an independent judge scores whether each claim is
354
+ supported by the retrieved passages — with the important lesson that
355
+ the judge must see the passages the answer was actually built from,
356
+ not a fresh retrieval)
357
+ - **User feedback** (thumbs up/down on every answer, logged to the same
358
+ evaluation loop)
359
+
360
+ Every change ships through the same gate — literally one command:
361
+ golden suite ×3 and the capability battery ×6 against the live service,
362
+ where any failed run fails the command. The suites also run in CI and
363
+ the service's cache versions flush on every retrieval change so no
364
+ answer is served from a superseded index. A humility note from the
365
+ measurement machinery itself: the capability battery's runner once
366
+ omitted to set a non-zero exit code on failure, so a failing battery
367
+ reported green through the gate — the same class of silent failure the
368
+ cache-version discipline exists to prevent. Exit codes are part of the
369
+ measurement contract now. The gate rejects as often as it accepts.
370
+ A candidate change that unioned additional retrieval candidates into a
371
+ rewritten query's pool — a plausible-looking "more evidence" improvement
372
+ — dropped recall@5 from 94.3% to 89.7%: topically close but wrong
373
+ documents outscored the right ones under the cross-encoder, and the
374
+ per-case diff named the three failing questions. The fix (additive
375
+ candidates may only replace an identical query, never dilute a rewritten
376
+ one) is now a comment in the code. A second candidate — grading
377
+ document-scoped retrieval more aggressively — was measured, found to buy
378
+ nothing, and reverted the same day. Additive is not free; the suite, not
379
+ the author, decides.
380
+
381
+ ## 10. Conclusions and outlook
382
+
383
+ A grounded, citation-linked, typed-rendering question-answering service
384
+ over the OIML corpus is measurable, economical, and — with the
385
+ producer-native ingestion path — increasingly maintained by the
386
+ documents' own structure rather than by scraping their renderings. The
387
+ open work is equally concrete: upstream bibliographic corrections for
388
+ the 36 families without derivable current editions; unit-level language
389
+ tagging so bilingual annexes inside English editions stop masquerading
390
+ as main text; a publisher-specific identifier flavor so citation joins
391
+ become exact; collection-level bundles for cross-document reasoning; and
392
+ interlingual unit alignment so the same clause can be answered in every
393
+ OIML language. The requirements behind several of these belong to the
394
+ authoring system itself, and we have filed them with the Metanorma
395
+ document model team as the producer side of an AI-native publishing
396
+ stack.
397
+
398
+ The service is live at ai.oimlsmart.org. Try it, and tell us when it is
399
+ wrong — the feedback button is the fastest path into the evaluation
400
+ loop that everything above runs on.
401
+
402
+ ---
403
+
404
+ ### References
405
+
406
+ 1. Al Masoud, A., Arazzi, M., Germani, S., Nocera, A. *Exploring
407
+ Structural Complexity in Normative RAG with Graph-based approaches: A
408
+ case study on the ETSI Standards.* arXiv:2604.09868 (2026).
409
+ 2. Guttal, P., et al. *Structure-Aware Chunking for Tabular Data in
410
+ Retrieval-Augmented Generation.* arXiv:2605.00318 (2026).
411
+ 3. Anthropic. *Contextual Retrieval.* Engineering blog (2024).
412
+ 4. Günther, M., et al. *Late Chunking: Contextual Chunk Embeddings Using
413
+ Long-Context Embedding Models.* arXiv:2409.04701 (2024).
414
+ 5. Lewis, P., et al. *Retrieval-Augmented Generation for
415
+ Knowledge-Intensive NLP Tasks.* NeurIPS (2020).
416
+ 6. Metanorma. *MN 116: Metanorma Knowledge Objects (MKO) machine
417
+ serialization format.* Metanorma documentation (2026).
418
+ 7. Xu, L., et al. *Equipping Retrieval-Augmented LLMs with Document
419
+ Structure Awareness.* arXiv:2510.04293 (2025).
@@ -0,0 +1,172 @@
1
+ # Product and branding options for the standards-intelligence engine
2
+
3
+ > Status: DECIDED (2026-09-12) — Option 1, a new high-level product.
4
+ > **The name is Konneal** (decided 2026-09-13; the GitHub org and the
5
+ > domain are secured — **konneal.org**). The name decodes on three levels, all of them
6
+ > the product's own: K = knowledge (K-onneal is knowledge annealment,
7
+ > the methodology this system invented and published); anneal stays
8
+ > fully legible inside the spelling, so the name carries the story in
9
+ > one step; and konne(ction) — the SDO connects to its members,
10
+ > questions connect to clauses, citations connect answers to the
11
+ > original document. Web collision scan found no company, product or
12
+ > brand using either Konneal or Konnea; formal trademark screening
13
+ > (classes 9/42, target jurisdictions) remains the one professional
14
+ > step before public marketing.
15
+ >
16
+ > Suite frame: *authored in Metanorma, modeled in Primmel, served by
17
+ > Konneal.* Deployments stay white-labeled per SDO — "[SDO] Answers,
18
+ > powered by Konneal" — as OIML SMART AI presents today. When the
19
+ > engine is extracted (migration step 5), it lands in the Konneal org;
20
+ > this repository becomes publisher-oiml, the reference profile.
21
+
22
+ ## The suite frame (applies to every option)
23
+
24
+ The pipeline has three layers, and the measured staircase prices them:
25
+
26
+ | Layer | Product role | What it gives the SDO |
27
+ |---|---|---|
28
+ | Metanorma | Authoring | The document as a model, not a rendering: clause structure, typed units, renderings |
29
+ | Primmel | Modeling | The standard as executable rules: constraints, calculations, sequences |
30
+ | The engine (this system) | Serving | Cited answers, verdicts, proofs of absence, edition awareness, on the SDO's own domain |
31
+
32
+ A deployment is white-labeled to the SDO (`ai.<sdo>.org`, their theme,
33
+ their identity provider), so the brand decision here concerns the
34
+ product the SDO buys — the engine and its integration contract — not the
35
+ name their users see. The staircase is the cross-sell narrative: plain
36
+ text answers 2 of 18 capability probes, Metanorma-authored content
37
+ unlocks typed retrieval, and Primmel models unlock execution.
38
+
39
+ ---
40
+
41
+ ## Option 1 — a new high-level product (recommended)
42
+
43
+ **Category.** A standards-intelligence platform for standards
44
+ organizations: grounded question answering, conformance checking by
45
+ execution, and corpus verification over the SDO's own publications.
46
+
47
+ **Promise.** Your members ask; your standards answer — every claim
48
+ cited to the clause, every conformance question decided by executing
49
+ your own rules, on your domain, under your brand.
50
+
51
+ **Buyer.** The SDO secretary general, publication director, or the
52
+ owner of the member-services programme. This buyer purchases member
53
+ value and organizational authority, not tooling.
54
+
55
+ **Name candidates** (the decision this document defends is the level,
56
+ not the word; candidates for the naming sprint):
57
+ - *Anneal* — coined from the project's own methodology (knowledge
58
+ annealment), family-fits Metanorma and Primmel as a coined single
59
+ word, and the methodology is already published in the whitepaper.
60
+ Cost: it needs one sentence of explanation, as Metanorma once did.
61
+ - *Verdict* — named for the flagship capability (conformance by
62
+ execution). Strong and concrete; crowded trademark space.
63
+ - *Plain descriptive* ("Standards Answers Platform") — fastest to
64
+ understand, weakest to own; workable as the category label whatever
65
+ the brand is named.
66
+
67
+ Deployment-level naming pattern: `[SDO] Answers`, powered by the
68
+ product — mirroring how OIML SMART AI presents today.
69
+
70
+ **Suite mechanics.** The product consumes Metanorma renderings and
71
+ Primmel packages as profile inputs. The measured staircase is the sales
72
+ tool: it shows an SDO exactly which capabilities their current content
73
+ unlocks and what the next layer buys. Each layer sells the next without
74
+ the next being mandatory.
75
+
76
+ **Go to market.** Direct to SDOs that already author in Metanorma
77
+ (shortest path to the full staircase), with OIML SMART AI as the
78
+ reference deployment and the whitepaper as the technical proof. The
79
+ publisher profile is the integration contract an SDO's team can
80
+ evaluate in an afternoon.
81
+
82
+ **Risks.** A new brand costs market education, and the product must
83
+ carry its own demand generation. Mitigated by the reference deployment
84
+ and by the suite story, which lets Metanorma's existing SDO
85
+ relationships do the introductions without lending the product
86
+ Metanorma's name.
87
+
88
+ ---
89
+
90
+ ## Option 2 — a Metanorma family product ("Metanorma Answers")
91
+
92
+ **Category.** The AI layer of the Metanorma toolchain: answers and
93
+ execution over Metanorma-authored corpora.
94
+
95
+ **Promise.** Standards you author in Metanorma become answerable and
96
+ executable, automatically.
97
+
98
+ **Buyer.** The existing Metanorma buyer: standards editors and
99
+ toolchain owners inside SDOs.
100
+
101
+ **Name.** Rides the family: Metanorma Answers, Metanorma AI.
102
+
103
+ **Suite mechanics.** Collapses the serving layer into the authoring
104
+ brand. The staircase still exists technically but is branded as one
105
+ product's capability tiers.
106
+
107
+ **Advantages.** Fastest to market: the brand exists, the relationships
108
+ exist, the story ("author it, then serve it") is one sentence.
109
+
110
+ **Risks — the reasons this is not the recommendation.**
111
+ - It implies the content must be Metanorma-authored, which the engine
112
+ does not require; SDOs with legacy corpora would read themselves out
113
+ of the market.
114
+ - It caps the product as a toolchain add-on in the buyer's mind, and
115
+ the buyer is wrong: the person who buys member-facing answers is not
116
+ the person who buys the authoring toolchain.
117
+ - It spends Metanorma's brand equity on a service with different
118
+ quality attributes (a wrong answer damages the authoring brand's
119
+ credibility by association).
120
+ - It forecloses the Primmel story: execution-the-top-of-the-staircase
121
+ deserves its own co-branding rather than being a feature of the
122
+ authoring tool.
123
+
124
+ ---
125
+
126
+ ## Option 3 — a Primmel family product
127
+
128
+ **Category.** The serving and execution surface of the Primmel model
129
+ world.
130
+
131
+ **Promise.** Ask your Primmel models anything; the answers are
132
+ computed, not retrieved.
133
+
134
+ **Buyer.** Modelers and the (currently small) Primmel community.
135
+
136
+ **Risks — disqualifying today.** Primmel is the youngest brand with the
137
+ least market recognition; the engine is not about Primmel (retrieval,
138
+ enrichment and citation machinery stand entirely apart from the model
139
+ plane); and equating the product with executable models under-sells the
140
+ 90 percent of the system that works on plain corpora. This route
141
+ becomes interesting only if Primmel itself becomes the strategic brand
142
+ of the estate, which is a larger decision than this one.
143
+
144
+ ---
145
+
146
+ ## Comparison and recommendation
147
+
148
+ | Criterion | New product | Metanorma sub-brand | Primmel sub-brand |
149
+ |---|---|---|---|
150
+ | Buyer fit | Secretary/publication director (the one with budget for member value) | Toolchain owner | Modeler |
151
+ | Implies Metanorma required | No | Yes | No (implies Primmel) |
152
+ | Brand risk to existing products | None | Answers' quality reflects on Metanorma | Ditto |
153
+ | Time to market | Slower (new brand) | Fast | Slow |
154
+ | Ownable position | The category itself | A feature of a toolchain | A feature of a modeling tool |
155
+ | White-label per SDO | Natural | Awkward (whose name does the SDO's user see?) | Awkward |
156
+
157
+ **Recommendation: Option 1.** The engine is architecturally,
158
+ commercially and reputationally a distinct product. Bring it to market
159
+ as a new high-level brand, marketed as part of the suite — *authored in
160
+ Metanorma, modeled in Primmel, served by [new brand]* — with each
161
+ deployment white-labeled to the SDO. Run the naming as its own short
162
+ sprint with proper trademark screening; "Anneal" is the internal
163
+ candidate with the strongest story.
164
+
165
+ **What would change this answer:**
166
+ - If the go-to-market constraint is the next two quarters and Metanorma
167
+ channel relationships are the only realistic demand source, Option 2
168
+ becomes the pragmatic bridge — ideally as "X, from the makers of
169
+ Metanorma" rather than "Metanorma X", preserving the exit path to a
170
+ standalone brand.
171
+ - If the estate decides Primmel is the strategic brand everything else
172
+ rides on, revisit Option 3 — but that is a portfolio-level decision.
@@ -0,0 +1,88 @@
1
+ # Projects — design (draft for review)
2
+
3
+ > The question: should chats group into **projects** that share a
4
+ > per-project contextual memory? Yes — and the parts already exist.
5
+ > This is the design; nothing is implemented yet.
6
+
7
+ ## 1. Prior art — what the best got right and wrong
8
+
9
+ | System | Model | What we take | What we avoid |
10
+ |---|---|---|---|
11
+ | **Claude Projects** | knowledge files + custom instructions scoped to a bucket of chats | per-project memory *files* (not a single prompt box); chats join/leave freely | no file-version awareness (stale PDFs silently ground answers) |
12
+ | **ChatGPT Projects** | pinned docs + project instructions; every chat in the project sees them | explicit membership: a chat is IN a project, visibly | memory edits don't re-trigger anything — old answers keep their old grounding, undiscoverably |
13
+ | **Notion/Linear "projects"** | a project is a *container* with its own views and defaults | defaults ride the project (scope, language), not the client | — |
14
+ | **GitHub repos** | shared context = files in the tree; issues/PRs reference them | memory files are addressable objects (ids), not ambient goo | — |
15
+
16
+ The failure modes to design against: **invisible grounding** (an answer
17
+ shaped by memory the user forgot was on), **stale memory** (edited file,
18
+ cached answers), and **membership confusion** (chat drifts between
19
+ projects).
20
+
21
+ ## 2. The design
22
+
23
+ A **project** is a container that owns: member-scoped memory files,
24
+ default database scope, and a set of conversations. Everything else in
25
+ the ask path already exists.
26
+
27
+ ### Data model (D1)
28
+
29
+ ```sql
30
+ projects (id, sub, name, default_datasets TEXT, created_at)
31
+ -- memory files move from user-level to project-level:
32
+ project_files (id, project_id, name, content, updated_at)
33
+ -- conversations gain an optional home:
34
+ ALTER TABLE conversations ADD COLUMN project_id TEXT NULL;
35
+ ```
36
+
37
+ Per-user today's `memories` table remains the *personal* tier; project
38
+ files are the *shared* tier. The ask accepts both; the note builder
39
+ labels them: "the user's personal memory" / "the project's shared
40
+ memory".
41
+
42
+ ### The three rules that make it honest
43
+
44
+ 1. **Selection is visible and per-question.** The memory chips the
45
+ answer USED are echoed in the response (`context_applied` gains a
46
+ `memory: [...]` block) and rendered on the message — an answer shaped
47
+ by memory says so, on its face. (The context-chip echo already works
48
+ this way.)
49
+ 2. **Memory salts the cache.** Already shipped for personal memories
50
+ (#185): the file-id set is part of the cache key. Project memory
51
+ joins the same salt — edit a project file, and every keyed answer
52
+ for that selection misses and regenerates.
53
+ 3. **Membership is a move, not a copy.** A conversation belongs to at
54
+ most one project (a nullable `project_id`); moving it re-scopes its
55
+ NEXT answer, never rewrites history. Old messages keep their
56
+ recorded grounding (the `context_applied` column already persists
57
+ per answer).
58
+
59
+ ### UX
60
+
61
+ - Sidebar: a **Projects** section above conversations — project rows
62
+ with their conversation counts; selecting one filters the list and
63
+ sets the composer's default scope/memory (visible as chips under the
64
+ composer, toggleable like datasets/memory rows — one grammar users
65
+ already know).
66
+ - Project settings drawer: name, default databases, shared memory files
67
+ (the same editor modal as #171), and the member list when sharing
68
+ arrives.
69
+ - A conversation outside any project behaves exactly as today.
70
+
71
+ ### What projects unlock next (the payoff)
72
+
73
+ - **Team tier**: `project_members` — shared memory becomes the lab's
74
+ institutional memory ("our instruments, our classes"), the natural
75
+ members'-side feature.
76
+ - **Cross-chat continuity**: a project summary (the existing history
77
+ summarizer, run over the project's conversations) grounds *new*
78
+ chats in what the project already established — continuity without
79
+ leaking between projects.
80
+ - **Reproducible scoped sessions**: default scope + shared memory +
81
+ the corpus generation stamp = an answer set a reviewer can re-run.
82
+
83
+ ## 3. Effort
84
+
85
+ Small: two D1 tables + one column, CRUD that mirrors #171's handlers,
86
+ an echo field, and sidebar/composer UI in the shipped grammar. The
87
+ expensive-looking parts (injection, salting, toggles, echo) already
88
+ exist from the dataset-scope and memory work.