@konneal/engine 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +29 -0
- package/README.md +13 -0
- package/dist/admin.d.ts +26 -0
- package/dist/ai.d.ts +6 -0
- package/dist/anchors.d.ts +6 -0
- package/dist/answercache.d.ts +22 -0
- package/dist/ask.d.ts +5 -0
- package/dist/auth.d.ts +12 -0
- package/dist/bubble.d.ts +14 -0
- package/dist/chunk-LLWPT2XV.js +49 -0
- package/dist/chunk-MB74PTRM.js +114 -0
- package/dist/chunk-WOGQM7DJ.js +197 -0
- package/dist/chunk-WWNCWKKC.js +42 -0
- package/dist/completion.d.ts +5 -0
- package/dist/config.d.ts +154 -0
- package/dist/config.js +37 -0
- package/dist/context.d.ts +115 -0
- package/dist/conversations.d.ts +5 -0
- package/dist/drafts.d.ts +129 -0
- package/dist/env.d.ts +57 -0
- package/dist/faithfulness.d.ts +5 -0
- package/dist/grader.d.ts +3 -0
- package/dist/graph.d.ts +13 -0
- package/dist/hybrid.d.ts +7 -0
- package/dist/index.d.ts +9 -0
- package/dist/index.js +5373 -0
- package/dist/internal_gateway.d.ts +14 -0
- package/dist/lexical.d.ts +7 -0
- package/dist/livedata.d.ts +77 -0
- package/dist/memories.d.ts +10 -0
- package/dist/modelplane.d.ts +61 -0
- package/dist/oidc.d.ts +73 -0
- package/dist/pipeline.d.ts +57 -0
- package/dist/profile.d.ts +2 -0
- package/dist/profile.gen.d.ts +70 -0
- package/dist/profile.js +8 -0
- package/dist/projects.d.ts +8 -0
- package/dist/prompts/conversational.md +8 -0
- package/dist/prompts/enrichment.md +3 -0
- package/dist/prompts/faithfulness.md +1 -0
- package/dist/prompts/grader.md +5 -0
- package/dist/prompts/listwise.md +3 -0
- package/dist/prompts/precision.md +1 -0
- package/dist/prompts/reflect.md +1 -0
- package/dist/prompts/relevancy.md +1 -0
- package/dist/prompts/research.md +10 -0
- package/dist/prompts/section-summary.md +5 -0
- package/dist/prompts/summarize.md +1 -0
- package/dist/prompts/system.md +18 -0
- package/dist/prompts/understanding.md +17 -0
- package/dist/quota.d.ts +13 -0
- package/dist/reflect.d.ts +5 -0
- package/dist/refs.d.ts +40 -0
- package/dist/refusal.d.ts +9 -0
- package/dist/refusal.js +9 -0
- package/dist/requestScope.d.ts +26 -0
- package/dist/requestScope.js +10 -0
- package/dist/research.d.ts +8 -0
- package/dist/search.d.ts +4 -0
- package/dist/selfquery.d.ts +7 -0
- package/dist/session.d.ts +1 -0
- package/dist/share.d.ts +2 -0
- package/dist/structural.d.ts +27 -0
- package/dist/tablecontext.d.ts +11 -0
- package/dist/understand.d.ts +11 -0
- package/dist/understandContract.d.ts +29 -0
- package/dist/verdict.d.ts +24 -0
- package/docs/API.md +451 -0
- package/docs/ARCHITECTURE.md +302 -0
- package/docs/AUDIT-2026-08-24.md +71 -0
- package/docs/CONTRIBUTOR-AUDIT-2026-08-25.md +147 -0
- package/docs/INGEST-ARCHITECTURE.md +158 -0
- package/docs/MCP.md +92 -0
- package/docs/METANORMA-AI-SERIALIZATION.md +247 -0
- package/docs/MKO-EXPORT-PIPELINE.md +147 -0
- package/docs/REDESIGN-NORMATIVE-RAG-ETSI.md +485 -0
- package/docs/RESEARCH-SOTA-2026.md +243 -0
- package/docs/ROADMAP-SOTA.md +130 -0
- package/docs/SOTA-STAGE-SPECS.md +509 -0
- package/docs/annealment/F1-verdict.md +27 -0
- package/docs/annealment/F10-notes.md +19 -0
- package/docs/annealment/F11-composition.md +17 -0
- package/docs/annealment/F12-passport.md +17 -0
- package/docs/annealment/F2-counterfactual.md +20 -0
- package/docs/annealment/F3-absence.md +21 -0
- package/docs/annealment/F4-instance.md +18 -0
- package/docs/annealment/F5-workflow.md +21 -0
- package/docs/annealment/F6-impact.md +21 -0
- package/docs/annealment/F7-editions.md +18 -0
- package/docs/annealment/F8-selfverify.md +19 -0
- package/docs/annealment/F9-projection-qa.md +17 -0
- package/docs/annealment/L0-locate.md +19 -0
- package/docs/annealment/L1-extract.md +18 -0
- package/docs/annealment/L2-nomenclature.md +22 -0
- package/docs/annealment/L3-geometry.md +23 -0
- package/docs/annealment/L4-composition.md +21 -0
- package/docs/annealment/L5-cross-standard.md +20 -0
- package/docs/annealment/L6-diachrony.md +21 -0
- package/docs/annealment/L7-perception.md +20 -0
- package/docs/annealment/L8-computation.md +22 -0
- package/docs/annealment/L9-instance-process.md +23 -0
- package/docs/annealment/README.md +10 -0
- package/docs/guidelines-metanorma-ai-programme.md +279 -0
- package/docs/identity-onboarding-rag.md +65 -0
- package/docs/identity-service.md +219 -0
- package/docs/knowledge-annealment.md +273 -0
- package/docs/konneal-extraction-plan.md +481 -0
- package/docs/metanorma-for-ai.md +270 -0
- package/docs/mirror-plan.md +36 -0
- package/docs/multi-sdo-architecture.md +191 -0
- package/docs/paper-annealment-comparison.md +259 -0
- package/docs/paper-assets/architecture.svg +94 -0
- package/docs/paper-assets/contract-v2.svg +94 -0
- package/docs/paper-assets/mko-ingest.svg +91 -0
- package/docs/paper-oiml-bulletin.md +402 -0
- package/docs/paper-oiml-bulletin.mdx +419 -0
- package/docs/product-branding-options.md +172 -0
- package/docs/projects-design.md +88 -0
- package/docs/sota-mechanisms.md +184 -0
- package/docs/spec-api.md +77 -0
- package/docs/spec-pipeline.md +126 -0
- package/docs/vector-adapter.md +88 -0
- package/package.json +70 -0
- package/profile/corpora.yaml +5 -0
- package/profile/datasets.yaml +14 -0
- package/profile/prompts.yaml +5 -0
- package/profile/publisher.yaml +17 -0
- package/profile/retrieval.yaml +1 -0
- package/profile/sources.yaml +5 -0
- package/profile/ui.yaml +7 -0
- package/scripts/gen_profile.mjs +33 -0
- package/workers/shared/ai.ts +21 -0
- package/workers/shared/auth.ts +16 -0
- package/workers/shared/chunk.ts +108 -0
- package/workers/shared/oidc.ts +312 -0
- package/workers/shared/router.ts +45 -0
- package/workers/shared/session.ts +104 -0
- package/workers/worker_internal/src/index.ts +157 -0
- package/workers/worker_internal/tsconfig.json +15 -0
- package/workers/worker_internal/wrangler.toml +32 -0
- package/workers/worker_mcp/src/index.ts +175 -0
- package/workers/worker_mcp/tsconfig.json +13 -0
- package/workers/worker_mcp/wrangler.toml +18 -0
- package/workers/worker_public/migrations/0002_conversations.sql +22 -0
- package/workers/worker_public/migrations/0003_shared_conversations.sql +9 -0
- package/workers/worker_public/migrations/0004_graph.sql +16 -0
- package/workers/worker_public/migrations/0005_documents.sql +19 -0
- package/workers/worker_public/migrations/0006_conversation_entities.sql +11 -0
- package/workers/worker_public/migrations/0007_chunks_fts.sql +43 -0
- package/workers/worker_public/migrations/0008_unit_payloads.sql +15 -0
- package/workers/worker_public/migrations/0009_chunks_unit.sql +8 -0
- package/workers/worker_public/migrations/0009_message_context.sql +7 -0
- package/workers/worker_public/migrations/0010_model_nodes.sql +33 -0
- package/workers/worker_public/migrations/0011_schema_union.sql +38 -0
- package/workers/worker_public/migrations/0012_memories.sql +15 -0
- package/workers/worker_public/migrations/0013_projects.sql +21 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.d.ts +16306 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.ts +16261 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.d.ts +16373 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.ts +16328 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.d.ts +16382 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.ts +16337 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.d.ts +16383 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.ts +16338 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.d.ts +16403 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.ts +16358 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.d.ts +16408 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.ts +16363 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.d.ts +16414 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.ts +16369 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.d.ts +16414 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.ts +16369 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/README.md +135 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/entrypoints.svg +53 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.d.ts +17095 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.ts +17050 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/index.d.ts +16306 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/index.ts +16261 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.d.ts +16447 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.ts +16402 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.d.ts +16306 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.ts +16261 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/package.json +11 -0
- package/workers/worker_public/package.json +13 -0
- package/workers/worker_public/prompts/conversational.md +8 -0
- package/workers/worker_public/prompts/enrichment.md +3 -0
- package/workers/worker_public/prompts/faithfulness.md +1 -0
- package/workers/worker_public/prompts/grader.md +5 -0
- package/workers/worker_public/prompts/listwise.md +3 -0
- package/workers/worker_public/prompts/precision.md +1 -0
- package/workers/worker_public/prompts/reflect.md +1 -0
- package/workers/worker_public/prompts/relevancy.md +1 -0
- package/workers/worker_public/prompts/research.md +10 -0
- package/workers/worker_public/prompts/section-summary.md +5 -0
- package/workers/worker_public/prompts/summarize.md +1 -0
- package/workers/worker_public/prompts/system.md +18 -0
- package/workers/worker_public/prompts/understanding.md +17 -0
- package/workers/worker_public/public/app.js +166 -0
- package/workers/worker_public/public/index.html +48 -0
- package/workers/worker_public/public/style.css +147 -0
- package/workers/worker_public/schema.sql +248 -0
- package/workers/worker_public/src/admin.ts +358 -0
- package/workers/worker_public/src/ai.ts +71 -0
- package/workers/worker_public/src/anchors.ts +41 -0
- package/workers/worker_public/src/answercache.ts +72 -0
- package/workers/worker_public/src/ask.ts +1094 -0
- package/workers/worker_public/src/auth.ts +252 -0
- package/workers/worker_public/src/bubble.ts +111 -0
- package/workers/worker_public/src/completion.ts +75 -0
- package/workers/worker_public/src/config.ts +238 -0
- package/workers/worker_public/src/context.ts +238 -0
- package/workers/worker_public/src/conversations.ts +162 -0
- package/workers/worker_public/src/drafts.ts +497 -0
- package/workers/worker_public/src/env.ts +90 -0
- package/workers/worker_public/src/faithfulness.ts +63 -0
- package/workers/worker_public/src/grader.ts +89 -0
- package/workers/worker_public/src/graph.ts +63 -0
- package/workers/worker_public/src/hybrid.ts +77 -0
- package/workers/worker_public/src/index.ts +441 -0
- package/workers/worker_public/src/internal_gateway.ts +41 -0
- package/workers/worker_public/src/lexical.ts +86 -0
- package/workers/worker_public/src/lib/hit.ts +4 -0
- package/workers/worker_public/src/lib/http.ts +83 -0
- package/workers/worker_public/src/lib/router.ts +4 -0
- package/workers/worker_public/src/livedata.ts +334 -0
- package/workers/worker_public/src/memories.ts +81 -0
- package/workers/worker_public/src/modelplane.ts +213 -0
- package/workers/worker_public/src/oidc.ts +333 -0
- package/workers/worker_public/src/pipeline.ts +377 -0
- package/workers/worker_public/src/ports/blobs.ts +7 -0
- package/workers/worker_public/src/ports/cloudflare/adapters.ts +177 -0
- package/workers/worker_public/src/ports/kv.ts +8 -0
- package/workers/worker_public/src/ports/model.ts +28 -0
- package/workers/worker_public/src/ports/runtime.ts +13 -0
- package/workers/worker_public/src/ports/store.ts +20 -0
- package/workers/worker_public/src/ports/vector.ts +26 -0
- package/workers/worker_public/src/profile.gen.ts +101 -0
- package/workers/worker_public/src/profile.ts +16 -0
- package/workers/worker_public/src/projects.ts +108 -0
- package/workers/worker_public/src/prompts.d.ts +6 -0
- package/workers/worker_public/src/quota.ts +54 -0
- package/workers/worker_public/src/reflect.ts +67 -0
- package/workers/worker_public/src/refs.ts +107 -0
- package/workers/worker_public/src/refusal.ts +65 -0
- package/workers/worker_public/src/requestScope.ts +71 -0
- package/workers/worker_public/src/research.ts +126 -0
- package/workers/worker_public/src/search.ts +56 -0
- package/workers/worker_public/src/selfquery.ts +25 -0
- package/workers/worker_public/src/session.ts +4 -0
- package/workers/worker_public/src/share.ts +53 -0
- package/workers/worker_public/src/stages/conceptGraph.ts +68 -0
- package/workers/worker_public/src/stages/conceptSteer.ts +39 -0
- package/workers/worker_public/src/stages/corpusScope.ts +25 -0
- package/workers/worker_public/src/stages/dedup.ts +10 -0
- package/workers/worker_public/src/stages/dense.ts +73 -0
- package/workers/worker_public/src/stages/diversity.ts +33 -0
- package/workers/worker_public/src/stages/editionCover.ts +63 -0
- package/workers/worker_public/src/stages/editionSteer.ts +88 -0
- package/workers/worker_public/src/stages/familyBoost.ts +22 -0
- package/workers/worker_public/src/stages/federate.ts +22 -0
- package/workers/worker_public/src/stages/glossary.ts +65 -0
- package/workers/worker_public/src/stages/graphLane.ts +31 -0
- package/workers/worker_public/src/stages/hyde.ts +29 -0
- package/workers/worker_public/src/stages/index.ts +69 -0
- package/workers/worker_public/src/stages/lexicalUnion.ts +21 -0
- package/workers/worker_public/src/stages/multiQuery.ts +57 -0
- package/workers/worker_public/src/stages/overviewDemote.ts +14 -0
- package/workers/worker_public/src/stages/poolOpen.ts +10 -0
- package/workers/worker_public/src/stages/propagate.ts +15 -0
- package/workers/worker_public/src/stages/rerank.ts +47 -0
- package/workers/worker_public/src/stages/seal.ts +16 -0
- package/workers/worker_public/src/stages/sectionDescent.ts +61 -0
- package/workers/worker_public/src/stages/stdRefNudge.ts +35 -0
- package/workers/worker_public/src/stages/subQuery.ts +42 -0
- package/workers/worker_public/src/stages/termNudge.ts +24 -0
- package/workers/worker_public/src/stages/typedPin.ts +131 -0
- package/workers/worker_public/src/stages/types.ts +112 -0
- package/workers/worker_public/src/stages/windowFloor.ts +23 -0
- package/workers/worker_public/src/structural.ts +171 -0
- package/workers/worker_public/src/tablecontext.ts +41 -0
- package/workers/worker_public/src/understand.ts +72 -0
- package/workers/worker_public/src/understandContract.ts +67 -0
- package/workers/worker_public/src/verdict.ts +255 -0
- package/workers/worker_public/tsconfig.json +18 -0
- package/workers/worker_public/wrangler.toml +104 -0
|
@@ -0,0 +1,419 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: "Grounding legal metrology: a retrieval-augmented question-answering service over the OIML publications corpus"
|
|
3
|
+
subtitle: "Draft article for the OIML Bulletin"
|
|
4
|
+
date: 2026-08-30
|
|
5
|
+
figures: paper-assets
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
export const Figure = ({ id, src, caption, width = 760 }) => (
|
|
9
|
+
<figure id={id} style={{ margin: "24px auto", textAlign: "center" }}>
|
|
10
|
+
<img src={src} alt={caption} width={width} style={{ maxWidth: "100%", height: "auto" }} />
|
|
11
|
+
<figcaption style={{ fontSize: "0.85em", color: "#475569", marginTop: 8 }}>{caption}</figcaption>
|
|
12
|
+
</figure>
|
|
13
|
+
);
|
|
14
|
+
|
|
15
|
+
*Draft article for the OIML Bulletin — 2026-08-30. Style: Bulletin
|
|
16
|
+
technical article (MS Word single-column on submission; this is the
|
|
17
|
+
authoring source). Numbers are production-measured unless noted.*
|
|
18
|
+
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## Abstract
|
|
22
|
+
|
|
23
|
+
The OIML publishes close to 900 documents — Recommendations, Basic
|
|
24
|
+
publications, Guides, Documents and Expert reports — that together form
|
|
25
|
+
the reference library of legal metrology. Finding what they say has
|
|
26
|
+
required knowing the corpus before querying it. We describe a
|
|
27
|
+
question-answering service, **OIML SMART AI** (ai.oimlsmart.org), that
|
|
28
|
+
answers natural-language questions from the indexed publications with
|
|
29
|
+
clause-level citations, verbatim quote anchors for normative values, and
|
|
30
|
+
typed renderings of the tables and equations the answers depend on — and
|
|
31
|
+
it reads figures and user-supplied photographs directly, from pixels. The
|
|
32
|
+
service measures every component it ships: a golden question set,
|
|
33
|
+
retrieval metrics in the tradition of the information-retrieval
|
|
34
|
+
literature, and judged faithfulness. We report the measured design
|
|
35
|
+
decisions, the corpus-structure findings that shaped them, and the open
|
|
36
|
+
path: producer-native ingestion of the Metanorma document model so that
|
|
37
|
+
any Metanorma-authored corpus can be indexed the same way.
|
|
38
|
+
|
|
39
|
+
---
|
|
40
|
+
|
|
41
|
+
## 1. Introduction
|
|
42
|
+
|
|
43
|
+
Legal metrology's knowledge lives in a large, cross-referential,
|
|
44
|
+
multi-edition corpus. A typical working question — *"what is the maximum
|
|
45
|
+
permissible error for a class III nonautomatic weighing instrument?"* —
|
|
46
|
+
requires knowing which Recommendation applies (OIML R 76-1), which
|
|
47
|
+
edition is current, which clause holds the table, and how to read the
|
|
48
|
+
table's rows against the instrument's verification scale intervals.
|
|
49
|
+
Every one of those steps assumes corpus knowledge the questioner may not
|
|
50
|
+
have.
|
|
51
|
+
|
|
52
|
+
Question answering over such corpora has been transformed by
|
|
53
|
+
retrieval-augmented generation (RAG): a language model is given passages
|
|
54
|
+
retrieved from a trusted corpus and may use only those. The engineering
|
|
55
|
+
challenge is no longer whether a model can write fluent prose — it can —
|
|
56
|
+
but whether the *right passages* are found, whether the answer stays
|
|
57
|
+
*faithful* to them, and whether the reader can *verify* the result. Our
|
|
58
|
+
service is built around those three properties.
|
|
59
|
+
|
|
60
|
+
This article describes the architecture as deployed, the measurements
|
|
61
|
+
that shaped it, and the corpus findings — including errors in the
|
|
62
|
+
bibliographic record itself — that surfaced because the system keeps
|
|
63
|
+
score. It closes with the producer-native ingestion path (Metanorma
|
|
64
|
+
Knowledge Objects) and what it makes possible for other standards bodies.
|
|
65
|
+
|
|
66
|
+
## 2. What the service does
|
|
67
|
+
|
|
68
|
+
A user asks a question in any language. The service:
|
|
69
|
+
|
|
70
|
+
1. **Understands the question** with a small language model — language,
|
|
71
|
+
named publication, edition, defined terms, complexity. There are no
|
|
72
|
+
keyword rules; the same model makes these judgements for every
|
|
73
|
+
question.
|
|
74
|
+
2. **Retrieves in parallel** — three lanes leave the question at the
|
|
75
|
+
same instant: the vector index is queried with the question's
|
|
76
|
+
embedding while understanding is still running (the dense lane's
|
|
77
|
+
results are reused unchanged when understanding adds nothing), a
|
|
78
|
+
full-corpus keyword index catches exact identifiers, part numbers
|
|
79
|
+
and defined terms that vector similarity alone misses, and
|
|
80
|
+
understanding's refinements — a named publication, an edition —
|
|
81
|
+
arrive as metadata filters on the dense query.
|
|
82
|
+
3. **Ranks** the fused candidates with a cross-encoder; complex
|
|
83
|
+
questions get a second, stronger re-ranking pass.
|
|
84
|
+
4. **Answers from the retrieved passages only**, under a contract that
|
|
85
|
+
requires inline citations (`[OIML R 60-1:2021 §4.4.2]`) and verbatim
|
|
86
|
+
quote anchors for normative values.
|
|
87
|
+
5. **Verifies** the answer after generation: a deterministic check that
|
|
88
|
+
every quoted phrase exists in a cited passage; a separate judge that
|
|
89
|
+
scores faithfulness; a correction-and-retry when either fails. An
|
|
90
|
+
answer that cannot be verified is never written to the answer cache.
|
|
91
|
+
|
|
92
|
+
<Figure
|
|
93
|
+
id="fig-architecture"
|
|
94
|
+
src="paper-assets/architecture.svg"
|
|
95
|
+
caption="Figure 1 — A question's path through the service. Three retrieval lanes run concurrently; every answer passes mechanical verification before it is served or cached. The edition registry derives publication status from supersession links, never from a status field."
|
|
96
|
+
/>
|
|
97
|
+
|
|
98
|
+
When the question names a publication, answers are steered to the
|
|
99
|
+
current edition by a registry derived from the bibliographic record's
|
|
100
|
+
supersession links (§5). When the answer depends on a table or
|
|
101
|
+
equation, the interface renders the *producer's* typed object — the
|
|
102
|
+
actual rows and columns — never the model's re-typing of it (§6).
|
|
103
|
+
|
|
104
|
+
If the corpus does not contain the answer, the system says so in one
|
|
105
|
+
canonical sentence and redirects; refusals are never cached, because a
|
|
106
|
+
refusal is a property of the moment, not of the question.
|
|
107
|
+
|
|
108
|
+
## 3. Corpus and structure
|
|
109
|
+
|
|
110
|
+
The indexed corpus comprises the English editions of the OIML
|
|
111
|
+
publications (~900 documents, ~9.3 million words). Documents are chunked
|
|
112
|
+
along clause boundaries — never arbitrary token windows — and each chunk
|
|
113
|
+
carries provenance: publication identifier, edition, language, clause
|
|
114
|
+
anchor, and bibliographic status.
|
|
115
|
+
|
|
116
|
+
Three corpus-level findings shaped the design:
|
|
117
|
+
|
|
118
|
+
**The bibliographic status field disagrees with its own record.** In 58
|
|
119
|
+
of 224 publication families, the machine-readable status field claims a
|
|
120
|
+
document is current while the same record's supersession links say
|
|
121
|
+
otherwise. The service therefore *derives* status from the links and
|
|
122
|
+
ignores the field. More consequentially, 36 of 224 families have no
|
|
123
|
+
recorded successor link at all, so the current edition cannot be derived
|
|
124
|
+
for them; the registry surfaces these gaps rather than guessing. Both
|
|
125
|
+
numbers are the worklist for upstream bibliographic corrections — the
|
|
126
|
+
kind that serving surfaces only because it keeps score.
|
|
127
|
+
|
|
128
|
+
**Identifiers are identity, not display.** A cross-referenced audit of
|
|
129
|
+
the dirty corpus found documents whose machine-readable identifier was
|
|
130
|
+
a placeholder ("OIML D 0:0000", "OIML D X") or polluted with language
|
|
131
|
+
markers — around 200 records — making them unreachable by
|
|
132
|
+
doc-scoped retrieval. Identifiers are now derived from slug identity
|
|
133
|
+
with placeholders never winning, and the index is reconciled against
|
|
134
|
+
the canonical chunk set: the first census measured 49,553 live vectors
|
|
135
|
+
against 31,512 canonical — 18,041 strays from every prior re-chunking,
|
|
136
|
+
deleted in place (upserts never delete). A retrieval service that never
|
|
137
|
+
enumerates its own index serves ghosts.
|
|
138
|
+
|
|
139
|
+
**Tables are where the normative values live.** Maximum permissible
|
|
140
|
+
errors, accuracy-class limits, verification interval bounds — the values
|
|
141
|
+
practitioners ask for — are tabular. Flattening tables into prose loses
|
|
142
|
+
the row/column geometry that makes them answerable. The index now treats
|
|
143
|
+
tables as atomic typed objects (§6).
|
|
144
|
+
|
|
145
|
+
**Structure is a retrieval signal.** Standards are hierarchically
|
|
146
|
+
organized with extensive cross-references; the system indexes the
|
|
147
|
+
section hierarchy and the citation graph as a queryable graph (7,128
|
|
148
|
+
nodes, 6,486 edges) alongside the text, so a defined term can be
|
|
149
|
+
resolved to the documents that define it even when the question's
|
|
150
|
+
vocabulary does not match the corpus's.
|
|
151
|
+
|
|
152
|
+
## 4. Retrieval: what the measurements showed
|
|
153
|
+
|
|
154
|
+
We evaluate retrieval with a golden question set (now 29 questions
|
|
155
|
+
including eight table-value questions with witness rows) using
|
|
156
|
+
recall@5, average precision@5, and mean reciprocal rank@5 — the protocol
|
|
157
|
+
of the standards-retrieval literature. Single runs proved noisy (±5
|
|
158
|
+
percentage points); all reported numbers are means over three runs on a
|
|
159
|
+
quiet account.
|
|
160
|
+
|
|
161
|
+
The single largest measured improvement came from making keyword
|
|
162
|
+
retrieval a **full-corpus first stage** rather than a re-scoring of the
|
|
163
|
+
vector results: recall@5 rose from 86% to 95%. The reason is structural
|
|
164
|
+
— standards vocabulary is exact ("n_LC", "R 60-3", "creep"), and dense
|
|
165
|
+
embeddings alone miss exact identifiers that a lexical scan of the whole
|
|
166
|
+
corpus recovers. Fusing both rankings (reciprocal rank fusion) lifted
|
|
167
|
+
precision and MRR without costing recall.
|
|
168
|
+
|
|
169
|
+
The second improvement came from **contextual enrichment**: a one-time
|
|
170
|
+
pass that prepends a short model-written preamble to every chunk ("this
|
|
171
|
+
clause defines the accuracy-class limits for load cells") before
|
|
172
|
+
embedding and indexing. This is the contextual-retrieval recipe
|
|
173
|
+
validated industry-wide; our corpus measurement confirmed it, and the
|
|
174
|
+
one-time cost (~$0.0017 per chunk) is amortized over every future
|
|
175
|
+
answer.
|
|
176
|
+
|
|
177
|
+
With the typed-table lane live (§6), retrieval lanes running
|
|
178
|
+
concurrently with query understanding, edition steering active (§5) and
|
|
179
|
+
multimodal generation on (§6), the current baseline is:
|
|
180
|
+
|
|
181
|
+
| Metric | Value |
|
|
182
|
+
|---|---|
|
|
183
|
+
| Recall@5 (mean of 3) | **94.3%** (range 93.1–96.6) |
|
|
184
|
+
| Average precision@5 | 0.82 |
|
|
185
|
+
| MRR@5 | 0.87 |
|
|
186
|
+
| End-to-end golden cases | 14/14 |
|
|
187
|
+
|
|
188
|
+
These numbers are on our own golden set, not on the benchmarks of the
|
|
189
|
+
works we build on — the honest comparison is per-technique, on the same
|
|
190
|
+
corpus, before and after. So read: the full-corpus lexical stage is the
|
|
191
|
+
ETSI study's recipe (ref 1) — adopting it lifted recall@5 here from 86%
|
|
192
|
+
to 95%; the contextual preambles are Anthropic's contextual retrieval
|
|
193
|
+
(ref 3) — confirmed on this corpus at ~$0.0017 per chunk one-time; the
|
|
194
|
+
clause-boundary, structure-preserving chunking follows the same
|
|
195
|
+
structure-aware line as refs 2 and 7, which our typed-unit pin extends
|
|
196
|
+
from "chunk better" to "guarantee the answering object a slot." Two
|
|
197
|
+
elements of the deployed system have no counterpart in the cited work:
|
|
198
|
+
the symbolic-reference contract, under which table data never passes
|
|
199
|
+
through the model at all (refs 2's error-reduction approach still
|
|
200
|
+
re-generates tables; we removed the corruption channel), and the
|
|
201
|
+
mechanical post-generation verification of every answer (quote anchors,
|
|
202
|
+
unit references, retyped-table detection) with a faithfulness judge that
|
|
203
|
+
sees the passages actually used.
|
|
204
|
+
|
|
205
|
+
Two structural additions complete the retrieval story. First, the
|
|
206
|
+
corpus is a tree — every chunk carries its clause anchor — and the
|
|
207
|
+
serving path uses the tree: a hit's score blends its neighbouring
|
|
208
|
+
clauses' scores, evidence is presented to the answer model in document
|
|
209
|
+
reading order, and a ranked section summary resolves to its quotable
|
|
210
|
+
child clauses (adapted from FABLE/BEAR, ref 7, at none of its index-time
|
|
211
|
+
cost, because Metanorma documents arrive as trees rather than needing
|
|
212
|
+
one inferred). Second, everyday words rarely match defined terms — the
|
|
213
|
+
one gap that failed for every corpus representation equally — so a
|
|
214
|
+
terminology index of the corpus's defined concepts links a question's
|
|
215
|
+
phrasing to candidate terms, and the answer model adjudicates and leads
|
|
216
|
+
with the corpus's own term, quoting its definition ("my output keeps
|
|
217
|
+
drifting" resolves to *span stability*, with the verbatim definition and
|
|
218
|
+
its defining clause cited).
|
|
219
|
+
|
|
220
|
+
## 5. Editions and trust
|
|
221
|
+
|
|
222
|
+
Which edition applies is as consequential as what it says. The service's
|
|
223
|
+
registry of publication families and editions is derived from
|
|
224
|
+
supersession links, not from status fields (§3). Citations carry the
|
|
225
|
+
edition and status of every source; superseded sources are marked.
|
|
226
|
+
Answers to questions that name a publication are steered to the current
|
|
227
|
+
edition; answers that must cite a superseded edition (the 2021 passages
|
|
228
|
+
do not reproduce a table the 2017 edition carries) say so explicitly.
|
|
229
|
+
|
|
230
|
+
Steering is measured, not assumed, and two mechanisms earned their place
|
|
231
|
+
by fixing observed failures. Superseded editions match archaic phrasing
|
|
232
|
+
strongly — a complex verification question once cited R 76-1:1992/1988
|
|
233
|
+
while the current 2006 edition was crowded out of the passage window —
|
|
234
|
+
so ranking now demotes older-edition chunks of a publication whenever a
|
|
235
|
+
newer edition of the same publication is present in the evidence
|
|
236
|
+
(cross-publication recency is deliberately untouched: a 1992 document
|
|
237
|
+
that is still current must not lose to unrelated 2024 ones). And because
|
|
238
|
+
query understanding can pin an edition for the wrong family, an edition
|
|
239
|
+
pin that the corpus itself fails to corroborate is dropped to the
|
|
240
|
+
document-level filter. After both fixes the probe cites R 76-1:2006 in
|
|
241
|
+
force, with one legitimate 1988 tail citation for a clause only that
|
|
242
|
+
edition carries; the window's measured average precision (0.85) and MRR
|
|
243
|
+
(0.88) are the best the service has recorded.
|
|
244
|
+
|
|
245
|
+
## 6. Tables, equations, and figures as typed objects
|
|
246
|
+
|
|
247
|
+
The most consequential rendering decision: the language model never
|
|
248
|
+
re-types normative data. When an answer depends on a table, the model
|
|
249
|
+
emits a symbolic reference to the table's unit (`[[u:table-1]]`); the
|
|
250
|
+
server validates the reference against the passages actually used and
|
|
251
|
+
resolves it to the producer's typed payload — columns, rows, units. The
|
|
252
|
+
interface renders the exact object. Small verbatim values in prose are
|
|
253
|
+
permitted and mechanically verified against the source. The payload
|
|
254
|
+
never passes through the model's output, so it cannot be corrupted by
|
|
255
|
+
generation.
|
|
256
|
+
|
|
257
|
+
The same principle extends from rendering to judgment: the model never
|
|
258
|
+
judges what a machine can compute. Where the corpus's machine-readable
|
|
259
|
+
models carry checkable objects — a constraint with an OCL expression, a
|
|
260
|
+
requirement with a threshold limit — a question naming one gets it
|
|
261
|
+
EXECUTED: the question's values bind to the rule's own parameters, the
|
|
262
|
+
check evaluates deterministically, and the verdict (pass, or the
|
|
263
|
+
standard's own violation word, with the arithmetic shown, or void naming
|
|
264
|
+
exactly what is missing) is attached to the answer as computed data.
|
|
265
|
+
Two sibling capabilities complete the set: provable absence (a topic
|
|
266
|
+
that a standard does not constrain returns an enumeration certificate
|
|
267
|
+
over its entire model plane — N nodes enumerated, zero matches — rather
|
|
268
|
+
than a refusal) and answer verification (any answer checkable against
|
|
269
|
+
the corpus: verbatim-quote containment, object-reference resolution, a
|
|
270
|
+
judged faithfulness score labeled as judged).
|
|
271
|
+
|
|
272
|
+
<Figure
|
|
273
|
+
id="fig-contract"
|
|
274
|
+
src="paper-assets/contract-v2.svg"
|
|
275
|
+
caption="Figure 2 — The answer contract. The model writes prose with inline citations and symbolic references; three mechanical checks run after generation; the server resolves surviving references to the producer's payloads and the client renders them."
|
|
276
|
+
/>
|
|
277
|
+
|
|
278
|
+
Figures are described by a vision-capable model that reads the actual
|
|
279
|
+
image pixels; equations are carried in their authoring-native formats
|
|
280
|
+
(AsciiMath and MathML) with a plain-language description. Both arrive as typed objects with the same validation.
|
|
281
|
+
|
|
282
|
+
Figures also flow the other way: when the retrieved passages contain
|
|
283
|
+
figure units, their actual images are attached to the answer model's
|
|
284
|
+
input, so descriptions and reasoning come from the drawing itself — the
|
|
285
|
+
model reads labels that exist only in the pixels (the pixel-label
|
|
286
|
+
probe — "what are the labeled example cases in the R 60-2 design-shapes
|
|
287
|
+
figure?" — is answered A, B and C from the drawing, stable across
|
|
288
|
+
repeated runs). Two measured invariants keep that true: assets must be
|
|
289
|
+
readable by vision pipelines (vector-sourced rasters that draw black on
|
|
290
|
+
transparent alpha flatten to a solid black rectangle inside a vision
|
|
291
|
+
model, whatever a browser shows — a corpus sweep detects and repairs
|
|
292
|
+
them), and images ride their own message in the generation call (long
|
|
293
|
+
passage text and image parts in one message triggers provider errors
|
|
294
|
+
that scale with payload). Users can likewise attach a photograph (an
|
|
295
|
+
instrument nameplate, a scale dial, a schematic) to their question; the
|
|
296
|
+
text still drives retrieval, and the image gives the model the visual
|
|
297
|
+
context, under the same citation contract.
|
|
298
|
+
|
|
299
|
+
## 7. Models, cost, and sovereignty
|
|
300
|
+
|
|
301
|
+
All models are open-weight and served on a single cloud platform
|
|
302
|
+
(Cloudflare Workers AI); no corpus data is sent to proprietary model
|
|
303
|
+
providers. The answer model (GLM-5.3 Flash, natively multimodal) serves
|
|
304
|
+
all tiers at approximately $0.001 per answer. The full one-time corpus
|
|
305
|
+
preparation — contextual enrichment of 42,000 chunks — cost
|
|
306
|
+
approximately $75. Daily serving at current traffic is under $5 per
|
|
307
|
+
month including infrastructure. The model policy is deliberately
|
|
308
|
+
cost-first on the serving path (every question pays it) and quality-first
|
|
309
|
+
on one-time work (enrichment, captioning, evaluation), where quality
|
|
310
|
+
persists into every future answer.
|
|
311
|
+
|
|
312
|
+
One lesson generalizes: read the model card before wiring a model. Three
|
|
313
|
+
separate live failures traced to defaults we never set — a
|
|
314
|
+
reasoning-effort parameter that silently defaults to maximum (starving
|
|
315
|
+
small output budgets), sampling defaults that let a thinking model loop
|
|
316
|
+
(repetition consuming the token budget that carried the structured
|
|
317
|
+
output), and a degraded non-thinking mode we were unknowingly paying
|
|
318
|
+
for. Every call site now states its reasoning mode, its sampling, and a
|
|
319
|
+
budget the reasoning cannot starve; the discipline costs nothing and
|
|
320
|
+
removed an entire class of silent failure.
|
|
321
|
+
|
|
322
|
+
## 8. The producer-native path: Metanorma Knowledge Objects
|
|
323
|
+
|
|
324
|
+
OIML publications are authored in Metanorma, a model-driven document
|
|
325
|
+
system: the source of truth is a typed document model, not any rendered
|
|
326
|
+
PDF or HTML. The service's newest ingestion path consumes that model
|
|
327
|
+
directly — one bundle per document containing typed units (clauses,
|
|
328
|
+
tables, terms, equations, figures, requirements), a section graph, a
|
|
329
|
+
native Glossarist glossary, and Relaton bibliographic objects — and
|
|
330
|
+
replaces HTML scraping end to end for the cleanly-authored portion of
|
|
331
|
+
the corpus. The format, Metanorma Knowledge Objects (MKO), is specified
|
|
332
|
+
as Metanorma note 116 with the consumer contract (symbolic unit
|
|
333
|
+
references and typed excerpts) contributed from this work.
|
|
334
|
+
|
|
335
|
+
<Figure
|
|
336
|
+
id="fig-mko"
|
|
337
|
+
src="paper-assets/mko-ingest.svg"
|
|
338
|
+
caption="Figure 3 — Producer-native ingestion. Each authored document exports as one MKO bundle; chunking, enrichment and indexing are derived from the bundle, and re-ingest is incremental by content hash."
|
|
339
|
+
/>
|
|
340
|
+
|
|
341
|
+
This path matters beyond OIML: any standards body whose publications are
|
|
342
|
+
authored in Metanorma gets structured, table-aware, graph-connected
|
|
343
|
+
question answering over its corpus without scraping renderings. The
|
|
344
|
+
guidelines for building such a service from scratch accompany this
|
|
345
|
+
article.
|
|
346
|
+
|
|
347
|
+
## 9. Evaluation as a discipline
|
|
348
|
+
|
|
349
|
+
The service's answers are evaluated on three axes, continuously:
|
|
350
|
+
|
|
351
|
+
- **Retrieval** (recall, precision, MRR against the golden set with
|
|
352
|
+
witness citations)
|
|
353
|
+
- **Faithfulness** (an independent judge scores whether each claim is
|
|
354
|
+
supported by the retrieved passages — with the important lesson that
|
|
355
|
+
the judge must see the passages the answer was actually built from,
|
|
356
|
+
not a fresh retrieval)
|
|
357
|
+
- **User feedback** (thumbs up/down on every answer, logged to the same
|
|
358
|
+
evaluation loop)
|
|
359
|
+
|
|
360
|
+
Every change ships through the same gate — literally one command:
|
|
361
|
+
golden suite ×3 and the capability battery ×6 against the live service,
|
|
362
|
+
where any failed run fails the command. The suites also run in CI and
|
|
363
|
+
the service's cache versions flush on every retrieval change so no
|
|
364
|
+
answer is served from a superseded index. A humility note from the
|
|
365
|
+
measurement machinery itself: the capability battery's runner once
|
|
366
|
+
omitted to set a non-zero exit code on failure, so a failing battery
|
|
367
|
+
reported green through the gate — the same class of silent failure the
|
|
368
|
+
cache-version discipline exists to prevent. Exit codes are part of the
|
|
369
|
+
measurement contract now. The gate rejects as often as it accepts.
|
|
370
|
+
A candidate change that unioned additional retrieval candidates into a
|
|
371
|
+
rewritten query's pool — a plausible-looking "more evidence" improvement
|
|
372
|
+
— dropped recall@5 from 94.3% to 89.7%: topically close but wrong
|
|
373
|
+
documents outscored the right ones under the cross-encoder, and the
|
|
374
|
+
per-case diff named the three failing questions. The fix (additive
|
|
375
|
+
candidates may only replace an identical query, never dilute a rewritten
|
|
376
|
+
one) is now a comment in the code. A second candidate — grading
|
|
377
|
+
document-scoped retrieval more aggressively — was measured, found to buy
|
|
378
|
+
nothing, and reverted the same day. Additive is not free; the suite, not
|
|
379
|
+
the author, decides.
|
|
380
|
+
|
|
381
|
+
## 10. Conclusions and outlook
|
|
382
|
+
|
|
383
|
+
A grounded, citation-linked, typed-rendering question-answering service
|
|
384
|
+
over the OIML corpus is measurable, economical, and — with the
|
|
385
|
+
producer-native ingestion path — increasingly maintained by the
|
|
386
|
+
documents' own structure rather than by scraping their renderings. The
|
|
387
|
+
open work is equally concrete: upstream bibliographic corrections for
|
|
388
|
+
the 36 families without derivable current editions; unit-level language
|
|
389
|
+
tagging so bilingual annexes inside English editions stop masquerading
|
|
390
|
+
as main text; a publisher-specific identifier flavor so citation joins
|
|
391
|
+
become exact; collection-level bundles for cross-document reasoning; and
|
|
392
|
+
interlingual unit alignment so the same clause can be answered in every
|
|
393
|
+
OIML language. The requirements behind several of these belong to the
|
|
394
|
+
authoring system itself, and we have filed them with the Metanorma
|
|
395
|
+
document model team as the producer side of an AI-native publishing
|
|
396
|
+
stack.
|
|
397
|
+
|
|
398
|
+
The service is live at ai.oimlsmart.org. Try it, and tell us when it is
|
|
399
|
+
wrong — the feedback button is the fastest path into the evaluation
|
|
400
|
+
loop that everything above runs on.
|
|
401
|
+
|
|
402
|
+
---
|
|
403
|
+
|
|
404
|
+
### References
|
|
405
|
+
|
|
406
|
+
1. Al Masoud, A., Arazzi, M., Germani, S., Nocera, A. *Exploring
|
|
407
|
+
Structural Complexity in Normative RAG with Graph-based approaches: A
|
|
408
|
+
case study on the ETSI Standards.* arXiv:2604.09868 (2026).
|
|
409
|
+
2. Guttal, P., et al. *Structure-Aware Chunking for Tabular Data in
|
|
410
|
+
Retrieval-Augmented Generation.* arXiv:2605.00318 (2026).
|
|
411
|
+
3. Anthropic. *Contextual Retrieval.* Engineering blog (2024).
|
|
412
|
+
4. Günther, M., et al. *Late Chunking: Contextual Chunk Embeddings Using
|
|
413
|
+
Long-Context Embedding Models.* arXiv:2409.04701 (2024).
|
|
414
|
+
5. Lewis, P., et al. *Retrieval-Augmented Generation for
|
|
415
|
+
Knowledge-Intensive NLP Tasks.* NeurIPS (2020).
|
|
416
|
+
6. Metanorma. *MN 116: Metanorma Knowledge Objects (MKO) machine
|
|
417
|
+
serialization format.* Metanorma documentation (2026).
|
|
418
|
+
7. Xu, L., et al. *Equipping Retrieval-Augmented LLMs with Document
|
|
419
|
+
Structure Awareness.* arXiv:2510.04293 (2025).
|
|
@@ -0,0 +1,172 @@
|
|
|
1
|
+
# Product and branding options for the standards-intelligence engine
|
|
2
|
+
|
|
3
|
+
> Status: DECIDED (2026-09-12) — Option 1, a new high-level product.
|
|
4
|
+
> **The name is Konneal** (decided 2026-09-13; the GitHub org and the
|
|
5
|
+
> domain are secured — **konneal.org**). The name decodes on three levels, all of them
|
|
6
|
+
> the product's own: K = knowledge (K-onneal is knowledge annealment,
|
|
7
|
+
> the methodology this system invented and published); anneal stays
|
|
8
|
+
> fully legible inside the spelling, so the name carries the story in
|
|
9
|
+
> one step; and konne(ction) — the SDO connects to its members,
|
|
10
|
+
> questions connect to clauses, citations connect answers to the
|
|
11
|
+
> original document. Web collision scan found no company, product or
|
|
12
|
+
> brand using either Konneal or Konnea; formal trademark screening
|
|
13
|
+
> (classes 9/42, target jurisdictions) remains the one professional
|
|
14
|
+
> step before public marketing.
|
|
15
|
+
>
|
|
16
|
+
> Suite frame: *authored in Metanorma, modeled in Primmel, served by
|
|
17
|
+
> Konneal.* Deployments stay white-labeled per SDO — "[SDO] Answers,
|
|
18
|
+
> powered by Konneal" — as OIML SMART AI presents today. When the
|
|
19
|
+
> engine is extracted (migration step 5), it lands in the Konneal org;
|
|
20
|
+
> this repository becomes publisher-oiml, the reference profile.
|
|
21
|
+
|
|
22
|
+
## The suite frame (applies to every option)
|
|
23
|
+
|
|
24
|
+
The pipeline has three layers, and the measured staircase prices them:
|
|
25
|
+
|
|
26
|
+
| Layer | Product role | What it gives the SDO |
|
|
27
|
+
|---|---|---|
|
|
28
|
+
| Metanorma | Authoring | The document as a model, not a rendering: clause structure, typed units, renderings |
|
|
29
|
+
| Primmel | Modeling | The standard as executable rules: constraints, calculations, sequences |
|
|
30
|
+
| The engine (this system) | Serving | Cited answers, verdicts, proofs of absence, edition awareness, on the SDO's own domain |
|
|
31
|
+
|
|
32
|
+
A deployment is white-labeled to the SDO (`ai.<sdo>.org`, their theme,
|
|
33
|
+
their identity provider), so the brand decision here concerns the
|
|
34
|
+
product the SDO buys — the engine and its integration contract — not the
|
|
35
|
+
name their users see. The staircase is the cross-sell narrative: plain
|
|
36
|
+
text answers 2 of 18 capability probes, Metanorma-authored content
|
|
37
|
+
unlocks typed retrieval, and Primmel models unlock execution.
|
|
38
|
+
|
|
39
|
+
---
|
|
40
|
+
|
|
41
|
+
## Option 1 — a new high-level product (recommended)
|
|
42
|
+
|
|
43
|
+
**Category.** A standards-intelligence platform for standards
|
|
44
|
+
organizations: grounded question answering, conformance checking by
|
|
45
|
+
execution, and corpus verification over the SDO's own publications.
|
|
46
|
+
|
|
47
|
+
**Promise.** Your members ask; your standards answer — every claim
|
|
48
|
+
cited to the clause, every conformance question decided by executing
|
|
49
|
+
your own rules, on your domain, under your brand.
|
|
50
|
+
|
|
51
|
+
**Buyer.** The SDO secretary general, publication director, or the
|
|
52
|
+
owner of the member-services programme. This buyer purchases member
|
|
53
|
+
value and organizational authority, not tooling.
|
|
54
|
+
|
|
55
|
+
**Name candidates** (the decision this document defends is the level,
|
|
56
|
+
not the word; candidates for the naming sprint):
|
|
57
|
+
- *Anneal* — coined from the project's own methodology (knowledge
|
|
58
|
+
annealment), family-fits Metanorma and Primmel as a coined single
|
|
59
|
+
word, and the methodology is already published in the whitepaper.
|
|
60
|
+
Cost: it needs one sentence of explanation, as Metanorma once did.
|
|
61
|
+
- *Verdict* — named for the flagship capability (conformance by
|
|
62
|
+
execution). Strong and concrete; crowded trademark space.
|
|
63
|
+
- *Plain descriptive* ("Standards Answers Platform") — fastest to
|
|
64
|
+
understand, weakest to own; workable as the category label whatever
|
|
65
|
+
the brand is named.
|
|
66
|
+
|
|
67
|
+
Deployment-level naming pattern: `[SDO] Answers`, powered by the
|
|
68
|
+
product — mirroring how OIML SMART AI presents today.
|
|
69
|
+
|
|
70
|
+
**Suite mechanics.** The product consumes Metanorma renderings and
|
|
71
|
+
Primmel packages as profile inputs. The measured staircase is the sales
|
|
72
|
+
tool: it shows an SDO exactly which capabilities their current content
|
|
73
|
+
unlocks and what the next layer buys. Each layer sells the next without
|
|
74
|
+
the next being mandatory.
|
|
75
|
+
|
|
76
|
+
**Go to market.** Direct to SDOs that already author in Metanorma
|
|
77
|
+
(shortest path to the full staircase), with OIML SMART AI as the
|
|
78
|
+
reference deployment and the whitepaper as the technical proof. The
|
|
79
|
+
publisher profile is the integration contract an SDO's team can
|
|
80
|
+
evaluate in an afternoon.
|
|
81
|
+
|
|
82
|
+
**Risks.** A new brand costs market education, and the product must
|
|
83
|
+
carry its own demand generation. Mitigated by the reference deployment
|
|
84
|
+
and by the suite story, which lets Metanorma's existing SDO
|
|
85
|
+
relationships do the introductions without lending the product
|
|
86
|
+
Metanorma's name.
|
|
87
|
+
|
|
88
|
+
---
|
|
89
|
+
|
|
90
|
+
## Option 2 — a Metanorma family product ("Metanorma Answers")
|
|
91
|
+
|
|
92
|
+
**Category.** The AI layer of the Metanorma toolchain: answers and
|
|
93
|
+
execution over Metanorma-authored corpora.
|
|
94
|
+
|
|
95
|
+
**Promise.** Standards you author in Metanorma become answerable and
|
|
96
|
+
executable, automatically.
|
|
97
|
+
|
|
98
|
+
**Buyer.** The existing Metanorma buyer: standards editors and
|
|
99
|
+
toolchain owners inside SDOs.
|
|
100
|
+
|
|
101
|
+
**Name.** Rides the family: Metanorma Answers, Metanorma AI.
|
|
102
|
+
|
|
103
|
+
**Suite mechanics.** Collapses the serving layer into the authoring
|
|
104
|
+
brand. The staircase still exists technically but is branded as one
|
|
105
|
+
product's capability tiers.
|
|
106
|
+
|
|
107
|
+
**Advantages.** Fastest to market: the brand exists, the relationships
|
|
108
|
+
exist, the story ("author it, then serve it") is one sentence.
|
|
109
|
+
|
|
110
|
+
**Risks — the reasons this is not the recommendation.**
|
|
111
|
+
- It implies the content must be Metanorma-authored, which the engine
|
|
112
|
+
does not require; SDOs with legacy corpora would read themselves out
|
|
113
|
+
of the market.
|
|
114
|
+
- It caps the product as a toolchain add-on in the buyer's mind, and
|
|
115
|
+
the buyer is wrong: the person who buys member-facing answers is not
|
|
116
|
+
the person who buys the authoring toolchain.
|
|
117
|
+
- It spends Metanorma's brand equity on a service with different
|
|
118
|
+
quality attributes (a wrong answer damages the authoring brand's
|
|
119
|
+
credibility by association).
|
|
120
|
+
- It forecloses the Primmel story: execution-the-top-of-the-staircase
|
|
121
|
+
deserves its own co-branding rather than being a feature of the
|
|
122
|
+
authoring tool.
|
|
123
|
+
|
|
124
|
+
---
|
|
125
|
+
|
|
126
|
+
## Option 3 — a Primmel family product
|
|
127
|
+
|
|
128
|
+
**Category.** The serving and execution surface of the Primmel model
|
|
129
|
+
world.
|
|
130
|
+
|
|
131
|
+
**Promise.** Ask your Primmel models anything; the answers are
|
|
132
|
+
computed, not retrieved.
|
|
133
|
+
|
|
134
|
+
**Buyer.** Modelers and the (currently small) Primmel community.
|
|
135
|
+
|
|
136
|
+
**Risks — disqualifying today.** Primmel is the youngest brand with the
|
|
137
|
+
least market recognition; the engine is not about Primmel (retrieval,
|
|
138
|
+
enrichment and citation machinery stand entirely apart from the model
|
|
139
|
+
plane); and equating the product with executable models under-sells the
|
|
140
|
+
90 percent of the system that works on plain corpora. This route
|
|
141
|
+
becomes interesting only if Primmel itself becomes the strategic brand
|
|
142
|
+
of the estate, which is a larger decision than this one.
|
|
143
|
+
|
|
144
|
+
---
|
|
145
|
+
|
|
146
|
+
## Comparison and recommendation
|
|
147
|
+
|
|
148
|
+
| Criterion | New product | Metanorma sub-brand | Primmel sub-brand |
|
|
149
|
+
|---|---|---|---|
|
|
150
|
+
| Buyer fit | Secretary/publication director (the one with budget for member value) | Toolchain owner | Modeler |
|
|
151
|
+
| Implies Metanorma required | No | Yes | No (implies Primmel) |
|
|
152
|
+
| Brand risk to existing products | None | Answers' quality reflects on Metanorma | Ditto |
|
|
153
|
+
| Time to market | Slower (new brand) | Fast | Slow |
|
|
154
|
+
| Ownable position | The category itself | A feature of a toolchain | A feature of a modeling tool |
|
|
155
|
+
| White-label per SDO | Natural | Awkward (whose name does the SDO's user see?) | Awkward |
|
|
156
|
+
|
|
157
|
+
**Recommendation: Option 1.** The engine is architecturally,
|
|
158
|
+
commercially and reputationally a distinct product. Bring it to market
|
|
159
|
+
as a new high-level brand, marketed as part of the suite — *authored in
|
|
160
|
+
Metanorma, modeled in Primmel, served by [new brand]* — with each
|
|
161
|
+
deployment white-labeled to the SDO. Run the naming as its own short
|
|
162
|
+
sprint with proper trademark screening; "Anneal" is the internal
|
|
163
|
+
candidate with the strongest story.
|
|
164
|
+
|
|
165
|
+
**What would change this answer:**
|
|
166
|
+
- If the go-to-market constraint is the next two quarters and Metanorma
|
|
167
|
+
channel relationships are the only realistic demand source, Option 2
|
|
168
|
+
becomes the pragmatic bridge — ideally as "X, from the makers of
|
|
169
|
+
Metanorma" rather than "Metanorma X", preserving the exit path to a
|
|
170
|
+
standalone brand.
|
|
171
|
+
- If the estate decides Primmel is the strategic brand everything else
|
|
172
|
+
rides on, revisit Option 3 — but that is a portfolio-level decision.
|
|
@@ -0,0 +1,88 @@
|
|
|
1
|
+
# Projects — design (draft for review)
|
|
2
|
+
|
|
3
|
+
> The question: should chats group into **projects** that share a
|
|
4
|
+
> per-project contextual memory? Yes — and the parts already exist.
|
|
5
|
+
> This is the design; nothing is implemented yet.
|
|
6
|
+
|
|
7
|
+
## 1. Prior art — what the best got right and wrong
|
|
8
|
+
|
|
9
|
+
| System | Model | What we take | What we avoid |
|
|
10
|
+
|---|---|---|---|
|
|
11
|
+
| **Claude Projects** | knowledge files + custom instructions scoped to a bucket of chats | per-project memory *files* (not a single prompt box); chats join/leave freely | no file-version awareness (stale PDFs silently ground answers) |
|
|
12
|
+
| **ChatGPT Projects** | pinned docs + project instructions; every chat in the project sees them | explicit membership: a chat is IN a project, visibly | memory edits don't re-trigger anything — old answers keep their old grounding, undiscoverably |
|
|
13
|
+
| **Notion/Linear "projects"** | a project is a *container* with its own views and defaults | defaults ride the project (scope, language), not the client | — |
|
|
14
|
+
| **GitHub repos** | shared context = files in the tree; issues/PRs reference them | memory files are addressable objects (ids), not ambient goo | — |
|
|
15
|
+
|
|
16
|
+
The failure modes to design against: **invisible grounding** (an answer
|
|
17
|
+
shaped by memory the user forgot was on), **stale memory** (edited file,
|
|
18
|
+
cached answers), and **membership confusion** (chat drifts between
|
|
19
|
+
projects).
|
|
20
|
+
|
|
21
|
+
## 2. The design
|
|
22
|
+
|
|
23
|
+
A **project** is a container that owns: member-scoped memory files,
|
|
24
|
+
default database scope, and a set of conversations. Everything else in
|
|
25
|
+
the ask path already exists.
|
|
26
|
+
|
|
27
|
+
### Data model (D1)
|
|
28
|
+
|
|
29
|
+
```sql
|
|
30
|
+
projects (id, sub, name, default_datasets TEXT, created_at)
|
|
31
|
+
-- memory files move from user-level to project-level:
|
|
32
|
+
project_files (id, project_id, name, content, updated_at)
|
|
33
|
+
-- conversations gain an optional home:
|
|
34
|
+
ALTER TABLE conversations ADD COLUMN project_id TEXT NULL;
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
Per-user today's `memories` table remains the *personal* tier; project
|
|
38
|
+
files are the *shared* tier. The ask accepts both; the note builder
|
|
39
|
+
labels them: "the user's personal memory" / "the project's shared
|
|
40
|
+
memory".
|
|
41
|
+
|
|
42
|
+
### The three rules that make it honest
|
|
43
|
+
|
|
44
|
+
1. **Selection is visible and per-question.** The memory chips the
|
|
45
|
+
answer USED are echoed in the response (`context_applied` gains a
|
|
46
|
+
`memory: [...]` block) and rendered on the message — an answer shaped
|
|
47
|
+
by memory says so, on its face. (The context-chip echo already works
|
|
48
|
+
this way.)
|
|
49
|
+
2. **Memory salts the cache.** Already shipped for personal memories
|
|
50
|
+
(#185): the file-id set is part of the cache key. Project memory
|
|
51
|
+
joins the same salt — edit a project file, and every keyed answer
|
|
52
|
+
for that selection misses and regenerates.
|
|
53
|
+
3. **Membership is a move, not a copy.** A conversation belongs to at
|
|
54
|
+
most one project (a nullable `project_id`); moving it re-scopes its
|
|
55
|
+
NEXT answer, never rewrites history. Old messages keep their
|
|
56
|
+
recorded grounding (the `context_applied` column already persists
|
|
57
|
+
per answer).
|
|
58
|
+
|
|
59
|
+
### UX
|
|
60
|
+
|
|
61
|
+
- Sidebar: a **Projects** section above conversations — project rows
|
|
62
|
+
with their conversation counts; selecting one filters the list and
|
|
63
|
+
sets the composer's default scope/memory (visible as chips under the
|
|
64
|
+
composer, toggleable like datasets/memory rows — one grammar users
|
|
65
|
+
already know).
|
|
66
|
+
- Project settings drawer: name, default databases, shared memory files
|
|
67
|
+
(the same editor modal as #171), and the member list when sharing
|
|
68
|
+
arrives.
|
|
69
|
+
- A conversation outside any project behaves exactly as today.
|
|
70
|
+
|
|
71
|
+
### What projects unlock next (the payoff)
|
|
72
|
+
|
|
73
|
+
- **Team tier**: `project_members` — shared memory becomes the lab's
|
|
74
|
+
institutional memory ("our instruments, our classes"), the natural
|
|
75
|
+
members'-side feature.
|
|
76
|
+
- **Cross-chat continuity**: a project summary (the existing history
|
|
77
|
+
summarizer, run over the project's conversations) grounds *new*
|
|
78
|
+
chats in what the project already established — continuity without
|
|
79
|
+
leaking between projects.
|
|
80
|
+
- **Reproducible scoped sessions**: default scope + shared memory +
|
|
81
|
+
the corpus generation stamp = an answer set a reviewer can re-run.
|
|
82
|
+
|
|
83
|
+
## 3. Effort
|
|
84
|
+
|
|
85
|
+
Small: two D1 tables + one column, CRUD that mirrors #171's handlers,
|
|
86
|
+
an echo field, and sidebar/composer UI in the shipped grammar. The
|
|
87
|
+
expensive-looking parts (injection, salting, toggles, echo) already
|
|
88
|
+
exist from the dataset-scope and memory work.
|