@konneal/engine 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +29 -0
- package/README.md +13 -0
- package/dist/admin.d.ts +26 -0
- package/dist/ai.d.ts +6 -0
- package/dist/anchors.d.ts +6 -0
- package/dist/answercache.d.ts +22 -0
- package/dist/ask.d.ts +5 -0
- package/dist/auth.d.ts +12 -0
- package/dist/bubble.d.ts +14 -0
- package/dist/chunk-LLWPT2XV.js +49 -0
- package/dist/chunk-MB74PTRM.js +114 -0
- package/dist/chunk-WOGQM7DJ.js +197 -0
- package/dist/chunk-WWNCWKKC.js +42 -0
- package/dist/completion.d.ts +5 -0
- package/dist/config.d.ts +154 -0
- package/dist/config.js +37 -0
- package/dist/context.d.ts +115 -0
- package/dist/conversations.d.ts +5 -0
- package/dist/drafts.d.ts +129 -0
- package/dist/env.d.ts +57 -0
- package/dist/faithfulness.d.ts +5 -0
- package/dist/grader.d.ts +3 -0
- package/dist/graph.d.ts +13 -0
- package/dist/hybrid.d.ts +7 -0
- package/dist/index.d.ts +9 -0
- package/dist/index.js +5373 -0
- package/dist/internal_gateway.d.ts +14 -0
- package/dist/lexical.d.ts +7 -0
- package/dist/livedata.d.ts +77 -0
- package/dist/memories.d.ts +10 -0
- package/dist/modelplane.d.ts +61 -0
- package/dist/oidc.d.ts +73 -0
- package/dist/pipeline.d.ts +57 -0
- package/dist/profile.d.ts +2 -0
- package/dist/profile.gen.d.ts +70 -0
- package/dist/profile.js +8 -0
- package/dist/projects.d.ts +8 -0
- package/dist/prompts/conversational.md +8 -0
- package/dist/prompts/enrichment.md +3 -0
- package/dist/prompts/faithfulness.md +1 -0
- package/dist/prompts/grader.md +5 -0
- package/dist/prompts/listwise.md +3 -0
- package/dist/prompts/precision.md +1 -0
- package/dist/prompts/reflect.md +1 -0
- package/dist/prompts/relevancy.md +1 -0
- package/dist/prompts/research.md +10 -0
- package/dist/prompts/section-summary.md +5 -0
- package/dist/prompts/summarize.md +1 -0
- package/dist/prompts/system.md +18 -0
- package/dist/prompts/understanding.md +17 -0
- package/dist/quota.d.ts +13 -0
- package/dist/reflect.d.ts +5 -0
- package/dist/refs.d.ts +40 -0
- package/dist/refusal.d.ts +9 -0
- package/dist/refusal.js +9 -0
- package/dist/requestScope.d.ts +26 -0
- package/dist/requestScope.js +10 -0
- package/dist/research.d.ts +8 -0
- package/dist/search.d.ts +4 -0
- package/dist/selfquery.d.ts +7 -0
- package/dist/session.d.ts +1 -0
- package/dist/share.d.ts +2 -0
- package/dist/structural.d.ts +27 -0
- package/dist/tablecontext.d.ts +11 -0
- package/dist/understand.d.ts +11 -0
- package/dist/understandContract.d.ts +29 -0
- package/dist/verdict.d.ts +24 -0
- package/docs/API.md +451 -0
- package/docs/ARCHITECTURE.md +302 -0
- package/docs/AUDIT-2026-08-24.md +71 -0
- package/docs/CONTRIBUTOR-AUDIT-2026-08-25.md +147 -0
- package/docs/INGEST-ARCHITECTURE.md +158 -0
- package/docs/MCP.md +92 -0
- package/docs/METANORMA-AI-SERIALIZATION.md +247 -0
- package/docs/MKO-EXPORT-PIPELINE.md +147 -0
- package/docs/REDESIGN-NORMATIVE-RAG-ETSI.md +485 -0
- package/docs/RESEARCH-SOTA-2026.md +243 -0
- package/docs/ROADMAP-SOTA.md +130 -0
- package/docs/SOTA-STAGE-SPECS.md +509 -0
- package/docs/annealment/F1-verdict.md +27 -0
- package/docs/annealment/F10-notes.md +19 -0
- package/docs/annealment/F11-composition.md +17 -0
- package/docs/annealment/F12-passport.md +17 -0
- package/docs/annealment/F2-counterfactual.md +20 -0
- package/docs/annealment/F3-absence.md +21 -0
- package/docs/annealment/F4-instance.md +18 -0
- package/docs/annealment/F5-workflow.md +21 -0
- package/docs/annealment/F6-impact.md +21 -0
- package/docs/annealment/F7-editions.md +18 -0
- package/docs/annealment/F8-selfverify.md +19 -0
- package/docs/annealment/F9-projection-qa.md +17 -0
- package/docs/annealment/L0-locate.md +19 -0
- package/docs/annealment/L1-extract.md +18 -0
- package/docs/annealment/L2-nomenclature.md +22 -0
- package/docs/annealment/L3-geometry.md +23 -0
- package/docs/annealment/L4-composition.md +21 -0
- package/docs/annealment/L5-cross-standard.md +20 -0
- package/docs/annealment/L6-diachrony.md +21 -0
- package/docs/annealment/L7-perception.md +20 -0
- package/docs/annealment/L8-computation.md +22 -0
- package/docs/annealment/L9-instance-process.md +23 -0
- package/docs/annealment/README.md +10 -0
- package/docs/guidelines-metanorma-ai-programme.md +279 -0
- package/docs/identity-onboarding-rag.md +65 -0
- package/docs/identity-service.md +219 -0
- package/docs/knowledge-annealment.md +273 -0
- package/docs/konneal-extraction-plan.md +481 -0
- package/docs/metanorma-for-ai.md +270 -0
- package/docs/mirror-plan.md +36 -0
- package/docs/multi-sdo-architecture.md +191 -0
- package/docs/paper-annealment-comparison.md +259 -0
- package/docs/paper-assets/architecture.svg +94 -0
- package/docs/paper-assets/contract-v2.svg +94 -0
- package/docs/paper-assets/mko-ingest.svg +91 -0
- package/docs/paper-oiml-bulletin.md +402 -0
- package/docs/paper-oiml-bulletin.mdx +419 -0
- package/docs/product-branding-options.md +172 -0
- package/docs/projects-design.md +88 -0
- package/docs/sota-mechanisms.md +184 -0
- package/docs/spec-api.md +77 -0
- package/docs/spec-pipeline.md +126 -0
- package/docs/vector-adapter.md +88 -0
- package/package.json +70 -0
- package/profile/corpora.yaml +5 -0
- package/profile/datasets.yaml +14 -0
- package/profile/prompts.yaml +5 -0
- package/profile/publisher.yaml +17 -0
- package/profile/retrieval.yaml +1 -0
- package/profile/sources.yaml +5 -0
- package/profile/ui.yaml +7 -0
- package/scripts/gen_profile.mjs +33 -0
- package/workers/shared/ai.ts +21 -0
- package/workers/shared/auth.ts +16 -0
- package/workers/shared/chunk.ts +108 -0
- package/workers/shared/oidc.ts +312 -0
- package/workers/shared/router.ts +45 -0
- package/workers/shared/session.ts +104 -0
- package/workers/worker_internal/src/index.ts +157 -0
- package/workers/worker_internal/tsconfig.json +15 -0
- package/workers/worker_internal/wrangler.toml +32 -0
- package/workers/worker_mcp/src/index.ts +175 -0
- package/workers/worker_mcp/tsconfig.json +13 -0
- package/workers/worker_mcp/wrangler.toml +18 -0
- package/workers/worker_public/migrations/0002_conversations.sql +22 -0
- package/workers/worker_public/migrations/0003_shared_conversations.sql +9 -0
- package/workers/worker_public/migrations/0004_graph.sql +16 -0
- package/workers/worker_public/migrations/0005_documents.sql +19 -0
- package/workers/worker_public/migrations/0006_conversation_entities.sql +11 -0
- package/workers/worker_public/migrations/0007_chunks_fts.sql +43 -0
- package/workers/worker_public/migrations/0008_unit_payloads.sql +15 -0
- package/workers/worker_public/migrations/0009_chunks_unit.sql +8 -0
- package/workers/worker_public/migrations/0009_message_context.sql +7 -0
- package/workers/worker_public/migrations/0010_model_nodes.sql +33 -0
- package/workers/worker_public/migrations/0011_schema_union.sql +38 -0
- package/workers/worker_public/migrations/0012_memories.sql +15 -0
- package/workers/worker_public/migrations/0013_projects.sql +21 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.d.ts +16306 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2021-11-03/index.ts +16261 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.d.ts +16373 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-01-31/index.ts +16328 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.d.ts +16382 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-03-21/index.ts +16337 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.d.ts +16383 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-08-04/index.ts +16338 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.d.ts +16403 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-10-31/index.ts +16358 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.d.ts +16408 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2022-11-30/index.ts +16363 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.d.ts +16414 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-03-01/index.ts +16369 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.d.ts +16414 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/2023-07-01/index.ts +16369 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/README.md +135 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/entrypoints.svg +53 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.d.ts +17095 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/experimental/index.ts +17050 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/index.d.ts +16306 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/index.ts +16261 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.d.ts +16447 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/latest/index.ts +16402 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.d.ts +16306 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/oldest/index.ts +16261 -0
- package/workers/worker_public/node_modules/@cloudflare/workers-types/package.json +11 -0
- package/workers/worker_public/package.json +13 -0
- package/workers/worker_public/prompts/conversational.md +8 -0
- package/workers/worker_public/prompts/enrichment.md +3 -0
- package/workers/worker_public/prompts/faithfulness.md +1 -0
- package/workers/worker_public/prompts/grader.md +5 -0
- package/workers/worker_public/prompts/listwise.md +3 -0
- package/workers/worker_public/prompts/precision.md +1 -0
- package/workers/worker_public/prompts/reflect.md +1 -0
- package/workers/worker_public/prompts/relevancy.md +1 -0
- package/workers/worker_public/prompts/research.md +10 -0
- package/workers/worker_public/prompts/section-summary.md +5 -0
- package/workers/worker_public/prompts/summarize.md +1 -0
- package/workers/worker_public/prompts/system.md +18 -0
- package/workers/worker_public/prompts/understanding.md +17 -0
- package/workers/worker_public/public/app.js +166 -0
- package/workers/worker_public/public/index.html +48 -0
- package/workers/worker_public/public/style.css +147 -0
- package/workers/worker_public/schema.sql +248 -0
- package/workers/worker_public/src/admin.ts +358 -0
- package/workers/worker_public/src/ai.ts +71 -0
- package/workers/worker_public/src/anchors.ts +41 -0
- package/workers/worker_public/src/answercache.ts +72 -0
- package/workers/worker_public/src/ask.ts +1094 -0
- package/workers/worker_public/src/auth.ts +252 -0
- package/workers/worker_public/src/bubble.ts +111 -0
- package/workers/worker_public/src/completion.ts +75 -0
- package/workers/worker_public/src/config.ts +238 -0
- package/workers/worker_public/src/context.ts +238 -0
- package/workers/worker_public/src/conversations.ts +162 -0
- package/workers/worker_public/src/drafts.ts +497 -0
- package/workers/worker_public/src/env.ts +90 -0
- package/workers/worker_public/src/faithfulness.ts +63 -0
- package/workers/worker_public/src/grader.ts +89 -0
- package/workers/worker_public/src/graph.ts +63 -0
- package/workers/worker_public/src/hybrid.ts +77 -0
- package/workers/worker_public/src/index.ts +441 -0
- package/workers/worker_public/src/internal_gateway.ts +41 -0
- package/workers/worker_public/src/lexical.ts +86 -0
- package/workers/worker_public/src/lib/hit.ts +4 -0
- package/workers/worker_public/src/lib/http.ts +83 -0
- package/workers/worker_public/src/lib/router.ts +4 -0
- package/workers/worker_public/src/livedata.ts +334 -0
- package/workers/worker_public/src/memories.ts +81 -0
- package/workers/worker_public/src/modelplane.ts +213 -0
- package/workers/worker_public/src/oidc.ts +333 -0
- package/workers/worker_public/src/pipeline.ts +377 -0
- package/workers/worker_public/src/ports/blobs.ts +7 -0
- package/workers/worker_public/src/ports/cloudflare/adapters.ts +177 -0
- package/workers/worker_public/src/ports/kv.ts +8 -0
- package/workers/worker_public/src/ports/model.ts +28 -0
- package/workers/worker_public/src/ports/runtime.ts +13 -0
- package/workers/worker_public/src/ports/store.ts +20 -0
- package/workers/worker_public/src/ports/vector.ts +26 -0
- package/workers/worker_public/src/profile.gen.ts +101 -0
- package/workers/worker_public/src/profile.ts +16 -0
- package/workers/worker_public/src/projects.ts +108 -0
- package/workers/worker_public/src/prompts.d.ts +6 -0
- package/workers/worker_public/src/quota.ts +54 -0
- package/workers/worker_public/src/reflect.ts +67 -0
- package/workers/worker_public/src/refs.ts +107 -0
- package/workers/worker_public/src/refusal.ts +65 -0
- package/workers/worker_public/src/requestScope.ts +71 -0
- package/workers/worker_public/src/research.ts +126 -0
- package/workers/worker_public/src/search.ts +56 -0
- package/workers/worker_public/src/selfquery.ts +25 -0
- package/workers/worker_public/src/session.ts +4 -0
- package/workers/worker_public/src/share.ts +53 -0
- package/workers/worker_public/src/stages/conceptGraph.ts +68 -0
- package/workers/worker_public/src/stages/conceptSteer.ts +39 -0
- package/workers/worker_public/src/stages/corpusScope.ts +25 -0
- package/workers/worker_public/src/stages/dedup.ts +10 -0
- package/workers/worker_public/src/stages/dense.ts +73 -0
- package/workers/worker_public/src/stages/diversity.ts +33 -0
- package/workers/worker_public/src/stages/editionCover.ts +63 -0
- package/workers/worker_public/src/stages/editionSteer.ts +88 -0
- package/workers/worker_public/src/stages/familyBoost.ts +22 -0
- package/workers/worker_public/src/stages/federate.ts +22 -0
- package/workers/worker_public/src/stages/glossary.ts +65 -0
- package/workers/worker_public/src/stages/graphLane.ts +31 -0
- package/workers/worker_public/src/stages/hyde.ts +29 -0
- package/workers/worker_public/src/stages/index.ts +69 -0
- package/workers/worker_public/src/stages/lexicalUnion.ts +21 -0
- package/workers/worker_public/src/stages/multiQuery.ts +57 -0
- package/workers/worker_public/src/stages/overviewDemote.ts +14 -0
- package/workers/worker_public/src/stages/poolOpen.ts +10 -0
- package/workers/worker_public/src/stages/propagate.ts +15 -0
- package/workers/worker_public/src/stages/rerank.ts +47 -0
- package/workers/worker_public/src/stages/seal.ts +16 -0
- package/workers/worker_public/src/stages/sectionDescent.ts +61 -0
- package/workers/worker_public/src/stages/stdRefNudge.ts +35 -0
- package/workers/worker_public/src/stages/subQuery.ts +42 -0
- package/workers/worker_public/src/stages/termNudge.ts +24 -0
- package/workers/worker_public/src/stages/typedPin.ts +131 -0
- package/workers/worker_public/src/stages/types.ts +112 -0
- package/workers/worker_public/src/stages/windowFloor.ts +23 -0
- package/workers/worker_public/src/structural.ts +171 -0
- package/workers/worker_public/src/tablecontext.ts +41 -0
- package/workers/worker_public/src/understand.ts +72 -0
- package/workers/worker_public/src/understandContract.ts +67 -0
- package/workers/worker_public/src/verdict.ts +255 -0
- package/workers/worker_public/tsconfig.json +18 -0
- package/workers/worker_public/wrangler.toml +104 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
# F5 — Certification workflow statefulness
|
|
2
|
+
|
|
3
|
+
**Powering objects:** `evaluation/processes.yaml` (layer-composed
|
|
4
|
+
process tree; `validate_provision` binds the campaign to requirement
|
|
5
|
+
URNs like /req/metrological/mpe), `evaluation/gateways.yaml`,
|
|
6
|
+
`execution/` (test-report forms, checklist), entities/performance-test-
|
|
7
|
+
evaluations, evaluation/sample-selection-rules.
|
|
8
|
+
|
|
9
|
+
**The win:** the assistant knows WHERE an evaluation stands and WHAT
|
|
10
|
+
GATES WHAT — not as prose but as traversable state:
|
|
11
|
+
- *"what must be true before the creep test?"* → the sequence rule
|
|
12
|
+
(MDLO baseline first) + the gateway that checks it;
|
|
13
|
+
- *"which requirements does this test campaign validate?"* → the
|
|
14
|
+
validate_provision URNs, exhaustively;
|
|
15
|
+
- agentic follow-through: checklist items flip state as evidence
|
|
16
|
+
lands; the assistant can say "3 of 62 tests outstanding, blocked on
|
|
17
|
+
the humidity chamber" — a document lane can quote the procedure but
|
|
18
|
+
cannot know the POSITION.
|
|
19
|
+
|
|
20
|
+
**Why documents can't follow:** procedures in prose are instructions;
|
|
21
|
+
here they are a state machine with recorded progress.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
# F6 — Impact analysis (committee tooling)
|
|
2
|
+
|
|
3
|
+
**Powering objects:** the binding graph — requirements ↔ conformance
|
|
4
|
+
tests ↔ formulas (`formulas-used.yaml` binds MDLO to conversion_factor_f,
|
|
5
|
+
E_L, E_R, C_M), tables (mpe_tiers feeds lookupMPE), aspects
|
|
6
|
+
(requirement-binding-targets), execution forms (E_R's report field).
|
|
7
|
+
|
|
8
|
+
**The win:** DRAFTING becomes queryable. *"If the committee raises the
|
|
9
|
+
class C limit_factor, what changes?"* → traverse the graph: the MPE
|
|
10
|
+
verdicts for class C, the tests whose verdicts derive from lookupMPE,
|
|
11
|
+
the R 60-3 report fields that record them, the instances already
|
|
12
|
+
evaluated. The answer is an IMPACT SET — machine-derived, exhaustive
|
|
13
|
+
(F3's completeness), with every hop's clause anchor.
|
|
14
|
+
|
|
15
|
+
**Why it matters:** this is the first capability whose customer is the
|
|
16
|
+
STANDARDS DEVELOPER, not the reader — the model pays for its own
|
|
17
|
+
maintenance by making revision risk computable.
|
|
18
|
+
|
|
19
|
+
**Why documents can't follow:** the binding "this test's verdict
|
|
20
|
+
consumes this table's factor" exists nowhere in any single document —
|
|
21
|
+
it is model knowledge spanning three parts.
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
# F7 — Semantic edition diffs + temporal jurisdiction
|
|
2
|
+
|
|
3
|
+
**Powering objects:** editions/2017/*.prl — a FULL prior-edition model
|
|
4
|
+
package (not a separate PDF) — plus edition lifecycle (status,
|
|
5
|
+
supersedes, validity.from).
|
|
6
|
+
|
|
7
|
+
**The win, beyond L6:** diffs become SEMANTIC, not textual:
|
|
8
|
+
*"what changed in creep between 2017 and 2021?"* → diff the two MODELS:
|
|
9
|
+
the constraint set (new OCL checks), attribute definitions (added
|
|
10
|
+
fields, moved scopes), test sequences (reordered steps), tier tables
|
|
11
|
+
(changed breakpoints) — each delta with both editions' clause anchors.
|
|
12
|
+
Text diffs of two PDFs produce noise; model diffs produce CHANGE RECORDS.
|
|
13
|
+
|
|
14
|
+
**Temporal jurisdiction:** validity windows answer retro-questions —
|
|
15
|
+
*"an evaluation performed 2019-11: which edition governed it, and is
|
|
16
|
+
its verdict still valid today?"* → 2017 governs; superseded-by chain
|
|
17
|
+
says whether re-evaluation is required. Legal-metrology gold: the
|
|
18
|
+
corpus knows TIME, not just content.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# F8 — Self-verification as a query (meta-grounding)
|
|
2
|
+
|
|
3
|
+
**Powering objects:** the model-linker rule set (quantity-coherence,
|
|
4
|
+
requirement-binding-targets) and the allowlist ledger whose history
|
|
5
|
+
reads "burned to empty — every formerly allowlisted site was fixed at
|
|
6
|
+
the data layer": the model is MACHINE-CHECKED consistent.
|
|
7
|
+
|
|
8
|
+
**The win:** for model-grounded claims, execution replaces the judge.
|
|
9
|
+
When the D lane answers "the class C limit factor is 0.35", it can
|
|
10
|
+
VERIFY that claim against the model (the value is read from the typed
|
|
11
|
+
table, not generated) — and the user can ask *"is this corpus
|
|
12
|
+
internally consistent?"* and receive the linker verdict. The
|
|
13
|
+
faithfulness judge (an LLM scoring prose) remains for synthesis; for
|
|
14
|
+
model-grounded facts, the check is arithmetic.
|
|
15
|
+
|
|
16
|
+
**Why it matters:** it closes the trust loop — the same mechanism that
|
|
17
|
+
answers also proves. The strongest possible grounding story for the
|
|
18
|
+
paper: not "the model was cited" but "the model was EXECUTED, and the
|
|
19
|
+
execution is checkable".
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
# F9 — Document-as-projection verification
|
|
2
|
+
|
|
3
|
+
**Powering objects:** documents/{1,2,3,a}/document.presentation.xml —
|
|
4
|
+
the model CARRIES its own rendered document projections per part.
|
|
5
|
+
|
|
6
|
+
**The win:** publishing QA becomes a query. *"Does the published
|
|
7
|
+
R 60-1 Table 4 match the model's mpe_tiers?"* → parse the projection's
|
|
8
|
+
table, compare cell-by-cell against the model's typed table → a
|
|
9
|
+
diff-and-verdict. The direction inverts: instead of the document being
|
|
10
|
+
the source the model was derived from, the MODEL becomes the source and
|
|
11
|
+
the document a VIEW to be verified. Forward: model → document
|
|
12
|
+
GENERATION (the view rendered FROM truth, drift structurally
|
|
13
|
+
impossible).
|
|
14
|
+
|
|
15
|
+
**Why it matters:** every standards body's dirty secret is render drift
|
|
16
|
+
(errata). Projection QA makes errata a computable diff. And it is the
|
|
17
|
+
on-ramp to the fully closed loop: author → model → verified document.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# L0 — LOCATE (primitives P1–P2)
|
|
2
|
+
|
|
3
|
+
**Definition:** find WHERE the corpus addresses a topic; return the clause
|
|
4
|
+
anchor. No value, no synthesis — presence and position only.
|
|
5
|
+
|
|
6
|
+
**Why it's the floor:** it isolates retrieval from understanding. A lane
|
|
7
|
+
failing L0 is mis-built (chunking/embedding bug), not "less annealed" —
|
|
8
|
+
the floor validates the harness before the ladder means anything.
|
|
9
|
+
|
|
10
|
+
**Example** — *"Where does R 60 address creep?"*
|
|
11
|
+
Witness: citation anchor ∈ {R 60-1 §3.x (definition), R 60-2 §2.11.x
|
|
12
|
+
(test)} — either is a pass; the anchor must EXIST and be normative.
|
|
13
|
+
Lane expectations: A/B/C/D/E all pass (sanity). Primmel additionally
|
|
14
|
+
binds the term to its `behaviors: [creep]` entry with stimulus/response
|
|
15
|
+
and the `load-cell` entity — same anchor, richer provenance.
|
|
16
|
+
|
|
17
|
+
**Harness notes:** 6 questions spanning definition-anchors, test-anchors
|
|
18
|
+
and annex-anchors; paraphrases use colloquial synonyms ("time drift under
|
|
19
|
+
load") so lexical lanes can't win by string match alone.
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
# L1 — EXTRACT (P1–P2)
|
|
2
|
+
|
|
3
|
+
**Definition:** return a normative value VERBATIM from prose — no
|
|
4
|
+
disambiguation beyond finding the sentence.
|
|
5
|
+
|
|
6
|
+
**Example** — *"Quote the note that governs which p_LC value the creep
|
|
7
|
+
MPE must use."*
|
|
8
|
+
Witness (real, from notes.yaml): "p_LC = 0.7" AND the qualifier
|
|
9
|
+
"regardless of any value declared by the manufacturer"
|
|
10
|
+
(note_creep_plc_always_0_7 — see also F10: this note is an OVERRIDE in
|
|
11
|
+
the model, prose everywhere else).
|
|
12
|
+
|
|
13
|
+
**Example** — *"What validity date does the current R 60 edition carry?"*
|
|
14
|
+
Witness: "2021-01-01" (standard.yaml edition.validity.from — P2
|
|
15
|
+
provenance makes this trivial in D, findable in C via the registry).
|
|
16
|
+
|
|
17
|
+
Lane expectations: all pass; separation is noise-level. L1 exists to
|
|
18
|
+
calibrate the harness's regex discipline before structure matters.
|
|
@@ -0,0 +1,22 @@
|
|
|
1
|
+
# L2 — NOMENCLATURE (P4)
|
|
2
|
+
|
|
3
|
+
**Definition:** resolve the user's vocabulary to the corpus's DEFINED
|
|
4
|
+
terms, return the authoritative definition and its vocabulary anchor.
|
|
5
|
+
|
|
6
|
+
**Sub-probes:**
|
|
7
|
+
- **a) corpus-local** — *"my load cell keeps drifting over months — what
|
|
8
|
+
does R 60 call that?"* → durability; witness: the definition text
|
|
9
|
+
("ability of a measuring instrument to maintain its performance
|
|
10
|
+
characteristics over a period of use") + anchor R 60-1 §3.1 with VIML
|
|
11
|
+
5.15 register (terminology.yaml: vocab_ref viml-2022 5.15).
|
|
12
|
+
- **b) cross-register** — *"which VIM/VIML concept does 'durability'
|
|
13
|
+
come from?"* → viml-2022 clause 5.15. Only lanes carrying vocab_ref
|
|
14
|
+
answer the REGISTER part (D natively; C via the glossary lane's
|
|
15
|
+
concept sources).
|
|
16
|
+
- **c) multilingual** — term spellings per language (Primmel
|
|
17
|
+
eng-Latn entries; multilingual spellings land with more languages).
|
|
18
|
+
|
|
19
|
+
Lane expectations: A/B luck-dependent (embedding proximity); C strong at
|
|
20
|
+
(a); D strong a–c. The colloquial gap ("drift"≠"durability") is the
|
|
21
|
+
measurement: without P4 the lanes must bridge vocabulary by embeddings
|
|
22
|
+
alone.
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
# L3 — TABULAR GEOMETRY (P3 + P8)
|
|
2
|
+
|
|
3
|
+
**Definition:** a value whose retrieval requires row×column
|
|
4
|
+
disambiguation; delivery requires the TYPED BLOCK (the producer's
|
|
5
|
+
object, never re-typed prose).
|
|
6
|
+
|
|
7
|
+
**Sub-probes:**
|
|
8
|
+
- **a) cell lookup** — *"minimum n_LC for class B"* → R 60-1 Table 3:
|
|
9
|
+
row B, lower-limit column. Witness: 5000 (family values: A 50000, B
|
|
10
|
+
5000, C 500, D 100 — verified in the corpus) + block artifact
|
|
11
|
+
`block:"table"`.
|
|
12
|
+
- **b) unit-aware cell** — *"what quantity does `load_min` carry in the
|
|
13
|
+
MPE tier table, and in what unit?"* → tables.yaml column
|
|
14
|
+
`{name: load_min, type: number, unit: v}` → "load in verification
|
|
15
|
+
intervals (v)". Only D answers the UNIT question structurally (P8);
|
|
16
|
+
C's payload has unit strings (metanorma-document#55 GAP-1).
|
|
17
|
+
- **c) derived cell** — *"the MPE limit factor for class C at a load
|
|
18
|
+
between the tier bounds"* → needs the tier BREAKPOINTS + factor: the
|
|
19
|
+
lookup operation (formula lookupMPE) — transitional to L8.
|
|
20
|
+
|
|
21
|
+
Lane expectations: A ≤40% (flattened geometry), B partial (pipes as
|
|
22
|
+
noise), C ≥90% at (a), D ≥90% a–c. **First hard separator; the
|
|
23
|
+
side-by-side demo on the site.**
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
# L4 — INTRA-DOCUMENT COMPOSITION (P2 + P5)
|
|
2
|
+
|
|
3
|
+
**Definition:** an answer that JOINS ≥2 clauses (often + a table) and
|
|
4
|
+
returns both anchors plus the derived relation.
|
|
5
|
+
|
|
6
|
+
**Sub-probes:**
|
|
7
|
+
- **a) intra-document** — *"initial-verification MPE vs in-service
|
|
8
|
+
limits for a class C load cell, and how they scale across the
|
|
9
|
+
range"* → R 60-1 MPE table (tier breakpoints, limit_factor) + the
|
|
10
|
+
applicability clauses. Witness: both anchors + the scaling statement
|
|
11
|
+
expressed in v-units.
|
|
12
|
+
- **b) cross-part** — *"which R 60-2 test validates the R 60-1
|
|
13
|
+
repeatability requirement, and where is its result recorded in
|
|
14
|
+
R 60-3?"* → requirement `/req/metrological/repeatability` (R 60-1) ↔
|
|
15
|
+
test `/conf/metrological-tests/...` (R 60-2) ↔ test-report form
|
|
16
|
+
(R 60-3 §2.1.3 E_R). Witness: the three part-anchors. In D the chain
|
|
17
|
+
is EDGES (formulas_used binds the test to
|
|
18
|
+
`repeatabilityError`); in C it must be co-retrieved.
|
|
19
|
+
|
|
20
|
+
Lane expectations: C/D strong at (a); (b) separates C (partial — needs
|
|
21
|
+
all three parts in the window) from D (edges). A/B partial.
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
# L5 — CROSS-STANDARD LINKING (P5)
|
|
2
|
+
|
|
3
|
+
**Definition:** answers that traverse a reference to ANOTHER standard,
|
|
4
|
+
returning the linked identity and clause.
|
|
5
|
+
|
|
6
|
+
**Sub-probes:**
|
|
7
|
+
- **a) cites-edge** — *"which ISO/IEC test method does R 60-2 invoke for
|
|
8
|
+
humidity?"* → witness: the ISO/IEC document id + clause from the
|
|
9
|
+
cites edges / references registry. ISOLATION: public lanes assert the
|
|
10
|
+
LINK only — ISO/IEC text never enters a public index; content probes
|
|
11
|
+
are member/internal-only, structurally enforced.
|
|
12
|
+
- **b) composition** — *"which CASCO vocabulary governs R 60's
|
|
13
|
+
certification activities, and which aspects of the load cell do its
|
|
14
|
+
requirements bind?"* → `uses: iso-iec-17000` + `iso-iec-17065` + the
|
|
15
|
+
aspects registry (markings, accompanying_document… via
|
|
16
|
+
requirement-binding-targets). Witness: both package ids + ≥1 aspect.
|
|
17
|
+
|
|
18
|
+
Lane expectations: D strong (composition is a first-class relation);
|
|
19
|
+
C partial at (a) (cites edges, no composition); A/B near-zero —
|
|
20
|
+
cross-standard vocabulary rarely co-embeds.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
# L6 — DIACHRONY / EDITIONS (P6)
|
|
2
|
+
|
|
3
|
+
**Definition:** answers that depend on WHICH edition, or on WHAT
|
|
4
|
+
CHANGED between editions.
|
|
5
|
+
|
|
6
|
+
**Sub-probes:**
|
|
7
|
+
- **a) current-edition selection** — *"which R 60 edition applies to a
|
|
8
|
+
type evaluation started this year?"* → 2021 (lifecycle: status
|
|
9
|
+
current, supersedes 2017). Witness: edition + status.
|
|
10
|
+
- **b) delta extraction** — *"what changed in the creep requirements
|
|
11
|
+
between 2017 and 2021?"* → D computes the diff between EDITION
|
|
12
|
+
PACKAGES (editions/2017/*.prl vs current) at the model level — the
|
|
13
|
+
witness is the delta fact (constraint added/changed, limits moved).
|
|
14
|
+
C can align units by hash but has no semantic diff (metanorma-document
|
|
15
|
+
#53 item 5). A/B near-zero.
|
|
16
|
+
- **c) temporal jurisdiction** — *"which edition governed an evaluation
|
|
17
|
+
performed in 2019, and under what validity rule?"* → 2017 (validity
|
|
18
|
+
windows: 2017 until 2021-01-01). Only D (validity.from on editions).
|
|
19
|
+
|
|
20
|
+
Lane expectations: D strong a–c; C partial (a, steering already ships);
|
|
21
|
+
A/B fail b/c. Deepened by frontier F7.
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
# L7 — MULTIMODAL PERCEPTION (P7)
|
|
2
|
+
|
|
3
|
+
**Definition:** answers whose evidence is in PIXELS, not text.
|
|
4
|
+
|
|
5
|
+
**Sub-probes:**
|
|
6
|
+
- **a) caption/asset** — *"what does the R 60-2 test setup figure
|
|
7
|
+
show?"* → figure unit + stored caption/vision description.
|
|
8
|
+
- **b) pixel-only content** — *"what are the three labeled cases in the
|
|
9
|
+
design-shapes example?"* → labels A/B/C that exist ONLY in the
|
|
10
|
+
drawing; witness: the label values (verified: the deployed system
|
|
11
|
+
reads them from pixels — attachFigureImages).
|
|
12
|
+
- **c) user-image grounding** — a nameplate photo → "which accuracy
|
|
13
|
+
class marking does this carry / which OIML R applies?" (image
|
|
14
|
+
questions, shipped). Witness: classification.
|
|
15
|
+
|
|
16
|
+
Lane expectations: C strong a–c (figure lane live + multimodal
|
|
17
|
+
generation); D partial (references figures in sequences —
|
|
18
|
+
`fig-2/fig-3` — asset binding pending); A/B zero (no units, no assets).
|
|
19
|
+
The most DEMO-VISIBLE separator: same question, lane C shows the image
|
|
20
|
+
and reads it, lane A shows prose soup.
|
|
@@ -0,0 +1,22 @@
|
|
|
1
|
+
# L8 — COMPUTATION (P9 + P8)
|
|
2
|
+
|
|
3
|
+
**Definition:** answers the documents DEFINE but never PRINT — values
|
|
4
|
+
that must be computed. Witnesses are COMPUTED OFFLINE from the model
|
|
5
|
+
(deterministic; no judge, no regex luck).
|
|
6
|
+
|
|
7
|
+
**Sub-probes:**
|
|
8
|
+
- **a) pure calculation** — *"the conversion factor f for these
|
|
9
|
+
indications"* → calculations.yaml conversionFactor (inputs:
|
|
10
|
+
avgIndicationAt75pct, indicationAtDmin; R 60-3 2.1.2.4). Witness: the
|
|
11
|
+
computed number.
|
|
12
|
+
- **b) constraint verdict** — *"is D_max = 0.8·E_max acceptable?"* → OCL
|
|
13
|
+
`dead_load_max_geometry` FAILS: d_max must lie in [0.9·E_max, E_max];
|
|
14
|
+
witness: verdict + the recorded violation_meaning ("the type
|
|
15
|
+
evaluation of this load cell is void") — F1's ladder form.
|
|
16
|
+
- **c) unit coherence** — *"are 3000 kgf and 30 kN the same force?"* →
|
|
17
|
+
the unit REGISTER converts (quantity_kind force; conversion to SI);
|
|
18
|
+
witness: the equality verdict. P8 pure.
|
|
19
|
+
|
|
20
|
+
Lane expectations: **D only, by construction** — A/B/C/E cannot compute
|
|
21
|
+
(they can at best quote a formula; E proves it's not prose quality).
|
|
22
|
+
The ladder's most objective rung: ground truth is arithmetic.
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
# L9 — INSTANCE & PROCESS (P10)
|
|
2
|
+
|
|
3
|
+
**Definition:** answers about PARTICULAR things and ORDERED activities —
|
|
4
|
+
knowledge that exists only where the model records instances and
|
|
5
|
+
sequences.
|
|
6
|
+
|
|
7
|
+
**Sub-probes:**
|
|
8
|
+
- **a) process order & consequence** — *"what is contaminated if creep
|
|
9
|
+
runs before MDLO on the same sample?"* → test_sequences.yaml
|
|
10
|
+
`mdlo-creep-dr`: MDLO is the BASELINE (order 1), creep FOLLOWS (order
|
|
11
|
+
2), DR reads against creep's D_max; witness: the ordering verdict +
|
|
12
|
+
"contaminates the baseline". The rationale is ENCODED as sequence
|
|
13
|
+
semantics — prose lanes only pass if the corpus happens to print it
|
|
14
|
+
(containment-checked; the printed form exists in R 60-2 — so the
|
|
15
|
+
probe asks the COUNTERFACTUAL form: "may I run them in my own
|
|
16
|
+
order?" → D answers from the model rule).
|
|
17
|
+
- **b) instance-grounded fact** — a sample instrument profile →
|
|
18
|
+
"what is ITS E_R budget?" — the model evaluates at the instance.
|
|
19
|
+
- **c) execution form** — *"where is E_R recorded in the test report?"*
|
|
20
|
+
→ execution/test-report.yaml + subforms (R 60-3 2.1.3).
|
|
21
|
+
|
|
22
|
+
Lane expectations: **D only.** The deepest binding: the corpus doesn't
|
|
23
|
+
describe the general — it RECORDED the particular and the ORDER.
|
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
# The annealment documentation set
|
|
2
|
+
|
|
3
|
+
One file per ladder rung (L0–L9) and per frontier win (F1–F12). The
|
|
4
|
+
instrument spec with the primitive taxonomy lives in
|
|
5
|
+
../knowledge-annealment.md; each file here ELABORATES its level with
|
|
6
|
+
worked examples, witnesses and lane expectations.
|
|
7
|
+
|
|
8
|
+
Ladder: [L0](L0-locate.md) · [L1](L1-extract.md) · [L2](L2-nomenclature.md) · [L3](L3-geometry.md) · [L4](L4-composition.md) · [L5](L5-cross-standard.md) · [L6](L6-diachrony.md) · [L7](L7-perception.md) · [L8](L8-computation.md) · [L9](L9-instance-process.md)
|
|
9
|
+
|
|
10
|
+
Frontier: [F1](F1-verdict.md) · [F2](F2-counterfactual.md) · [F3](F3-absence.md) · [F4](F4-instance.md) · [F5](F5-workflow.md) · [F6](F6-impact.md) · [F7](F7-editions.md) · [F8](F8-selfverify.md) · [F9](F9-projection-qa.md) · [F10](F10-notes.md) · [F11](F11-composition.md) · [F12](F12-passport.md)
|
|
@@ -0,0 +1,279 @@
|
|
|
1
|
+
# Metanorma AI — Building a Grounded RAG Service from Metanorma Sources
|
|
2
|
+
|
|
3
|
+
*Programme guidelines: how any standards body or Metanorma user builds a
|
|
4
|
+
citation-grounded question-answering service over their own corpus,
|
|
5
|
+
from scratch, using the Metanorma document model as the source of
|
|
6
|
+
truth. Companion to the OIML Bulletin article (2026-08-30). Everything
|
|
7
|
+
below is extracted from a production deployment — the OIML SMART AI
|
|
8
|
+
service — and is reproducible.*
|
|
9
|
+
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
## Who this is for
|
|
13
|
+
|
|
14
|
+
You have documents authored in Metanorma (or convertible to it) and you
|
|
15
|
+
want people to *ask questions* and get *verifiable, cited answers* —
|
|
16
|
+
not a search box over PDFs. The path below assumes no prior RAG
|
|
17
|
+
experience; every stage names the decision, the cost, and the
|
|
18
|
+
measurement that proves it works.
|
|
19
|
+
|
|
20
|
+
## Principles (the short list)
|
|
21
|
+
|
|
22
|
+
1. **Grounding is the product.** Fluency is free; trust is engineered.
|
|
23
|
+
Every claim cites the exact publication, edition and clause. When
|
|
24
|
+
the corpus doesn't contain the answer, the system says so — one
|
|
25
|
+
canonical refusal sentence, never cached.
|
|
26
|
+
2. **Facts from the source model, never from renderings.** Metanorma
|
|
27
|
+
documents are typed object models; consuming the model (via MKO)
|
|
28
|
+
beats scraping HTML or PDF on every axis that matters.
|
|
29
|
+
3. **Input = full fidelity; output = references.** The model *reads*
|
|
30
|
+
tables, equations, figures; it *never re-types* them. It writes a
|
|
31
|
+
symbolic reference; your renderer draws the producer's payload.
|
|
32
|
+
4. **Measure everything, ship nothing unmeasured.** A golden question
|
|
33
|
+
set with witness citations, retrieval metrics over repeated runs,
|
|
34
|
+
and a faithfulness judge that sees the passages the answer was
|
|
35
|
+
actually built from. CI runs the suites; a change without a number
|
|
36
|
+
is a guess.
|
|
37
|
+
5. **Cost-first on the hot path, quality-first on one-time work.**
|
|
38
|
+
Every question pays the serving model; enrichment, captioning and
|
|
39
|
+
evaluation are one-time costs whose quality persists.
|
|
40
|
+
|
|
41
|
+
---
|
|
42
|
+
|
|
43
|
+
## Stage 0 — Decide your corpus and its identity backbone
|
|
44
|
+
|
|
45
|
+
**What:** the documents you will index, and the canonical identifier
|
|
46
|
+
that makes two spellings of the same publication collapse into one.
|
|
47
|
+
|
|
48
|
+
- Establish a registry: publication families → editions → status
|
|
49
|
+
(current / superseded / withdrawn). **Derive status from
|
|
50
|
+
supersession links, never from a status field** — fields lie; edges
|
|
51
|
+
are structure.
|
|
52
|
+
- Surface the gaps: families with no derivable current edition are your
|
|
53
|
+
upstream bibliographic worklist, not something to hide.
|
|
54
|
+
- Record language and edition metadata on every document *before*
|
|
55
|
+
chunking; it becomes the filter your retrieval uses.
|
|
56
|
+
|
|
57
|
+
**Cost:** $0. **Measurement:** registry completeness (families with a
|
|
58
|
+
derivable current edition / total families).
|
|
59
|
+
|
|
60
|
+
## Stage 1 — Ingest the Metanorma model (MKO), not renderings
|
|
61
|
+
|
|
62
|
+
**What:** export each document as a Metanorma Knowledge Objects bundle
|
|
63
|
+
(MN 116): typed units (clauses, tables, terms, equations, figures,
|
|
64
|
+
requirements), a section graph (part_of, cites, defines edges), a
|
|
65
|
+
native glossary, and native bibliographic objects.
|
|
66
|
+
|
|
67
|
+
```
|
|
68
|
+
metanorma compile document.adoc -x mko # (or the Ruby exporter today)
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
**Key decisions:**
|
|
72
|
+
- Chunks follow **clause boundaries** — the unit a human cites — never
|
|
73
|
+
token windows. Tables, equations and figures are **atomic units**
|
|
74
|
+
(never split; cite their parent clause).
|
|
75
|
+
- Each chunk carries: docidentifier, edition, language, clause anchor,
|
|
76
|
+
status, and its **unit id** (stable content hash) for incremental
|
|
77
|
+
re-ingest.
|
|
78
|
+
- Write a **data-quality report** on every build: counts by type,
|
|
79
|
+
missing anchors, oversized payloads. Serving keeps score; upstream
|
|
80
|
+
fixes follow.
|
|
81
|
+
|
|
82
|
+
**Cost:** $0 (local compute). **Measurement:** unit counts per document
|
|
83
|
+
against a hand-check; zero empty chunks for preface-only documents
|
|
84
|
+
(a real producer bug we hit).
|
|
85
|
+
|
|
86
|
+
## Stage 2 — Contextual enrichment (the highest-ROI one-time spend)
|
|
87
|
+
|
|
88
|
+
**What:** a one-time pass where a strong model writes a 2–3 sentence
|
|
89
|
+
preamble situating each chunk in its document ("This clause defines the
|
|
90
|
+
accuracy-class limits for load cells in R 60-1:2017 §4.1.2"), then you
|
|
91
|
+
embed **preamble + chunk text** and index that.
|
|
92
|
+
|
|
93
|
+
- Cache preambles by chunk id (KV or object store); unchanged chunks
|
|
94
|
+
are free on re-runs. Invalidate by content hash.
|
|
95
|
+
- Budget: ~260 in / ~210 out tokens per chunk. At our models that is
|
|
96
|
+
≈$0.0017 per chunk; a 40,000-chunk corpus is ≈$70 one-time.
|
|
97
|
+
- Expect a measurable retrieval lift and, more visibly, the recovery of
|
|
98
|
+
paraphrased questions that share no vocabulary with the corpus.
|
|
99
|
+
|
|
100
|
+
**Cost:** ≈$0.002 per chunk, one-time. **Measurement:** retrieval
|
|
101
|
+
recall@5 before/after on a paraphrase probe set.
|
|
102
|
+
|
|
103
|
+
## Stage 3 — Hybrid retrieval (dense + full-corpus lexical)
|
|
104
|
+
|
|
105
|
+
**What:** two retrieval lanes run in parallel on every question:
|
|
106
|
+
|
|
107
|
+
1. **Dense** — the question is embedded (same embedding model as the
|
|
108
|
+
index; this is mandatory) and queried against the vector index.
|
|
109
|
+
2. **Lexical** — a full-corpus BM25/FTS index over the *enriched* chunk
|
|
110
|
+
text, catching exact identifiers, part numbers and defined terms
|
|
111
|
+
that dense similarity misses.
|
|
112
|
+
|
|
113
|
+
Fuse with reciprocal rank fusion (RRF, k=60). Then a cross-encoder
|
|
114
|
+
reranks the fused candidates; optionally a stronger listwise pass for
|
|
115
|
+
complex questions.
|
|
116
|
+
|
|
117
|
+
**The lesson that cost us a baseline:** lexical retrieval must scan the
|
|
118
|
+
*whole corpus* as a first stage. Re-scoring only what dense already
|
|
119
|
+
found cannot recover what dense missed. On our corpus this single
|
|
120
|
+
change lifted recall@5 from 86% to 95%.
|
|
121
|
+
|
|
122
|
+
**Cost:** embedding ≈$0.00001 per query; lexical ≈free (D1/FTS5).
|
|
123
|
+
**Measurement:** recall@5, AP@5, MRR@5 on the golden set, **mean over
|
|
124
|
+
≥3 runs** (single runs are noise).
|
|
125
|
+
|
|
126
|
+
## Stage 4 — Query understanding (LLM-decided, zero rules)
|
|
127
|
+
|
|
128
|
+
**What:** a small model reads the question first: language, named
|
|
129
|
+
publication, edition, defined terms, complexity, whether it's
|
|
130
|
+
conversational. There are no keyword rules and no regex on user input —
|
|
131
|
+
the same model judges every question, in every language.
|
|
132
|
+
|
|
133
|
+
- Its output drives **metadata filters** (doc number, edition) on the
|
|
134
|
+
dense query — the exact-identifier query class embeddings handle
|
|
135
|
+
poorly.
|
|
136
|
+
- A **terminology graph** (defined term → defining documents) resolves
|
|
137
|
+
colloquial phrasing to the corpus's own vocabulary.
|
|
138
|
+
- Run the understanding call **in parallel** with the query embedding;
|
|
139
|
+
they have no dependency when no filter is emitted.
|
|
140
|
+
|
|
141
|
+
**Cost:** ≈$0.0002 per question. **Measurement:** filter precision
|
|
142
|
+
(named publication → correct doc_number).
|
|
143
|
+
|
|
144
|
+
## Stage 5 — The answer contract
|
|
145
|
+
|
|
146
|
+
**What:** the model generates from the retrieved passages only, under a
|
|
147
|
+
data-file prompt contract:
|
|
148
|
+
|
|
149
|
+
- **Inline citations** on every claim: `[OIML R 60-1:2021 §4.4.2]`
|
|
150
|
+
- **Verbatim quote anchors** for normative values — the quoted phrase
|
|
151
|
+
must exist, character-for-character, in a cited passage
|
|
152
|
+
- **Symbolic unit references** (`[[u:table-1]]`) when presenting a
|
|
153
|
+
whole table/equation/figure — the model *points*, your renderer
|
|
154
|
+
draws the producer's payload
|
|
155
|
+
- **One canonical refusal sentence** for off-corpus questions
|
|
156
|
+
|
|
157
|
+
**Verification (all mechanical, all post-generation):**
|
|
158
|
+
|
|
159
|
+
| Check | Method | On failure |
|
|
160
|
+
|---|---|---|
|
|
161
|
+
| Quote anchors | deterministic text match against used passages | one corrective regeneration; never cache the unverified answer |
|
|
162
|
+
| Unit references | every `[[u:…]]` must exist in the used passages | dropped, never rendered |
|
|
163
|
+
| Table retyping | markdown table in output while a typed unit was available | corrective regeneration with a targeted note |
|
|
164
|
+
| Faithfulness | independent judge scores claims vs the passages used | answer withheld or regenerated |
|
|
165
|
+
|
|
166
|
+
The crucial correctness rule: **the faithfulness judge sees the
|
|
167
|
+
passages the answer was actually built from** (return them in the
|
|
168
|
+
response), not a fresh retrieval — otherwise you are measuring a
|
|
169
|
+
different question's evidence.
|
|
170
|
+
|
|
171
|
+
**Cost:** judge ≈$0.0003 per answer. **Measurement:** faithfulness
|
|
172
|
+
score distribution; corrective-regen rate; zero cached violations.
|
|
173
|
+
|
|
174
|
+
## Stage 6 — Serving, caching, and cost control
|
|
175
|
+
|
|
176
|
+
- **Exact answer cache** (keyed by query hash + index version) and a
|
|
177
|
+
**semantic cache** (leading-dimension signature, cosine ≥0.97) serve
|
|
178
|
+
near-duplicate questions instantly. Refusals and conversational turns
|
|
179
|
+
are never cached. `fresh=true` bypasses *every* cache (we shipped a
|
|
180
|
+
bug where it missed the semantic one — test this).
|
|
181
|
+
- **Cache versioning:** bump the version on every retrieval or prompt
|
|
182
|
+
change; old answers expire by key.
|
|
183
|
+
- **Model policy:** all tiers can share one good model if it's cheap
|
|
184
|
+
enough (ours: ≈$0.001/answer). Hot-path understanding stays on a
|
|
185
|
+
smaller model; only the *answer* needs the quality.
|
|
186
|
+
|
|
187
|
+
## Stage 7 — Evaluation as a CI gate
|
|
188
|
+
|
|
189
|
+
Build these suites **before** you need them:
|
|
190
|
+
|
|
191
|
+
1. **Golden set** — 20–30 questions with expected citations (regex over
|
|
192
|
+
docidentifier) and witness values for table questions. Run on every
|
|
193
|
+
change; report R/AP/MRR means over ≥3 runs.
|
|
194
|
+
2. **Paraphrase probes** — the same question asked 8 ways must hit the
|
|
195
|
+
same sources.
|
|
196
|
+
3. **End-to-end behavioural cases** — refusal correctness, language
|
|
197
|
+
fidelity, edition steering, block rendering, auth tiers.
|
|
198
|
+
4. **Faithfulness battery** — judged scores over the golden set,
|
|
199
|
+
re-run after any generation-model change.
|
|
200
|
+
|
|
201
|
+
Our suites caught: a corpus lane invisible to filtered retrieval (an
|
|
202
|
+
identity-parse bug), a mangled CSS rule that blanked dark mode, a cache
|
|
203
|
+
path that served stale answers to regeneration requests, and a
|
|
204
|
+
diversity cap that structurally evicted typed tables. None were visible
|
|
205
|
+
in code review; all were visible in a suite.
|
|
206
|
+
|
|
207
|
+
## Stage 8 — The typed-block pipeline (tables, equations, figures)
|
|
208
|
+
|
|
209
|
+
This is the stage that separates a chatbot from a *standards*
|
|
210
|
+
assistant:
|
|
211
|
+
|
|
212
|
+
1. **Typed payloads in a serving store** (D1/KV table: unit_id →
|
|
213
|
+
payload), written at ingest, validated against the MN 116 schemas.
|
|
214
|
+
2. **Passage headers declare their units** so the model can reference
|
|
215
|
+
them: `[3] OIML R 60-1:2021 §5.1.2 unit u:table-1 (table) — …`
|
|
216
|
+
3. **Server-side reference resolution**: validate against used
|
|
217
|
+
passages, fetch the payload, return a `blocks` array alongside the
|
|
218
|
+
prose.
|
|
219
|
+
4. **Client rendering**: real `<table>` from columns/rows, KaTeX from
|
|
220
|
+
LaTeX, `<img>` from the asset store, term cards from designations +
|
|
221
|
+
definition. Unknown block types render nothing (forward
|
|
222
|
+
compatibility).
|
|
223
|
+
5. **Figures**: upload image assets to immutable URLs; a vision-capable
|
|
224
|
+
model describes each from pixels (one-time, ≈cents per figure);
|
|
225
|
+
store the description so text-only tiers can explain figures too.
|
|
226
|
+
|
|
227
|
+
## Stage 9 — Operations
|
|
228
|
+
|
|
229
|
+
- **Incremental re-ingest:** diff bundles by (document, unit_id) +
|
|
230
|
+
content hash; a wording change re-processes only the affected chunks.
|
|
231
|
+
Our corpus-proven case: a synonym change diffed exactly 34 of 3,053
|
|
232
|
+
chunks; a no-change re-ingest reports zero.
|
|
233
|
+
- **Spend ledger:** log per-request model spend to a queryable table;
|
|
234
|
+
review weekly against the model policy.
|
|
235
|
+
- **Upstream loop:** every corpus defect the service surfaces (missing
|
|
236
|
+
successor links, malformed table serializations, preface-only
|
|
237
|
+
documents) goes back as an upstream PR or issue — the corpus gets
|
|
238
|
+
better because serving keeps score.
|
|
239
|
+
|
|
240
|
+
---
|
|
241
|
+
|
|
242
|
+
## Reference stack (as deployed)
|
|
243
|
+
|
|
244
|
+
| Component | Choice | Notes |
|
|
245
|
+
|---|---|---|
|
|
246
|
+
| Platform | Cloudflare Workers only | minimal accounts, no proprietary model providers |
|
|
247
|
+
| Answer model | GLM-5.3 Flash (open-weight, multimodal) | ≈$0.001/answer, all tiers |
|
|
248
|
+
| Understanding | Qwen3-30B-A3B | cost-first hot path |
|
|
249
|
+
| Embeddings | Qwen3-Embedding-0.6B | same model both sides (mandatory) |
|
|
250
|
+
| Reranker | bge-reranker-base | ≈$0.00003/query |
|
|
251
|
+
| Vector index | Cloudflare Vectorize | dense lane |
|
|
252
|
+
| Lexical index | D1 + FTS5 (porter) | full-corpus BM25 |
|
|
253
|
+
| Registry + graph | D1 (documents, graph_nodes/edges) | derived status |
|
|
254
|
+
| Caches | KV (answer + semantic) | versioned by index |
|
|
255
|
+
| Typed payloads | D1 (unit_payloads) | MN 116 schemas |
|
|
256
|
+
| Assets | R2 (unit-keyed, immutable) | /assets/u:<id>.<ext> |
|
|
257
|
+
| Ingest | Python (pydantic models) | MKO → chunks → enrich → embed → upsert |
|
|
258
|
+
| Eval | Node (golden, paraphrase, e2e, RAGAS-style) | CI-gated |
|
|
259
|
+
|
|
260
|
+
## What not to do (each one cost us a measurement)
|
|
261
|
+
|
|
262
|
+
- Don't scrape rendered HTML when the document model exists.
|
|
263
|
+
- Don't let the model re-type tables; it will, and the values will be
|
|
264
|
+
wrong in ways that pass review.
|
|
265
|
+
- Don't trust a status field over the record's own edges.
|
|
266
|
+
- Don't judge faithfulness against a fresh retrieval.
|
|
267
|
+
- Don't report single-run metrics; the noise band is ±5pp.
|
|
268
|
+
- Don't run your heavy eval/enrichment traffic against the same account
|
|
269
|
+
that serves users — you'll throttle yourself and mis-measure
|
|
270
|
+
everything.
|
|
271
|
+
- Don't skip the refusal path in your test suite; it's the contract
|
|
272
|
+
users trust most.
|
|
273
|
+
|
|
274
|
+
---
|
|
275
|
+
|
|
276
|
+
*Production reference: ai.oimlsmart.org — OIML SMART AI. Format spec:
|
|
277
|
+
Metanorma MN 116 (Metanorma Knowledge Objects). This guide is extracted
|
|
278
|
+
from the system's own documentation set and can be adapted per corpus;
|
|
279
|
+
the measurement discipline transfers unchanged.*
|