@hviana/sema 0.7.3 → 0.7.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (46) hide show
  1. package/AGENTS.md +95 -843
  2. package/README.md +11 -11
  3. package/dist/src/mind/graph-search.d.ts +32 -22
  4. package/dist/src/mind/graph-search.js +97 -54
  5. package/dist/src/mind/mind.js +14 -0
  6. package/dist/src/store-sqlite.js +17 -0
  7. package/dist/src/store.d.ts +18 -0
  8. package/dist/src/store.js +10 -0
  9. package/docs/INDEX.md +71 -0
  10. package/docs/INVARIANTS.md +19 -0
  11. package/docs/architecture/bounded-reads.md +85 -0
  12. package/docs/architecture/caches.md +89 -0
  13. package/docs/architecture/commonality.md +45 -0
  14. package/docs/architecture/cost-model.md +71 -0
  15. package/docs/architecture/determinism.md +73 -0
  16. package/docs/architecture/exact-vs-approximate.md +47 -0
  17. package/docs/architecture/factored-machinery.md +28 -0
  18. package/docs/architecture/fold-contract.md +87 -0
  19. package/docs/architecture/halo-sketch.md +99 -0
  20. package/docs/architecture/match-project.md +62 -0
  21. package/docs/architecture/mechanism-market.md +95 -0
  22. package/docs/architecture/memoization.md +96 -0
  23. package/docs/architecture/meter.md +55 -0
  24. package/docs/architecture/saturation.md +92 -0
  25. package/docs/architecture/store.md +79 -0
  26. package/docs/architecture/thresholds.md +79 -0
  27. package/docs/failures/tempting-but-wrong.md +144 -0
  28. package/docs/harness/gates.md +56 -0
  29. package/docs/mechanisms/alu.md +75 -0
  30. package/docs/mechanisms/cast.md +75 -0
  31. package/docs/mechanisms/confluence.md +36 -0
  32. package/docs/mechanisms/cover.md +54 -0
  33. package/docs/mechanisms/extraction.md +53 -0
  34. package/docs/mechanisms/prefix-completion.md +54 -0
  35. package/docs/mechanisms/recall.md +69 -0
  36. package/docs/mechanisms/reference.md +58 -0
  37. package/jsr.json +1 -1
  38. package/package.json +1 -1
  39. package/src/mind/graph-search.ts +101 -55
  40. package/src/mind/mind.ts +14 -0
  41. package/src/store-sqlite.ts +19 -0
  42. package/src/store.ts +22 -0
  43. package/test/89-completion-recursion.test.mjs +30 -10
  44. package/test/97-store-seed.test.mjs +105 -0
  45. package/test/98-completion-chaining.test.mjs +140 -0
  46. package/HOW_IT_WORKS.md +0 -5836
@@ -0,0 +1,62 @@
1
+ # Match → Project → Gate
2
+
3
+ Every grounding mechanism is a configuration of one shared operation in
4
+ `src/mind/match.ts`. The family is defined once and imported many times;
5
+ duplicating it forks the corpus contract, moving it hides who owns the gate.
6
+
7
+ ## The shared family in `mind/match.ts`
8
+
9
+ The match layer locates structure, the project layer moves along the store, and
10
+ the gate layer decides whether the shape licences voicing. All three are pure
11
+ functions over bytes and the store — no mechanism owns a private copy.
12
+
13
+ ## The triple
14
+
15
+ | Role | Symbols | What it does |
16
+ | ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------- |
17
+ | **Match** (locate structure) | `locate` (exact → halo → gist ladder), `alignRuns` (literal W-gram weave), `alignGraded` (literal + halo gaps), `alignAround` / `frameSlots` (seeded frame with contracted gaps), `bestHaloMate` (in-list halo), `analogyStrength` / `sharedFrameStrength` (distributional + structural analogy) | Finds where a query sits in a learnt form. |
18
+ | **Project** (direction) | `follow` (forward to fixpoint, first hop may `conceptHop`), `reverseContext` (reverse to context), `project` (forward else reverse), `conceptHop` (halo sibling with edge) | Moves along the store from the match — forward toward answers, reverse toward contexts. |
19
+ | **Gate** (structural licence) | `isSpanShaped` (sparse subsequence — open reading), `carriesFillers` (substitution carriage — strict voicing licence) | Decides whether the shape licences voicing. |
20
+
21
+ Mechanisms declare only `(matcher, direction, gate)`. Thresholds behind gates
22
+ live in `src/geometry.ts` — the match layer never invents a cutoff.
23
+
24
+ The graded ladder inside `locate` is exact → distributional → geometric:
25
+ content-addressed identity first, halo similarity second, gist resonance last.
26
+ Reordering the ladder or letting an approximate score override an exact hit is a
27
+ correctness bug (see `exact-vs-approximate.md`).
28
+
29
+ ## Frame reading — matcher reports, gate judges, inventory elects nothing
30
+
31
+ `frameSlots` is the shared frame reader. It contracts every gap via
32
+ `contractGap` to its varying core, tags it `substitution` / `insertion` /
33
+ `deletion`, and attaches `covered` — the bytes the frame accounts for. It
34
+ applies no gate; it reports.
35
+
36
+ `carriesFillers` is the substitution gate. It judges byte-exactly:
37
+
38
+ ```
39
+ substituteAll(contA, fillersA → fillersB) == contB
40
+ ```
41
+
42
+ If the equality holds, voicing through the slot is a derivation; if not, the
43
+ slot cannot carry. This is the only place that decision is made.
44
+
45
+ `Precomputed.frames` is the inventory. It enumerates every frame pairing the
46
+ match layer finds and elects nothing — ranking and refusal belong to the
47
+ consumer.
48
+
49
+ ## Voicing gates belong to the consumer
50
+
51
+ The shared layer never refuses on a consumer's behalf. Reference owns its four
52
+ gates: frame dominates the query, each slot reaches `W` on both sides, no
53
+ insertion/deletion, fillers pairwise distinct — plus `carriesFillers` on the
54
+ chosen pair. CAST, recall, and cover each apply their own gate over the same
55
+ shared inventory. Moving a consumer's gate into `match.ts` would hide who is
56
+ responsible for the refusal.
57
+
58
+ ## Pins
59
+
60
+ - `test/47` — frame reading split (matcher vs gate vs inventory).
61
+ - `test/50` — CAST / reference voicing via `carriesFillers`.
62
+ - `test/24` / `test/76` — match/project family and span-shape readings.
@@ -0,0 +1,95 @@
1
+ # Mechanism Market — The Free-Will Architecture
2
+
3
+ Every grounding mechanism — including the ALU and user extensions — speaks one
4
+ interface (`mind/pipeline-mechanism.ts`). The decider in `mind/pipeline.ts`
5
+ (`think`) holds a plain list and never branches on which mechanism it holds.
6
+
7
+ ## Interface
8
+
9
+ ```ts
10
+ interface PipelineMechanism {
11
+ parse?(query: Uint8Array): Promise<ComputedSpan[]>; // authoritative spans
12
+ floor(ctx, query, pre, worthRunning): Promise<number | null>; // bound or null
13
+ run(ctx, query, pre): Promise<MechanismResult[]>; // candidates
14
+ }
15
+ interface MechanismResult {
16
+ bytes: Uint8Array;
17
+ accounted: [number, number][];
18
+ moves: number;
19
+ unexplained: string;
20
+ scaffolding?: number;
21
+ complete?: boolean;
22
+ }
23
+ ```
24
+
25
+ - `parse` is optional; all results are collected into `Precomputed.computed`
26
+ before any `floor`/`run`.
27
+ - `floor` returns `null` when structurally impossible, otherwise an admissible
28
+ lower bound (never overstates cost).
29
+ - `run` returns candidates with travelling evidence (below).
30
+
31
+ ## Decider
32
+
33
+ `think` in `mind/pipeline.ts` iterates `defaultMechanisms` in list order:
34
+
35
+ ```
36
+ defaultMechanisms = [cover, cast, confluence, extraction, reference, recall,
37
+ prefix-completion] + ALU (`aluToMechanism`) + extensions
38
+ ```
39
+
40
+ Weight is one currency: `weight = moves + PASS · unaccountedBytes` where
41
+ `unaccountedBytes = unexplainedSpans(query.length, accounted)`. Comparison is at
42
+ `STEP` grade (`grade = floor(weight/STEP)`); equal grade prefers fewer
43
+ `scaffolding` bytes, then list order.
44
+
45
+ ## Four constraints
46
+
47
+ 1. **Decoupled** — zero cross-imports between `mind/mechanisms/*`. Adding one
48
+ never touches another; no mechanism asks what already decided.
49
+ 2. **Declared competence** — binary structural gates inside `floor`/`run` (query
50
+ length, anchor shape, weave existence). Never a learned score; rationale
51
+ states exactly why a mechanism abstained.
52
+ 3. **Visible budget** — every corpus-scale loop is capped at a named constant:
53
+ `√N` via `hubBound`/`hubCap` and `k = 2·recallQueryK` (`Precomputed.k`).
54
+ Enforced at the store level.
55
+ 4. **Evidence travels** — every candidate carries `accounted` (query spans
56
+ explained), `moves` (priced on `MICRO/STEP/CONCEPT/PASS`), `unexplained`
57
+ (diagnostic label); optionally `scaffolding` (answer bytes from unrecognised
58
+ spans — equal-grade tie-break) and `complete` (trained-form continuation
59
+ reached via identity; post-grounding must not extend). The decider honours
60
+ both without knowing who set them.
61
+
62
+ ## Two disciplines
63
+
64
+ - **Admissible-floor pruning.** `floor` runs for every mechanism in list order
65
+ before any `run`. `run` fires only if `worthRunning(floor)` where
66
+ `worthRunning = (floor) => best === null || grade(floor) < grade(best.weight)`.
67
+ Cover runs first so a near-zero-cost computed span prunes the rest through the
68
+ same mechanism — not a special case.
69
+
70
+ - **Investment discipline.** `worthRunning` is passed _into_ `floor`. A floor
71
+ that would first-touch an expensive shared analysis (`pre.attention()` climb,
72
+ `pre.weave()`, `pre.resonance()`) checks `worthRunning(cheapestBound)` first
73
+ and returns the uninvested bound when it already loses. Never compute a shared
74
+ analysis just to discard it. `cast.ts`/`extraction.ts` are the references.
75
+
76
+ ## Accounting
77
+
78
+ - **Extraction:** located frames are always evidence; the span between them
79
+ counts only when _both_ borders were located. An open-ended read is priced by
80
+ exclusion (`PASS`/byte).
81
+ - **Reverse reading:** `reverseContext` produces bytes but explains nothing
82
+ forward: `accounted = []`, weight ≈ `PASS·|query|` — last resort by
83
+ arithmetic, not rule.
84
+ - **Paid acts are accounted:** the bridge's corroborated substitutions cost
85
+ `CONCEPT` each in `moves`, so their spans must be `accounted`; otherwise the
86
+ same act is charged twice (`PASS`/byte dominates).
87
+
88
+ `accounted` is a cost-ladder quantity; `cover.ts` leaves masked computed spans
89
+ out of it so `PASS`-bridged bytes are still charged. `unexplained`,
90
+ `narrowDecision`, `thinGrounding` are observational only.
91
+
92
+ ## Pins
93
+
94
+ - `test/01-floor` — floor geometry.
95
+ - `test/04-think` — decider, admissible pruning, investment discipline.
@@ -0,0 +1,96 @@
1
+ # Memoization — Shared Evidence Without Duplication
2
+
3
+ > **Law:** asking never writes, so structural reads are pure during one
4
+ > response. Memoization elides probes, not evidence.
5
+
6
+ Two layers: `Precomputed` (response-scoped shared analyses) and `Mind`
7
+ per-response memos. Both are accelerators that must not change what inference
8
+ computes.
9
+
10
+ ## Precomputed — one response, one container
11
+
12
+ `Precomputed` (`src/mind/pipeline-mechanism.ts`) is the sole place a response's
13
+ shared evidence lives. Created by `think` (`src/mind/pipeline.ts`) before the
14
+ mechanism loop.
15
+
16
+ ### Eager — populated before any `floor`/`run`
17
+
18
+ - `rec: Recognition` — structural + canonical decomposition (`recognise`)
19
+ - `computed: ComputedSpan[]` — `parse()` results from all mechanisms (e.g. ALU)
20
+ - `guide: Vec` — query gist, the response-wide disambiguation guide
21
+ - `k: number` — `cfg.recallQueryK * 2`, the breadth for resonance/weave/climb
22
+
23
+ ### Lazy — computed on first touch, cached by promise
24
+
25
+ Expensive analyses are `async` and cached by promise: the first caller starts
26
+ the work, every later caller awaits the same promise.
27
+
28
+ - `attention()` — `climbAttentionAll` (roots + ranked anchors)
29
+ - `weave()` — `alignGraded` over top-k anchors
30
+ - `resonance()` — `store.resonate(guide, k)` (single ANN query)
31
+ - `frames()` — `frameSlots` inventory from resonance
32
+ - `spanShapedOf(anchor)` / `spanShapedAll()` — per-anchor `skillExemplar`,
33
+ memoised per id
34
+ - `queryWindows` / `queryResolved` / `windowsOf(anchor)` — W-window identities
35
+ - `reachMemo` — `sharedReachMemo(ctx)` (ancestor reach, § below)
36
+
37
+ A mechanism that never asks pays nothing; two mechanisms asking the same
38
+ question pay once. `floor()` must gate on `worthRunning` before first-touching
39
+ an expensive analysis.
40
+
41
+ ## Mind memos — `beginResponse` → `endResponse`
42
+
43
+ `Mind` (`src/mind/mind.ts:beginResponse`/`endResponse`) swaps per-response state
44
+ for each inference call. `respond` takes fresh maps; `respondTurn` reuses the
45
+ conversation's persistent ones (content-keyed, cross-turn).
46
+
47
+ | Memo | Key | Scope |
48
+ | ------------------- | ----------------------------------------- | -------------------------------------------- |
49
+ | `perceiveMemo` | `perceiveKey(bytes)` (latin1) | response / conversation |
50
+ | `recogniseMemo` | `latin1Key(bytes)` | response / conversation |
51
+ | `climbMemo` | `latin1Key(bytes)` | response / conversation |
52
+ | `canonMemo` | `latin1Key(bytes)` | response (when `canon` set) |
53
+ | `_resolvedSubtrees` | `WeakMap<Sema, {id,len}>` (node identity) | response / conversation |
54
+ | `_edgeChoice` | `Map<nodeId, pick>` | response only — **cleared** in `endResponse` |
55
+ | `_gistCache` | `BoundedMap<nodeId, Vec>` 32 MB | **session-lifetime** (not per-response) |
56
+
57
+ `_gistCache` (≈ 8K gists at D=1024) survives across responses; all others are
58
+ dropped or cleared at `endResponse`. `_resolvedSubtrees` elides store probes
59
+ when `visit` is absent; with a visitor it still walks in full (see
60
+ `src/mind/primitives.ts:foldTree`).
61
+
62
+ ## Trace boundary — what is bypassed
63
+
64
+ Traced responses must emit every step, but must not change the answer.
65
+
66
+ - **Bypassed:** `_edgeChoice` via `guidedNext`
67
+ (`src/mind/traverse.ts:guidedNext`) and `sharedReachMemo`
68
+ (`src/mind/traverse.ts:sharedReachMemo`). Both return fresh empty maps when
69
+ `ctx.trace !== null`; `chooseNext` recomputes identically (pure over store +
70
+ guide).
71
+ - **Always consulted:** `perceiveMemo`, `recogniseMemo`, `climbMemo` (and their
72
+ underlying `perceive`/`recognise`/`climbAttention` caches). Bypassing breaks
73
+ idempotence.
74
+
75
+ `foldTree`'s subtree fast path is taken only when no `visit` is supplied. With a
76
+ visitor (recognition, attention) the walk still descends; the cache elides only
77
+ probes. Bypassing `recogniseMemo` under trace re-ran `recogniseImpl` with a warm
78
+ `_resolvedSubtrees` and emitted fewer sites (observed 31 → 5) — a correctness
79
+ change, not just a slowdown.
80
+
81
+ ## Meter — charge work to itself
82
+
83
+ Shared analyses are charged to their own phase (`meter.time(phase, fn)` in
84
+ `Precomputed.shared`), not to the mechanism that first touched them
85
+ (`src/meter.ts:PhaseCost`). Without this, the profile reads "cast.floor costs 2
86
+ s" when the cost was the consensus climb cast paid for on everyone's behalf.
87
+
88
+ ## Adding a shared analysis
89
+
90
+ Add one lazy method to `Precomputed`. No new memo map elsewhere. Gate it behind
91
+ `worthRunning` in `floor()`.
92
+
93
+ ## Pins
94
+
95
+ - `test/42` — recognition idempotence under trace: traced and untraced
96
+ `recognise` return the same site count and cached object.
@@ -0,0 +1,55 @@
1
+ # Meter — Work Accounting
2
+
3
+ `src/meter.ts` is the one computational-usage accounting surface. It counts what
4
+ inference _cost_ so a slow response can be attributed instead of guessed at. The
5
+ rationale says why an answer was chosen; the meter says what it cost to choose
6
+ it. Harness: `bench/profile-inference.mjs`.
7
+
8
+ ## Five contracts
9
+
10
+ 1. **Write-only from inference.** No counter reaches a decision, a threshold, or
11
+ an ordering. Determinism survives only because the meter is observed, never
12
+ consulted. Every call site is `meter?.x++` on a nullable field.
13
+
14
+ 2. **Counters vs hints.** Counters are deterministic and diffable between runs;
15
+ the same query on the same store meters identically, so a regression is
16
+ visible in a diff. Millisecond fields (`elapsedMs`, per-phase `ms`) are
17
+ non-deterministic hints reported separately — never use them to gate
18
+ behaviour.
19
+
20
+ 3. **Phases nest, they do not partition.** `think` contains every mechanism
21
+ phase; a mechanism's `floor` contains whatever shared analysis it
22
+ first-touched; `recall.run` contains `substitutionBridge`. Read a phase as
23
+ inclusive wall-clock — never sum phases and expect the total.
24
+ `CostReport.elapsedMs` is the only whole.
25
+
26
+ 4. **Count once.** Off by default and free when off
27
+ (`new Mind({ profile:
28
+ true })` to attach). A layer that wants to be
29
+ visible bumps a field in `meter.ts` — it does not grow a private counter.
30
+ (Legacy `danglingReads` / `compactFailures` in `store.ts` are health
31
+ counters, not per-response work.)
32
+
33
+ 5. **Shared analyses charged to themselves.** The first toucher pays the wall
34
+ clock, but every later consumer gets the result free. Attribution follows the
35
+ analysis, not the mechanism that first triggered it — otherwise the profile
36
+ misreads which work is expensive (e.g. the consensus climb billed through
37
+ whichever mechanism happened to need it first).
38
+
39
+ ## Where counters live
40
+
41
+ `src/meter.ts:Meter` is the only definition of a counter name. Phases are
42
+ charged via `meter.time(phase, fn)` / `meter.timeSync(phase, fn)`, which
43
+ snapshot counters on entry and attribute the delta to the phase. The sync/async
44
+ seam is load-bearing: synchronous layers (perception, recognition, graph search)
45
+ must use `timeSync` so the profiled path does not await where the unprofiled
46
+ path does not.
47
+
48
+ `CostReport` is plain JSON (`version`, `elapsedMs`, `queryBytes`, `counters`,
49
+ `phases`). Zero-valued counters are dropped; `formatReport` renders the three
50
+ heaviest counters per phase.
51
+
52
+ ## Pins
53
+
54
+ - `test/55` — `Meter`, `CostReport`, `searchPops` / `searchPushes`, phase
55
+ nesting.
@@ -0,0 +1,92 @@
1
+ # Saturation — Named Stop, Not Cap
2
+
3
+ > **Law:** every walk names a deciding saturation beside its cap. The cap is a
4
+ > safety net; saturation is the derived stop that decides and terminates.
5
+
6
+ A walk with only a cap drifts to the cap. A walk with a saturation stops the
7
+ moment the answer (saturated vs. decided) is known — bounded, exact below the
8
+ bound, and named in the trace.
9
+
10
+ ## Cap vs. saturation
11
+
12
+ | Role | Value | Nature |
13
+ | ---------- | ------------------------------------------------------------ | ------------------------------------------------ |
14
+ | Cap | `hubBound = ceil(sqrt(N))` per read; `hubBound * W` per walk | Safety net — prevents corpus-proportional work |
15
+ | Saturation | Named derived stop (`SaturationReason`) | Decision — proves continuing cannot discriminate |
16
+
17
+ `N = corpusN = max(2, edgeSourceCount())`, `W = maxGroup`. Defined once in
18
+ `mind/traverse.ts` (`corpusN`, `hubBound`, `boundFor`); never spelled inline.
19
+
20
+ ## edgeAncestors — EXPAND-UNTIL-DECIDED
21
+
22
+ `mind/traverse.ts:edgeAncestors` is the model: it climbs the structural DAG
23
+ (parents + containment) until one of five saturations decides, and is exact
24
+ below every one of them.
25
+
26
+ 1. **Predecessor fan-in** — `prevCount(node) > bound` via `store.prevCount`. One
27
+ indexed count; no read. Proves `> bound` distinct contexts on its own.
28
+ 2. **Distinct-context limit** — `ctxSeen.size > bound` after
29
+ `prevFirst(node, bound)`. The accumulated set of distinct contexts reachable
30
+ from roots visited so far exceeds `sqrt(N)`.
31
+ 3. **Parent fan-out** — `parentsFirst(node, bound+1).length > bound`. One
32
+ `LIMIT bound+1` read distinguishes "exactly bound parents" from "hub". The
33
+ node itself is not expanded; saturated reaches are never voted.
34
+ 4. **Lateral-cone cumulative** — `lateral > bound`, where `lateral` sums
35
+ `fresh-1` over every expanded node's extra parents beyond its first. The
36
+ per-node guard catches concentration at one node; this catches the same
37
+ commonness distributed across the cone. A deep chain in one structure accrues
38
+ zero laterals and still reaches its root at any depth.
39
+ 5. **Byte-atom commonality** — `atomIsHub(N,W)` when
40
+ `atomReach(N,W) = max(1, ceil(N*W/256)) > bound`. Atoms carry no kid/contain
41
+ rows, so containment is unmeasurable; the uniform-expectation floor replaces
42
+ it. Above the scale the atom abstains as voter (edges remain traversable for
43
+ tier-0 recall).
44
+
45
+ Below every threshold the walk is exact — `prevFirst(bound)` is the full list,
46
+ `parentsFirst(bound+1)` is the full list, `containersSlice` pages are walked in
47
+ full — identical to the unbounded climb. Work is `O(bound)` contexts times local
48
+ structure, never `O(N)`. Container seeding is streamed in `bound`-sized pages
49
+ for the same reason.
50
+
51
+ Trace records the first deciding stop as
52
+ `SaturationStop { reason, node, observed, limit }` plus `visited`/`maxDepth`;
53
+ absent when unsaturated or untraced.
54
+
55
+ ## pivotInto — longest-wins
56
+
57
+ `mind/resonance.ts:pivotInto` ranks candidates by `contentLen(id, answerLen+1)`
58
+ descending (first-inserted tie-break). The byte score is length itself, so the
59
+ scan is decided at the first candidate that passes every filter — a shorter
60
+ candidate can never outscore it. At most one winner's bytes are reconstructed;
61
+ every shorter proposal is skipped without a read. Saturation, not cap.
62
+
63
+ ## Junction — hub guards vs. budget
64
+
65
+ `mind/junction.ts:junctionContainersFrom` has three disciplines:
66
+
67
+ - **Phrase-scale reads** — `bytesPrefix(maxContainer+1)` per visit; a node
68
+ beyond the cap prunes its branch.
69
+ - **Per-node hub guards** (real saturations) — `parentsFirst(bound+1) > bound`
70
+ not expanded; one `containersSlice(bound+1)` page beyond `bound` not expanded.
71
+ Each is exact below `bound`.
72
+ - **Expansion budget** — at most `bound * W` pops total (shared across a tier's
73
+ walks). Budget exhaustion is an abstention (`junctionBudgetExhausted`) that
74
+ falls through to the resonance tier — a net, not a saturation.
75
+
76
+ Refuted tightening: applying `edgeAncestors`' cumulative lateral-cone limit here
77
+ would discard half the successful junctions (measured lateral 1425/1426 at
78
+ `bound` ~ 570).
79
+
80
+ Refuted early-stop: **one-cone-exhausted** — stopping when one side's upward
81
+ cone empties — is wrong in both hub-guarded and hub-flagged forms. A junction
82
+ can be reachable from only one side when that side's seed is a fold sub-node of
83
+ the container (`test/16`: "cold or hot" reached from window "cold" while 3-byte
84
+ "hot" cone is empty; `test/34` n-ary binding fails the same way). Exhausting one
85
+ cone never proves no junction remains; the walk must keep the `bound*W` net
86
+ after per-node saturations.
87
+
88
+ ## Pins
89
+
90
+ - `test/16` — one-cone-exhausted refutation (bridge junction).
91
+ - `test/34` — one-cone-exhausted refutation (n-ary cross-region binding).
92
+ - `test/27` — saturation-drop gate (leading/trailing saturated intervals).
@@ -0,0 +1,79 @@
1
+ # Store — AbstractStore Owns the Domain, Adapters Own the Wires
2
+
3
+ > **Law:** `AbstractStore` (`src/store.ts`) owns every domain decision — dedup,
4
+ > near-dedup, gist/halo indexing, containment, batching, LRU, compaction.
5
+ > `SQliteStore` (`src/store-sqlite.ts`) implements only `_db*`/`_vec*` thin
6
+ > wrappers.
7
+
8
+ ## Template method
9
+
10
+ ```
11
+ AbstractStore all logic: caches, merge gates, halo schedule,
12
+ buffers, chain transparency, length walks
13
+ └─ SQliteStore SQL + VectorDatabase glue — one statement per method
14
+ └─ <NewBackend> same contract — subclass AbstractStore only
15
+ ```
16
+
17
+ A new backend subclasses `AbstractStore`; never re-implements dedup or batching.
18
+
19
+ ## IDs and leaves
20
+
21
+ Branch ids are dense non-negative `0,1,2,…` (`_nextId`, never deleted).
22
+ Single-byte leaves are **implicit** negative ids `-256..-1` (`-(byte+1)`), never
23
+ a row. `has(id)` is `id < 0 || id < _nextId`.
24
+
25
+ ## Flat branches and bytes
26
+
27
+ A branch whose kids are all leaves is **flat** — stored as raw bytes in `leaf`
28
+ with an empty `kids` blob as marker (`flatKidsBytes`/`flatBytesKids`). Dedup
29
+ probes hash then verify: `hashOf`→`h`→`LIMIT 1` fetch→byte compare (bloom
30
+ negative filter first).
31
+
32
+ `bytes(id)`/`bytesPrefix(id, cap)` are shared with `BoundedMap` caches — callers
33
+ must **never mutate** the returned buffer. `contentLen(id, cap)` walks with
34
+ memo; when `cap` is given it saturates (`>= cap` without finishing) so one huge
35
+ root never costs a full walk.
36
+
37
+ ## Gist, halo, dedup
38
+
39
+ On `put*`, content dedup (`hashOf`→probe→mint) gates first. `DedupKey` caches
40
+ short keys (`DEDUP_KEY_MAX` bypass). Near-dedup merges by `mergeThreshold(D)` on
41
+ unit gist cosine. Gists sit in `_pendingGist` (byte-budgeted `BoundedMap`);
42
+ `indexSubtree` & `pourHalo` promote via `_vecContentUpsert`/`_vecHaloUpsert` in
43
+ `batchSize` batches. Buffers flush on cadence, `commit()`, and close. Halo mass
44
+ re-indexes geometrically (`mass<=4 || powerOfTwo`) and encodes 2-bit quantized.
45
+ Canon index is optional: `canonAdd`/ `canonFind`/`canonCount` over 32-bit
46
+ canonical hashes, caller verifies bytes.
47
+
48
+ ## Containment, batching, LRU
49
+
50
+ `addContainer(child,parent)` buffers per child; flush appends via
51
+ `_dbAppendContain` (packed pages, geometric merge) — never rewrite the whole
52
+ list. `containersSlice` pages through it. Edges and kids write through the same
53
+ deferred transaction.
54
+
55
+ Every in-memory cache is a `BoundedMap` with byte accounting and eviction (`lru`
56
+ vs `smallest` + `clock`/`reorder` recency). ANN reads
57
+ (`resonate`/`resonateHalo`) are content-addressed (`vecKey`) and dropped on any
58
+ index mutation; `RESonate_CACHE_MAX=4096`.
59
+
60
+ ## Maintenance (incremental)
61
+
62
+ - `compactContentIndex(minParents)` — scans only entries since last watermark
63
+ (`_vecContentEntriesSince`), removes indexed-but-isolated nodes (`<minParents`
64
+ parents, no edges/halos), compacts the vector DB.
65
+ - `repairContentIndex(regenerateGist)` — walks `_dbEdgeOrHaloIds()` candidates
66
+ only, re-inserts missing bridge nodes whose gists were evicted before
67
+ indexing.
68
+ - `buildCanonIndex` (`_buildCanonIndex`) — iterates `eachContent(fromId)` and
69
+ `canonAdd`s; `fromId` makes refresh incremental.
70
+
71
+ Full scans (`parents()`, `next()`, `containers()`) are maintenance-only — hot
72
+ paths use `LIMIT`ed probes (`parentsFirst`/`nextFirst`/`prevFirst`), `has*`,
73
+ `prevCount`.
74
+
75
+ ## Adding a backend
76
+
77
+ Implement every `protected abstract _db*`/`_vec*` in `src/store.ts` as a thin
78
+ wrapper around your storage. Keep `_dbGet*First`/`Slice` as real `LIMIT` queries
79
+ and `has*`/`COUNT` as point probes — never materialise-then-slice.
@@ -0,0 +1,79 @@
1
+ # Thresholds — derived, never tuned
2
+
3
+ Every decision cutoff is a formula over `D` (vector dimension), `W` (`maxGroup`,
4
+ perception window), or `N` (corpus size). No threshold is tuned or added to
5
+ `src/config.ts`.
6
+
7
+ ## Source of truth
8
+
9
+ - `src/geometry.ts` — all similarity/decision thresholds.
10
+ - `src/mind/traverse.ts` — corpus-scale readings that parameterise bounded
11
+ walks.
12
+ - `src/sema.ts` — positional coordinate algebra.
13
+ - `src/config.ts` — capacities and budgets only (cache byte budgets, batch
14
+ sizes, index parameters, query `k`, ALU precision, seed). Never a threshold.
15
+
16
+ ## Geometry thresholds (`src/geometry.ts`)
17
+
18
+ | Symbol | Definition | Formula |
19
+ | ----------------------- | ---------------------------------------------------------------------------- | ----------------------------------- |
20
+ | `mergeThreshold(D)` | Store identity bar — cosine at which `intern` treats two gists as same node | `1 - 1/√D` |
21
+ | `identityBar(D,W,len)` | Scale-aware whole-span identity claim | `max(mergeThreshold(D), 1 - W/len)` |
22
+ | `reachThreshold(W)` | Recall confidence floor — half a river quantum | `1 - 1/(2·W)` |
23
+ | `estimatorNoise(D)` | RaBitQ noise floor — 1σ of random cosine | `1/√D` |
24
+ | `significanceBar(D)` | Whole-query relatedness — 3σ above chance | `3/√D` |
25
+ | `conceptThreshold(D)` | Halo concept sharing — structural midpoint + ½σ | `0.5 + 0.5/√D` |
26
+ | `dominates(part,whole)` | Half-dominance predicate | `part*2 > whole` |
27
+ | `profileCapacity(D)` | Superposition capacity — terms before readout collapses | `floor(√D)` (min 1) |
28
+ | `consensusFloor(N)` | Pooled-vote significance floor | `ln(N) + ½` |
29
+ | `coverageBar(_,D)` | Reach-index gating (currently unused hot-path; batch compaction replaces it) | `conceptThreshold(D)` |
30
+
31
+ `N` in `consensusFloor` is `corpusN` (edge-source count, floored at 2).
32
+
33
+ ## Corpus-scale readings (`src/mind/traverse.ts`)
34
+
35
+ | Symbol | Definition | Formula |
36
+ | ---------------- | ------------------------------------------------------ | --------------------------- |
37
+ | `corpusN` | Distinct learnt contexts, floored | `max(2, edgeSourceCount())` |
38
+ | `hubBound` | Hub bound, used for every `LIMIT` read | `ceil(√max(2,N))` |
39
+ | `hubCap(ids)` | Fan-out cap — list-side reading of `hubBound` | `ids.slice(0, hubBound)` |
40
+ | `atomReach(N,W)` | Uniform-expectation floor on a byte atom's commonality | `max(1, ceil(N·W/256))` |
41
+ | `atomIsHub(N,W)` | Whether atom abstains as consensus voter | `atomReach > hubBound` |
42
+
43
+ `hubBound` is enforced at the store level (`nextFirst`, `parentsFirst`,
44
+ `containersSlice`, `hasNext`/`hasParents`, `bytesPrefix`, `chainRun`).
45
+ `atomReach` is the honest floor for atoms — they carry no kid/contain rows, so
46
+ their reach is unmeasurable and must not default to "maximally rare".
47
+
48
+ ## Seat algebra (`src/sema.ts`)
49
+
50
+ `twoEndedSeat(seatCount, size, index)` — the one positional-coordinate algebra
51
+ shared by perception, `fold`, and every synthetic/canonical fold. First half
52
+ uses low seats, second half uses high seats:
53
+ `index < (size+1)/2 ? index : seatCount-size+index`.
54
+
55
+ ## Two derivations that bite
56
+
57
+ **1. `identityBar` is scale-aware.** A fixed cosine `1-1/√D` over a `4·√D`-byte
58
+ span tolerates four whole windows of foreign bytes while still claiming
59
+ "near-identical". An identity claim may tolerate at most one window `W` (the
60
+ perception quantum, same budget as `differsByOneWindow` in near-dedup), so the
61
+ bar must be `1-W/len` floored at `mergeThreshold`. Reusing `mergeThreshold` for
62
+ a whole-span claim silently widens the byte budget with span length.
63
+
64
+ **2. `consensusFloor` is priced for pooled climb votes, not for support
65
+ counts.** Each region contributes at most `ln(N/c) ≤ ln(N)`; `ln(N)+½` demands
66
+ corroboration beyond one maximally-specific region. `chooseNext`'s `bestSupport`
67
+ (`prevCount` of one destination) is N-invariant — bounded by retellings of that
68
+ fact, not by `N`. Gating it against `consensusFloor` guarantees failure once `N`
69
+ is large enough (observed: 2-vs-1-1-1 corroboration refused at N≈325K, falling
70
+ back to a noisy concept-hop).
71
+
72
+ ## Pins
73
+
74
+ - `test/40-choosenext-scale-guard` — `consensusFloor` must not gate `chooseNext`
75
+ support counts.
76
+ - `test/64-two-ended-thresholds` — `mergeThreshold` / `identityBar` /
77
+ `reachThreshold` derivations and the scale-aware floor.
78
+ - `test/78-atom-hub-recognition-cliff` — `atomReach` / `atomIsHub` hub
79
+ abstention at scale.