@hviana/sema 0.9.4 → 0.9.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/AGENTS.md +16 -7
  2. package/dist/src/mind/learning.js +11 -12
  3. package/dist/src/mind/mind.js +4 -4
  4. package/dist/src/mind/types.d.ts +10 -15
  5. package/dist/src/store.d.ts +10 -10
  6. package/dist/src/store.js +5 -5
  7. package/docs/INDEX.md +62 -63
  8. package/docs/INVARIANTS.md +36 -18
  9. package/docs/architecture/bounded-reads.md +31 -71
  10. package/docs/architecture/caches.md +61 -81
  11. package/docs/architecture/closure.md +88 -107
  12. package/docs/architecture/commonality.md +47 -38
  13. package/docs/architecture/cost-model.md +57 -79
  14. package/docs/architecture/determinism.md +43 -55
  15. package/docs/architecture/evidence.md +158 -235
  16. package/docs/architecture/exact-vs-approximate.md +41 -35
  17. package/docs/architecture/factored-machinery.md +34 -20
  18. package/docs/architecture/fold-contract.md +110 -118
  19. package/docs/architecture/halo-sketch.md +105 -96
  20. package/docs/architecture/match-project.md +51 -42
  21. package/docs/architecture/mechanism-market.md +87 -91
  22. package/docs/architecture/memoization.md +60 -74
  23. package/docs/architecture/meter.md +37 -47
  24. package/docs/architecture/saturation.md +75 -101
  25. package/docs/architecture/store.md +118 -99
  26. package/docs/architecture/thresholds.md +66 -73
  27. package/docs/failures/tempting-but-wrong.md +139 -165
  28. package/docs/harness/gates.md +27 -32
  29. package/docs/mechanisms/alu.md +22 -69
  30. package/docs/mechanisms/cast.md +76 -71
  31. package/docs/mechanisms/confluence.md +22 -29
  32. package/docs/mechanisms/cover.md +58 -66
  33. package/docs/mechanisms/extraction.md +33 -37
  34. package/docs/mechanisms/prefix-completion.md +36 -39
  35. package/docs/mechanisms/recall.md +60 -53
  36. package/docs/mechanisms/reference.md +63 -49
  37. package/jsr.json +1 -1
  38. package/package.json +1 -1
  39. package/src/alu/README.md +90 -298
  40. package/src/derive/README.md +94 -256
  41. package/src/mind/learning.ts +11 -12
  42. package/src/mind/mind.ts +4 -4
  43. package/src/mind/types.ts +10 -15
  44. package/src/rabitq-ivf/README.md +11 -8
  45. package/src/store.ts +5 -5
@@ -1,29 +1,43 @@
1
1
  # Factored Machinery — One Definition, Many Consumers
2
2
 
3
- Every shared operation is defined once and imported many times. Duplicating it
4
- forks the corpus contract; moving it hides who owns the gate.
3
+ > **Law:** every operation that more than one layer relies on is defined once
4
+ > and imported everywhere. A second copy is a write/read drift waiting to
5
+ > happen. A copy moved into a consumer hides who owns the gate.
5
6
 
6
- Siblings: `match-project.md`, `commonality.md`, `meter.md`.
7
+ Each row is a contract that broke at least once when it existed twice:
8
+
9
+ - `contentBoundaries` carried its own hash loop;
10
+ - the ingest cache re-implemented deposit;
11
+ - recall spelled its own fragment test.
7
12
 
8
13
  ## Single-definition contracts
9
14
 
10
- | Symbol | Defined in | One fact |
11
- | ------------------------------------------------------------- | ------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
12
- | `contentLevels` | `src/geometry.ts` | Single boundary rule: cuts + levels from one rolling hash pass; every segmentation reads it. |
13
- | `canonicalWindows` / `chainReach` / `leafIdRun` / `windowIds` | `src/mind/canonical.ts` | Write/read contract: training interns `W-1,W` windows, reading chains to `W²` and probes `W`-windows — drift silences recognition. |
14
- | `junction.ts` + `WalkCache` | `src/mind/junction.ts` | Shared junction ascent (parents + containers) with bounded `√N·W` walk; `WalkCache` memoizes capped reads/parents/containers per response; bridge and attention share it. |
15
- | `joinWithBridge` | `src/mind/resonance.ts` | One out-of-search assembly: `bridge(left,right)` or bare concat with `bridgeMiss` trace. |
16
- | `dismissedKnownContent` | `src/mind/bridge.ts` | Pure attestation: any unaccounted `W`-window that resolves as known content — shared gap guard for substitution and CAST. |
17
- | `sharedReachMemo` | `src/mind/traverse.ts` | One response-scoped `AncestorReach` memo (cleared on write and for traces); every `reachOf`/`edgeAncestors` consumer shares it. |
18
- | `guidedFirst` | `src/mind/traverse.ts` | Guided-or-first answer bytes: `guidedNext` else first-inserted edge (`LIMIT 1`). |
19
- | `leadsSomewhere` | `src/store.ts` (raw), `src/mind/traverse.ts` (memoised) | Admission predicate: `hasNext` or `hasHalo`, defined once on the store; traverse memoises the edge tier per response. |
20
- | `isChunk` | `src/sema.ts` | `kids !== null && kids.every(k=>k.kids===null)` — smallest grouped unit; governs regions, seams, indexing. |
21
- | `twoEndedSeat` | `src/sema.ts` | One seat algebra: first half low seats, second half high seats; shared by perception, `fold`, and canonical folds. |
22
- | `closed`/`admissible`/`advance`/`closeOver` | `src/mind/derivation.ts` | One admission for every derivation step, the readings every tier asks, and the engine that walks the post-grounding layers. |
23
- | `exactNode` | `src/mind/primitives.ts` | The exact content-addressed lookup (`resolve`, `canonResolve`): the identity fold (`contentIdentity`, geometry.ts) — a miss decided by the segment filter, a hit named bottom-up, no vectors. |
15
+ | Symbol | Home | The one fact it owns |
16
+ | ---------------------------------------------------------------------- | ------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- |
17
+ | `contentLevels` | `geometry.ts` | the boundary rule: cuts and levels from one rolling hash (`fold-contract.md`) |
18
+ | `groupByLevel` / `foldSlice` | `geometry.ts` | the tree's shape, written once over two algebras (vector and identity) |
19
+ | `twoEndedSeat` | `sema.ts` | the seat algebra: low seats for the first half, high seats for the second |
20
+ | `isChunk` | `sema.ts` | the smallest grouped unit, a node whose children are all leaves |
21
+ | `exactNode` / `branchNaming` | `mind/primitives.ts` | exact lookup: the identity fold, naming a branch exactly as `intern` does |
22
+ | `canonicalWindows` / `chainReach` / `leafIdRun` / `windowIds` | `mind/canonical.ts` | the window contract: deposits intern windows of `W−1` and `W`, and reads chain them to `W²` |
23
+ | `leadsSomewhere` | `store.ts`, memoized in `traverse.ts` | admission: `hasNext \|\| hasHalo` |
24
+ | `corpusN` / `hubBound` / `hubCap` | `mind/traverse.ts` | the scale and its cap (`bounded-reads.md`) |
25
+ | `sharedReachMemo` | `mind/traverse.ts` | one `AncestorReach` memo for every `reachOf`/`edgeAncestors` consumer |
26
+ | `answersOtherQuestions` | `mind/traverse.ts` | the fragment predicate: a form inside other forms, with several continuations, that the question leaves partly outside (`evidence.md`) |
27
+ | `junctionContainersFrom` + `WalkCache` | `mind/junction.ts` | the content-addressed ascent shared by the bridge and attention |
28
+ | `joinWithBridge` | `mind/resonance.ts` | joining two results outside the search: `bridge(left, right)`, or a bare join traced as `bridgeMiss` |
29
+ | `dismissedKnownContent` | `mind/bridge.ts` | the gap guard shared by substitution and CAST: no dismissed window may resolve as known content |
30
+ | `locate` / `alignGraded` / `frameSlots` / `project` | `mind/match.ts` | the matcher family; gates belong to the consumers (`match-project.md`) |
31
+ | `witness` / `windowIndex` | `mind/evidence.ts` | order-free W-window coverage, the evidence reading (`evidence.md`) |
32
+ | `closed` / `admissible` / `advance` / `closeOver` / `unexplainedSpans` | `mind/derivation.ts` | the closure law, its span algebra and its engine (`closure.md`) |
33
+ | `Precomputed` | `mind/pipeline-mechanism.ts` | the per-response shared analyses (`memoization.md`) |
34
+ | `weigh` | `mind/pipeline.ts` | the price, `moves + PASS·unaccounted` (`cost-model.md`) |
35
+ | `Meter` | `meter.ts` | every counter name (`meter.md`) |
24
36
 
25
37
  ## Pins
26
38
 
27
- - `test/47` — frame reading split (matcher vs gate vs inventory).
28
- - `test/50` — CAST analog / consensus floor (dismissed content, `MIN_WEAVE` /
29
- `dominates` frame, `carriesFillers` refusal).
39
+ - `test/137` — the law lives once and below: no hand-built transition outside
40
+ `advance`, no inline re-spelling of a law reading, and no dead export.
41
+ - `test/47` — the frame reading split into matcher, inventory and gate.
42
+ - `test/50` — CAST's shared guards: dismissed content, the frame from
43
+ `MIN_WEAVE` and `dominates`, and the `carriesFillers` refusal.
@@ -1,137 +1,129 @@
1
- # Fold Contract — One Tree For The Same Bytes
1
+ # Fold Contract — One Tree for the Same Bytes
2
+
3
+ > **Law:** perception is a pure function of the bytes. A deposit
4
+ > (`perceiveDeposit`) and a question (`perceive`) fold the same bytes into the
5
+ > same tree, and the read side names that tree with the node the write side
6
+ > interned. Nothing absent from the bytes may shape the tree: turn boundaries,
7
+ > offsets and index-aligned grids are excluded.
8
+
9
+ Identity is content (`store.md`), so a question finds what memory holds only if
10
+ it is cut exactly as the deposit was. When the two sides disagreed, alignment
11
+ went quadratic (5.2M cells on a 476-byte context, against 0 when they agree),
12
+ and cumulative contexts stopped resolving to what they were trained as.
13
+ Conversation state (`ConversationState`, `answeredSpans`, `currentTurnStart`) is
14
+ API metadata and never reaches the geometry. Passing turn boundaries into the
15
+ fold is a correctness bug, not a tuning choice.
16
+
17
+ ## The boundary rule — `contentLevels`
18
+
19
+ `contentLevels` is the one boundary rule, and `contentBoundaries` is its
20
+ projection.
21
+
22
+ - **Cuts.** A rolling hash runs over a bounded window of the bytes, and a cut
23
+ falls where it vanishes. A cut is level `L` when the hash vanishes mod
24
+ `W^(L+1)`, so level-`L` cuts nest inside level-`(L−1)` cuts, and a level-`L`
25
+ span averages `W^(L+1)` bytes.
26
+ - **Segments.** A segment runs from `W−1` bytes to `seats.length` bytes, the
27
+ most one flat node folds, with a forced cut at the maximum.
28
+ - **Grouping.** `groupByLevel` groups segments by their level, never by count.
29
+
30
+ **Why content, not position.** A change moves only the cut it falls inside.
31
+ After shifts of 1–7 bytes, 99.7% of cuts on real deposits survive, against 14.3%
32
+ for a fixed grid. Grouping by index (a stride, a tile, a fixed-arity row such as
33
+ `riverFold`'s `W`-ary grouping from byte 0) makes the same run a different
34
+ subtree at a different offset.
35
+
36
+ **The distribution is load-bearing.** Each of these has been changed
37
+ experimentally, and each change broke tests:
38
+
39
+ | Change | Tests broken |
40
+ | ----------------------- | --------------- |
41
+ | cut rate | 15–18 |
42
+ | bits read | 19–21 |
43
+ | normalized chunking | 5–6 |
44
+ | relaxing the forced cut | the same suites |
45
+
46
+ The forced cut accounts for 32% of all cuts. Every mechanism downstream is
47
+ fitted to this distribution, so any change here must be re-measured on the whole
48
+ suite. A second copy of the hash loop, which `contentBoundaries` once carried,
49
+ lets the write side and the read side drift without a type error.
2
50
 
3
- > **Law:** `perceiveDeposit` ≡ `perceive` — same bytes ⇒ same tree and same node
4
- > id. Deposit imposes nothing (no boundaries, no turn convention). Geometry
5
- > never sees conversation metadata.
51
+ ## One shape, two algebras
6
52
 
7
- ## The identity
53
+ `groupByLevel` and `foldSlice` are written once, over a fold algebra (`join`,
54
+ `key`), and run over two algebras:
8
55
 
9
- Perception is a pure function of the bytes. The deposit path and the inference
10
- path compute the same content-defined fold for the same input, so a trained
11
- context node and `resolve(query)` reach the same node. When the two sides
12
- disagreed, alignment went quadratic (measured 5.2M cells on a 476-byte context
13
- vs 0 when they agree) and cumulative contexts stopped resolving to what they
14
- were trained as.
56
+ - **The vector fold (`perceive`)** produces gists.
57
+ - **The identity fold (`contentIdentity`)** names the node a stream folds to by
58
+ asking the store bottom-up as it folds.
15
59
 
16
- ## Deposit imposes nothing
60
+ Inside an over-long row, grouping also reads each item's `itemKey`: eight raw
61
+ gist coordinates. The identity fold computes only those coordinates, lazily,
62
+ with the same float32 additions in the same order. `exactNode` is the identity
63
+ fold, so resolving a span builds no `D`-dimensional gist. On the 31.7M-node
64
+ store, one join query went from 185 s and 4.4 GB retained to 40 s and 81 MB,
65
+ with the same answer. Written once, the two folds cannot disagree about the
66
+ tree.
17
67
 
18
- No boundaries, no turn convention, nothing read out of the bytes. Conversational
19
- turn offsets are API metadata — they feed `ConversationState`, `answeredSpans`
20
- and `currentTurnStart`; the geometry never sees them. Passing turn boundaries
21
- into the fold is a correctness bug, not a tuning choice.
68
+ ## The read side names as the write side names
22
69
 
23
- ## Boundaries vs reuse — two problems
70
+ The store's `intern` names a branch by its children first. When the children
71
+ name nothing, it reuses the flat node over the same bytes ("same bytes, same
72
+ node"). Deposits intern flat nodes, so a later deposit's branch is often stored
73
+ as an earlier deposit's window: `ver` + `!` is stored as the window `ver!`.
74
+ `branchNaming` (`mind/primitives.ts`) applies the same order on the read side,
75
+ in `exactNode` and `foldTree`, and an unnamed child does not settle the
76
+ question. On the 31.7M-node store, 7 of 80 dialogue turns asked verbatim used to
77
+ resolve to nothing, costing 26–41 s each on the composition path. They now
78
+ resolve to their own context.
24
79
 
25
- `contentFoldIncremental` and `stablePrefixFold` solve different problems;
26
- conflating them is what once put an imposed boundary set on the inference path.
80
+ A name found only through the bytes is a flat entry, not a learnt structure. In
81
+ that case `resolve` also asks the canonical class (`exactNaming`'s `byBytes`)
82
+ for the member that leads somewhere. That class is an optional capability
83
+ (`store.md`), and its candidates are hash-then-verified.
27
84
 
28
- | Mechanism | What it buys | Cost / shape |
29
- | ------------------------ | --------------------------------------------------------- | ----------------------------------------------------------------------------------------------- |
30
- | `contentFoldIncremental` | Transparent segment reuse (cost only) | Imposes nothing; tree identical to the cold fold |
31
- | `stablePrefixFold` | Caller-supplied cuts left-nested for prefix-ROOT identity | One prefix-ROOT per cut becomes an identical subtree (and same node id) inside the grown stream |
85
+ Known limit: a context stored with cuts at the ends of earlier turns does not
86
+ resolve. No current deposit path produces that shape, and reproducing it would
87
+ mean guessing turn boundaries, which this contract forbids.
32
88
 
33
- Both carry the same precondition: `prev` must be a fold of a byte-identical
34
- prefix — reuse is keyed on `[start,end)` offsets, which cannot witness byte
35
- agreement. A mismatched `prev` produced a wrong tree on 336 of 400 random
36
- streams. `perceiveDeposit` discharges this via the prefix bytes as cache key; a
37
- conversation advances only by append. A caller that cannot prove the prefix must
38
- pass no `prev`.
89
+ ## Canonical windows — finding a span whatever the fold did
39
90
 
40
- ## Identity must not depend on W or absolute offset
91
+ Besides its tree, every deposit interns flat nodes:
41
92
 
42
- `contentLevels` is the single boundary rule (rolling hash over a bounded
43
- window). Any grouping by index — stride, tile, fixed-arity row — reintroduces
44
- the grid's phase bug: `riverFold` groups `W`-ary from byte 0, so the same byte
45
- run is a different subtree at a different offset. The content-defined hash
46
- removes this: a change upstream moves only the cut it falls inside; downstream
47
- cuts and segments are unchanged (99.7% cuts preserved on real deposits after
48
- shifts of 1..7 bytes vs 14.3% for the grid). The `groupByLevel` above the
49
- segments splits by content level, not by count.
93
+ - its whole stream;
94
+ - each window of `W−1` and `W` bytes, contained by the chunks it overlaps
95
+ (`canonicalWindows`, `indexSubSpans`).
50
96
 
51
- ## `contentLevels` is single source; `contentBoundaries` is projection
97
+ Recognition's canonical reading chains these windows up to `chainReach = W²`
98
+ leaf ids from a position. This write/read pair finds a form embedded at an
99
+ offset where the question's own cuts do not line up with it (`canonical.ts`).
52
100
 
53
- `contentBoundaries(space, bytes)` is `contentLevels(space, bytes).cuts`. It once
54
- carried its own rolling-hash loop, which is how a write side and a read side
55
- drift without a type error. Levels are read from the hash the cut was accepted
56
- at — level `L` when `h` vanishes mod `W^(L+1)` — so level-`L` cuts nest inside
57
- level-`(L-1)` and expected span is `W^(L+1)` bytes.
101
+ ## Incremental folds — reuse is not boundaries
58
102
 
59
- ## One shape, two algebras
103
+ `contentFoldIncremental` reuses already-folded segments of a byte-identical
104
+ prefix. The result is the cold fold's tree, so the reuse changes cost only. It
105
+ is the only fold any internal path computes: `perceiveDeposit` keys the reuse on
106
+ the prefix bytes, and a conversation advances only by appending.
60
107
 
61
- Which items group under which parent is decided by the cut levels, the keyring
62
- and — inside an over-long row — each item's `itemKey` (eight raw gist
63
- coordinates). `groupByLevel`/`foldSlice` are written ONCE over a fold algebra
64
- (`join`, `key`) and run over two of them: the vector fold (`perceive`) and the
65
- identity fold (`contentIdentity`), which names the node a stream folds to by
66
- asking the store bottom-up and reads only the coordinates `itemKey` hashes,
67
- lazily, with the same float32 additions in the same order. `exactNode` is the
68
- identity fold, so resolving a span builds no D-dimensional gist and leaves
69
- nothing in the perception memo (measured on the 31.7M-node store: one join query
70
- went 185 s / 4.4 GB retained → 40 s / 81 MB, same answer). A grouping rule
71
- written twice would be a write/read drift waiting to happen; written once, the
72
- two folds cannot disagree about the tree.
73
-
74
- ## Same tree is not enough — the read side names as the write side names
75
-
76
- The store's `intern` names a branch by its children, and when they name none it
77
- looks up the FLAT node over the same bytes and REUSES it (step 1b, "same bytes,
78
- same node"). Every deposit interns flat nodes — its whole input and each
79
- canonical window — so a later deposit's branch is often stored as an earlier
80
- deposit's window (`ver` + `!` stored as the window `ver!`). The same tree is
81
- then named by a node no child lookup can reach. `exactNode` and `foldTree` name
82
- a branch exactly as `intern` does (`branchNaming`, `src/mind/primitives.ts`):
83
- the children first, the flat node over the same bytes when they name nothing —
84
- and an unnamed child does not settle it, since the write side minted that child
85
- and still reached step 1b. Measured on the 31.7M-node store: 7 of 80 dialogue
86
- turns asked verbatim had resolved to nothing and fell to the composition path
87
- (26–41 s); they now resolve to their own context.
88
-
89
- A name found only through the bytes is where the exact lookup used to MISS, and
90
- a flat index entry is not a learnt structure, so `resolve` still asks the
91
- canonical class there (`exactNaming`'s `byBytes`): it holds the learnt member
92
- that leads somewhere, when there is one. The store holds such byte-only names
93
- from an earlier deposit path (today's deposits always intern the structure
94
- beside the flat copy), which is why `test/152` builds the state through the
95
- store's write API.
96
-
97
- One stored turn of the same set still does not resolve: its context was stored
98
- with cuts at the ends of the conversation's EARLIER contexts (103 and 227 bytes
99
- into a 257-byte context) — a boundary-imposed shape neither today's deposit path
100
- nor the training-time one produces from these bytes, and that the read side
101
- could reproduce only by guessing turn boundaries, which this contract forbids.
102
-
103
- ## Optional canonical capability
104
-
105
- `canonAdd`/`canonFind` (`src/store.ts` — `canonCount`/`eachContent`) is an
106
- optional backend capability. A backend may omit all four; resolution then has no
107
- equivalence fallback. The store never learns the equivalence — the canonicalizer
108
- (`Canon` in `src/canon.ts`, e.g. `textCanon`) is injected by the caller and
109
- every candidate is hash-then-verified (re-canonicalize stored bytes, compare). A
110
- hash collision costs a read, never a wrong id.
111
-
112
- ## Cost of changing the cut distribution
113
-
114
- The cut rate, which bits are read, `minLen`/`maxLen`, and the forced cut at
115
- `maxLen` set the segment distribution every downstream mechanism is fitted to.
116
- Each has been changed experimentally and cost 5–21 tests (rate: 15–18, bits:
117
- 19–21, normalized chunking: 5–6). `W-1` is the minimum (one window minus one);
118
- `seats.length` is the maximum (one flat node folds exactly one segment). The
119
- forced cut is load-bearing — relaxing it to reduce the current 32% forced rate
120
- looked like a tidying but broke the same suites. Re-measure the whole suite for
121
- any change here.
108
+ `stablePrefixFold` is a separate, public geometry capability. It takes cuts
109
+ supplied by the caller and nests them to the left, so that each prefix root is
110
+ an identical subtree inside the grown stream. No path inside the mind supplies
111
+ such cuts, and confusing the two capabilities is how an imposed boundary set
112
+ once reached inference.
113
+
114
+ Both capabilities require `prev` to be a fold of a byte-identical prefix. Reuse
115
+ is keyed on offsets, and offsets cannot witness that the bytes agree: a
116
+ mismatched `prev` produced a wrong tree on 336 of 400 random streams. A caller
117
+ that cannot prove its prefix must pass no `prev`.
122
118
 
123
119
  ## Pins
124
120
 
125
- - `test/59` — shift invariance floors (content-defined cuts preserved over
126
- random binary and prose).
127
- - `test/63` — offset/W invariance and `contentLevels` distribution expectations.
128
- - `test/148` — the identity fold names exactly what the vector fold names (every
129
- sub-span of corpus and noise), and groups exactly as it does over long
130
- low-entropy streams that force the `itemKey` split.
121
+ - `test/59` — shift invariance: content-defined cuts survive shifts in random
122
+ binary and in prose.
123
+ - `test/63` — offset and `W` invariance, and the expected `contentLevels`
124
+ distribution.
125
+ - `test/148` — the identity fold names exactly what the vector fold names, and
126
+ groups as it does over long low-entropy streams that force the `itemKey`
127
+ split.
131
128
  - `test/152` — a deposit stored through an earlier deposit's flat window
132
- resolves to its own context, both folds name it, and it is answered on the
133
- exact path.
134
-
135
- See:
136
- `src/geometry.ts:contentLevels`/`contentBoundaries`/`contentFoldIncremental`/`stablePrefixFold`/`contentIdentity`;
137
- `AGENTS.md` §2 invariants.
129
+ resolves to its own context, and is answered on the exact path.
@@ -1,99 +1,108 @@
1
- # Halo & Sketch — Distributional Memory
2
-
3
- A node's **halo** is its distributional signature: the superposition of
4
- identity-bound company signatures poured from every episode it participated in.
5
- Its **gist** is the VSA fold of its own bytes — content, not company.
6
-
7
- ## Two vectors per node, two indexes
8
-
9
- | Vector | Encodes | Index | Query |
10
- | ------ | ----------------------------- | ------------- | -------------- |
11
- | gist | what the node _is made of_ | content index | `resonate` |
12
- | halo | what _company_ the node keeps | halo index | `resonateHalo` |
13
-
14
- Both indexes are RaBitQ-IVF (`src/rabitq-ivf/`) — 1-bit ANN over the same
15
- vectors; scores are estimates, never identity. Halos are also persisted durably
16
- (see below).
17
-
18
- ## Quantization — 2-bit on disk, float in session
19
-
20
- A halo is a superposition of quasi-orthogonal signatures, so coordinates are
21
- Gaussian. Storage exploits this:
22
-
23
- - **In-session accumulator** (`_haloExact` in `src/store.ts`): `Float32Array` —
24
- exact, additive, incremented by `pourHalo`.
25
- - **Durable row** (`_dbUpsertHalo`): 2-bit Lloyd–Max quantizer — decision at
26
- ±0.9816σ, levels at ±0.4528σ / ±1.5104σ, σ derived from the stored norm
27
- (`norm/√D`). Header is the 4-byte norm; body is 2 bits/coordinate. Keeps ≥0.88
28
- correlation with the exact vector.
29
- - **ANN index**: 1-bit RaBitQ, irreversible — answers only "which halos are near
30
- this query?"
31
-
32
- Re-indexing is geometric: a halo re-enters the ANN when mass is small or crosses
33
- a power of two (`geometricMass`), so index writes are O(log mass).
34
-
35
- ## Bottom-k sketch & company profile
36
-
37
- A whole-partner signature alone records tokens, not types — halos of genuine
38
- synonyms would be quasi-orthogonal. `companyProfile` (`src/mind/learning.ts`)
39
- superposes:
40
-
41
- 1. the partner's own identity signature, plus
42
- 2. the bottom-k **constituent sketch** — the
43
- `k = profileCapacity(D) = floor(√D)` minimal units of its subtree with
44
- smallest `unitPriority`, deduped.
45
-
46
- The sketch is composable (bottom-k of a union = bottom-k of children's
47
- sketches), durable derived state via `sketchGet`/`sketchPut`, and bounded: at
48
- most `k` constituents are classified, each by one `LIMIT`ed parent read
49
- (`hubBound`). Beyond `√D` terms a single constituent contributes less than
50
- `1/√D` — below RaBitQ noise — and extra terms shrink every accepted one; the cap
51
- is a correctness limit.
52
-
53
- ## Gist vectors
54
-
55
- Folded by the river (`src/geometry.ts`): leaves are alphabet vectors, groups
56
- bind by two-ended seats, intermediate gists stay unnormalized (magnitude ∝
57
- √len), only the root is normalized. Gist resonance reads byte-proportional
58
- overlap; halo resonance reads distributional overlap — the two are independent.
59
-
60
- ## Thresholds & gating
61
-
62
- All bars are derived in `src/geometry.ts`; no tunable constant:
63
-
64
- | Symbol | Formula | Use |
65
- | ------------------ | -------------- | ------------------------------------------------------------------------------------------------- |
66
- | `estimatorNoise` | `1/√D` | 1σ RaBitQ noise; contrastive margin must clear it |
67
- | `significanceBar` | `3/√D` | whole-query relatedness — 3σ above chance; gates consensus climb and `analogyStrength` |
68
- | `conceptThreshold` | `0.5 + 0.5/√D` | halo concept sharing — structural midpoint + ½σ; gates `haloSiblings`, concept hops, articulation |
69
-
70
- The significance bar gates the whole query; `conceptThreshold` gates per-pair
71
- halo cosine.
72
-
73
- ## Probes: `haloMass` and `hasHalo`
74
-
75
- - `haloMass(id)` — count of poured episodes; evidence weight, tie-breaker in
76
- `chooseAmong`.
77
- - `hasHalo(id)` — existence probe (indexed point check, no vector decode);
78
- mirrors `halo(id) !== null`. One tier of the `leadsSomewhere` admission
79
- predicate (with `hasNext`/`hasParents`).
80
-
81
- Both are `meter`-counted probes, not full decodes — `halo(id)` is the bounded
82
- vector read; `resonateHalo` is the IVF ANN query.
83
-
84
- ## Relation to invariants
85
-
86
- - **Derived thresholds** — all bars above live in `geometry.ts`.
87
- - **Exact decides / approximate proposes** — halo scores rank and gate; identity
88
- is content-addressed. The graded ladder is exact → halo → gist
89
- (`mind/match.ts`).
90
- - **Bounded reads** — `hasHalo`/`haloMass` are point probes; `resonateHalo` is
91
- capped ANN; constituent classification uses `LIMIT hubBound+1` reads. No
92
- per-query scan grows with corpus.
1
+ # Halo — Distributional Memory
2
+
3
+ > **Law:** a node's halo is the superposition of the company it was deposited
4
+ > in. It measures nearness of _use_, independently of the gist, which measures
5
+ > nearness of _form_. It is a statistic over the record: it may propose, never
6
+ > decide (`exact-vs-approximate.md`).
7
+
8
+ The gist puts `colour` near `colours`. Only the halo can put `colour` near
9
+ `hue`, two words whose bytes have nothing in common.
10
+
11
+ ## How a halo is poured
12
+
13
+ Each deposited fact `(context → continuation)` makes two pours (`ingestPair`):
14
+
15
+ - every part of the context newly seen by this deposit receives the
16
+ continuation's profile, bound to seat 1, _what it led to_;
17
+ - the continuation receives each such part's profile, bound to seat 0, _what led
18
+ to it_.
19
+
20
+ Pours add. Each profile is normalized, so one episode pours one unit of mass
21
+ (`haloMass` counts episodes). Addition forgets order and keeps proportion, which
22
+ is why the halo is a statistic rather than a record.
23
+
24
+ ## Company by type, not by token — the profile
25
+
26
+ A partner's own signature (`companySignature`, a random unit vector seeded by
27
+ its node id) records a token: "occurred near node #4711992". Whole deposits
28
+ almost never recur (whole-span dedup on the trained store is 0.98×), so the
29
+ halos of genuine synonyms came out quasi-orthogonal. The best distributional
30
+ sibling of `Eiffel Tower` scored 0.146 against a bar of 0.516.
31
+
32
+ So `companyProfile` superposes the partner's own signature with the signatures
33
+ of its **bottom-k constituent sketch**: the `k = profileCapacity(D) = ⌊√D⌋`
34
+ minimal units of its subtree with the smallest identity-keyed priority.
35
+
36
+ - **The sketch descends to every depth.** Content-defined cuts depend on the
37
+ surrounding bytes, so a shared unit such as `Paris` is often not a top-level
38
+ chunk of either partner. A depth-1 read measured 0.0319 against 0.0416 for an
39
+ unrelated control: no signal.
40
+ - **The sketch is independent of arrival order.** A stop rule like "descend
41
+ until a unit is attested twice" depends on which partner was trained first,
42
+ and its pairs never met (0.0165 against 0.0375). The priority is a hash of the
43
+ unit's id, so a unit shared by two partners is kept by both or by neither.
44
+ - **Hubs are excluded, but still descended into.** A unit with more than `√N`
45
+ parents (`is`, `the`) would put one shared term into every profile, all halos
46
+ would correlate, and the concept bar's null model would collapse. Byte atoms
47
+ are skipped for the same reason. So is a unit that dominates the partner,
48
+ because that unit is the partner itself.
49
+ - **The sketch is composable and durable.** Bottom-k of a union is bottom-k of
50
+ the children's sketches. Sketches are stored (`sketchGet`/`sketchPut`), so a
51
+ partner met again costs `O(k)`.
52
+
53
+ Every term is a seeded function of a node identity, never of a gist, so
54
+ resemblance of spelling cannot leak into resemblance of use. Two partners
55
+ sharing `j` of `k` discriminating constituents meet at `j/(1+k)`. A profile is
56
+ fixed given the corpus seen so far. As `N` grows, a term can move only from
57
+ contributing to excluded-as-hub.
58
+
59
+ `k = √D` is a correctness limit, not a budget. Beyond it, one more term
60
+ contributes less than `1/√D`, below the estimator's noise, and dilutes every
61
+ accepted term.
62
+
63
+ ## Storage — exact in session, quantized at rest
64
+
65
+ | Layer | Form |
66
+ | --------- | ----------------------------------------------------------------------------------------------------------------------- |
67
+ | session | `Float32Array` accumulator, exact and additive (`pourHalo`) |
68
+ | durable | 2-bit Lloyd–Max quantizer: decision at ±0.9816σ, levels ±0.4528σ / ±1.5104σ, σ = norm/√D. Correlation ≥ 0.88 with exact |
69
+ | ANN index | 1-bit RaBitQ, answering only "which halos are near this one?" |
70
+
71
+ A halo re-enters the index when its mass is at most 4 or a power of two, so
72
+ index writes are `O(log mass)`.
73
+
74
+ Halo comparisons do not depend on the configured seed, because signatures are
75
+ keyed on node ids. Gist comparisons do: a Mind built with a seed other than its
76
+ store's makes every gist comparison meaningless, while its halo comparisons
77
+ remain valid.
78
+
79
+ ## Where halos are read
80
+
81
+ | Use | Gate |
82
+ | ---------------------------------------------------------------------------------------------- | ------------------ |
83
+ | admission: a form _leads somewhere_ if `hasNext \|\| hasHalo` (`store.md`) | existence probe |
84
+ | `locate`'s second tier (halo role), `alignGraded`'s halo-matched sites | best halo mate |
85
+ | synonym tiers: junction bridge, `crossRegionVotes`, substitution bridge | `conceptThreshold` |
86
+ | concept hops in `cover`: an edge-less form borrows a halo sibling's continuation, at `CONCEPT` | `conceptThreshold` |
87
+ | CAST's analogy strength | `significanceBar` |
88
+ | `chooseNext`: halo mass breaks ties after `prevCount` | ordering only |
89
+
90
+ The bars are derived in `geometry.ts` (`thresholds.md`):
91
+
92
+ - `estimatorNoise = 1/√D`;
93
+ - `significanceBar = 3/√D`;
94
+ - `conceptThreshold = 0.5 + 0.5/√D`, the structural midpoint plus half a noise
95
+ unit.
96
+
97
+ `haloMass` and `hasHalo` are point probes. `halo(id)` is the bounded vector
98
+ read, and `resonateHalo` is the capped ANN query.
93
99
 
94
100
  ## Pins
95
101
 
96
- - `test/08 storage halo` — halo persistence, 2-bit round-trip,
97
- `haloMass`/`hasHalo` contract, index survival across reopen.
98
- - `test/35 ivf` — RaBitQ-IVF contract (recall vs brute force, sublinear
99
- scaling); covers the halo index's own layer.
102
+ - `test/08` — halo persistence, the 2-bit round trip, the `haloMass`/`hasHalo`
103
+ contract, and the index surviving a reopen.
104
+ - `test/76-type-level-company` — company shared by type: synonyms meet through
105
+ their constituents.
106
+ - `test/29` C1 — CAST's halo-tier analogy evidence. A profile polluted by byte
107
+ atoms silenced it.
108
+ - `test/35-ivf` — the RaBitQ-IVF contract under the halo index.