@hviana/sema 0.9.3 → 0.9.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (46) hide show
  1. package/AGENTS.md +16 -7
  2. package/dist/src/mind/learning.js +11 -12
  3. package/dist/src/mind/mind.js +4 -4
  4. package/dist/src/mind/types.d.ts +10 -15
  5. package/dist/src/store.d.ts +10 -10
  6. package/dist/src/store.js +5 -5
  7. package/docs/INDEX.md +62 -59
  8. package/docs/INVARIANTS.md +36 -18
  9. package/docs/PHILOSOPHY.md +317 -0
  10. package/docs/architecture/bounded-reads.md +31 -71
  11. package/docs/architecture/caches.md +61 -81
  12. package/docs/architecture/closure.md +88 -107
  13. package/docs/architecture/commonality.md +47 -38
  14. package/docs/architecture/cost-model.md +57 -79
  15. package/docs/architecture/determinism.md +43 -55
  16. package/docs/architecture/evidence.md +158 -235
  17. package/docs/architecture/exact-vs-approximate.md +41 -35
  18. package/docs/architecture/factored-machinery.md +34 -20
  19. package/docs/architecture/fold-contract.md +110 -118
  20. package/docs/architecture/halo-sketch.md +105 -96
  21. package/docs/architecture/match-project.md +51 -42
  22. package/docs/architecture/mechanism-market.md +87 -91
  23. package/docs/architecture/memoization.md +60 -74
  24. package/docs/architecture/meter.md +37 -47
  25. package/docs/architecture/saturation.md +75 -101
  26. package/docs/architecture/store.md +118 -99
  27. package/docs/architecture/thresholds.md +66 -73
  28. package/docs/failures/tempting-but-wrong.md +139 -165
  29. package/docs/harness/gates.md +27 -32
  30. package/docs/mechanisms/alu.md +22 -69
  31. package/docs/mechanisms/cast.md +76 -71
  32. package/docs/mechanisms/confluence.md +22 -29
  33. package/docs/mechanisms/cover.md +58 -66
  34. package/docs/mechanisms/extraction.md +33 -37
  35. package/docs/mechanisms/prefix-completion.md +36 -39
  36. package/docs/mechanisms/recall.md +60 -53
  37. package/docs/mechanisms/reference.md +63 -49
  38. package/jsr.json +1 -1
  39. package/package.json +1 -1
  40. package/src/alu/README.md +90 -298
  41. package/src/derive/README.md +94 -256
  42. package/src/mind/learning.ts +11 -12
  43. package/src/mind/mind.ts +4 -4
  44. package/src/mind/types.ts +10 -15
  45. package/src/rabitq-ivf/README.md +11 -8
  46. package/src/store.ts +5 -5
@@ -1,104 +1,78 @@
1
- # Saturation — Named Stop, Not Cap
2
-
3
- > **Law:** every walk names a deciding saturation beside its cap. The cap is a
4
- > safety net; saturation is the derived stop that decides and terminates.
5
-
6
- A walk with only a cap drifts to the cap. A walk with a saturation stops the
7
- moment the answer (saturated vs. decided) is known — bounded, exact below the
8
- bound, and named in the trace.
9
-
10
- ## Cap vs. saturation
11
-
12
- | Role | Value | Nature |
13
- | ---------- | ------------------------------------------------------------ | ------------------------------------------------ |
14
- | Cap | `hubBound = ceil(sqrt(N))` per read; `hubBound * W` per walk | Safety net — prevents corpus-proportional work |
15
- | Saturation | Named derived stop (`SaturationReason`) | Decision — proves continuing cannot discriminate |
16
-
17
- `N = corpusN = max(2, edgeSourceCount())`, `W = maxGroup`. Defined once in
18
- `mind/traverse.ts` (`corpusN`, `hubBound`, `boundFor`); never spelled inline.
19
-
20
- ## edgeAncestors — EXPAND-UNTIL-DECIDED
21
-
22
- `mind/traverse.ts:edgeAncestors` is the model: it climbs the structural DAG
23
- (parents + containment) until one of five saturations decides, and is exact
24
- below every one of them.
25
-
26
- 1. **Predecessor fan-in** — `prevCount(node) > bound` via `store.prevCount`. One
27
- indexed count; no read. Proves `> bound` distinct contexts on its own.
28
- 2. **Distinct-context limit** — `ctxSeen.size > bound` after
29
- `prevFirst(node, bound)`. The accumulated set of distinct contexts reachable
30
- from roots visited so far exceeds `sqrt(N)`.
31
- 3. **Parent fan-out** — `parentsFirst(node, bound+1).length > bound`. One
32
- `LIMIT bound+1` read distinguishes "exactly bound parents" from "hub". The
33
- node itself is not expanded; saturated reaches are never voted.
34
- 4. **Lateral-cone cumulative** — `lateral > bound`, where `lateral` sums
35
- `fresh-1` over every expanded node's extra parents beyond its first. The
36
- per-node guard catches concentration at one node; this catches the same
37
- commonness distributed across the cone. A deep chain in one structure accrues
38
- zero laterals and still reaches its root at any depth.
39
- 5. **Byte-atom commonality** — `atomIsHub(N,W)` when
40
- `atomReach(N,W) = max(1, ceil(N*W/256)) > bound`. Atoms carry no kid/contain
41
- rows, so containment is unmeasurable; the uniform-expectation floor replaces
42
- it. Above the scale the atom abstains as voter (edges remain traversable for
43
- tier-0 recall).
44
-
45
- Below every threshold the walk is exact — `prevFirst(bound)` is the full list,
46
- `parentsFirst(bound+1)` is the full list, `containersSlice` pages are walked in
47
- full — identical to the unbounded climb. Work is `O(bound)` contexts times local
48
- structure, never `O(N)`. Container seeding is streamed in `bound`-sized pages
49
- for the same reason.
50
-
51
- Refuted tightening: **container fan-out as a hub** — deciding a containment seed
52
- saturated when its first `containersSlice(bound+1)` page overflows, the reading
53
- `junction.ts` applies to the same links. It conflates the PLACES a window occurs
54
- with the CONTEXTS it reaches, and saturation is defined over contexts. Proved on
55
- `test/49`'s fixture (N = 5, `sqrt(N)` = 3): `fran` sits in 4 containers yet its
56
- full climb reaches 2 contexts — discriminative — and the rule called it
57
- saturated, as it did `of f`, `fra`, `is`, `ranc`, `ance`. On the 31.7M-node
58
- store no such window climbed unsaturated (18 composition queries; those climbs
59
- were 38% of all visits), because at `sqrt(N)` = 1,560 a full page of containers
60
- rarely converges on few contexts — but the rule is the same wrong reading at
61
- both scales. The streamed seed stays.
62
-
63
- Trace records the first deciding stop as
64
- `SaturationStop { reason, node, observed, limit }` plus `visited`/`maxDepth`;
65
- absent when unsaturated or untraced.
66
-
67
- ## pivotInto — longest-wins
68
-
69
- `mind/resonance.ts:pivotInto` ranks candidates by `contentLen(id, answerLen+1)`
70
- descending (first-inserted tie-break). The byte score is length itself, so the
71
- scan is decided at the first candidate that passes every filter — a shorter
72
- candidate can never outscore it. At most one winner's bytes are reconstructed;
73
- every shorter proposal is skipped without a read. Saturation, not cap.
74
-
75
- ## Junction — hub guards vs. budget
76
-
77
- `mind/junction.ts:junctionContainersFrom` has three disciplines:
78
-
79
- - **Phrase-scale reads** — `bytesPrefix(maxContainer+1)` per visit; a node
80
- beyond the cap prunes its branch.
81
- - **Per-node hub guards** (real saturations) — `parentsFirst(bound+1) > bound`
82
- not expanded; one `containersSlice(bound+1)` page beyond `bound` not expanded.
83
- Each is exact below `bound`.
84
- - **Expansion budget** — at most `bound * W` pops total (shared across a tier's
85
- walks). Budget exhaustion is an abstention (`junctionBudgetExhausted`) that
86
- falls through to the resonance tier — a net, not a saturation.
87
-
88
- Refuted tightening: applying `edgeAncestors`' cumulative lateral-cone limit here
89
- would discard half the successful junctions (measured lateral 1425/1426 at
90
- `bound` ~ 570).
91
-
92
- Refuted early-stop: **one-cone-exhausted** — stopping when one side's upward
93
- cone empties — is wrong in both hub-guarded and hub-flagged forms. A junction
94
- can be reachable from only one side when that side's seed is a fold sub-node of
95
- the container (`test/16`: "cold or hot" reached from window "cold" while 3-byte
96
- "hot" cone is empty; `test/34` n-ary binding fails the same way). Exhausting one
97
- cone never proves no junction remains; the walk must keep the `bound*W` net
98
- after per-node saturations.
1
+ # Saturation — A Named Stop, Not a Cap
2
+
3
+ > **Law:** every walk names the saturation that decides when it stops, beside
4
+ > its cap. The cap (`bounded-reads.md`) is a safety net. The saturation is a
5
+ > derived proof that continuing cannot discriminate, and it is recorded in the
6
+ > trace.
7
+
8
+ **Why.** A walk with only a cap drifts to the cap: it spends the whole budget
9
+ and stops at an arbitrary place. A walk with a saturation stops the moment the
10
+ outcome is known, is exact below every stop, and says why it stopped. Deflate's
11
+ match finder truncates long chains arbitrarily. Sema may cap, but it must still
12
+ justify every stop.
13
+
14
+ The trace records the first stop that decides as
15
+ `SaturationStop { reason, node, observed, limit }`, together with `visited` and
16
+ `maxDepth`.
17
+
18
+ ## `edgeAncestors` — expand until decided (`mind/traverse.ts`)
19
+
20
+ This is the model walk. It climbs parents and containers from a node to the
21
+ learnt contexts it reaches, and stops at the first of five saturations. Each is
22
+ exact below its threshold (`bound = √N`):
23
+
24
+ 1. **Predecessor fan-in.** `prevCount(node) > bound`: one indexed count proves
25
+ more than `bound` contexts.
26
+ 2. **Distinct contexts.** `ctxSeen.size > bound` after `prevFirst(node, bound)`.
27
+ 3. **Parent fan-out.** `parentsFirst(node, bound + 1)` returns more than
28
+ `bound`, and the node is not expanded.
29
+ 4. **Lateral cone.** The extra parents summed over every expanded node exceed
30
+ `bound`. Commonness spread across the cone is caught as surely as commonness
31
+ concentrated at one node, while a deep chain inside one structure accrues no
32
+ laterals and reaches its root at any depth.
33
+ 5. **Byte atoms.** Atoms carry no containment rows, so their reach is
34
+ unmeasurable, and the uniform expectation `atomReach(N, W)` replaces it. Past
35
+ `bound`, an atom abstains as a voter (`atomIsHub`), though its edges remain
36
+ traversable.
37
+
38
+ Below every stop, the walk equals the unbounded climb, and its work is
39
+ `O(bound)` contexts. Container seeds are streamed in `bound`-sized pages.
40
+
41
+ **Refuted:** treating container fan-out as a hub. It confuses the _places_ a
42
+ window occurs with the _contexts_ it reaches, and saturation is defined over
43
+ contexts. On `test/49`, `fran` sits in 4 containers yet reaches 2 contexts, and
44
+ the rule called it saturated.
45
+
46
+ ## `pivotInto` — the longest wins (`mind/resonance.ts`)
47
+
48
+ Candidates rank by `contentLen(id, answerLen + 1)`, descending, with
49
+ first-inserted breaking ties. The score is length itself, so the first candidate
50
+ that passes every filter wins, and no shorter candidate's bytes are ever read.
51
+
52
+ ## Junction ascent — guards against a budget (`mind/junction.ts`)
53
+
54
+ `junctionContainersFrom` keeps three disciplines apart:
55
+
56
+ - **Phrase-scale reads.** `bytesPrefix(maxContainer + 1)` per visit; a node past
57
+ the cap prunes its branch.
58
+ - **Hub guards, which are true saturations.** A node is not expanded when
59
+ `parentsFirst(bound + 1)` or one `containersSlice` page exceeds `bound`.
60
+ - **The expansion budget, a net.** At most `bound · W` pops, shared across a
61
+ tier's walks. Running out is an abstention (`junctionBudgetExhausted`), and
62
+ the search falls through to the resonance tier.
63
+
64
+ Two tightenings were refuted:
65
+
66
+ - **The lateral-cone limit from `edgeAncestors`** would discard half the
67
+ successful junctions (lateral 1425 of 1426 at `bound` ≈ 570).
68
+ - **Stopping when one side's cone is exhausted** is wrong. A junction can be
69
+ reachable from only one side when that side's seed is a fold sub-node of the
70
+ container. In `test/16`, `cold or hot` is reached from `cold` while the cone
71
+ of `hot` is empty. Exhausting one cone never proves that no junction remains.
99
72
 
100
73
  ## Pins
101
74
 
102
- - `test/16` — one-cone-exhausted refutation (bridge junction).
103
- - `test/34` — one-cone-exhausted refutation (n-ary cross-region binding).
104
- - `test/27` — saturation-drop gate (leading/trailing saturated intervals).
75
+ - `test/16`, `test/34` — refute stopping when one cone is exhausted, in the
76
+ bridge junction and in n-ary cross-region binding.
77
+ - `test/27` — dropping saturated leading and trailing intervals from the climb.
78
+ - `test/49` — refutes treating container fan-out as a hub.
@@ -1,102 +1,121 @@
1
- # Store — AbstractStore Owns the Domain, Adapters Own the Wires
2
-
3
- > **Law:** `AbstractStore` (`src/store.ts`) owns every domain decision — dedup,
4
- > near-dedup, gist/halo indexing, containment, batching, LRU, compaction.
5
- > `SQliteStore` (`src/store-sqlite.ts`) implements only `_db*`/`_vec*` thin
6
- > wrappers.
7
-
8
- ## Template method
9
-
10
- ```
11
- AbstractStore all logic: caches, merge gates, halo schedule,
12
- buffers, chain transparency, length walks
13
- └─ SQliteStore SQL + VectorDatabase glue — one statement per method
14
- └─ <NewBackend> same contract — subclass AbstractStore only
15
- ```
16
-
17
- A new backend subclasses `AbstractStore`; never re-implements dedup or batching.
18
-
19
- ## IDs and leaves
20
-
21
- Branch ids are dense non-negative `0,1,2,…` (`_nextId`, never deleted).
22
- Single-byte leaves are **implicit** negative ids `-256..-1` (`-(byte+1)`), never
23
- a row. `has(id)` is `id < 0 || id < _nextId`.
24
-
25
- ## Flat branches and bytes
26
-
27
- A branch whose kids are all leaves is **flat** — stored as raw bytes in `leaf`
28
- with an empty `kids` blob as marker (`flatKidsBytes`/`flatBytesKids`). Dedup
29
- probes hash then verify: `hashOf`→`h`→`LIMIT 1` fetch→byte compare (bloom
30
- negative filter first). `flatBranchMayExist(bytes)` is the filter alone — its
31
- `false` is exact, its `true` means "look it up" — for a caller that will verify
32
- anyway: `exactNode` (primitives.ts) refuses a stream whose level-0 segment
33
- cannot exist before naming any node, so a resolve miss costs hashing.
34
- `findFlatBranch` memoizes HITS (`_flatKey`, keyed by the bytes themselves) and
35
- never misses — the filter answers first, so a span that is not stored builds no
36
- key, while the segments the identity fold names, asked by every span that
37
- contains them, cost a map hit. `flatSpans(bytes)` returns a prober over ONE
38
- buffer that answers exactly `findFlatBranch(bytes.subarray(s, e))`. `hashOf` is
39
- FNV-1a, a left fold, so the prober extends the hash each start was last probed
40
- at, and keeps one byte-short copy for the edge scan's trimmed retry. A span
41
- scanner sweeping ends upward pays O(1) per probe instead of O(span).
42
- Recognition's interior pass was hashing O(n·reach²) bytes: 251,660,406 for one
43
- response on the 31.7M-node store (`test/153.4`).
44
-
45
- `leadsSomewhere(id)` — `hasNext || hasHalo` — is the admission predicate's ONE
46
- raw definition; `traverse.ts` memoises its edge tier per response.
47
-
48
- `bytes(id)`/`bytesPrefix(id, cap)` are shared with `BoundedMap` caches — callers
49
- must **never mutate** the returned buffer. `contentLen(id, cap)` walks with
50
- memo; when `cap` is given it saturates (`>= cap` without finishing) so one huge
51
- root never costs a full walk.
52
-
53
- ## Gist, halo, dedup
54
-
55
- On `put*`, content dedup (`hashOf`→probe→mint) gates first. Short keys are
56
- cached (`DEDUP_KEY_MAX` bypass). Near-dedup: `identityBar(D, W, len)`, one
57
- window apart. Gists sit in `_pendingGist` (byte-budgeted `BoundedMap`);
58
- `indexSubtree` & `pourHalo` promote via `_vecContentUpsert`/`_vecHaloUpsert` in
59
- `batchSize` batches. Buffers flush on cadence, `commit()`, and close. Halo mass
60
- re-indexes geometrically (`mass<=4 || powerOfTwo`) and encodes 2-bit quantized.
61
- Canon index is optional: `canonAdd`/ `canonFind`/`canonCount` over 32-bit
62
- canonical hashes, caller verifies bytes. The SQLite backend keeps a negative
63
- filter over the canon hashes too (kept exact on `canonAdd`; the table is never
64
- deleted from), because recognition and the join probe it once per span and
65
- almost every answer is "no such key". It is PERSISTED (`canon_bloom`) in the
66
- same transaction as the canon rows it covers, stamped with that commit's
67
- `canon.upto`; an open whose meta disagrees with the stamp (rows written without
68
- it) rebuilds it by one scan of the h column — seconds on a trained store, paid
69
- once instead of per process (`test/36`).
70
-
71
- ## Containment, batching, LRU
72
-
73
- `addContainer(child,parent)` buffers per child; flush appends via
74
- `_dbAppendContain` (packed pages, geometric merge) — never rewrite the whole
75
- list. `containersSlice` pages through it. Edges and kids write through the same
76
- deferred transaction.
77
-
78
- Every in-memory cache is a `BoundedMap` with byte accounting and eviction (`lru`
79
- vs `smallest` + `clock`/`reorder` recency). ANN reads
80
- (`resonate`/`resonateHalo`) are content-addressed (`vecKey`) and dropped on any
81
- index mutation; `RESONATE_CACHE_MAX=4096`.
82
-
83
- ## Maintenance (incremental)
84
-
85
- - `compactContentIndex(minParents)` — scans only entries since last watermark
86
- (`_vecContentEntriesSince`), removes indexed-but-isolated nodes (`<minParents`
87
- parents, no edges/halos), compacts the vector DB.
88
- - `repairContentIndex(regenerateGist)` — walks `_dbEdgeOrHaloIds()` candidates
89
- only, re-inserts missing bridge nodes whose gists were evicted before
90
- indexing.
91
- - `buildCanonIndex` (`_buildCanonIndex`) — iterates `eachContent(fromId)` and
92
- `canonAdd`s; `fromId` makes refresh incremental.
93
-
94
- Full scans (`parents()`, `next()`, `containers()`) are maintenance-only — hot
95
- paths use `LIMIT`ed probes (`parentsFirst`/`nextFirst`/`prevFirst`), `has*`,
96
- `prevCount`.
1
+ # Store — A Content-Addressed DAG, One Owner of Its Logic
2
+
3
+ > **Law:** a node is named by its content, and equal content is stored once.
4
+ > `AbstractStore` (`src/store.ts`) owns every domain decision: dedup, merge,
5
+ > indexing, containment, batching, caching and maintenance. A backend
6
+ > (`store-sqlite.ts`) implements only thin `_db*`/`_vec*` wrappers.
7
+
8
+ ## Nodes
9
+
10
+ | Kind | Id | Stored as |
11
+ | ----------- | ----------------------------------------------- | -------------------------------------------------------- |
12
+ | byte leaf | implicit, `-(byte+1)` in `−256…−1` | nothing — it always exists |
13
+ | branch | dense `0,1,2,…`, minted in order, never deleted | its ordered child ids |
14
+ | flat branch | as a branch | raw bytes, with an empty `kids` marker (`flatKidsBytes`) |
15
+
16
+ An id is an arbitrary mark; content decides which mark a thing gets. A subtree
17
+ shared by a thousand deposits is one node with a thousand parents, so storage
18
+ grows with distinct content, not with volume, and every span the fold produced
19
+ can be addressed.
20
+
21
+ ## Interning — dedup, then same bytes, then near, then mint
22
+
23
+ `intern` runs four steps, in order:
24
+
25
+ 1. **Exact dedup.** `findLeaf` or `findBranch` looks for equal content: a cache
26
+ first, then a durable probe, so dedup survives a cold cache or a resumed run.
27
+ 2. **Same bytes, same node.** A branch whose children name nothing reuses the
28
+ flat node over the same bytes. Its gist is re-captured, because the same
29
+ bytes fold differently standing alone and embedded, and the index must hold
30
+ the gist that a direct query will present (`fold-contract.md`).
31
+ 3. **Near dedup.** This applies to branches only, and only against whole
32
+ experiences still in the write buffer. The gist proposes a candidate, which
33
+ must reach the branch's own `identityBar`. The bytes then decide: the two
34
+ must be identical except for one span of at most `W` bytes
35
+ (`differsByOneWindow`). This is the only place two different contents share
36
+ an id.
37
+ 4. **Mint** a new id.
38
+
39
+ A refuted alternative: probing the flushed ANN index for near targets. It was
40
+ the dominant training cost, and it was wrong. The 1-bit code ranked a
41
+ byte-distinct branch as nearest and collapsed two different subtrees onto one id
42
+ (`test/02`).
43
+
44
+ ## Relations on the nodes
45
+
46
+ - **Continuation edges.** An edge is a unique `(src, dst)` pair, so
47
+ `prevCount(dst)` counts distinct establishing contexts, never repetitions.
48
+ - **Containment.** `addContainer(child, parent)` buffers per child, and the
49
+ flush appends packed pages (`_dbAppendContain`) without rewriting the list.
50
+ `containersSlice` pages through it.
51
+ - **Gist index** (content) and **halo index** (company). Both are RaBitQ-IVF ANN
52
+ over node ids. Gists wait in `_pendingGist` until `indexSubtree` promotes them
53
+ in batches. Halos accumulate exactly in session and persist 2-bit quantized
54
+ (`halo-sketch.md`).
55
+ - **Sketches.** `sketchGet`/`sketchPut` hold durable derived state: a miss costs
56
+ time, never correctness.
57
+
58
+ `leadsSomewhere(id) = hasNext || hasHalo` is the one raw admission predicate: a
59
+ form is knowledge only if it continues somewhere or keeps company.
60
+
61
+ ## Probes that do not grow with N
62
+
63
+ Hot paths read through `LIMIT` variants, point probes and capped byte reads, and
64
+ full scans are for maintenance only (`bounded-reads.md`). Flat lookups are
65
+ hashed, then verified:
66
+
67
+ - `flatBranchMayExist` is a negative filter, so its `false` is exact.
68
+ - `findFlatBranch` memoizes hits only, and a miss builds no key.
69
+ - `flatSpans(bytes)` extends one FNV-1a hash per start position, so sweeping a
70
+ span's ends costs O(1) per probe instead of O(span). Recognition's interior
71
+ pass once hashed 251,660,406 bytes for a single response (`test/153.4`).
72
+
73
+ `bytes(id)` and `bytesPrefix(id, cap)` return buffers shared with the caches.
74
+ Callers must never mutate them.
75
+
76
+ ## The canonical index — optional, injected, verified
77
+
78
+ `canonAdd`, `canonFind`, `canonCount` and `eachContent` are an optional backend
79
+ capability. Without it, `resolve` has no equivalence fallback.
80
+
81
+ - **The store never learns the equivalence.** The caller injects the
82
+ canonicalizer (`Canon`, for example `textCanon`).
83
+ - **Every candidate is hash-then-verified:** the stored bytes are
84
+ re-canonicalized and compared, so a collision costs a read, never a wrong id.
85
+ - **`buildCanonIndex` runs after training** and is incremental from a watermark.
86
+ A fixture store that skips it under-reports every case variant.
87
+ - **SQLite keeps a negative filter over the canonical hashes** (`canon_bloom`).
88
+ It is persisted in the same transaction as the rows it covers and stamped with
89
+ `canon.upto`. An open whose stamp disagrees rebuilds the filter once
90
+ (`test/36-bloom`).
91
+
92
+ ## Caches and maintenance
93
+
94
+ Every in-memory cache is a byte-budgeted `BoundedMap` (`caches.md`). ANN read
95
+ caches are keyed by content (`vecKey`), dropped on any index mutation, and
96
+ cleared at `RESONATE_CACHE_MAX`.
97
+
98
+ Maintenance is incremental:
99
+
100
+ - `compactContentIndex` removes isolated index entries written since its last
101
+ watermark.
102
+ - `repairContentIndex` re-inserts bridge nodes whose gists were evicted before
103
+ they were indexed.
104
+ - `buildCanonIndex` extends the canonical index from a given id.
97
105
 
98
106
  ## Adding a backend
99
107
 
100
- Implement every `protected abstract _db*`/`_vec*` in `src/store.ts` as a thin
101
- wrapper around your storage. Keep `_dbGet*First`/`Slice` as real `LIMIT` queries
102
- and `has*`/`COUNT` as point probes — never materialise-then-slice.
108
+ 1. Subclass `AbstractStore` and implement every
109
+ `protected abstract _db*`/`_vec*` method as a thin wrapper. Never
110
+ re-implement dedup or batching.
111
+ 2. Keep the `*First`/`*Slice` methods as real `LIMIT ?` queries, and keep `has*`
112
+ and counts as point probes. Never materialise a list and then slice it.
113
+ 3. Run the whole suite with your store substituted.
114
+
115
+ ## Pins
116
+
117
+ - `test/08` — storage: the node layout, halo persistence and 2-bit round trip,
118
+ the `haloMass`/`hasHalo` contract, and survival across a reopen.
119
+ - `test/02` — exact round trip: near dedup never merges distinct bytes.
120
+ - `test/36-bloom` — the filters' exactness and persistence.
121
+ - `test/153.4` — the span prober answers exactly what `findFlatBranch` answers.
@@ -1,79 +1,72 @@
1
- # Thresholds — derived, never tuned
2
-
3
- Every decision cutoff is a formula over `D` (vector dimension), `W` (`maxGroup`,
4
- perception window), or `N` (corpus size). No threshold is tuned or added to
5
- `src/config.ts`.
6
-
7
- ## Source of truth
8
-
9
- - `src/geometry.ts` — all similarity/decision thresholds.
10
- - `src/mind/traverse.ts` — corpus-scale readings that parameterise bounded
11
- walks.
12
- - `src/sema.ts` — positional coordinate algebra.
13
- - `src/config.ts` — capacities and budgets only (cache byte budgets, batch
14
- sizes, index parameters, query `k`, ALU precision, seed). Never a threshold.
15
-
16
- ## Geometry thresholds (`src/geometry.ts`)
17
-
18
- | Symbol | Definition | Formula |
19
- | ----------------------- | ---------------------------------------------------------------------------- | ----------------------------------- |
20
- | `mergeThreshold(D)` | Cosine below which two gists are near enough to consider merging | `1 - 1/√D` |
21
- | `identityBar(D,W,len)` | Scale-aware whole-span identity claim | `max(mergeThreshold(D), 1 - W/len)` |
22
- | `reachThreshold(W)` | Recall confidence floor — half a river quantum | `1 - 1/(2·W)` |
23
- | `estimatorNoise(D)` | RaBitQ noise floor — 1σ of random cosine | `1/√D` |
24
- | `significanceBar(D)` | Whole-query relatedness — 3σ above chance | `3/√D` |
25
- | `conceptThreshold(D)` | Halo concept sharing — structural midpoint + ½σ | `0.5 + 0.5/√D` |
26
- | `dominates(part,whole)` | Half-dominance predicate | `part*2 > whole` |
27
- | `profileCapacity(D)` | Superposition capacity — terms before readout collapses | `floor(√D)` (min 1) |
28
- | `consensusFloor(N)` | Pooled-vote significance floor | `ln(N) + ½` |
29
- | `coverageBar(_,D)` | Reach-index gating (currently unused hot-path; batch compaction replaces it) | `conceptThreshold(D)` |
30
-
31
- `N` in `consensusFloor` is `corpusN` (edge-source count, floored at 2).
32
-
33
- ## Corpus-scale readings (`src/mind/traverse.ts`)
34
-
35
- | Symbol | Definition | Formula |
36
- | ---------------- | ------------------------------------------------------ | --------------------------- |
37
- | `corpusN` | Distinct learnt contexts, floored | `max(2, edgeSourceCount())` |
38
- | `hubBound` | Hub bound, used for every `LIMIT` read | `ceil(√max(2,N))` |
39
- | `hubCap(ids)` | Fan-out cap — list-side reading of `hubBound` | `ids.slice(0, hubBound)` |
40
- | `atomReach(N,W)` | Uniform-expectation floor on a byte atom's commonality | `max(1, ceil(N·W/256))` |
41
- | `atomIsHub(N,W)` | Whether atom abstains as consensus voter | `atomReach > hubBound` |
42
-
43
- `hubBound` is enforced at the store level (`nextFirst`, `parentsFirst`,
44
- `containersSlice`, `hasNext`/`hasParents`, `bytesPrefix`, `chainRun`).
45
- `atomReach` is the honest floor for atoms — they carry no kid/contain rows, so
46
- their reach is unmeasurable and must not default to "maximally rare".
47
-
48
- ## Seat algebra (`src/sema.ts`)
49
-
50
- `twoEndedSeat(seatCount, size, index)` — the one positional-coordinate algebra
51
- shared by perception, `fold`, and every synthetic/canonical fold. First half
52
- uses low seats, second half uses high seats:
53
- `index < (size+1)/2 ? index : seatCount-size+index`.
1
+ # Thresholds — Derived, Never Tuned
2
+
3
+ > **Law:** every decision cutoff is a formula over the vector dimension `D`, the
4
+ > perception window `W` (`maxGroup`) or the corpus size `N`. `config.ts` holds
5
+ > capacities and budgets only: cache bytes, batch sizes, index parameters, query
6
+ > `k`, ALU precision and the seed. It never holds a threshold.
7
+
8
+ **Why.** A kept memory grows, and a fixed number silently changes meaning as
9
+ `N`, `D` or a span's length changes. A derived bar is measured against the
10
+ medium's own chance and scale, so there is no development set to overfit and no
11
+ calibration that expires with new data. With nothing fitted, a wrong answer
12
+ indicts a law, never a parameter.
13
+
14
+ ## Three sources of scale
15
+
16
+ **`D`: chance.** Random unit vectors have cosine about `N(0, 1/D)`, so one noise
17
+ unit is `1/√D`.
18
+
19
+ - `estimatorNoise = 1/√D`
20
+ - `significanceBar = 3/√D`, three noise units above chance
21
+ - `conceptThreshold = 0.5 + 0.5/√D`, for halo concept sharing: the structural
22
+ midpoint plus half a noise unit
23
+ - `mergeThreshold = 1 − 1/√D`
24
+ - `profileCapacity = ⌊√D⌋`, the terms a superposition holds before each falls
25
+ below noise
26
+
27
+ **`W`: the perception quantum.** Below one window, a byte overlap is chance.
28
+
29
+ - `identityBar(D, W, len) = max(mergeThreshold, 1 − W/len)`
30
+ - `reachThreshold = 1 − 1/(2W)`
31
+ - canonical windows of `W−1` and `W` bytes, and `chainReach = W²`
32
+
33
+ **`N`: commonness,** with `corpusN = max(2, edge sources)`.
34
+
35
+ - `hubBound = ⌈√N⌉`, the bound on every `LIMIT` read (`hubCap` on lists)
36
+ - `consensusFloor = ln N + ½`, for pooled climb votes
37
+ - `atomReach = max(1, ⌈N·W/256⌉)`; an atom is a hub when `atomReach > hubBound`
38
+
39
+ `dominates(part, whole) = 2·part > whole` is the one scale-free bar: a strict
40
+ majority.
41
+
42
+ The vector bars are compared against RaBitQ estimates, not exact cosines. That
43
+ is benign for inequality gates over broad regions, and it is why no bar ever
44
+ decides identity (`exact-vs-approximate.md`).
45
+
46
+ The formulas live in `src/geometry.ts`, `src/mind/traverse.ts` (corpus scale)
47
+ and `src/mind/canonical.ts` (windows). A new bar is added there, never to
48
+ `config.ts`.
54
49
 
55
50
  ## Two derivations that bite
56
51
 
57
- **1. `identityBar` is scale-aware.** A fixed cosine `1-1/√D` over a `4·√D`-byte
58
- span tolerates four whole windows of foreign bytes while still claiming
59
- "near-identical". An identity claim may tolerate at most one window `W` (the
60
- perception quantum, same budget as `differsByOneWindow` in near-dedup), so the
61
- bar must be `1-W/len` floored at `mergeThreshold`. Reusing `mergeThreshold` for
62
- a whole-span claim silently widens the byte budget with span length.
63
-
64
- **2. `consensusFloor` is priced for pooled climb votes, not for support
65
- counts.** Each region contributes at most `ln(N/c) ≤ ln(N)`; `ln(N)+½` demands
66
- corroboration beyond one maximally-specific region. `chooseNext`'s `bestSupport`
67
- (`prevCount` of one destination) is N-invariant — bounded by retellings of that
68
- fact, not by `N`. Gating it against `consensusFloor` guarantees failure once `N`
69
- is large enough (observed: 2-vs-1-1-1 corroboration refused at N≈325K, falling
70
- back to a noisy concept-hop).
52
+ **`identityBar` must scale with length.** A fixed cosine of `1 − 1/√D` over a
53
+ `4·√D`-byte span tolerates four whole windows of foreign bytes while still
54
+ claiming near-identity. An identity claim may tolerate at most one window (`W`),
55
+ so the bar is `1 − W/len`, floored at `mergeThreshold`. Reusing `mergeThreshold`
56
+ for a whole-span claim silently widens the tolerance as the span grows.
57
+
58
+ **A bar is priced for one quantity only.** `consensusFloor` is calibrated for
59
+ pooled climb votes, where each region contributes at most `ln(N/c) ≤ ln N`, so
60
+ `ln N + ½` asks for corroboration beyond one maximally specific region.
61
+ `chooseNext`'s support count is different: it is bounded by retellings of one
62
+ fact and does not grow with `N`. Gating that count by `consensusFloor` must
63
+ eventually fail, and it did: at N≈325K, a 2-against-1-1-1 corroboration was
64
+ refused, and the answer fell back to a noisy concept hop.
71
65
 
72
66
  ## Pins
73
67
 
74
- - `test/40-choosenext-scale-guard` — `consensusFloor` must not gate `chooseNext`
75
- support counts.
76
- - `test/64-two-ended-thresholds` — `mergeThreshold` / `identityBar` /
77
- `reachThreshold` derivations and the scale-aware floor.
78
- - `test/78-atom-hub-recognition-cliff` — `atomReach` / `atomIsHub` hub
79
- abstention at scale.
68
+ - `test/64` — `mergeThreshold`, `identityBar` and `reachThreshold`, and the
69
+ scale-aware floor.
70
+ - `test/40` — `consensusFloor` must not gate `chooseNext`'s support counts.
71
+ - `test/78` — `atomReach`/`atomIsHub`: atoms abstain once the corpus makes them
72
+ hubs.