@hviana/sema 0.7.3 → 0.7.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (46) hide show
  1. package/AGENTS.md +95 -843
  2. package/README.md +11 -11
  3. package/dist/src/mind/graph-search.d.ts +32 -22
  4. package/dist/src/mind/graph-search.js +97 -54
  5. package/dist/src/mind/mind.js +14 -0
  6. package/dist/src/store-sqlite.js +17 -0
  7. package/dist/src/store.d.ts +18 -0
  8. package/dist/src/store.js +10 -0
  9. package/docs/INDEX.md +71 -0
  10. package/docs/INVARIANTS.md +19 -0
  11. package/docs/architecture/bounded-reads.md +85 -0
  12. package/docs/architecture/caches.md +89 -0
  13. package/docs/architecture/commonality.md +45 -0
  14. package/docs/architecture/cost-model.md +71 -0
  15. package/docs/architecture/determinism.md +73 -0
  16. package/docs/architecture/exact-vs-approximate.md +47 -0
  17. package/docs/architecture/factored-machinery.md +28 -0
  18. package/docs/architecture/fold-contract.md +87 -0
  19. package/docs/architecture/halo-sketch.md +99 -0
  20. package/docs/architecture/match-project.md +62 -0
  21. package/docs/architecture/mechanism-market.md +95 -0
  22. package/docs/architecture/memoization.md +96 -0
  23. package/docs/architecture/meter.md +55 -0
  24. package/docs/architecture/saturation.md +92 -0
  25. package/docs/architecture/store.md +79 -0
  26. package/docs/architecture/thresholds.md +79 -0
  27. package/docs/failures/tempting-but-wrong.md +144 -0
  28. package/docs/harness/gates.md +56 -0
  29. package/docs/mechanisms/alu.md +75 -0
  30. package/docs/mechanisms/cast.md +75 -0
  31. package/docs/mechanisms/confluence.md +36 -0
  32. package/docs/mechanisms/cover.md +54 -0
  33. package/docs/mechanisms/extraction.md +53 -0
  34. package/docs/mechanisms/prefix-completion.md +54 -0
  35. package/docs/mechanisms/recall.md +69 -0
  36. package/docs/mechanisms/reference.md +58 -0
  37. package/jsr.json +1 -1
  38. package/package.json +1 -1
  39. package/src/mind/graph-search.ts +101 -55
  40. package/src/mind/mind.ts +14 -0
  41. package/src/store-sqlite.ts +19 -0
  42. package/src/store.ts +22 -0
  43. package/test/89-completion-recursion.test.mjs +30 -10
  44. package/test/97-store-seed.test.mjs +105 -0
  45. package/test/98-completion-chaining.test.mjs +140 -0
  46. package/HOW_IT_WORKS.md +0 -5836
@@ -0,0 +1,85 @@
1
+ # Bounded Reads — No Per-Query Read Grows With the Corpus
2
+
3
+ > **Law:** the cost of one query is proportional to the query, not to how much
4
+ > was learned. No per-query read may grow with corpus size N.
5
+
6
+ Every fan-out, walk, and disambiguation is capped at `hubBound` —
7
+ `ceil(sqrt(N))` — derived once from `corpusN` and floored at 2 so `sqrt` and
8
+ `ln` stay meaningful on a near-empty store. There is no second convention; do
9
+ not invent one.
10
+
11
+ ## Scale
12
+
13
+ ```
14
+ corpusN(ctx) = max(2, store.edgeSourceCount()) // distinct learnt contexts
15
+ hubBound(ctx) = ceil(sqrt(corpusN(ctx))) // >= 2, the store cap
16
+ hubCap(ctx, ids) = ids.slice(0, hubBound(ctx)) // list-side reading
17
+ boundFor(n) = ceil(sqrt(max(2, n))) // ctx-free reading
18
+ ```
19
+
20
+ Defined once in `mind/traverse.ts` (`corpusN`, `hubBound`, `hubCap`,
21
+ `boundFor`). Every consumer imports them; never spell `Math.sqrt` inline.
22
+
23
+ ## Enforcement at the store level
24
+
25
+ The cap is not advisory — adapters must make bounded reads bounded in SQL.
26
+
27
+ ### 1. LIMITed reads — real `LIMIT ?`
28
+
29
+ `nextFirst(id, limit)`, `prevFirst(id, limit)`, `parentsFirst(id, limit)`,
30
+ `containersSlice(child, offset, limit)`.
31
+
32
+ Same statement and `ORDER BY` as the full read, with `LIMIT ?`. Never
33
+ "materialise then slice". Reading `hubBound + 1` parents decides "hub or not"
34
+ exactly without reading the rest. Implemented as thin wrappers in
35
+ `store-sqlite.ts` over `AbstractStore` in `store.ts`.
36
+
37
+ ### 2. Existence probes — indexed point probes
38
+
39
+ `hasNext(id)`, `hasParents(id)`, `hasContainers(child)`, `hasHalo(id)`,
40
+ `prevCount(id)`.
41
+
42
+ One indexed `EXISTS` / `COUNT` probe that never decodes vectors or unpacks
43
+ blobs. Use them for every "does this lead anywhere?" question instead of
44
+ `next(id).length > 0` or `prev(id).length`. `prevCount` is the reverse-edge
45
+ support count for `chooseNext`/`chooseAmong`; `hasNext`/`hasHalo` gate the
46
+ `leadsSomewhere` admission predicate in `mind/traverse.ts`.
47
+
48
+ ### 3. Prefix-capped reads — reject without reconstruction
49
+
50
+ `bytesPrefix(id, cap)` and `contentLen(id, cap)`.
51
+
52
+ `contentLen` under a cap returns an exact length below the cap and `>= cap`
53
+ otherwise — an indexed memo walk that stops early, never a full subtree walk.
54
+ `bytesPrefix` stops after `cap` bytes. A candidate exceeding the cap is rejected
55
+ on the length probe alone; the weave, junction walks, and bridge all read this
56
+ way. Uncapped reads there cost seconds per query on a large store.
57
+
58
+ ### 4. Transparent scaffolding — one bounded read
59
+
60
+ `chainRun(id)` climbs a run of transparent nodes (exactly one parent, no edges)
61
+ in a single recursive CTE, cached for the store lifetime and dropped on writes
62
+ that break transparency. The climber hops the whole run where a node-at-a-time
63
+ ascent would pay three probes per node.
64
+
65
+ ## Maintenance only
66
+
67
+ The full materialising reads — `next(id)`, `prev(id)`, `parents(id)`,
68
+ `containers(child)` — exist for inspection, repair, and compaction only
69
+ (`compactContentIndex`, `repairContentIndex`). Keep them off hot paths.
70
+
71
+ ## Adding a walk
72
+
73
+ Any new fan-out walk uses `hubBound`/`hubCap`. Do not call `edgeSourceCount()`
74
+ or `Math.ceil(Math.sqrt(...))` inline, and do not invent a per-walk limit. The
75
+ walk's saturation decision (when to stop) is separate from the cap (the safety
76
+ net); a walk with only a cap drifts to the cap.
77
+
78
+ ## Pins
79
+
80
+ - `test/14` — sublinear inference in corpus size and constant-rate in input
81
+ length; training throughput floor; exact recall at scale.
82
+ - `test/89` — completion recursion stays output-sensitive (nested searches/pops
83
+ sublinear); guards the count of reads, not just per-read size.
84
+ - `test/90` — connector probe (`resolveConnectors`) reads by the query length
85
+ (`QUERY.length + 1`), not by the learnt continuation; per-read size bound.
@@ -0,0 +1,89 @@
1
+ # Caches — Every Acceleration Is a BoundedMap
2
+
3
+ > **Law:** every acceleration is a `BoundedMap` with a byte budget. A miss
4
+ > re-derives from durable state. Degradation order is speed/reach lost, never
5
+ > identity.
6
+
7
+ No cache may change what is stored, what is resolved, or what tree is folded.
8
+ Eviction costs a re-read, a re-walk, or a narrower reach — never a wrong answer
9
+ or a wrong tree.
10
+
11
+ ## `BoundedMap` — the one cache primitive
12
+
13
+ `src/store.ts:BoundedMap<K,V>` — LRU with byte accounting (`maxBytes`, `sizeOf`,
14
+ `evict`, `recency`).
15
+
16
+ - `evict: "lru"` — uniform-cost entries (dedup, vectors, records).
17
+ - `evict: "smallest"` — variable-cost reconstruction (`_bytesCache`): protects
18
+ expensive large branches over cheap leaves.
19
+ - `recency: "reorder"` (default) — exact LRU via `delete+set`; required when
20
+ eviction choice is load-bearing (`_depositTrees` — 8 entries, victim changes
21
+ fold).
22
+ - `recency: "clock"` — bit instead of reorder; only for transparent caches where
23
+ wrong victim costs a re-read (`_bytesCache`, `_recCache`). Measured: same
24
+ entries cached, hot-path time 55% → bit.
25
+
26
+ Persistent cursor over V8 insertion order makes eviction amortised O(1);
27
+ candidate window for `"smallest"` never rescans from front.
28
+
29
+ ## Store caches — budgets in `src/config.ts:StoreConfig`
30
+
31
+ | Cache | Field | Budget | `sizeOf` | Eviction |
32
+ | ------------------- | -------------------------- | ---------------------------- | ------------------- | -------------- |
33
+ | dedup leaf/branch | `_leafKey` / `_branchKey` | `dedupCacheMax` 1M entries | 1 | lru |
34
+ | reconstructed bytes | `_bytesCache` | `bytesCacheMax` 20 MB | `byteLength` | smallest+clock |
35
+ | content length | `_lenCache` | `bytesCacheMax` | 16 | lru |
36
+ | node records | `_recCache` | `recCacheBytes` 10 MB | leaf+4·kids+12 | lru+clock |
37
+ | pending gists | `_pendingGist` | `pendingGistBytes` 16 MB | `byteLength` (D·4) | lru |
38
+ | halo exact / norm | `_haloExact` / `_haloNorm` | `haloCacheBytes` 16 MB each | `byteLength` | lru |
39
+ | skipped interiors | `_coveredIds` | `coveredIdsMax` 100K entries | 1 | lru |
40
+ | indexed ids | `_indexedIds` | `coveredIdsMax` | 1 | lru |
41
+ | transparent chains | `_chainMemo` | `chainCacheBytes` 16 MB | 4·len+32 | lru |
42
+ | ingest memo | `CachedIngest._memo` | `ingestCacheBytes` 50 MB | vector+ids+keyBytes | lru |
43
+
44
+ ANN read caches (`_resonateCache`, `_resonateHaloCache`) are `Map<string,Hit[]>`
45
+ keyed by `vecKey(v)+":"+k`, dropped on any index mutation.
46
+ `vectorCacheMb`/`sqliteCacheMb` are pure page-cache latency knobs.
47
+
48
+ `_bytesCache` only caches complete reconstructions — `bytesPrefix(id,cap)` with
49
+ `got < cap`; a truncated prefix is never stored. `_chainMemo` is dropped on any
50
+ write that could break transparency; `_pendingGist` eviction falls back to DAG
51
+ climb; halo eviction re-decodes the durable 2-bit row.
52
+
53
+ ## Mind caches — session and per-response
54
+
55
+ | Cache | Location | Budget | Scope / invalidation |
56
+ | ------------------- | ------------------------------------------------ | --------- | ------------------------------------------------------------------ |
57
+ | `_gistCache` | `Mind._gistCache` | 32 MB | session-lifetime, never invalidated (perception pure) |
58
+ | `_depositTrees` | `Mind._depositTrees` | 8 entries | session; `perceiveDeposit` only when `conversational` |
59
+ | `_depositLens` | `Mind._depositLens` | — | byte lengths for prefix probes; cleared with map when >64 |
60
+ | `_internIds` | `Mind._internIds: WeakMap<Sema,number>` | — | Mind lifetime; ids permanent |
61
+ | `_resolvedSubtrees` | `Mind._resolvedSubtrees: WeakMap<Sema,{id,len}>` | — | per-response/conversation; fast path only when `visit===undefined` |
62
+
63
+ `REACH_MEMO_MAX` / `STRUCT_MEMO_MAX` 100K (`src/mind/traverse.ts`) — whole-climb
64
+ and per-node structural probes (`hasNext`/`prevCount`/`hasParents`); cleared on
65
+ write or when cap reached. `reachMemo`/`structCaches` keyed by `_structMemoKey`,
66
+ bypassed under trace.
67
+
68
+ ## Deposit caches — offset-keyed, caller-discharged
69
+
70
+ `_depositTrees`/`_depositLens`/`_internIds`/`_resolvedSubtrees` key **offsets**,
71
+ not bytes, for O(1) reuse. Offsets alone cannot witness byte agreement — caller
72
+ must discharge it.
73
+
74
+ - Correct: conversation append — each turn extends the prior cumulative context
75
+ by its own bytes; longest cached proper prefix hit (`L < bytes.length`) reuses
76
+ `contentFoldIncremental` segments bit-identically.
77
+ - Wrong: mismatched `prev` reused by offset produced wrong tree (336 vs 400
78
+ bytes) — a coincidental prefix length aliased unrelated content.
79
+ - Now: `perceiveDeposit` keys by `latin1Key(bytes.subarray(0,L))` (prefix
80
+ bytes), probes longest cached proper prefix first; `_depositTrees` populated
81
+ only for conversational deposits (budget discipline), otherwise cold path
82
+ always correct.
83
+
84
+ ## Pins
85
+
86
+ - `test/91` — `chainRun` via capped `_prefix`: bounded transparent-chain hop,
87
+ not per-node probes.
88
+ - `test/96` — `_bytesCache` is byte-accounted `BoundedMap` that evicts; miss
89
+ re-derives.
@@ -0,0 +1,45 @@
1
+ # Two Measures of Commonality
2
+
3
+ Sema needs "what is shared" in two different populations. One is corpus-global
4
+ (how widely a structure is reused), the other is weave-local (what a local
5
+ cohort of overlapping forms agrees on). They use different data and different
6
+ formulas and must not be conflated.
7
+
8
+ ## Corpus-global — `reachOf` + `dominates`
9
+
10
+ _Defined in `src/mind/traverse.ts` + `src/geometry.ts`; used by climb,
11
+ containment, IDF pooling._
12
+
13
+ For a node id, `reachOf(id, N)` counts how many learnt contexts contain it
14
+ (ancestor reach via capped graph walks, memoised per response in
15
+ `sharedReachMemo`). `dominates(reach, N)` then asks whether that reach is above
16
+ the corpus-determined majority threshold (derived in `geometry.ts` over `N`).
17
+ Intuition: minority reach discriminates (a filler), majority reach is
18
+ scaffolding. Powers the consensus climb, edge following, and vote pooling.
19
+
20
+ ## Weave-local — `depth[]` + `MIN_WEAVE` + `dominates`
21
+
22
+ _Defined in `src/mind/match.ts` (`depth[]`, `MIN_WEAVE`, `frame`) and gated in
23
+ `src/mind/match.ts:frame`; used by CAST._
24
+
25
+ For an alignment weave, `depth[i]` counts how many aligned structures cover byte
26
+ `i` of the query. `MIN_WEAVE = 2` requires agreement beyond a pair (pair columns
27
+ are ambiguous with insertions/deletions), and `dominates(depth[i], aligned)`
28
+ requires agreement by a majority of the aligned cohort:
29
+
30
+ ```
31
+ frame(i) ⇔ depth[i] > MIN_WEAVE ∧ dominates(depth[i], aligned)
32
+ ```
33
+
34
+ This powers CAST's frame gate: what the local cohort shares vs what
35
+ differentiates one member. It never consults corpus reach.
36
+
37
+ The two measures answer different questions over different populations; CAST's
38
+ frame must not be replaced by a reach check and the climb must not be driven by
39
+ weave depth.
40
+
41
+ ## Pins
42
+
43
+ - `test/17` — weave-local frame / `MIN_WEAVE` / `dominates` vs corpus-global
44
+ reach.
45
+ - `test/34` — containment and reach-driven disambiguation.
@@ -0,0 +1,71 @@
1
+ # Cost Model — One Currency
2
+
3
+ Every mechanism and every byte competes on one cost ladder defined in
4
+ `src/mind/graph-search.ts`. GraphSearch and `pipeline.ts:think` use the same
5
+ units, so a mechanism-level choice and a byte-level choice are the same kind of
6
+ decision: a lightest derivation.
7
+
8
+ ## Ladder (`src/mind/graph-search.ts`)
9
+
10
+ | Cost | Value | Meaning |
11
+ | --------- | ------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
12
+ | `MICRO` | `1e-3` | Recognised advance (one `rec` bridge); per-byte unit of the A\* heuristic. A recomposed form's onward edge is also `MICRO`. |
13
+ | `STEP` | `1` | Every edge hop (first or fifth), every computed result, every projection. Charging every hop makes the lightest derivation the shortest chain. |
14
+ | `CONCEPT` | `10` | Halo-mediated act (synonym hop, consensus climb) and abandoning an edge chain early (`CONCEPT` above chain cost — genuine fixpoint at `+0` always beats giving up at same depth). |
15
+ | `PASS` | `1000` / byte | Carrying a byte nothing explains. Dominates everything so the search always prefers to recognise. |
16
+
17
+ Only the **ordering** `MICRO < STEP < CONCEPT < PASS` matters; any constants
18
+ with that order give the same derivations.
19
+
20
+ ## Pipeline weighing (`src/mind/pipeline.ts:think`)
21
+
22
+ Mechanism candidates are weighed in the same ladder:
23
+
24
+ ```
25
+ weight = moves + PASS * unaccounted_bytes
26
+ grade = floor(weight / STEP)
27
+ ```
28
+
29
+ `unaccounted` is the query bytes no `accounted` span covers. Comparison is at
30
+ `STEP` resolution: lowest `grade` wins. At equal grade the candidate with fewer
31
+ `scaffolding` bytes (answer bytes lifted from unrecognised spans) wins; only
32
+ then does mechanism list order decide.
33
+
34
+ ## Two semirings
35
+
36
+ - **(min, +) tropical** — lightest derivation in `GraphSearch` via `src/derive`
37
+ (`lightestDerivation`). Cost accumulates with `+`, choice selects `min`.
38
+ Powers `cover`/`form`/`out`, edge following, fusing, and the A\* agenda
39
+ (`g + h`).
40
+
41
+ - **(+, +) arithmetic** — evidence pooling in `src/mind/attention.ts:poolVotes`.
42
+ Each region's vote is an axiom; rules carry `Rule.combine = 'sum'` so costs to
43
+ the same anchor **add** rather than minimise. Powers IDF-weighted consensus,
44
+ `votes`/`votesIdf`/`support`, and `regionSupport`/`regionPeak`.
45
+
46
+ ## Admissibility
47
+
48
+ The A\* heuristic is admissible and consistent:
49
+
50
+ ```
51
+ h(it) = (queryLen - right) * MICRO
52
+ ```
53
+
54
+ `right` is `p` for `cover(p)` or `j` for `form/out [i,j)`. `MICRO` is the
55
+ minimum per-byte cost in the ladder (every real per-byte cost is `>= MICRO`,
56
+ including `PASS`), and only the suffix past `right` is counted, so `h` never
57
+ exceeds the true remaining cost.
58
+
59
+ ## Policy is not cost
60
+
61
+ "Computation always wins" is **not** priced into the ladder (a computed result
62
+ costs `STEP`, same as a learned edge). It is enforced by masking: `pipeline.ts`
63
+ removes recognised sites overlapped by a `ComputedResult` so the computation is
64
+ the sole completion there. Keep policy in callers; keep the engine neutral.
65
+
66
+ ## Pins
67
+
68
+ - `test/52` — climb consensus instrumentation
69
+ - `test/53` — cross-region probe instrumentation
70
+ - `test/54` — evidence `k` instrumentation
71
+ - `test/55` — cost meter (`Meter`, `CostReport`, `searchPops`/`searchPushes`)
@@ -0,0 +1,73 @@
1
+ # Determinism — Same Seed + Same Deposits + Same Query ⇒ Same Bytes
2
+
3
+ ## The law
4
+
5
+ > Same `seed` + same deposit order + same query ⇒ byte-identical answer.
6
+
7
+ Determinism is the product. Every code path that can reach output must be
8
+ deterministic given `(seed, store contents, query bytes)`.
9
+
10
+ ## Forbidden
11
+
12
+ No `Math.random` or `Date.now` in behaviour, and no iteration over unordered
13
+ collections where order can reach output. Example-only uses
14
+ (`example/train_base`) are outside the library contract. If a test becomes
15
+ flaky, the contract was broken, not the test.
16
+
17
+ ## All randomness flows from `seed`
18
+
19
+ `MindConfig.seed` (`src/config.ts:resolveConfig`, `DEFAULT_CONFIG`) is the sole
20
+ entropy root. Subsystems derive deterministically:
21
+
22
+ - **Alphabet** — `Alphabet` (`src/alphabet.ts`) via `rng` (`src/vec.ts:rng`)
23
+ seeded as `seed ^ seedMask`; builds 16→64→256 vectors by refinement.
24
+ - **Keyring / Space** — `Space.seats` (`src/sema.ts:Space`) via `makeKeyring`
25
+ (`src/vec.ts:makeKeyring`) and `rng` seeded from `seed` in `Mind`
26
+ (`src/mind/mind.ts`); `fold`/`twoEndedSeat`/`companySignature` are pure over
27
+ `Space`.
28
+ - **Vector indexes** — `VectorDatabase` (`src/rabitq-ivf/src/database.ts`) and
29
+ `Prng` (`src/rabitq-ivf/src/rabitq.ts`) seeded from config; insertion order is
30
+ the stored order, not a random choice.
31
+
32
+ No other PRNG source may affect grounding. Thresholds in `geometry.ts` are
33
+ derived from `D`/`W`/`N`, not sampled.
34
+
35
+ ## Tie-breaks are corpus-determined
36
+
37
+ Every choice among equals bottoms out in a fixed ordering — insertion order or
38
+ lowest node id. The universal no-evidence fallback is **first-inserted**:
39
+
40
+ - `guidedFirst` (`src/mind/traverse.ts:guidedFirst`) — guided pick via
41
+ `chooseNext` else first-inserted edge (`nextFirst` LIMIT 1).
42
+ - `chooseNext` (`src/mind/traverse.ts:chooseNext`) — capped `nextFirst` read
43
+ (`hubBound`), ranked by `prevCount` then `haloMass`; equal ⇒ first-inserted.
44
+ - `chooseAmong` (`src/mind/traverse.ts:chooseAmong`) — `hubCap` + `argmaxCosine`
45
+ over `candidateGist`; first-inserted on tie via stable scan.
46
+ - `companySignature` (`src/sema.ts:companySignature`) — `rng(id ^ 0x9e3779b9)`,
47
+ i.e. seeded by node id, not observation order.
48
+
49
+ Last-inserted was once used in one place; it was a bug. Never reintroduce it.
50
+
51
+ ## Memoization and trace must not break identity
52
+
53
+ Per-response memos (`Precomputed`, `perceiveMemo`, `recogniseMemo`, `climbMemo`,
54
+ `_resolvedSubtrees` via `foldTree`, `_edgeChoice`, `_gistCache` in
55
+ `src/mind/mind.ts` / `src/mind/pipeline-mechanism.ts` /
56
+ `src/mind/primitives.ts`) are sound because asking never writes. Only
57
+ `guidedNext`/`sharedReachMemo` are trace-bypassed;
58
+ `perceiveMemo`/`recogniseMemo`/`climbMemo` are always consulted — `foldTree`'s
59
+ subtree fast path skips `visit` (and thus site emission) for cached subtrees, so
60
+ bypassing makes `recognise` non-idempotent.
61
+
62
+ ## Follow it
63
+
64
+ When you add any choice among equals, name the tie-break explicitly and make it
65
+ corpus-determined. Thread new randomness through `seed`-derived `rng`; never
66
+ call `Math.random`/`Date.now` on a behavioural path.
67
+
68
+ ## Pins
69
+
70
+ - `test/42` pins recognition idempotence under trace — traced and untraced
71
+ `recognise` must return the same cached object and site count.
72
+ - Determinism suites — `test/03`, `test/04`, `test/08`, `test/20` and others
73
+ assert same seed + same training ⇒ byte-identical answers and stores.
@@ -0,0 +1,47 @@
1
+ # Exact vs Approximate — The Law and Its Five Ladders
2
+
3
+ Vector scores (`resonate` / `resonateHalo`) are RaBitQ **estimates**. They rank
4
+ candidates and gate broad regions; they never decide identity. Identity is
5
+ decided only by content-addressed lookup — `resolve` / `findLeaf` / `findBranch`
6
+ / `canonResolve` — and by re-folding bytes to verify.
7
+
8
+ ## The law
9
+
10
+ > Scores propose, bytes dispose.
11
+
12
+ Even recall's echo decision re-folds the top hit's bytes rather than trusting
13
+ the estimate it already has. No `score >= threshold` path may mint an identity
14
+ claim; thresholds derived in `geometry.ts` gate search breadth, not truth.
15
+
16
+ ## Graded evidence ladders
17
+
18
+ Five subsystems share one shape — **exact → distributional → geometric** — with
19
+ earlier tiers strictly preferred. Never reorder tiers; never let an approximate
20
+ tier override an exact one.
21
+
22
+ | # | Site | Ladder (strong → weak) | File |
23
+ | - | ------------------ | ------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------- |
24
+ | 1 | `resolve` | exact content-addressed fold → `canonResolve` (equivalence class, hash-then-verify) | `mind/primitives.ts` |
25
+ | 2 | `locate` | exact bytes → halo role → gist | `mind/match.ts` |
26
+ | 3 | `alignGraded` | literal W-gram runs → halo-matched sites + climb proposals (weave) | `mind/match.ts` / `pipeline-mechanism.ts` |
27
+ | 4 | `bridge` | junction containers → edge → synonym → whole-gist | `mind/resonance.ts` |
28
+ | 5 | `crossRegionVotes` | exact containers → single synonym → double → `structuralResonance` (synthetic gist, gated hardest — no byte containment) | `mind/attention.ts` |
29
+
30
+ ## Asymmetries (attention)
31
+
32
+ Two rules in `attention.ts` encode "exact decides" and must not be flattened:
33
+
34
+ - Only the **EXACT** tier may explain ordinary votes away.
35
+ - Only **container-backed** evidence may consume its endpoints.
36
+
37
+ ## Pins
38
+
39
+ - `test/51` pins the cross-region tier ladder and its gating.
40
+ - Recognition idempotence under trace (`test/42`) depends on exact identity
41
+ remaining byte-determined, not score-determined.
42
+
43
+ ## Adding a matcher
44
+
45
+ Add a tier to the shared family in `mind/match.ts` with a derived gate
46
+ (`geometry.ts`), never a private `score >= k` check. A new mechanism is a
47
+ `(matcher, direction, gate)` configuration over that family (§2.5).
@@ -0,0 +1,28 @@
1
+ # Factored Machinery — One Definition, Many Consumers
2
+
3
+ Every shared operation is defined once and imported many times. Duplicating it
4
+ forks the corpus contract; moving it hides who owns the gate.
5
+
6
+ For the match → project → gate family see `match-project.md`; for the two
7
+ commonality measures see `commonality.md`; for work accounting see `meter.md`.
8
+
9
+ ## Single-definition contracts
10
+
11
+ | Symbol | Defined in | One fact |
12
+ | ------------------------------------------------------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
13
+ | `contentLevels` | `src/geometry.ts` | Single boundary rule: cuts + levels from one rolling hash pass; every segmentation reads it. |
14
+ | `canonicalWindows` / `chainReach` / `leafIdRun` / `windowIds` | `src/mind/canonical.ts` | Write/read contract: training interns `W-1,W` windows, reading chains to `W²` and probes `W`-windows — drift silences recognition. |
15
+ | `junction.ts` + `WalkCache` | `src/mind/junction.ts` | Shared junction ascent (parents + containers) with bounded `√N·W` walk; `WalkCache` memoizes capped reads/parents/containers per response; bridge and attention share it. |
16
+ | `joinWithBridge` | `src/mind/resonance.ts` | One out-of-search assembly: `bridge(left,right)` or bare concat with `bridgeMiss` trace. |
17
+ | `dismissedKnownContent` | `src/mind/bridge.ts` | Pure attestation: any unaccounted `W`-window that resolves as known content — shared gap guard for substitution and CAST. |
18
+ | `sharedReachMemo` | `src/mind/traverse.ts` | One response-scoped `AncestorReach` memo (cleared on write and for traces); every `reachOf`/`edgeAncestors` consumer shares it. |
19
+ | `guidedFirst` | `src/mind/traverse.ts` | Guided-or-first answer bytes: `guidedNext` else first-inserted edge (`LIMIT 1`). |
20
+ | `leadsSomewhere` | `src/mind/traverse.ts` | Admission predicate: `hasNext` (cached) or `hasHalo`; sites that lead nowhere contribute no derivation. |
21
+ | `isChunk` | `src/sema.ts` | `kids !== null && kids.every(k=>k.kids===null)` — smallest grouped unit; governs regions, seams, indexing. |
22
+ | `twoEndedSeat` | `src/sema.ts` | One seat algebra: first half low seats, second half high seats; shared by perception, `fold`, and canonical folds. |
23
+
24
+ ## Pins
25
+
26
+ - `test/47` — frame reading split (matcher vs gate vs inventory).
27
+ - `test/50` — CAST analog / consensus floor (dismissed content, `MIN_WEAVE` /
28
+ `dominates` frame, `carriesFillers` refusal).
@@ -0,0 +1,87 @@
1
+ # Fold Contract — One Tree For The Same Bytes
2
+
3
+ > **Law:** `perceiveDeposit` ≡ `perceive` — same bytes ⇒ same tree and same node
4
+ > id. Deposit imposes nothing (no boundaries, no turn convention). Geometry
5
+ > never sees conversation metadata.
6
+
7
+ ## The identity
8
+
9
+ Perception is a pure function of the bytes. The deposit path and the inference
10
+ path compute the same content-defined fold for the same input, so a trained
11
+ context node and `resolve(query)` reach the same node. When the two sides
12
+ disagreed, alignment went quadratic (measured 5.2M cells on a 476-byte context
13
+ vs 0 when they agree) and cumulative contexts stopped resolving to what they
14
+ were trained as.
15
+
16
+ ## Deposit imposes nothing
17
+
18
+ No boundaries, no turn convention, nothing read out of the bytes. Conversational
19
+ turn offsets are API metadata — they feed `ConversationState`, `answeredSpans`
20
+ and `currentTurnStart`; the geometry never sees them. Passing turn boundaries
21
+ into the fold is a correctness bug, not a tuning choice.
22
+
23
+ ## Boundaries vs reuse — two problems
24
+
25
+ `contentFoldIncremental` and `stablePrefixFold` solve different problems;
26
+ conflating them is what once put an imposed boundary set on the inference path.
27
+
28
+ | Mechanism | What it buys | Cost / shape |
29
+ | ------------------------ | --------------------------------------------------------- | ----------------------------------------------------------------------------------------------- |
30
+ | `contentFoldIncremental` | Transparent segment reuse (cost only) | Imposes nothing; tree identical to the cold fold |
31
+ | `stablePrefixFold` | Caller-supplied cuts left-nested for prefix-ROOT identity | One prefix-ROOT per cut becomes an identical subtree (and same node id) inside the grown stream |
32
+
33
+ Both carry the same precondition: `prev` must be a fold of a byte-identical
34
+ prefix — reuse is keyed on `[start,end)` offsets, which cannot witness byte
35
+ agreement. A mismatched `prev` produced a wrong tree on 336 of 400 random
36
+ streams. `perceiveDeposit` discharges this via the prefix bytes as cache key; a
37
+ conversation advances only by append. A caller that cannot prove the prefix must
38
+ pass no `prev`.
39
+
40
+ ## Identity must not depend on W or absolute offset
41
+
42
+ `contentLevels` is the single boundary rule (rolling hash over a bounded
43
+ window). Any grouping by index — stride, tile, fixed-arity row — reintroduces
44
+ the grid's phase bug: `riverFold` groups `W`-ary from byte 0, so the same byte
45
+ run is a different subtree at a different offset. The content-defined hash
46
+ removes this: a change upstream moves only the cut it falls inside; downstream
47
+ cuts and segments are unchanged (99.7% cuts preserved on real deposits after
48
+ shifts of 1..7 bytes vs 14.3% for the grid). The `groupByLevel` above the
49
+ segments splits by content level, not by count.
50
+
51
+ ## `contentLevels` is single source; `contentBoundaries` is projection
52
+
53
+ `contentBoundaries(space, bytes)` is `contentLevels(space, bytes).cuts`. It once
54
+ carried its own rolling-hash loop, which is how a write side and a read side
55
+ drift without a type error. Levels are read from the hash the cut was accepted
56
+ at — level `L` when `h` vanishes mod `W^(L+1)` — so level-`L` cuts nest inside
57
+ level-`(L-1)` and expected span is `W^(L+1)` bytes.
58
+
59
+ ## Optional canonical capability
60
+
61
+ `canonAdd`/`canonFind` (`src/store.ts` — `canonCount`/`eachContent`) is an
62
+ optional backend capability. A backend may omit all four; resolution then has no
63
+ equivalence fallback. The store never learns the equivalence — the canonicalizer
64
+ (`Canon` in `src/canon.ts`, e.g. `textCanon`) is injected by the caller and
65
+ every candidate is hash-then-verified (re-canonicalize stored bytes, compare). A
66
+ hash collision costs a read, never a wrong id.
67
+
68
+ ## Cost of changing the cut distribution
69
+
70
+ The cut rate, which bits are read, `minLen`/`maxLen`, and the forced cut at
71
+ `maxLen` set the segment distribution every downstream mechanism is fitted to.
72
+ Each has been changed experimentally and cost 5–21 tests (rate: 15–18, bits:
73
+ 19–21, normalized chunking: 5–6). `W-1` is the minimum (one window minus one);
74
+ `seats.length` is the maximum (one flat node folds exactly one segment). The
75
+ forced cut is load-bearing — relaxing it to reduce the current 32% forced rate
76
+ looked like a tidying but broke the same suites. Re-measure the whole suite for
77
+ any change here.
78
+
79
+ ## Pins
80
+
81
+ - `test/59` — shift invariance floors (content-defined cuts preserved over
82
+ random binary and prose).
83
+ - `test/63` — offset/W invariance and `contentLevels` distribution expectations.
84
+
85
+ See:
86
+ `src/geometry.ts:contentLevels`/`contentBoundaries`/`contentFoldIncremental`/`stablePrefixFold`;
87
+ `AGENTS.md` bootloader invariants.
@@ -0,0 +1,99 @@
1
+ # Halo & Sketch — Distributional Memory
2
+
3
+ A node's **halo** is its distributional signature: the superposition of
4
+ identity-bound company signatures poured from every episode it participated in.
5
+ Its **gist** is the VSA fold of its own bytes — content, not company.
6
+
7
+ ## Two vectors per node, two indexes
8
+
9
+ | Vector | Encodes | Index | Query |
10
+ | ------ | ----------------------------- | ------------- | -------------- |
11
+ | gist | what the node _is made of_ | content index | `resonate` |
12
+ | halo | what _company_ the node keeps | halo index | `resonateHalo` |
13
+
14
+ Both indexes are RaBitQ-IVF (`src/rabitq-ivf/`) — 1-bit ANN over the same
15
+ vectors; scores are estimates, never identity. Halos are also persisted durably
16
+ (see below).
17
+
18
+ ## Quantization — 2-bit on disk, float in session
19
+
20
+ A halo is a superposition of quasi-orthogonal signatures, so coordinates are
21
+ Gaussian. Storage exploits this:
22
+
23
+ - **In-session accumulator** (`_haloExact` in `src/store.ts`): `Float32Array` —
24
+ exact, additive, incremented by `pourHalo`.
25
+ - **Durable row** (`_dbUpsertHalo`): 2-bit Lloyd–Max quantizer — decision at
26
+ ±0.9816σ, levels at ±0.4528σ / ±1.5104σ, σ derived from the stored norm
27
+ (`norm/√D`). Header is the 4-byte norm; body is 2 bits/coordinate. Keeps ≥0.88
28
+ correlation with the exact vector.
29
+ - **ANN index**: 1-bit RaBitQ, irreversible — answers only "which halos are near
30
+ this query?"
31
+
32
+ Re-indexing is geometric: a halo re-enters the ANN when mass is small or crosses
33
+ a power of two (`geometricMass`), so index writes are O(log mass).
34
+
35
+ ## Bottom-k sketch & company profile
36
+
37
+ A whole-partner signature alone records tokens, not types — halos of genuine
38
+ synonyms would be quasi-orthogonal. `companyProfile` (`src/mind/learning.ts`)
39
+ superposes:
40
+
41
+ 1. the partner's own identity signature, plus
42
+ 2. the bottom-k **constituent sketch** — the
43
+ `k = profileCapacity(D) = floor(√D)` minimal units of its subtree with
44
+ smallest `unitPriority`, deduped.
45
+
46
+ The sketch is composable (bottom-k of a union = bottom-k of children's
47
+ sketches), durable derived state via `sketchGet`/`sketchPut`, and bounded: at
48
+ most `k` constituents are classified, each by one `LIMIT`ed parent read
49
+ (`hubBound`). Beyond `√D` terms a single constituent contributes less than
50
+ `1/√D` — below RaBitQ noise — and extra terms shrink every accepted one; the cap
51
+ is a correctness limit.
52
+
53
+ ## Gist vectors
54
+
55
+ Folded by the river (`src/geometry.ts`): leaves are alphabet vectors, groups
56
+ bind by two-ended seats, intermediate gists stay unnormalized (magnitude ∝
57
+ √len), only the root is normalized. Gist resonance reads byte-proportional
58
+ overlap; halo resonance reads distributional overlap — the two are independent.
59
+
60
+ ## Thresholds & gating
61
+
62
+ All bars are derived in `src/geometry.ts`; no tunable constant:
63
+
64
+ | Symbol | Formula | Use |
65
+ | ------------------ | -------------- | ------------------------------------------------------------------------------------------------- |
66
+ | `estimatorNoise` | `1/√D` | 1σ RaBitQ noise; contrastive margin must clear it |
67
+ | `significanceBar` | `3/√D` | whole-query relatedness — 3σ above chance; gates consensus climb and `analogyStrength` |
68
+ | `conceptThreshold` | `0.5 + 0.5/√D` | halo concept sharing — structural midpoint + ½σ; gates `haloSiblings`, concept hops, articulation |
69
+
70
+ The significance bar gates the whole query; `conceptThreshold` gates per-pair
71
+ halo cosine.
72
+
73
+ ## Probes: `haloMass` and `hasHalo`
74
+
75
+ - `haloMass(id)` — count of poured episodes; evidence weight, tie-breaker in
76
+ `chooseAmong`.
77
+ - `hasHalo(id)` — existence probe (indexed point check, no vector decode);
78
+ mirrors `halo(id) !== null`. One tier of the `leadsSomewhere` admission
79
+ predicate (with `hasNext`/`hasParents`).
80
+
81
+ Both are `meter`-counted probes, not full decodes — `halo(id)` is the bounded
82
+ vector read; `resonateHalo` is the IVF ANN query.
83
+
84
+ ## Relation to invariants
85
+
86
+ - **Derived thresholds** — all bars above live in `geometry.ts`.
87
+ - **Exact decides / approximate proposes** — halo scores rank and gate; identity
88
+ is content-addressed. The graded ladder is exact → halo → gist
89
+ (`mind/match.ts`).
90
+ - **Bounded reads** — `hasHalo`/`haloMass` are point probes; `resonateHalo` is
91
+ capped ANN; constituent classification uses `LIMIT hubBound+1` reads. No
92
+ per-query scan grows with corpus.
93
+
94
+ ## Pins
95
+
96
+ - `test/08 storage halo` — halo persistence, 2-bit round-trip,
97
+ `haloMass`/`hasHalo` contract, index survival across reopen.
98
+ - `test/35 ivf` — RaBitQ-IVF contract (recall vs brute force, sublinear
99
+ scaling); covers the halo index's own layer.