@hviana/sema 0.7.2 → 0.7.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (43) hide show
  1. package/AGENTS.md +95 -843
  2. package/README.md +11 -11
  3. package/dist/src/mind/mind.js +14 -0
  4. package/dist/src/store-sqlite.js +17 -0
  5. package/dist/src/store.d.ts +51 -4
  6. package/dist/src/store.js +81 -14
  7. package/docs/INDEX.md +71 -0
  8. package/docs/INVARIANTS.md +19 -0
  9. package/docs/architecture/bounded-reads.md +85 -0
  10. package/docs/architecture/caches.md +89 -0
  11. package/docs/architecture/commonality.md +45 -0
  12. package/docs/architecture/cost-model.md +71 -0
  13. package/docs/architecture/determinism.md +73 -0
  14. package/docs/architecture/exact-vs-approximate.md +47 -0
  15. package/docs/architecture/factored-machinery.md +28 -0
  16. package/docs/architecture/fold-contract.md +87 -0
  17. package/docs/architecture/halo-sketch.md +99 -0
  18. package/docs/architecture/match-project.md +62 -0
  19. package/docs/architecture/mechanism-market.md +95 -0
  20. package/docs/architecture/memoization.md +96 -0
  21. package/docs/architecture/meter.md +55 -0
  22. package/docs/architecture/saturation.md +92 -0
  23. package/docs/architecture/store.md +79 -0
  24. package/docs/architecture/thresholds.md +79 -0
  25. package/docs/failures/tempting-but-wrong.md +144 -0
  26. package/docs/harness/gates.md +56 -0
  27. package/docs/mechanisms/alu.md +75 -0
  28. package/docs/mechanisms/cast.md +75 -0
  29. package/docs/mechanisms/confluence.md +36 -0
  30. package/docs/mechanisms/cover.md +54 -0
  31. package/docs/mechanisms/extraction.md +53 -0
  32. package/docs/mechanisms/prefix-completion.md +54 -0
  33. package/docs/mechanisms/recall.md +69 -0
  34. package/docs/mechanisms/reference.md +58 -0
  35. package/jsr.json +1 -1
  36. package/package.json +1 -1
  37. package/src/mind/mind.ts +14 -0
  38. package/src/store-sqlite.ts +19 -0
  39. package/src/store.ts +92 -16
  40. package/test/89-completion-recursion.test.mjs +30 -10
  41. package/test/96-bytes-walk-termination.test.mjs +115 -0
  42. package/test/97-store-seed.test.mjs +105 -0
  43. package/HOW_IT_WORKS.md +0 -5836
@@ -0,0 +1,96 @@
1
+ # Memoization — Shared Evidence Without Duplication
2
+
3
+ > **Law:** asking never writes, so structural reads are pure during one
4
+ > response. Memoization elides probes, not evidence.
5
+
6
+ Two layers: `Precomputed` (response-scoped shared analyses) and `Mind`
7
+ per-response memos. Both are accelerators that must not change what inference
8
+ computes.
9
+
10
+ ## Precomputed — one response, one container
11
+
12
+ `Precomputed` (`src/mind/pipeline-mechanism.ts`) is the sole place a response's
13
+ shared evidence lives. Created by `think` (`src/mind/pipeline.ts`) before the
14
+ mechanism loop.
15
+
16
+ ### Eager — populated before any `floor`/`run`
17
+
18
+ - `rec: Recognition` — structural + canonical decomposition (`recognise`)
19
+ - `computed: ComputedSpan[]` — `parse()` results from all mechanisms (e.g. ALU)
20
+ - `guide: Vec` — query gist, the response-wide disambiguation guide
21
+ - `k: number` — `cfg.recallQueryK * 2`, the breadth for resonance/weave/climb
22
+
23
+ ### Lazy — computed on first touch, cached by promise
24
+
25
+ Expensive analyses are `async` and cached by promise: the first caller starts
26
+ the work, every later caller awaits the same promise.
27
+
28
+ - `attention()` — `climbAttentionAll` (roots + ranked anchors)
29
+ - `weave()` — `alignGraded` over top-k anchors
30
+ - `resonance()` — `store.resonate(guide, k)` (single ANN query)
31
+ - `frames()` — `frameSlots` inventory from resonance
32
+ - `spanShapedOf(anchor)` / `spanShapedAll()` — per-anchor `skillExemplar`,
33
+ memoised per id
34
+ - `queryWindows` / `queryResolved` / `windowsOf(anchor)` — W-window identities
35
+ - `reachMemo` — `sharedReachMemo(ctx)` (ancestor reach, § below)
36
+
37
+ A mechanism that never asks pays nothing; two mechanisms asking the same
38
+ question pay once. `floor()` must gate on `worthRunning` before first-touching
39
+ an expensive analysis.
40
+
41
+ ## Mind memos — `beginResponse` → `endResponse`
42
+
43
+ `Mind` (`src/mind/mind.ts:beginResponse`/`endResponse`) swaps per-response state
44
+ for each inference call. `respond` takes fresh maps; `respondTurn` reuses the
45
+ conversation's persistent ones (content-keyed, cross-turn).
46
+
47
+ | Memo | Key | Scope |
48
+ | ------------------- | ----------------------------------------- | -------------------------------------------- |
49
+ | `perceiveMemo` | `perceiveKey(bytes)` (latin1) | response / conversation |
50
+ | `recogniseMemo` | `latin1Key(bytes)` | response / conversation |
51
+ | `climbMemo` | `latin1Key(bytes)` | response / conversation |
52
+ | `canonMemo` | `latin1Key(bytes)` | response (when `canon` set) |
53
+ | `_resolvedSubtrees` | `WeakMap<Sema, {id,len}>` (node identity) | response / conversation |
54
+ | `_edgeChoice` | `Map<nodeId, pick>` | response only — **cleared** in `endResponse` |
55
+ | `_gistCache` | `BoundedMap<nodeId, Vec>` 32 MB | **session-lifetime** (not per-response) |
56
+
57
+ `_gistCache` (≈ 8K gists at D=1024) survives across responses; all others are
58
+ dropped or cleared at `endResponse`. `_resolvedSubtrees` elides store probes
59
+ when `visit` is absent; with a visitor it still walks in full (see
60
+ `src/mind/primitives.ts:foldTree`).
61
+
62
+ ## Trace boundary — what is bypassed
63
+
64
+ Traced responses must emit every step, but must not change the answer.
65
+
66
+ - **Bypassed:** `_edgeChoice` via `guidedNext`
67
+ (`src/mind/traverse.ts:guidedNext`) and `sharedReachMemo`
68
+ (`src/mind/traverse.ts:sharedReachMemo`). Both return fresh empty maps when
69
+ `ctx.trace !== null`; `chooseNext` recomputes identically (pure over store +
70
+ guide).
71
+ - **Always consulted:** `perceiveMemo`, `recogniseMemo`, `climbMemo` (and their
72
+ underlying `perceive`/`recognise`/`climbAttention` caches). Bypassing breaks
73
+ idempotence.
74
+
75
+ `foldTree`'s subtree fast path is taken only when no `visit` is supplied. With a
76
+ visitor (recognition, attention) the walk still descends; the cache elides only
77
+ probes. Bypassing `recogniseMemo` under trace re-ran `recogniseImpl` with a warm
78
+ `_resolvedSubtrees` and emitted fewer sites (observed 31 → 5) — a correctness
79
+ change, not just a slowdown.
80
+
81
+ ## Meter — charge work to itself
82
+
83
+ Shared analyses are charged to their own phase (`meter.time(phase, fn)` in
84
+ `Precomputed.shared`), not to the mechanism that first touched them
85
+ (`src/meter.ts:PhaseCost`). Without this, the profile reads "cast.floor costs 2
86
+ s" when the cost was the consensus climb cast paid for on everyone's behalf.
87
+
88
+ ## Adding a shared analysis
89
+
90
+ Add one lazy method to `Precomputed`. No new memo map elsewhere. Gate it behind
91
+ `worthRunning` in `floor()`.
92
+
93
+ ## Pins
94
+
95
+ - `test/42` — recognition idempotence under trace: traced and untraced
96
+ `recognise` return the same site count and cached object.
@@ -0,0 +1,55 @@
1
+ # Meter — Work Accounting
2
+
3
+ `src/meter.ts` is the one computational-usage accounting surface. It counts what
4
+ inference _cost_ so a slow response can be attributed instead of guessed at. The
5
+ rationale says why an answer was chosen; the meter says what it cost to choose
6
+ it. Harness: `bench/profile-inference.mjs`.
7
+
8
+ ## Five contracts
9
+
10
+ 1. **Write-only from inference.** No counter reaches a decision, a threshold, or
11
+ an ordering. Determinism survives only because the meter is observed, never
12
+ consulted. Every call site is `meter?.x++` on a nullable field.
13
+
14
+ 2. **Counters vs hints.** Counters are deterministic and diffable between runs;
15
+ the same query on the same store meters identically, so a regression is
16
+ visible in a diff. Millisecond fields (`elapsedMs`, per-phase `ms`) are
17
+ non-deterministic hints reported separately — never use them to gate
18
+ behaviour.
19
+
20
+ 3. **Phases nest, they do not partition.** `think` contains every mechanism
21
+ phase; a mechanism's `floor` contains whatever shared analysis it
22
+ first-touched; `recall.run` contains `substitutionBridge`. Read a phase as
23
+ inclusive wall-clock — never sum phases and expect the total.
24
+ `CostReport.elapsedMs` is the only whole.
25
+
26
+ 4. **Count once.** Off by default and free when off
27
+ (`new Mind({ profile:
28
+ true })` to attach). A layer that wants to be
29
+ visible bumps a field in `meter.ts` — it does not grow a private counter.
30
+ (Legacy `danglingReads` / `compactFailures` in `store.ts` are health
31
+ counters, not per-response work.)
32
+
33
+ 5. **Shared analyses charged to themselves.** The first toucher pays the wall
34
+ clock, but every later consumer gets the result free. Attribution follows the
35
+ analysis, not the mechanism that first triggered it — otherwise the profile
36
+ misreads which work is expensive (e.g. the consensus climb billed through
37
+ whichever mechanism happened to need it first).
38
+
39
+ ## Where counters live
40
+
41
+ `src/meter.ts:Meter` is the only definition of a counter name. Phases are
42
+ charged via `meter.time(phase, fn)` / `meter.timeSync(phase, fn)`, which
43
+ snapshot counters on entry and attribute the delta to the phase. The sync/async
44
+ seam is load-bearing: synchronous layers (perception, recognition, graph search)
45
+ must use `timeSync` so the profiled path does not await where the unprofiled
46
+ path does not.
47
+
48
+ `CostReport` is plain JSON (`version`, `elapsedMs`, `queryBytes`, `counters`,
49
+ `phases`). Zero-valued counters are dropped; `formatReport` renders the three
50
+ heaviest counters per phase.
51
+
52
+ ## Pins
53
+
54
+ - `test/55` — `Meter`, `CostReport`, `searchPops` / `searchPushes`, phase
55
+ nesting.
@@ -0,0 +1,92 @@
1
+ # Saturation — Named Stop, Not Cap
2
+
3
+ > **Law:** every walk names a deciding saturation beside its cap. The cap is a
4
+ > safety net; saturation is the derived stop that decides and terminates.
5
+
6
+ A walk with only a cap drifts to the cap. A walk with a saturation stops the
7
+ moment the answer (saturated vs. decided) is known — bounded, exact below the
8
+ bound, and named in the trace.
9
+
10
+ ## Cap vs. saturation
11
+
12
+ | Role | Value | Nature |
13
+ | ---------- | ------------------------------------------------------------ | ------------------------------------------------ |
14
+ | Cap | `hubBound = ceil(sqrt(N))` per read; `hubBound * W` per walk | Safety net — prevents corpus-proportional work |
15
+ | Saturation | Named derived stop (`SaturationReason`) | Decision — proves continuing cannot discriminate |
16
+
17
+ `N = corpusN = max(2, edgeSourceCount())`, `W = maxGroup`. Defined once in
18
+ `mind/traverse.ts` (`corpusN`, `hubBound`, `boundFor`); never spelled inline.
19
+
20
+ ## edgeAncestors — EXPAND-UNTIL-DECIDED
21
+
22
+ `mind/traverse.ts:edgeAncestors` is the model: it climbs the structural DAG
23
+ (parents + containment) until one of five saturations decides, and is exact
24
+ below every one of them.
25
+
26
+ 1. **Predecessor fan-in** — `prevCount(node) > bound` via `store.prevCount`. One
27
+ indexed count; no read. Proves `> bound` distinct contexts on its own.
28
+ 2. **Distinct-context limit** — `ctxSeen.size > bound` after
29
+ `prevFirst(node, bound)`. The accumulated set of distinct contexts reachable
30
+ from roots visited so far exceeds `sqrt(N)`.
31
+ 3. **Parent fan-out** — `parentsFirst(node, bound+1).length > bound`. One
32
+ `LIMIT bound+1` read distinguishes "exactly bound parents" from "hub". The
33
+ node itself is not expanded; saturated reaches are never voted.
34
+ 4. **Lateral-cone cumulative** — `lateral > bound`, where `lateral` sums
35
+ `fresh-1` over every expanded node's extra parents beyond its first. The
36
+ per-node guard catches concentration at one node; this catches the same
37
+ commonness distributed across the cone. A deep chain in one structure accrues
38
+ zero laterals and still reaches its root at any depth.
39
+ 5. **Byte-atom commonality** — `atomIsHub(N,W)` when
40
+ `atomReach(N,W) = max(1, ceil(N*W/256)) > bound`. Atoms carry no kid/contain
41
+ rows, so containment is unmeasurable; the uniform-expectation floor replaces
42
+ it. Above the scale the atom abstains as voter (edges remain traversable for
43
+ tier-0 recall).
44
+
45
+ Below every threshold the walk is exact — `prevFirst(bound)` is the full list,
46
+ `parentsFirst(bound+1)` is the full list, `containersSlice` pages are walked in
47
+ full — identical to the unbounded climb. Work is `O(bound)` contexts times local
48
+ structure, never `O(N)`. Container seeding is streamed in `bound`-sized pages
49
+ for the same reason.
50
+
51
+ Trace records the first deciding stop as
52
+ `SaturationStop { reason, node, observed, limit }` plus `visited`/`maxDepth`;
53
+ absent when unsaturated or untraced.
54
+
55
+ ## pivotInto — longest-wins
56
+
57
+ `mind/resonance.ts:pivotInto` ranks candidates by `contentLen(id, answerLen+1)`
58
+ descending (first-inserted tie-break). The byte score is length itself, so the
59
+ scan is decided at the first candidate that passes every filter — a shorter
60
+ candidate can never outscore it. At most one winner's bytes are reconstructed;
61
+ every shorter proposal is skipped without a read. Saturation, not cap.
62
+
63
+ ## Junction — hub guards vs. budget
64
+
65
+ `mind/junction.ts:junctionContainersFrom` has three disciplines:
66
+
67
+ - **Phrase-scale reads** — `bytesPrefix(maxContainer+1)` per visit; a node
68
+ beyond the cap prunes its branch.
69
+ - **Per-node hub guards** (real saturations) — `parentsFirst(bound+1) > bound`
70
+ not expanded; one `containersSlice(bound+1)` page beyond `bound` not expanded.
71
+ Each is exact below `bound`.
72
+ - **Expansion budget** — at most `bound * W` pops total (shared across a tier's
73
+ walks). Budget exhaustion is an abstention (`junctionBudgetExhausted`) that
74
+ falls through to the resonance tier — a net, not a saturation.
75
+
76
+ Refuted tightening: applying `edgeAncestors`' cumulative lateral-cone limit here
77
+ would discard half the successful junctions (measured lateral 1425/1426 at
78
+ `bound` ~ 570).
79
+
80
+ Refuted early-stop: **one-cone-exhausted** — stopping when one side's upward
81
+ cone empties — is wrong in both hub-guarded and hub-flagged forms. A junction
82
+ can be reachable from only one side when that side's seed is a fold sub-node of
83
+ the container (`test/16`: "cold or hot" reached from window "cold" while 3-byte
84
+ "hot" cone is empty; `test/34` n-ary binding fails the same way). Exhausting one
85
+ cone never proves no junction remains; the walk must keep the `bound*W` net
86
+ after per-node saturations.
87
+
88
+ ## Pins
89
+
90
+ - `test/16` — one-cone-exhausted refutation (bridge junction).
91
+ - `test/34` — one-cone-exhausted refutation (n-ary cross-region binding).
92
+ - `test/27` — saturation-drop gate (leading/trailing saturated intervals).
@@ -0,0 +1,79 @@
1
+ # Store — AbstractStore Owns the Domain, Adapters Own the Wires
2
+
3
+ > **Law:** `AbstractStore` (`src/store.ts`) owns every domain decision — dedup,
4
+ > near-dedup, gist/halo indexing, containment, batching, LRU, compaction.
5
+ > `SQliteStore` (`src/store-sqlite.ts`) implements only `_db*`/`_vec*` thin
6
+ > wrappers.
7
+
8
+ ## Template method
9
+
10
+ ```
11
+ AbstractStore all logic: caches, merge gates, halo schedule,
12
+ buffers, chain transparency, length walks
13
+ └─ SQliteStore SQL + VectorDatabase glue — one statement per method
14
+ └─ <NewBackend> same contract — subclass AbstractStore only
15
+ ```
16
+
17
+ A new backend subclasses `AbstractStore`; never re-implements dedup or batching.
18
+
19
+ ## IDs and leaves
20
+
21
+ Branch ids are dense non-negative `0,1,2,…` (`_nextId`, never deleted).
22
+ Single-byte leaves are **implicit** negative ids `-256..-1` (`-(byte+1)`), never
23
+ a row. `has(id)` is `id < 0 || id < _nextId`.
24
+
25
+ ## Flat branches and bytes
26
+
27
+ A branch whose kids are all leaves is **flat** — stored as raw bytes in `leaf`
28
+ with an empty `kids` blob as marker (`flatKidsBytes`/`flatBytesKids`). Dedup
29
+ probes hash then verify: `hashOf`→`h`→`LIMIT 1` fetch→byte compare (bloom
30
+ negative filter first).
31
+
32
+ `bytes(id)`/`bytesPrefix(id, cap)` are shared with `BoundedMap` caches — callers
33
+ must **never mutate** the returned buffer. `contentLen(id, cap)` walks with
34
+ memo; when `cap` is given it saturates (`>= cap` without finishing) so one huge
35
+ root never costs a full walk.
36
+
37
+ ## Gist, halo, dedup
38
+
39
+ On `put*`, content dedup (`hashOf`→probe→mint) gates first. `DedupKey` caches
40
+ short keys (`DEDUP_KEY_MAX` bypass). Near-dedup merges by `mergeThreshold(D)` on
41
+ unit gist cosine. Gists sit in `_pendingGist` (byte-budgeted `BoundedMap`);
42
+ `indexSubtree` & `pourHalo` promote via `_vecContentUpsert`/`_vecHaloUpsert` in
43
+ `batchSize` batches. Buffers flush on cadence, `commit()`, and close. Halo mass
44
+ re-indexes geometrically (`mass<=4 || powerOfTwo`) and encodes 2-bit quantized.
45
+ Canon index is optional: `canonAdd`/ `canonFind`/`canonCount` over 32-bit
46
+ canonical hashes, caller verifies bytes.
47
+
48
+ ## Containment, batching, LRU
49
+
50
+ `addContainer(child,parent)` buffers per child; flush appends via
51
+ `_dbAppendContain` (packed pages, geometric merge) — never rewrite the whole
52
+ list. `containersSlice` pages through it. Edges and kids write through the same
53
+ deferred transaction.
54
+
55
+ Every in-memory cache is a `BoundedMap` with byte accounting and eviction (`lru`
56
+ vs `smallest` + `clock`/`reorder` recency). ANN reads
57
+ (`resonate`/`resonateHalo`) are content-addressed (`vecKey`) and dropped on any
58
+ index mutation; `RESonate_CACHE_MAX=4096`.
59
+
60
+ ## Maintenance (incremental)
61
+
62
+ - `compactContentIndex(minParents)` — scans only entries since last watermark
63
+ (`_vecContentEntriesSince`), removes indexed-but-isolated nodes (`<minParents`
64
+ parents, no edges/halos), compacts the vector DB.
65
+ - `repairContentIndex(regenerateGist)` — walks `_dbEdgeOrHaloIds()` candidates
66
+ only, re-inserts missing bridge nodes whose gists were evicted before
67
+ indexing.
68
+ - `buildCanonIndex` (`_buildCanonIndex`) — iterates `eachContent(fromId)` and
69
+ `canonAdd`s; `fromId` makes refresh incremental.
70
+
71
+ Full scans (`parents()`, `next()`, `containers()`) are maintenance-only — hot
72
+ paths use `LIMIT`ed probes (`parentsFirst`/`nextFirst`/`prevFirst`), `has*`,
73
+ `prevCount`.
74
+
75
+ ## Adding a backend
76
+
77
+ Implement every `protected abstract _db*`/`_vec*` in `src/store.ts` as a thin
78
+ wrapper around your storage. Keep `_dbGet*First`/`Slice` as real `LIMIT` queries
79
+ and `has*`/`COUNT` as point probes — never materialise-then-slice.
@@ -0,0 +1,79 @@
1
+ # Thresholds — derived, never tuned
2
+
3
+ Every decision cutoff is a formula over `D` (vector dimension), `W` (`maxGroup`,
4
+ perception window), or `N` (corpus size). No threshold is tuned or added to
5
+ `src/config.ts`.
6
+
7
+ ## Source of truth
8
+
9
+ - `src/geometry.ts` — all similarity/decision thresholds.
10
+ - `src/mind/traverse.ts` — corpus-scale readings that parameterise bounded
11
+ walks.
12
+ - `src/sema.ts` — positional coordinate algebra.
13
+ - `src/config.ts` — capacities and budgets only (cache byte budgets, batch
14
+ sizes, index parameters, query `k`, ALU precision, seed). Never a threshold.
15
+
16
+ ## Geometry thresholds (`src/geometry.ts`)
17
+
18
+ | Symbol | Definition | Formula |
19
+ | ----------------------- | ---------------------------------------------------------------------------- | ----------------------------------- |
20
+ | `mergeThreshold(D)` | Store identity bar — cosine at which `intern` treats two gists as same node | `1 - 1/√D` |
21
+ | `identityBar(D,W,len)` | Scale-aware whole-span identity claim | `max(mergeThreshold(D), 1 - W/len)` |
22
+ | `reachThreshold(W)` | Recall confidence floor — half a river quantum | `1 - 1/(2·W)` |
23
+ | `estimatorNoise(D)` | RaBitQ noise floor — 1σ of random cosine | `1/√D` |
24
+ | `significanceBar(D)` | Whole-query relatedness — 3σ above chance | `3/√D` |
25
+ | `conceptThreshold(D)` | Halo concept sharing — structural midpoint + ½σ | `0.5 + 0.5/√D` |
26
+ | `dominates(part,whole)` | Half-dominance predicate | `part*2 > whole` |
27
+ | `profileCapacity(D)` | Superposition capacity — terms before readout collapses | `floor(√D)` (min 1) |
28
+ | `consensusFloor(N)` | Pooled-vote significance floor | `ln(N) + ½` |
29
+ | `coverageBar(_,D)` | Reach-index gating (currently unused hot-path; batch compaction replaces it) | `conceptThreshold(D)` |
30
+
31
+ `N` in `consensusFloor` is `corpusN` (edge-source count, floored at 2).
32
+
33
+ ## Corpus-scale readings (`src/mind/traverse.ts`)
34
+
35
+ | Symbol | Definition | Formula |
36
+ | ---------------- | ------------------------------------------------------ | --------------------------- |
37
+ | `corpusN` | Distinct learnt contexts, floored | `max(2, edgeSourceCount())` |
38
+ | `hubBound` | Hub bound, used for every `LIMIT` read | `ceil(√max(2,N))` |
39
+ | `hubCap(ids)` | Fan-out cap — list-side reading of `hubBound` | `ids.slice(0, hubBound)` |
40
+ | `atomReach(N,W)` | Uniform-expectation floor on a byte atom's commonality | `max(1, ceil(N·W/256))` |
41
+ | `atomIsHub(N,W)` | Whether atom abstains as consensus voter | `atomReach > hubBound` |
42
+
43
+ `hubBound` is enforced at the store level (`nextFirst`, `parentsFirst`,
44
+ `containersSlice`, `hasNext`/`hasParents`, `bytesPrefix`, `chainRun`).
45
+ `atomReach` is the honest floor for atoms — they carry no kid/contain rows, so
46
+ their reach is unmeasurable and must not default to "maximally rare".
47
+
48
+ ## Seat algebra (`src/sema.ts`)
49
+
50
+ `twoEndedSeat(seatCount, size, index)` — the one positional-coordinate algebra
51
+ shared by perception, `fold`, and every synthetic/canonical fold. First half
52
+ uses low seats, second half uses high seats:
53
+ `index < (size+1)/2 ? index : seatCount-size+index`.
54
+
55
+ ## Two derivations that bite
56
+
57
+ **1. `identityBar` is scale-aware.** A fixed cosine `1-1/√D` over a `4·√D`-byte
58
+ span tolerates four whole windows of foreign bytes while still claiming
59
+ "near-identical". An identity claim may tolerate at most one window `W` (the
60
+ perception quantum, same budget as `differsByOneWindow` in near-dedup), so the
61
+ bar must be `1-W/len` floored at `mergeThreshold`. Reusing `mergeThreshold` for
62
+ a whole-span claim silently widens the byte budget with span length.
63
+
64
+ **2. `consensusFloor` is priced for pooled climb votes, not for support
65
+ counts.** Each region contributes at most `ln(N/c) ≤ ln(N)`; `ln(N)+½` demands
66
+ corroboration beyond one maximally-specific region. `chooseNext`'s `bestSupport`
67
+ (`prevCount` of one destination) is N-invariant — bounded by retellings of that
68
+ fact, not by `N`. Gating it against `consensusFloor` guarantees failure once `N`
69
+ is large enough (observed: 2-vs-1-1-1 corroboration refused at N≈325K, falling
70
+ back to a noisy concept-hop).
71
+
72
+ ## Pins
73
+
74
+ - `test/40-choosenext-scale-guard` — `consensusFloor` must not gate `chooseNext`
75
+ support counts.
76
+ - `test/64-two-ended-thresholds` — `mergeThreshold` / `identityBar` /
77
+ `reachThreshold` derivations and the scale-aware floor.
78
+ - `test/78-atom-hub-recognition-cliff` — `atomReach` / `atomIsHub` hub
79
+ abstention at scale.
@@ -0,0 +1,144 @@
1
+ # Tempting but Wrong — 12 Traps
2
+
3
+ Twelve shortcuts that look plausible and break an invariant. Each states what
4
+ not to do, why it fails, and what to do instead.
5
+
6
+ ### 1. `score >= threshold` decides identity
7
+
8
+ - **WRONG:** Treat a RaBitQ cosine above a cutoff as proof the bytes are the
9
+ same node.
10
+ - **WHY:** Scores are estimates that rank and gate; they never decide identity
11
+ (`AGENTS §2` Invariant 3 — Exact decides / approximate proposes).
12
+ - **CORRECT:** Gate with the score, decide with
13
+ `resolve`/`findLeaf`/`canonResolve` and re-fold verification. Pinned by
14
+ `test/51-structural-resonance-ladder.test.mjs` and
15
+ `test/56-bridge-identity-admission.test.mjs`.
16
+
17
+ ### 2. `Math.random()` / `Date.now()` on a behavioural path
18
+
19
+ - **WRONG:** Sample randomness or wall-clock time in grounding, indexing, or
20
+ tie-breaking.
21
+ - **WHY:** Determinism is the product: same seed + deposit order + query ⇒
22
+ identical bytes (`AGENTS §2` Invariant 1).
23
+ - **CORRECT:** Derive all randomness from `MindConfig.seed` via `rng`/`Prng`;
24
+ keep `example/train_base` as the only non-library exception. Pinned by
25
+ `test/20-stability.test.mjs` and
26
+ `test/42-recognise-trace-idempotence.test.mjs`.
27
+
28
+ ### 3. Last-inserted tie-break
29
+
30
+ - **WRONG:** Break equal-rank ties by picking the most recently inserted
31
+ edge/node.
32
+ - **WHY:** Tie-breaks must be corpus-determined and stable; last-inserted is
33
+ recency-dependent and was fixed as a bug (`AGENTS §2` Invariant 1 —
34
+ first-inserted fallback).
35
+ - **CORRECT:** `guidedFirst`/`chooseNext`/`chooseAmong`: rank then
36
+ first-inserted (lowest node id / `LIMIT 1` insertion order). Pinned by
37
+ `test/03-recall.test.mjs` determinism suites.
38
+
39
+ ### 4. Tunable threshold in `config.ts`
40
+
41
+ - **WRONG:** Add a new `threshold: number` to `src/config.ts` and tune it.
42
+ - **WHY:** Every cutoff is a formula over `D`, `W`, or `N` in `src/geometry.ts`;
43
+ config holds only capacities/budgets (`AGENTS §2` Invariant 2 — Derived
44
+ thresholds).
45
+ - **CORRECT:** Add `mergeThreshold`/`identityBar`/`significanceBar` etc.
46
+ derivation in `geometry.ts`; `traverse.ts:hubBound` for scale caps. Pinned by
47
+ `test/64-two-ended-thresholds.test.mjs` and
48
+ `test/40-choosenext-scale-guard.test.mjs`.
49
+
50
+ ### 5. Tuning `PASS` to encode policy
51
+
52
+ - **WRONG:** Raise/lower `PASS` (1000/byte) so "computation always wins" or
53
+ another preference falls out of pricing.
54
+ - **WHY:** The ladder's order `MICRO < STEP < CONCEPT < PASS` is the contract;
55
+ policy is enforced by masking, not pricing (`AGENTS §2` Invariant 4 — One cost
56
+ currency; `docs/architecture/cost-model.md` § Policy is not cost).
57
+ - **CORRECT:** Keep `PASS` dominating; enforce precedence in the caller (e.g.
58
+ `pipeline.ts` masks recognised sites overlapped by `ComputedResult`). Pinned
59
+ by `test/04-think.test.mjs` and `test/55-cost-meter.test.mjs`.
60
+
61
+ ### 6. Reimplementing `locate`/`align` inside a mechanism
62
+
63
+ - **WRONG:** Copy-paste matching logic into `mind/mechanisms/*.ts` with a
64
+ private gate.
65
+ - **WHY:** Match/project/gate is factored once in `mind/match.ts` (`AGENTS §2`
66
+ Cross-cutting contracts; `AGENTS §3` — Where things live).
67
+ - **CORRECT:** Configure the shared family:
68
+ `locate`/`alignRuns`/`alignGraded`/`frameSlots` + `follow`/`reverseContext` +
69
+ `isSpanShaped`/`carriesFillers` with a `geometry.ts` gate. Pinned by
70
+ `test/50-cast-analog-consensus-floor.test.mjs`.
71
+
72
+ ### 7. Putting voicing gates in `frameSlots`
73
+
74
+ - **WRONG:** Make `frameSlots` refuse pairings that fail `carriesFillers` or
75
+ reference's four voicing conditions.
76
+ - **WHY:** `frameSlots` reports (contracted gaps tagged
77
+ substitution/insertion/deletion); `carriesFillers` judges;
78
+ `Precomputed.frames` inventories — elects nothing (`AGENTS §2` Cross-cutting
79
+ contracts — `docs/architecture/factored-machinery.md` § Frame reading).
80
+ - **CORRECT:** Report everything in the shared layer; apply
81
+ `substituteAll(contA, fillersA→fillersB)==contB` and
82
+ frame-dominance/`W`-reach/distinctness in the consumer (reference). Pinned by
83
+ `test/47-cast-comparison-coverage.test.mjs`.
84
+
85
+ ### 8. Swapping corpus-global and weave-local commonality
86
+
87
+ - **WRONG:** Use `reachOf`/`dominates(reach,N)` to decide CAST's frame, or
88
+ `depth[i]`/`dominates(depth, aligned)` to decide climb/IDF.
89
+ - **WHY:** They measure different things: global reach (minority discriminates,
90
+ powers climb/pooling) vs weave-local depth with `MIN_WEAVE=2` (what the local
91
+ cohort shares, powers CAST) (`AGENTS §2` Cross-cutting contracts — Two
92
+ measures of commonality).
93
+ - **CORRECT:** Climb/attention uses corpus-global;
94
+ `frame(i) ⇔ depth[i]>MIN_WEAVE ∧ dominates(depth[i],aligned)` for CAST. Pinned
95
+ by `test/50-cast-analog-consensus-floor.test.mjs` and
96
+ `test/67-climb-anchor-breadth.test.mjs`.
97
+
98
+ ### 9. Materialise-then-slice instead of `LIMIT ?`
99
+
100
+ - **WRONG:** `store.next(id).slice(0, k)` or `parents(id).length` to cap a
101
+ fan-out.
102
+ - **WHY:** Per-query reads must not grow with `N`; caps are enforced in SQL as
103
+ `LIMIT ?` / `EXISTS` probes (`AGENTS §2` Invariant 5 — Bounded reads).
104
+ - **CORRECT:** `nextFirst`/`parentsFirst`/`containersSlice` with `hubBound`
105
+ (`ceil(sqrt(N))`), `hasNext`/`hasParents`/`hasHalo` probes,
106
+ `bytesPrefix`/`contentLen` caps, `chainRun` CTE. Pinned by
107
+ `test/14-scaling.test.mjs` and `test/90-connector-read-cap.test.mjs`.
108
+
109
+ ### 10. Bypassing `recogniseMemo` under trace
110
+
111
+ - **WRONG:** Skip `recogniseMemo`/`perceiveMemo`/`climbMemo` when
112
+ `ctx.trace !== null` to "emit more steps."
113
+ - **WHY:** Only `guidedNext`/`sharedReachMemo` are trace-bypassed; bypassing
114
+ recognition re-runs `recogniseImpl` with a warm cache and changes site count —
115
+ 31→5 observed (`AGENTS §2` Cross-cutting — `Precomputed` owns memoization;
116
+ `docs/architecture/memoization.md`).
117
+ - **CORRECT:** Always consult `recogniseMemo`; `foldTree` already descends fully
118
+ when `visit` is present. Pinned by
119
+ `test/42-recognise-trace-idempotence.test.mjs`.
120
+
121
+ ### 11. Stopping junction ascent when one cone is exhausted
122
+
123
+ - **WRONG:** Terminate the `junction.ts` walk as soon as parents or containers
124
+ run out.
125
+ - **WHY:** Junction ascent climbs both cones within a bounded `√N·W` walk via
126
+ `WalkCache`; exhausting one cone does not imply the other is exhausted —
127
+ stopping early misses the shared ancestor (`AGENTS §2` Cross-cutting contracts
128
+ — `junction.ts` is the shared ascent; `AGENTS §3` — `WalkCache`).
129
+ - **CORRECT:** Continue the live cone until the walk budget is spent or a
130
+ meeting point is found; cap reads with `hubBound`. Pinned by
131
+ `test/34-cross-region.test.mjs` and
132
+ `test/52-climb-consensus-instrumentation.test.mjs`.
133
+
134
+ ### 12. Imposing turn boundaries on `fold`
135
+
136
+ - **WRONG:** Cut the byte stream at conversation turn edges before folding, so
137
+ deposits and queries fold differently.
138
+ - **WHY:** Perception is a pure function of the bytes; deposit and inference
139
+ must compute the same tree for the same input (`AGENTS §1` Orientation — Hard
140
+ facts).
141
+ - **CORRECT:** Fold content-defined cuts (`contentLevels` in `geometry.ts` +
142
+ `twoEndedSeat`); turns are API state in `mind/mind.ts`, not segmentation.
143
+ Pinned by `test/59-fold-invariance.test.mjs` and
144
+ `test/63-fold-invariants.test.mjs`.
@@ -0,0 +1,56 @@
1
+ # Gates
2
+
3
+ Four executable gates. Each: run the command, check what it guards, follow its
4
+ §.
5
+
6
+ ## 1 — Correctness (all 87 suites)
7
+
8
+ ```bash
9
+ npm test
10
+ ```
11
+
12
+ Guards honest silence, determinism, and every pinned contract. Silence:
13
+ unrelated queries ground to nothing (`test/28`, `50`, `56`, `67`, `76`, `84`).
14
+ Determinism: same seed + deposit order + query gives byte-identical answer
15
+ (`test/20`). Every invariant is pinned — a simplification that fails a test is
16
+ wrong until the test is shown wrong. §14–25 (pipeline), §8 (derived thresholds),
17
+ AGENTS.md §2 invariants 1–5.
18
+
19
+ ## 2 — Work accounting (profiler)
20
+
21
+ ```bash
22
+ node bench/profile-inference.mjs # add [n] to limit probes
23
+ node bench/profile-inference.mjs --trace # trace is a debugging aid, not product
24
+ ```
25
+
26
+ Guards without trace: counters deterministic and diffable between runs; phases
27
+ nest (not disjoint — `think` contains every mechanism phase); shared analyses
28
+ charged to themselves, not to the first toucher; millisecond fields are
29
+ non-deterministic hints only. With `--trace`, recognition idempotence still
30
+ holds (`test/42`). `src/meter.ts`, `docs/architecture/meter.md`, §26, AGENTS.md
31
+ §2 invariant meter/cost.
32
+
33
+ ## 3 — Dependency footprint
34
+
35
+ ```bash
36
+ node --test test/88-dependency-footprint.test.mjs
37
+ ```
38
+
39
+ Guards `dist/src` imports only `node:` + relative paths, and `package.json`
40
+ declares no `dependencies` (examples use `devDependencies` lazily). The
41
+ near-zero footprint is a product feature. AGENTS.md §6, §3 (store has one
42
+ runtime dep: `node:sqlite`).
43
+
44
+ ## 4 — Fold invariance and sublinear scaling
45
+
46
+ ```bash
47
+ node --test test/59-fold-invariance.test.mjs test/63-fold-invariants.test.mjs
48
+ node --test test/14-scaling.test.mjs
49
+ ```
50
+
51
+ Guards: `59+63` — segmentation is content-defined (`contentBoundaries`), not
52
+ positional; grid regression (14.3% survival) cannot pass. `14` — inference cost
53
+ is sublinear in corpus size (power-law exponent ≪ 1) and constant-rate in input
54
+ length; measured on independent disjoint corpora via log–log slope.
55
+ `src/geometry.ts` (`contentLevels`), `docs/architecture/fold-contract.md` +
56
+ `bounded-reads.md`, §10, §29.