@hviana/sema 0.7.3 → 0.7.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +95 -843
- package/README.md +11 -11
- package/dist/src/mind/graph-search.d.ts +32 -22
- package/dist/src/mind/graph-search.js +97 -54
- package/dist/src/mind/mind.js +14 -0
- package/dist/src/store-sqlite.js +17 -0
- package/dist/src/store.d.ts +18 -0
- package/dist/src/store.js +10 -0
- package/docs/INDEX.md +71 -0
- package/docs/INVARIANTS.md +19 -0
- package/docs/architecture/bounded-reads.md +85 -0
- package/docs/architecture/caches.md +89 -0
- package/docs/architecture/commonality.md +45 -0
- package/docs/architecture/cost-model.md +71 -0
- package/docs/architecture/determinism.md +73 -0
- package/docs/architecture/exact-vs-approximate.md +47 -0
- package/docs/architecture/factored-machinery.md +28 -0
- package/docs/architecture/fold-contract.md +87 -0
- package/docs/architecture/halo-sketch.md +99 -0
- package/docs/architecture/match-project.md +62 -0
- package/docs/architecture/mechanism-market.md +95 -0
- package/docs/architecture/memoization.md +96 -0
- package/docs/architecture/meter.md +55 -0
- package/docs/architecture/saturation.md +92 -0
- package/docs/architecture/store.md +79 -0
- package/docs/architecture/thresholds.md +79 -0
- package/docs/failures/tempting-but-wrong.md +144 -0
- package/docs/harness/gates.md +56 -0
- package/docs/mechanisms/alu.md +75 -0
- package/docs/mechanisms/cast.md +75 -0
- package/docs/mechanisms/confluence.md +36 -0
- package/docs/mechanisms/cover.md +54 -0
- package/docs/mechanisms/extraction.md +53 -0
- package/docs/mechanisms/prefix-completion.md +54 -0
- package/docs/mechanisms/recall.md +69 -0
- package/docs/mechanisms/reference.md +58 -0
- package/jsr.json +1 -1
- package/package.json +1 -1
- package/src/mind/graph-search.ts +101 -55
- package/src/mind/mind.ts +14 -0
- package/src/store-sqlite.ts +19 -0
- package/src/store.ts +22 -0
- package/test/89-completion-recursion.test.mjs +30 -10
- package/test/97-store-seed.test.mjs +105 -0
- package/test/98-completion-chaining.test.mjs +140 -0
- package/HOW_IT_WORKS.md +0 -5836
|
@@ -0,0 +1,85 @@
|
|
|
1
|
+
# Bounded Reads — No Per-Query Read Grows With the Corpus
|
|
2
|
+
|
|
3
|
+
> **Law:** the cost of one query is proportional to the query, not to how much
|
|
4
|
+
> was learned. No per-query read may grow with corpus size N.
|
|
5
|
+
|
|
6
|
+
Every fan-out, walk, and disambiguation is capped at `hubBound` —
|
|
7
|
+
`ceil(sqrt(N))` — derived once from `corpusN` and floored at 2 so `sqrt` and
|
|
8
|
+
`ln` stay meaningful on a near-empty store. There is no second convention; do
|
|
9
|
+
not invent one.
|
|
10
|
+
|
|
11
|
+
## Scale
|
|
12
|
+
|
|
13
|
+
```
|
|
14
|
+
corpusN(ctx) = max(2, store.edgeSourceCount()) // distinct learnt contexts
|
|
15
|
+
hubBound(ctx) = ceil(sqrt(corpusN(ctx))) // >= 2, the store cap
|
|
16
|
+
hubCap(ctx, ids) = ids.slice(0, hubBound(ctx)) // list-side reading
|
|
17
|
+
boundFor(n) = ceil(sqrt(max(2, n))) // ctx-free reading
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
Defined once in `mind/traverse.ts` (`corpusN`, `hubBound`, `hubCap`,
|
|
21
|
+
`boundFor`). Every consumer imports them; never spell `Math.sqrt` inline.
|
|
22
|
+
|
|
23
|
+
## Enforcement at the store level
|
|
24
|
+
|
|
25
|
+
The cap is not advisory — adapters must make bounded reads bounded in SQL.
|
|
26
|
+
|
|
27
|
+
### 1. LIMITed reads — real `LIMIT ?`
|
|
28
|
+
|
|
29
|
+
`nextFirst(id, limit)`, `prevFirst(id, limit)`, `parentsFirst(id, limit)`,
|
|
30
|
+
`containersSlice(child, offset, limit)`.
|
|
31
|
+
|
|
32
|
+
Same statement and `ORDER BY` as the full read, with `LIMIT ?`. Never
|
|
33
|
+
"materialise then slice". Reading `hubBound + 1` parents decides "hub or not"
|
|
34
|
+
exactly without reading the rest. Implemented as thin wrappers in
|
|
35
|
+
`store-sqlite.ts` over `AbstractStore` in `store.ts`.
|
|
36
|
+
|
|
37
|
+
### 2. Existence probes — indexed point probes
|
|
38
|
+
|
|
39
|
+
`hasNext(id)`, `hasParents(id)`, `hasContainers(child)`, `hasHalo(id)`,
|
|
40
|
+
`prevCount(id)`.
|
|
41
|
+
|
|
42
|
+
One indexed `EXISTS` / `COUNT` probe that never decodes vectors or unpacks
|
|
43
|
+
blobs. Use them for every "does this lead anywhere?" question instead of
|
|
44
|
+
`next(id).length > 0` or `prev(id).length`. `prevCount` is the reverse-edge
|
|
45
|
+
support count for `chooseNext`/`chooseAmong`; `hasNext`/`hasHalo` gate the
|
|
46
|
+
`leadsSomewhere` admission predicate in `mind/traverse.ts`.
|
|
47
|
+
|
|
48
|
+
### 3. Prefix-capped reads — reject without reconstruction
|
|
49
|
+
|
|
50
|
+
`bytesPrefix(id, cap)` and `contentLen(id, cap)`.
|
|
51
|
+
|
|
52
|
+
`contentLen` under a cap returns an exact length below the cap and `>= cap`
|
|
53
|
+
otherwise — an indexed memo walk that stops early, never a full subtree walk.
|
|
54
|
+
`bytesPrefix` stops after `cap` bytes. A candidate exceeding the cap is rejected
|
|
55
|
+
on the length probe alone; the weave, junction walks, and bridge all read this
|
|
56
|
+
way. Uncapped reads there cost seconds per query on a large store.
|
|
57
|
+
|
|
58
|
+
### 4. Transparent scaffolding — one bounded read
|
|
59
|
+
|
|
60
|
+
`chainRun(id)` climbs a run of transparent nodes (exactly one parent, no edges)
|
|
61
|
+
in a single recursive CTE, cached for the store lifetime and dropped on writes
|
|
62
|
+
that break transparency. The climber hops the whole run where a node-at-a-time
|
|
63
|
+
ascent would pay three probes per node.
|
|
64
|
+
|
|
65
|
+
## Maintenance only
|
|
66
|
+
|
|
67
|
+
The full materialising reads — `next(id)`, `prev(id)`, `parents(id)`,
|
|
68
|
+
`containers(child)` — exist for inspection, repair, and compaction only
|
|
69
|
+
(`compactContentIndex`, `repairContentIndex`). Keep them off hot paths.
|
|
70
|
+
|
|
71
|
+
## Adding a walk
|
|
72
|
+
|
|
73
|
+
Any new fan-out walk uses `hubBound`/`hubCap`. Do not call `edgeSourceCount()`
|
|
74
|
+
or `Math.ceil(Math.sqrt(...))` inline, and do not invent a per-walk limit. The
|
|
75
|
+
walk's saturation decision (when to stop) is separate from the cap (the safety
|
|
76
|
+
net); a walk with only a cap drifts to the cap.
|
|
77
|
+
|
|
78
|
+
## Pins
|
|
79
|
+
|
|
80
|
+
- `test/14` — sublinear inference in corpus size and constant-rate in input
|
|
81
|
+
length; training throughput floor; exact recall at scale.
|
|
82
|
+
- `test/89` — completion recursion stays output-sensitive (nested searches/pops
|
|
83
|
+
sublinear); guards the count of reads, not just per-read size.
|
|
84
|
+
- `test/90` — connector probe (`resolveConnectors`) reads by the query length
|
|
85
|
+
(`QUERY.length + 1`), not by the learnt continuation; per-read size bound.
|
|
@@ -0,0 +1,89 @@
|
|
|
1
|
+
# Caches — Every Acceleration Is a BoundedMap
|
|
2
|
+
|
|
3
|
+
> **Law:** every acceleration is a `BoundedMap` with a byte budget. A miss
|
|
4
|
+
> re-derives from durable state. Degradation order is speed/reach lost, never
|
|
5
|
+
> identity.
|
|
6
|
+
|
|
7
|
+
No cache may change what is stored, what is resolved, or what tree is folded.
|
|
8
|
+
Eviction costs a re-read, a re-walk, or a narrower reach — never a wrong answer
|
|
9
|
+
or a wrong tree.
|
|
10
|
+
|
|
11
|
+
## `BoundedMap` — the one cache primitive
|
|
12
|
+
|
|
13
|
+
`src/store.ts:BoundedMap<K,V>` — LRU with byte accounting (`maxBytes`, `sizeOf`,
|
|
14
|
+
`evict`, `recency`).
|
|
15
|
+
|
|
16
|
+
- `evict: "lru"` — uniform-cost entries (dedup, vectors, records).
|
|
17
|
+
- `evict: "smallest"` — variable-cost reconstruction (`_bytesCache`): protects
|
|
18
|
+
expensive large branches over cheap leaves.
|
|
19
|
+
- `recency: "reorder"` (default) — exact LRU via `delete+set`; required when
|
|
20
|
+
eviction choice is load-bearing (`_depositTrees` — 8 entries, victim changes
|
|
21
|
+
fold).
|
|
22
|
+
- `recency: "clock"` — bit instead of reorder; only for transparent caches where
|
|
23
|
+
wrong victim costs a re-read (`_bytesCache`, `_recCache`). Measured: same
|
|
24
|
+
entries cached, hot-path time 55% → bit.
|
|
25
|
+
|
|
26
|
+
Persistent cursor over V8 insertion order makes eviction amortised O(1);
|
|
27
|
+
candidate window for `"smallest"` never rescans from front.
|
|
28
|
+
|
|
29
|
+
## Store caches — budgets in `src/config.ts:StoreConfig`
|
|
30
|
+
|
|
31
|
+
| Cache | Field | Budget | `sizeOf` | Eviction |
|
|
32
|
+
| ------------------- | -------------------------- | ---------------------------- | ------------------- | -------------- |
|
|
33
|
+
| dedup leaf/branch | `_leafKey` / `_branchKey` | `dedupCacheMax` 1M entries | 1 | lru |
|
|
34
|
+
| reconstructed bytes | `_bytesCache` | `bytesCacheMax` 20 MB | `byteLength` | smallest+clock |
|
|
35
|
+
| content length | `_lenCache` | `bytesCacheMax` | 16 | lru |
|
|
36
|
+
| node records | `_recCache` | `recCacheBytes` 10 MB | leaf+4·kids+12 | lru+clock |
|
|
37
|
+
| pending gists | `_pendingGist` | `pendingGistBytes` 16 MB | `byteLength` (D·4) | lru |
|
|
38
|
+
| halo exact / norm | `_haloExact` / `_haloNorm` | `haloCacheBytes` 16 MB each | `byteLength` | lru |
|
|
39
|
+
| skipped interiors | `_coveredIds` | `coveredIdsMax` 100K entries | 1 | lru |
|
|
40
|
+
| indexed ids | `_indexedIds` | `coveredIdsMax` | 1 | lru |
|
|
41
|
+
| transparent chains | `_chainMemo` | `chainCacheBytes` 16 MB | 4·len+32 | lru |
|
|
42
|
+
| ingest memo | `CachedIngest._memo` | `ingestCacheBytes` 50 MB | vector+ids+keyBytes | lru |
|
|
43
|
+
|
|
44
|
+
ANN read caches (`_resonateCache`, `_resonateHaloCache`) are `Map<string,Hit[]>`
|
|
45
|
+
keyed by `vecKey(v)+":"+k`, dropped on any index mutation.
|
|
46
|
+
`vectorCacheMb`/`sqliteCacheMb` are pure page-cache latency knobs.
|
|
47
|
+
|
|
48
|
+
`_bytesCache` only caches complete reconstructions — `bytesPrefix(id,cap)` with
|
|
49
|
+
`got < cap`; a truncated prefix is never stored. `_chainMemo` is dropped on any
|
|
50
|
+
write that could break transparency; `_pendingGist` eviction falls back to DAG
|
|
51
|
+
climb; halo eviction re-decodes the durable 2-bit row.
|
|
52
|
+
|
|
53
|
+
## Mind caches — session and per-response
|
|
54
|
+
|
|
55
|
+
| Cache | Location | Budget | Scope / invalidation |
|
|
56
|
+
| ------------------- | ------------------------------------------------ | --------- | ------------------------------------------------------------------ |
|
|
57
|
+
| `_gistCache` | `Mind._gistCache` | 32 MB | session-lifetime, never invalidated (perception pure) |
|
|
58
|
+
| `_depositTrees` | `Mind._depositTrees` | 8 entries | session; `perceiveDeposit` only when `conversational` |
|
|
59
|
+
| `_depositLens` | `Mind._depositLens` | — | byte lengths for prefix probes; cleared with map when >64 |
|
|
60
|
+
| `_internIds` | `Mind._internIds: WeakMap<Sema,number>` | — | Mind lifetime; ids permanent |
|
|
61
|
+
| `_resolvedSubtrees` | `Mind._resolvedSubtrees: WeakMap<Sema,{id,len}>` | — | per-response/conversation; fast path only when `visit===undefined` |
|
|
62
|
+
|
|
63
|
+
`REACH_MEMO_MAX` / `STRUCT_MEMO_MAX` 100K (`src/mind/traverse.ts`) — whole-climb
|
|
64
|
+
and per-node structural probes (`hasNext`/`prevCount`/`hasParents`); cleared on
|
|
65
|
+
write or when cap reached. `reachMemo`/`structCaches` keyed by `_structMemoKey`,
|
|
66
|
+
bypassed under trace.
|
|
67
|
+
|
|
68
|
+
## Deposit caches — offset-keyed, caller-discharged
|
|
69
|
+
|
|
70
|
+
`_depositTrees`/`_depositLens`/`_internIds`/`_resolvedSubtrees` key **offsets**,
|
|
71
|
+
not bytes, for O(1) reuse. Offsets alone cannot witness byte agreement — caller
|
|
72
|
+
must discharge it.
|
|
73
|
+
|
|
74
|
+
- Correct: conversation append — each turn extends the prior cumulative context
|
|
75
|
+
by its own bytes; longest cached proper prefix hit (`L < bytes.length`) reuses
|
|
76
|
+
`contentFoldIncremental` segments bit-identically.
|
|
77
|
+
- Wrong: mismatched `prev` reused by offset produced wrong tree (336 vs 400
|
|
78
|
+
bytes) — a coincidental prefix length aliased unrelated content.
|
|
79
|
+
- Now: `perceiveDeposit` keys by `latin1Key(bytes.subarray(0,L))` (prefix
|
|
80
|
+
bytes), probes longest cached proper prefix first; `_depositTrees` populated
|
|
81
|
+
only for conversational deposits (budget discipline), otherwise cold path
|
|
82
|
+
always correct.
|
|
83
|
+
|
|
84
|
+
## Pins
|
|
85
|
+
|
|
86
|
+
- `test/91` — `chainRun` via capped `_prefix`: bounded transparent-chain hop,
|
|
87
|
+
not per-node probes.
|
|
88
|
+
- `test/96` — `_bytesCache` is byte-accounted `BoundedMap` that evicts; miss
|
|
89
|
+
re-derives.
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
# Two Measures of Commonality
|
|
2
|
+
|
|
3
|
+
Sema needs "what is shared" in two different populations. One is corpus-global
|
|
4
|
+
(how widely a structure is reused), the other is weave-local (what a local
|
|
5
|
+
cohort of overlapping forms agrees on). They use different data and different
|
|
6
|
+
formulas and must not be conflated.
|
|
7
|
+
|
|
8
|
+
## Corpus-global — `reachOf` + `dominates`
|
|
9
|
+
|
|
10
|
+
_Defined in `src/mind/traverse.ts` + `src/geometry.ts`; used by climb,
|
|
11
|
+
containment, IDF pooling._
|
|
12
|
+
|
|
13
|
+
For a node id, `reachOf(id, N)` counts how many learnt contexts contain it
|
|
14
|
+
(ancestor reach via capped graph walks, memoised per response in
|
|
15
|
+
`sharedReachMemo`). `dominates(reach, N)` then asks whether that reach is above
|
|
16
|
+
the corpus-determined majority threshold (derived in `geometry.ts` over `N`).
|
|
17
|
+
Intuition: minority reach discriminates (a filler), majority reach is
|
|
18
|
+
scaffolding. Powers the consensus climb, edge following, and vote pooling.
|
|
19
|
+
|
|
20
|
+
## Weave-local — `depth[]` + `MIN_WEAVE` + `dominates`
|
|
21
|
+
|
|
22
|
+
_Defined in `src/mind/match.ts` (`depth[]`, `MIN_WEAVE`, `frame`) and gated in
|
|
23
|
+
`src/mind/match.ts:frame`; used by CAST._
|
|
24
|
+
|
|
25
|
+
For an alignment weave, `depth[i]` counts how many aligned structures cover byte
|
|
26
|
+
`i` of the query. `MIN_WEAVE = 2` requires agreement beyond a pair (pair columns
|
|
27
|
+
are ambiguous with insertions/deletions), and `dominates(depth[i], aligned)`
|
|
28
|
+
requires agreement by a majority of the aligned cohort:
|
|
29
|
+
|
|
30
|
+
```
|
|
31
|
+
frame(i) ⇔ depth[i] > MIN_WEAVE ∧ dominates(depth[i], aligned)
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
This powers CAST's frame gate: what the local cohort shares vs what
|
|
35
|
+
differentiates one member. It never consults corpus reach.
|
|
36
|
+
|
|
37
|
+
The two measures answer different questions over different populations; CAST's
|
|
38
|
+
frame must not be replaced by a reach check and the climb must not be driven by
|
|
39
|
+
weave depth.
|
|
40
|
+
|
|
41
|
+
## Pins
|
|
42
|
+
|
|
43
|
+
- `test/17` — weave-local frame / `MIN_WEAVE` / `dominates` vs corpus-global
|
|
44
|
+
reach.
|
|
45
|
+
- `test/34` — containment and reach-driven disambiguation.
|
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
# Cost Model — One Currency
|
|
2
|
+
|
|
3
|
+
Every mechanism and every byte competes on one cost ladder defined in
|
|
4
|
+
`src/mind/graph-search.ts`. GraphSearch and `pipeline.ts:think` use the same
|
|
5
|
+
units, so a mechanism-level choice and a byte-level choice are the same kind of
|
|
6
|
+
decision: a lightest derivation.
|
|
7
|
+
|
|
8
|
+
## Ladder (`src/mind/graph-search.ts`)
|
|
9
|
+
|
|
10
|
+
| Cost | Value | Meaning |
|
|
11
|
+
| --------- | ------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
12
|
+
| `MICRO` | `1e-3` | Recognised advance (one `rec` bridge); per-byte unit of the A\* heuristic. A recomposed form's onward edge is also `MICRO`. |
|
|
13
|
+
| `STEP` | `1` | Every edge hop (first or fifth), every computed result, every projection. Charging every hop makes the lightest derivation the shortest chain. |
|
|
14
|
+
| `CONCEPT` | `10` | Halo-mediated act (synonym hop, consensus climb) and abandoning an edge chain early (`CONCEPT` above chain cost — genuine fixpoint at `+0` always beats giving up at same depth). |
|
|
15
|
+
| `PASS` | `1000` / byte | Carrying a byte nothing explains. Dominates everything so the search always prefers to recognise. |
|
|
16
|
+
|
|
17
|
+
Only the **ordering** `MICRO < STEP < CONCEPT < PASS` matters; any constants
|
|
18
|
+
with that order give the same derivations.
|
|
19
|
+
|
|
20
|
+
## Pipeline weighing (`src/mind/pipeline.ts:think`)
|
|
21
|
+
|
|
22
|
+
Mechanism candidates are weighed in the same ladder:
|
|
23
|
+
|
|
24
|
+
```
|
|
25
|
+
weight = moves + PASS * unaccounted_bytes
|
|
26
|
+
grade = floor(weight / STEP)
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
`unaccounted` is the query bytes no `accounted` span covers. Comparison is at
|
|
30
|
+
`STEP` resolution: lowest `grade` wins. At equal grade the candidate with fewer
|
|
31
|
+
`scaffolding` bytes (answer bytes lifted from unrecognised spans) wins; only
|
|
32
|
+
then does mechanism list order decide.
|
|
33
|
+
|
|
34
|
+
## Two semirings
|
|
35
|
+
|
|
36
|
+
- **(min, +) tropical** — lightest derivation in `GraphSearch` via `src/derive`
|
|
37
|
+
(`lightestDerivation`). Cost accumulates with `+`, choice selects `min`.
|
|
38
|
+
Powers `cover`/`form`/`out`, edge following, fusing, and the A\* agenda
|
|
39
|
+
(`g + h`).
|
|
40
|
+
|
|
41
|
+
- **(+, +) arithmetic** — evidence pooling in `src/mind/attention.ts:poolVotes`.
|
|
42
|
+
Each region's vote is an axiom; rules carry `Rule.combine = 'sum'` so costs to
|
|
43
|
+
the same anchor **add** rather than minimise. Powers IDF-weighted consensus,
|
|
44
|
+
`votes`/`votesIdf`/`support`, and `regionSupport`/`regionPeak`.
|
|
45
|
+
|
|
46
|
+
## Admissibility
|
|
47
|
+
|
|
48
|
+
The A\* heuristic is admissible and consistent:
|
|
49
|
+
|
|
50
|
+
```
|
|
51
|
+
h(it) = (queryLen - right) * MICRO
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
`right` is `p` for `cover(p)` or `j` for `form/out [i,j)`. `MICRO` is the
|
|
55
|
+
minimum per-byte cost in the ladder (every real per-byte cost is `>= MICRO`,
|
|
56
|
+
including `PASS`), and only the suffix past `right` is counted, so `h` never
|
|
57
|
+
exceeds the true remaining cost.
|
|
58
|
+
|
|
59
|
+
## Policy is not cost
|
|
60
|
+
|
|
61
|
+
"Computation always wins" is **not** priced into the ladder (a computed result
|
|
62
|
+
costs `STEP`, same as a learned edge). It is enforced by masking: `pipeline.ts`
|
|
63
|
+
removes recognised sites overlapped by a `ComputedResult` so the computation is
|
|
64
|
+
the sole completion there. Keep policy in callers; keep the engine neutral.
|
|
65
|
+
|
|
66
|
+
## Pins
|
|
67
|
+
|
|
68
|
+
- `test/52` — climb consensus instrumentation
|
|
69
|
+
- `test/53` — cross-region probe instrumentation
|
|
70
|
+
- `test/54` — evidence `k` instrumentation
|
|
71
|
+
- `test/55` — cost meter (`Meter`, `CostReport`, `searchPops`/`searchPushes`)
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
# Determinism — Same Seed + Same Deposits + Same Query ⇒ Same Bytes
|
|
2
|
+
|
|
3
|
+
## The law
|
|
4
|
+
|
|
5
|
+
> Same `seed` + same deposit order + same query ⇒ byte-identical answer.
|
|
6
|
+
|
|
7
|
+
Determinism is the product. Every code path that can reach output must be
|
|
8
|
+
deterministic given `(seed, store contents, query bytes)`.
|
|
9
|
+
|
|
10
|
+
## Forbidden
|
|
11
|
+
|
|
12
|
+
No `Math.random` or `Date.now` in behaviour, and no iteration over unordered
|
|
13
|
+
collections where order can reach output. Example-only uses
|
|
14
|
+
(`example/train_base`) are outside the library contract. If a test becomes
|
|
15
|
+
flaky, the contract was broken, not the test.
|
|
16
|
+
|
|
17
|
+
## All randomness flows from `seed`
|
|
18
|
+
|
|
19
|
+
`MindConfig.seed` (`src/config.ts:resolveConfig`, `DEFAULT_CONFIG`) is the sole
|
|
20
|
+
entropy root. Subsystems derive deterministically:
|
|
21
|
+
|
|
22
|
+
- **Alphabet** — `Alphabet` (`src/alphabet.ts`) via `rng` (`src/vec.ts:rng`)
|
|
23
|
+
seeded as `seed ^ seedMask`; builds 16→64→256 vectors by refinement.
|
|
24
|
+
- **Keyring / Space** — `Space.seats` (`src/sema.ts:Space`) via `makeKeyring`
|
|
25
|
+
(`src/vec.ts:makeKeyring`) and `rng` seeded from `seed` in `Mind`
|
|
26
|
+
(`src/mind/mind.ts`); `fold`/`twoEndedSeat`/`companySignature` are pure over
|
|
27
|
+
`Space`.
|
|
28
|
+
- **Vector indexes** — `VectorDatabase` (`src/rabitq-ivf/src/database.ts`) and
|
|
29
|
+
`Prng` (`src/rabitq-ivf/src/rabitq.ts`) seeded from config; insertion order is
|
|
30
|
+
the stored order, not a random choice.
|
|
31
|
+
|
|
32
|
+
No other PRNG source may affect grounding. Thresholds in `geometry.ts` are
|
|
33
|
+
derived from `D`/`W`/`N`, not sampled.
|
|
34
|
+
|
|
35
|
+
## Tie-breaks are corpus-determined
|
|
36
|
+
|
|
37
|
+
Every choice among equals bottoms out in a fixed ordering — insertion order or
|
|
38
|
+
lowest node id. The universal no-evidence fallback is **first-inserted**:
|
|
39
|
+
|
|
40
|
+
- `guidedFirst` (`src/mind/traverse.ts:guidedFirst`) — guided pick via
|
|
41
|
+
`chooseNext` else first-inserted edge (`nextFirst` LIMIT 1).
|
|
42
|
+
- `chooseNext` (`src/mind/traverse.ts:chooseNext`) — capped `nextFirst` read
|
|
43
|
+
(`hubBound`), ranked by `prevCount` then `haloMass`; equal ⇒ first-inserted.
|
|
44
|
+
- `chooseAmong` (`src/mind/traverse.ts:chooseAmong`) — `hubCap` + `argmaxCosine`
|
|
45
|
+
over `candidateGist`; first-inserted on tie via stable scan.
|
|
46
|
+
- `companySignature` (`src/sema.ts:companySignature`) — `rng(id ^ 0x9e3779b9)`,
|
|
47
|
+
i.e. seeded by node id, not observation order.
|
|
48
|
+
|
|
49
|
+
Last-inserted was once used in one place; it was a bug. Never reintroduce it.
|
|
50
|
+
|
|
51
|
+
## Memoization and trace must not break identity
|
|
52
|
+
|
|
53
|
+
Per-response memos (`Precomputed`, `perceiveMemo`, `recogniseMemo`, `climbMemo`,
|
|
54
|
+
`_resolvedSubtrees` via `foldTree`, `_edgeChoice`, `_gistCache` in
|
|
55
|
+
`src/mind/mind.ts` / `src/mind/pipeline-mechanism.ts` /
|
|
56
|
+
`src/mind/primitives.ts`) are sound because asking never writes. Only
|
|
57
|
+
`guidedNext`/`sharedReachMemo` are trace-bypassed;
|
|
58
|
+
`perceiveMemo`/`recogniseMemo`/`climbMemo` are always consulted — `foldTree`'s
|
|
59
|
+
subtree fast path skips `visit` (and thus site emission) for cached subtrees, so
|
|
60
|
+
bypassing makes `recognise` non-idempotent.
|
|
61
|
+
|
|
62
|
+
## Follow it
|
|
63
|
+
|
|
64
|
+
When you add any choice among equals, name the tie-break explicitly and make it
|
|
65
|
+
corpus-determined. Thread new randomness through `seed`-derived `rng`; never
|
|
66
|
+
call `Math.random`/`Date.now` on a behavioural path.
|
|
67
|
+
|
|
68
|
+
## Pins
|
|
69
|
+
|
|
70
|
+
- `test/42` pins recognition idempotence under trace — traced and untraced
|
|
71
|
+
`recognise` must return the same cached object and site count.
|
|
72
|
+
- Determinism suites — `test/03`, `test/04`, `test/08`, `test/20` and others
|
|
73
|
+
assert same seed + same training ⇒ byte-identical answers and stores.
|
|
@@ -0,0 +1,47 @@
|
|
|
1
|
+
# Exact vs Approximate — The Law and Its Five Ladders
|
|
2
|
+
|
|
3
|
+
Vector scores (`resonate` / `resonateHalo`) are RaBitQ **estimates**. They rank
|
|
4
|
+
candidates and gate broad regions; they never decide identity. Identity is
|
|
5
|
+
decided only by content-addressed lookup — `resolve` / `findLeaf` / `findBranch`
|
|
6
|
+
/ `canonResolve` — and by re-folding bytes to verify.
|
|
7
|
+
|
|
8
|
+
## The law
|
|
9
|
+
|
|
10
|
+
> Scores propose, bytes dispose.
|
|
11
|
+
|
|
12
|
+
Even recall's echo decision re-folds the top hit's bytes rather than trusting
|
|
13
|
+
the estimate it already has. No `score >= threshold` path may mint an identity
|
|
14
|
+
claim; thresholds derived in `geometry.ts` gate search breadth, not truth.
|
|
15
|
+
|
|
16
|
+
## Graded evidence ladders
|
|
17
|
+
|
|
18
|
+
Five subsystems share one shape — **exact → distributional → geometric** — with
|
|
19
|
+
earlier tiers strictly preferred. Never reorder tiers; never let an approximate
|
|
20
|
+
tier override an exact one.
|
|
21
|
+
|
|
22
|
+
| # | Site | Ladder (strong → weak) | File |
|
|
23
|
+
| - | ------------------ | ------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------- |
|
|
24
|
+
| 1 | `resolve` | exact content-addressed fold → `canonResolve` (equivalence class, hash-then-verify) | `mind/primitives.ts` |
|
|
25
|
+
| 2 | `locate` | exact bytes → halo role → gist | `mind/match.ts` |
|
|
26
|
+
| 3 | `alignGraded` | literal W-gram runs → halo-matched sites + climb proposals (weave) | `mind/match.ts` / `pipeline-mechanism.ts` |
|
|
27
|
+
| 4 | `bridge` | junction containers → edge → synonym → whole-gist | `mind/resonance.ts` |
|
|
28
|
+
| 5 | `crossRegionVotes` | exact containers → single synonym → double → `structuralResonance` (synthetic gist, gated hardest — no byte containment) | `mind/attention.ts` |
|
|
29
|
+
|
|
30
|
+
## Asymmetries (attention)
|
|
31
|
+
|
|
32
|
+
Two rules in `attention.ts` encode "exact decides" and must not be flattened:
|
|
33
|
+
|
|
34
|
+
- Only the **EXACT** tier may explain ordinary votes away.
|
|
35
|
+
- Only **container-backed** evidence may consume its endpoints.
|
|
36
|
+
|
|
37
|
+
## Pins
|
|
38
|
+
|
|
39
|
+
- `test/51` pins the cross-region tier ladder and its gating.
|
|
40
|
+
- Recognition idempotence under trace (`test/42`) depends on exact identity
|
|
41
|
+
remaining byte-determined, not score-determined.
|
|
42
|
+
|
|
43
|
+
## Adding a matcher
|
|
44
|
+
|
|
45
|
+
Add a tier to the shared family in `mind/match.ts` with a derived gate
|
|
46
|
+
(`geometry.ts`), never a private `score >= k` check. A new mechanism is a
|
|
47
|
+
`(matcher, direction, gate)` configuration over that family (§2.5).
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Factored Machinery — One Definition, Many Consumers
|
|
2
|
+
|
|
3
|
+
Every shared operation is defined once and imported many times. Duplicating it
|
|
4
|
+
forks the corpus contract; moving it hides who owns the gate.
|
|
5
|
+
|
|
6
|
+
For the match → project → gate family see `match-project.md`; for the two
|
|
7
|
+
commonality measures see `commonality.md`; for work accounting see `meter.md`.
|
|
8
|
+
|
|
9
|
+
## Single-definition contracts
|
|
10
|
+
|
|
11
|
+
| Symbol | Defined in | One fact |
|
|
12
|
+
| ------------------------------------------------------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
13
|
+
| `contentLevels` | `src/geometry.ts` | Single boundary rule: cuts + levels from one rolling hash pass; every segmentation reads it. |
|
|
14
|
+
| `canonicalWindows` / `chainReach` / `leafIdRun` / `windowIds` | `src/mind/canonical.ts` | Write/read contract: training interns `W-1,W` windows, reading chains to `W²` and probes `W`-windows — drift silences recognition. |
|
|
15
|
+
| `junction.ts` + `WalkCache` | `src/mind/junction.ts` | Shared junction ascent (parents + containers) with bounded `√N·W` walk; `WalkCache` memoizes capped reads/parents/containers per response; bridge and attention share it. |
|
|
16
|
+
| `joinWithBridge` | `src/mind/resonance.ts` | One out-of-search assembly: `bridge(left,right)` or bare concat with `bridgeMiss` trace. |
|
|
17
|
+
| `dismissedKnownContent` | `src/mind/bridge.ts` | Pure attestation: any unaccounted `W`-window that resolves as known content — shared gap guard for substitution and CAST. |
|
|
18
|
+
| `sharedReachMemo` | `src/mind/traverse.ts` | One response-scoped `AncestorReach` memo (cleared on write and for traces); every `reachOf`/`edgeAncestors` consumer shares it. |
|
|
19
|
+
| `guidedFirst` | `src/mind/traverse.ts` | Guided-or-first answer bytes: `guidedNext` else first-inserted edge (`LIMIT 1`). |
|
|
20
|
+
| `leadsSomewhere` | `src/mind/traverse.ts` | Admission predicate: `hasNext` (cached) or `hasHalo`; sites that lead nowhere contribute no derivation. |
|
|
21
|
+
| `isChunk` | `src/sema.ts` | `kids !== null && kids.every(k=>k.kids===null)` — smallest grouped unit; governs regions, seams, indexing. |
|
|
22
|
+
| `twoEndedSeat` | `src/sema.ts` | One seat algebra: first half low seats, second half high seats; shared by perception, `fold`, and canonical folds. |
|
|
23
|
+
|
|
24
|
+
## Pins
|
|
25
|
+
|
|
26
|
+
- `test/47` — frame reading split (matcher vs gate vs inventory).
|
|
27
|
+
- `test/50` — CAST analog / consensus floor (dismissed content, `MIN_WEAVE` /
|
|
28
|
+
`dominates` frame, `carriesFillers` refusal).
|
|
@@ -0,0 +1,87 @@
|
|
|
1
|
+
# Fold Contract — One Tree For The Same Bytes
|
|
2
|
+
|
|
3
|
+
> **Law:** `perceiveDeposit` ≡ `perceive` — same bytes ⇒ same tree and same node
|
|
4
|
+
> id. Deposit imposes nothing (no boundaries, no turn convention). Geometry
|
|
5
|
+
> never sees conversation metadata.
|
|
6
|
+
|
|
7
|
+
## The identity
|
|
8
|
+
|
|
9
|
+
Perception is a pure function of the bytes. The deposit path and the inference
|
|
10
|
+
path compute the same content-defined fold for the same input, so a trained
|
|
11
|
+
context node and `resolve(query)` reach the same node. When the two sides
|
|
12
|
+
disagreed, alignment went quadratic (measured 5.2M cells on a 476-byte context
|
|
13
|
+
vs 0 when they agree) and cumulative contexts stopped resolving to what they
|
|
14
|
+
were trained as.
|
|
15
|
+
|
|
16
|
+
## Deposit imposes nothing
|
|
17
|
+
|
|
18
|
+
No boundaries, no turn convention, nothing read out of the bytes. Conversational
|
|
19
|
+
turn offsets are API metadata — they feed `ConversationState`, `answeredSpans`
|
|
20
|
+
and `currentTurnStart`; the geometry never sees them. Passing turn boundaries
|
|
21
|
+
into the fold is a correctness bug, not a tuning choice.
|
|
22
|
+
|
|
23
|
+
## Boundaries vs reuse — two problems
|
|
24
|
+
|
|
25
|
+
`contentFoldIncremental` and `stablePrefixFold` solve different problems;
|
|
26
|
+
conflating them is what once put an imposed boundary set on the inference path.
|
|
27
|
+
|
|
28
|
+
| Mechanism | What it buys | Cost / shape |
|
|
29
|
+
| ------------------------ | --------------------------------------------------------- | ----------------------------------------------------------------------------------------------- |
|
|
30
|
+
| `contentFoldIncremental` | Transparent segment reuse (cost only) | Imposes nothing; tree identical to the cold fold |
|
|
31
|
+
| `stablePrefixFold` | Caller-supplied cuts left-nested for prefix-ROOT identity | One prefix-ROOT per cut becomes an identical subtree (and same node id) inside the grown stream |
|
|
32
|
+
|
|
33
|
+
Both carry the same precondition: `prev` must be a fold of a byte-identical
|
|
34
|
+
prefix — reuse is keyed on `[start,end)` offsets, which cannot witness byte
|
|
35
|
+
agreement. A mismatched `prev` produced a wrong tree on 336 of 400 random
|
|
36
|
+
streams. `perceiveDeposit` discharges this via the prefix bytes as cache key; a
|
|
37
|
+
conversation advances only by append. A caller that cannot prove the prefix must
|
|
38
|
+
pass no `prev`.
|
|
39
|
+
|
|
40
|
+
## Identity must not depend on W or absolute offset
|
|
41
|
+
|
|
42
|
+
`contentLevels` is the single boundary rule (rolling hash over a bounded
|
|
43
|
+
window). Any grouping by index — stride, tile, fixed-arity row — reintroduces
|
|
44
|
+
the grid's phase bug: `riverFold` groups `W`-ary from byte 0, so the same byte
|
|
45
|
+
run is a different subtree at a different offset. The content-defined hash
|
|
46
|
+
removes this: a change upstream moves only the cut it falls inside; downstream
|
|
47
|
+
cuts and segments are unchanged (99.7% cuts preserved on real deposits after
|
|
48
|
+
shifts of 1..7 bytes vs 14.3% for the grid). The `groupByLevel` above the
|
|
49
|
+
segments splits by content level, not by count.
|
|
50
|
+
|
|
51
|
+
## `contentLevels` is single source; `contentBoundaries` is projection
|
|
52
|
+
|
|
53
|
+
`contentBoundaries(space, bytes)` is `contentLevels(space, bytes).cuts`. It once
|
|
54
|
+
carried its own rolling-hash loop, which is how a write side and a read side
|
|
55
|
+
drift without a type error. Levels are read from the hash the cut was accepted
|
|
56
|
+
at — level `L` when `h` vanishes mod `W^(L+1)` — so level-`L` cuts nest inside
|
|
57
|
+
level-`(L-1)` and expected span is `W^(L+1)` bytes.
|
|
58
|
+
|
|
59
|
+
## Optional canonical capability
|
|
60
|
+
|
|
61
|
+
`canonAdd`/`canonFind` (`src/store.ts` — `canonCount`/`eachContent`) is an
|
|
62
|
+
optional backend capability. A backend may omit all four; resolution then has no
|
|
63
|
+
equivalence fallback. The store never learns the equivalence — the canonicalizer
|
|
64
|
+
(`Canon` in `src/canon.ts`, e.g. `textCanon`) is injected by the caller and
|
|
65
|
+
every candidate is hash-then-verified (re-canonicalize stored bytes, compare). A
|
|
66
|
+
hash collision costs a read, never a wrong id.
|
|
67
|
+
|
|
68
|
+
## Cost of changing the cut distribution
|
|
69
|
+
|
|
70
|
+
The cut rate, which bits are read, `minLen`/`maxLen`, and the forced cut at
|
|
71
|
+
`maxLen` set the segment distribution every downstream mechanism is fitted to.
|
|
72
|
+
Each has been changed experimentally and cost 5–21 tests (rate: 15–18, bits:
|
|
73
|
+
19–21, normalized chunking: 5–6). `W-1` is the minimum (one window minus one);
|
|
74
|
+
`seats.length` is the maximum (one flat node folds exactly one segment). The
|
|
75
|
+
forced cut is load-bearing — relaxing it to reduce the current 32% forced rate
|
|
76
|
+
looked like a tidying but broke the same suites. Re-measure the whole suite for
|
|
77
|
+
any change here.
|
|
78
|
+
|
|
79
|
+
## Pins
|
|
80
|
+
|
|
81
|
+
- `test/59` — shift invariance floors (content-defined cuts preserved over
|
|
82
|
+
random binary and prose).
|
|
83
|
+
- `test/63` — offset/W invariance and `contentLevels` distribution expectations.
|
|
84
|
+
|
|
85
|
+
See:
|
|
86
|
+
`src/geometry.ts:contentLevels`/`contentBoundaries`/`contentFoldIncremental`/`stablePrefixFold`;
|
|
87
|
+
`AGENTS.md` bootloader invariants.
|
|
@@ -0,0 +1,99 @@
|
|
|
1
|
+
# Halo & Sketch — Distributional Memory
|
|
2
|
+
|
|
3
|
+
A node's **halo** is its distributional signature: the superposition of
|
|
4
|
+
identity-bound company signatures poured from every episode it participated in.
|
|
5
|
+
Its **gist** is the VSA fold of its own bytes — content, not company.
|
|
6
|
+
|
|
7
|
+
## Two vectors per node, two indexes
|
|
8
|
+
|
|
9
|
+
| Vector | Encodes | Index | Query |
|
|
10
|
+
| ------ | ----------------------------- | ------------- | -------------- |
|
|
11
|
+
| gist | what the node _is made of_ | content index | `resonate` |
|
|
12
|
+
| halo | what _company_ the node keeps | halo index | `resonateHalo` |
|
|
13
|
+
|
|
14
|
+
Both indexes are RaBitQ-IVF (`src/rabitq-ivf/`) — 1-bit ANN over the same
|
|
15
|
+
vectors; scores are estimates, never identity. Halos are also persisted durably
|
|
16
|
+
(see below).
|
|
17
|
+
|
|
18
|
+
## Quantization — 2-bit on disk, float in session
|
|
19
|
+
|
|
20
|
+
A halo is a superposition of quasi-orthogonal signatures, so coordinates are
|
|
21
|
+
Gaussian. Storage exploits this:
|
|
22
|
+
|
|
23
|
+
- **In-session accumulator** (`_haloExact` in `src/store.ts`): `Float32Array` —
|
|
24
|
+
exact, additive, incremented by `pourHalo`.
|
|
25
|
+
- **Durable row** (`_dbUpsertHalo`): 2-bit Lloyd–Max quantizer — decision at
|
|
26
|
+
±0.9816σ, levels at ±0.4528σ / ±1.5104σ, σ derived from the stored norm
|
|
27
|
+
(`norm/√D`). Header is the 4-byte norm; body is 2 bits/coordinate. Keeps ≥0.88
|
|
28
|
+
correlation with the exact vector.
|
|
29
|
+
- **ANN index**: 1-bit RaBitQ, irreversible — answers only "which halos are near
|
|
30
|
+
this query?"
|
|
31
|
+
|
|
32
|
+
Re-indexing is geometric: a halo re-enters the ANN when mass is small or crosses
|
|
33
|
+
a power of two (`geometricMass`), so index writes are O(log mass).
|
|
34
|
+
|
|
35
|
+
## Bottom-k sketch & company profile
|
|
36
|
+
|
|
37
|
+
A whole-partner signature alone records tokens, not types — halos of genuine
|
|
38
|
+
synonyms would be quasi-orthogonal. `companyProfile` (`src/mind/learning.ts`)
|
|
39
|
+
superposes:
|
|
40
|
+
|
|
41
|
+
1. the partner's own identity signature, plus
|
|
42
|
+
2. the bottom-k **constituent sketch** — the
|
|
43
|
+
`k = profileCapacity(D) = floor(√D)` minimal units of its subtree with
|
|
44
|
+
smallest `unitPriority`, deduped.
|
|
45
|
+
|
|
46
|
+
The sketch is composable (bottom-k of a union = bottom-k of children's
|
|
47
|
+
sketches), durable derived state via `sketchGet`/`sketchPut`, and bounded: at
|
|
48
|
+
most `k` constituents are classified, each by one `LIMIT`ed parent read
|
|
49
|
+
(`hubBound`). Beyond `√D` terms a single constituent contributes less than
|
|
50
|
+
`1/√D` — below RaBitQ noise — and extra terms shrink every accepted one; the cap
|
|
51
|
+
is a correctness limit.
|
|
52
|
+
|
|
53
|
+
## Gist vectors
|
|
54
|
+
|
|
55
|
+
Folded by the river (`src/geometry.ts`): leaves are alphabet vectors, groups
|
|
56
|
+
bind by two-ended seats, intermediate gists stay unnormalized (magnitude ∝
|
|
57
|
+
√len), only the root is normalized. Gist resonance reads byte-proportional
|
|
58
|
+
overlap; halo resonance reads distributional overlap — the two are independent.
|
|
59
|
+
|
|
60
|
+
## Thresholds & gating
|
|
61
|
+
|
|
62
|
+
All bars are derived in `src/geometry.ts`; no tunable constant:
|
|
63
|
+
|
|
64
|
+
| Symbol | Formula | Use |
|
|
65
|
+
| ------------------ | -------------- | ------------------------------------------------------------------------------------------------- |
|
|
66
|
+
| `estimatorNoise` | `1/√D` | 1σ RaBitQ noise; contrastive margin must clear it |
|
|
67
|
+
| `significanceBar` | `3/√D` | whole-query relatedness — 3σ above chance; gates consensus climb and `analogyStrength` |
|
|
68
|
+
| `conceptThreshold` | `0.5 + 0.5/√D` | halo concept sharing — structural midpoint + ½σ; gates `haloSiblings`, concept hops, articulation |
|
|
69
|
+
|
|
70
|
+
The significance bar gates the whole query; `conceptThreshold` gates per-pair
|
|
71
|
+
halo cosine.
|
|
72
|
+
|
|
73
|
+
## Probes: `haloMass` and `hasHalo`
|
|
74
|
+
|
|
75
|
+
- `haloMass(id)` — count of poured episodes; evidence weight, tie-breaker in
|
|
76
|
+
`chooseAmong`.
|
|
77
|
+
- `hasHalo(id)` — existence probe (indexed point check, no vector decode);
|
|
78
|
+
mirrors `halo(id) !== null`. One tier of the `leadsSomewhere` admission
|
|
79
|
+
predicate (with `hasNext`/`hasParents`).
|
|
80
|
+
|
|
81
|
+
Both are `meter`-counted probes, not full decodes — `halo(id)` is the bounded
|
|
82
|
+
vector read; `resonateHalo` is the IVF ANN query.
|
|
83
|
+
|
|
84
|
+
## Relation to invariants
|
|
85
|
+
|
|
86
|
+
- **Derived thresholds** — all bars above live in `geometry.ts`.
|
|
87
|
+
- **Exact decides / approximate proposes** — halo scores rank and gate; identity
|
|
88
|
+
is content-addressed. The graded ladder is exact → halo → gist
|
|
89
|
+
(`mind/match.ts`).
|
|
90
|
+
- **Bounded reads** — `hasHalo`/`haloMass` are point probes; `resonateHalo` is
|
|
91
|
+
capped ANN; constituent classification uses `LIMIT hubBound+1` reads. No
|
|
92
|
+
per-query scan grows with corpus.
|
|
93
|
+
|
|
94
|
+
## Pins
|
|
95
|
+
|
|
96
|
+
- `test/08 storage halo` — halo persistence, 2-bit round-trip,
|
|
97
|
+
`haloMass`/`hasHalo` contract, index survival across reopen.
|
|
98
|
+
- `test/35 ivf` — RaBitQ-IVF contract (recall vs brute force, sublinear
|
|
99
|
+
scaling); covers the halo index's own layer.
|