@hviana/sema 0.9.3 → 0.9.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +16 -7
- package/dist/src/mind/learning.js +11 -12
- package/dist/src/mind/mind.js +4 -4
- package/dist/src/mind/types.d.ts +10 -15
- package/dist/src/store.d.ts +10 -10
- package/dist/src/store.js +5 -5
- package/docs/INDEX.md +62 -59
- package/docs/INVARIANTS.md +36 -18
- package/docs/PHILOSOPHY.md +317 -0
- package/docs/architecture/bounded-reads.md +31 -71
- package/docs/architecture/caches.md +61 -81
- package/docs/architecture/closure.md +88 -107
- package/docs/architecture/commonality.md +47 -38
- package/docs/architecture/cost-model.md +57 -79
- package/docs/architecture/determinism.md +43 -55
- package/docs/architecture/evidence.md +158 -235
- package/docs/architecture/exact-vs-approximate.md +41 -35
- package/docs/architecture/factored-machinery.md +34 -20
- package/docs/architecture/fold-contract.md +110 -118
- package/docs/architecture/halo-sketch.md +105 -96
- package/docs/architecture/match-project.md +51 -42
- package/docs/architecture/mechanism-market.md +87 -91
- package/docs/architecture/memoization.md +60 -74
- package/docs/architecture/meter.md +37 -47
- package/docs/architecture/saturation.md +75 -101
- package/docs/architecture/store.md +118 -99
- package/docs/architecture/thresholds.md +66 -73
- package/docs/failures/tempting-but-wrong.md +139 -165
- package/docs/harness/gates.md +27 -32
- package/docs/mechanisms/alu.md +22 -69
- package/docs/mechanisms/cast.md +76 -71
- package/docs/mechanisms/confluence.md +22 -29
- package/docs/mechanisms/cover.md +58 -66
- package/docs/mechanisms/extraction.md +33 -37
- package/docs/mechanisms/prefix-completion.md +36 -39
- package/docs/mechanisms/recall.md +60 -53
- package/docs/mechanisms/reference.md +63 -49
- package/jsr.json +1 -1
- package/package.json +1 -1
- package/src/alu/README.md +90 -298
- package/src/derive/README.md +94 -256
- package/src/mind/learning.ts +11 -12
- package/src/mind/mind.ts +4 -4
- package/src/mind/types.ts +10 -15
- package/src/rabitq-ivf/README.md +11 -8
- package/src/store.ts +5 -5
|
@@ -1,104 +1,78 @@
|
|
|
1
|
-
# Saturation — Named Stop, Not Cap
|
|
2
|
-
|
|
3
|
-
> **Law:** every walk names
|
|
4
|
-
>
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
`mind/traverse.ts`
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
`
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
`
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
candidate can never outscore it. At most one winner's bytes are reconstructed;
|
|
73
|
-
every shorter proposal is skipped without a read. Saturation, not cap.
|
|
74
|
-
|
|
75
|
-
## Junction — hub guards vs. budget
|
|
76
|
-
|
|
77
|
-
`mind/junction.ts:junctionContainersFrom` has three disciplines:
|
|
78
|
-
|
|
79
|
-
- **Phrase-scale reads** — `bytesPrefix(maxContainer+1)` per visit; a node
|
|
80
|
-
beyond the cap prunes its branch.
|
|
81
|
-
- **Per-node hub guards** (real saturations) — `parentsFirst(bound+1) > bound`
|
|
82
|
-
not expanded; one `containersSlice(bound+1)` page beyond `bound` not expanded.
|
|
83
|
-
Each is exact below `bound`.
|
|
84
|
-
- **Expansion budget** — at most `bound * W` pops total (shared across a tier's
|
|
85
|
-
walks). Budget exhaustion is an abstention (`junctionBudgetExhausted`) that
|
|
86
|
-
falls through to the resonance tier — a net, not a saturation.
|
|
87
|
-
|
|
88
|
-
Refuted tightening: applying `edgeAncestors`' cumulative lateral-cone limit here
|
|
89
|
-
would discard half the successful junctions (measured lateral 1425/1426 at
|
|
90
|
-
`bound` ~ 570).
|
|
91
|
-
|
|
92
|
-
Refuted early-stop: **one-cone-exhausted** — stopping when one side's upward
|
|
93
|
-
cone empties — is wrong in both hub-guarded and hub-flagged forms. A junction
|
|
94
|
-
can be reachable from only one side when that side's seed is a fold sub-node of
|
|
95
|
-
the container (`test/16`: "cold or hot" reached from window "cold" while 3-byte
|
|
96
|
-
"hot" cone is empty; `test/34` n-ary binding fails the same way). Exhausting one
|
|
97
|
-
cone never proves no junction remains; the walk must keep the `bound*W` net
|
|
98
|
-
after per-node saturations.
|
|
1
|
+
# Saturation — A Named Stop, Not a Cap
|
|
2
|
+
|
|
3
|
+
> **Law:** every walk names the saturation that decides when it stops, beside
|
|
4
|
+
> its cap. The cap (`bounded-reads.md`) is a safety net. The saturation is a
|
|
5
|
+
> derived proof that continuing cannot discriminate, and it is recorded in the
|
|
6
|
+
> trace.
|
|
7
|
+
|
|
8
|
+
**Why.** A walk with only a cap drifts to the cap: it spends the whole budget
|
|
9
|
+
and stops at an arbitrary place. A walk with a saturation stops the moment the
|
|
10
|
+
outcome is known, is exact below every stop, and says why it stopped. Deflate's
|
|
11
|
+
match finder truncates long chains arbitrarily. Sema may cap, but it must still
|
|
12
|
+
justify every stop.
|
|
13
|
+
|
|
14
|
+
The trace records the first stop that decides as
|
|
15
|
+
`SaturationStop { reason, node, observed, limit }`, together with `visited` and
|
|
16
|
+
`maxDepth`.
|
|
17
|
+
|
|
18
|
+
## `edgeAncestors` — expand until decided (`mind/traverse.ts`)
|
|
19
|
+
|
|
20
|
+
This is the model walk. It climbs parents and containers from a node to the
|
|
21
|
+
learnt contexts it reaches, and stops at the first of five saturations. Each is
|
|
22
|
+
exact below its threshold (`bound = √N`):
|
|
23
|
+
|
|
24
|
+
1. **Predecessor fan-in.** `prevCount(node) > bound`: one indexed count proves
|
|
25
|
+
more than `bound` contexts.
|
|
26
|
+
2. **Distinct contexts.** `ctxSeen.size > bound` after `prevFirst(node, bound)`.
|
|
27
|
+
3. **Parent fan-out.** `parentsFirst(node, bound + 1)` returns more than
|
|
28
|
+
`bound`, and the node is not expanded.
|
|
29
|
+
4. **Lateral cone.** The extra parents summed over every expanded node exceed
|
|
30
|
+
`bound`. Commonness spread across the cone is caught as surely as commonness
|
|
31
|
+
concentrated at one node, while a deep chain inside one structure accrues no
|
|
32
|
+
laterals and reaches its root at any depth.
|
|
33
|
+
5. **Byte atoms.** Atoms carry no containment rows, so their reach is
|
|
34
|
+
unmeasurable, and the uniform expectation `atomReach(N, W)` replaces it. Past
|
|
35
|
+
`bound`, an atom abstains as a voter (`atomIsHub`), though its edges remain
|
|
36
|
+
traversable.
|
|
37
|
+
|
|
38
|
+
Below every stop, the walk equals the unbounded climb, and its work is
|
|
39
|
+
`O(bound)` contexts. Container seeds are streamed in `bound`-sized pages.
|
|
40
|
+
|
|
41
|
+
**Refuted:** treating container fan-out as a hub. It confuses the _places_ a
|
|
42
|
+
window occurs with the _contexts_ it reaches, and saturation is defined over
|
|
43
|
+
contexts. On `test/49`, `fran` sits in 4 containers yet reaches 2 contexts, and
|
|
44
|
+
the rule called it saturated.
|
|
45
|
+
|
|
46
|
+
## `pivotInto` — the longest wins (`mind/resonance.ts`)
|
|
47
|
+
|
|
48
|
+
Candidates rank by `contentLen(id, answerLen + 1)`, descending, with
|
|
49
|
+
first-inserted breaking ties. The score is length itself, so the first candidate
|
|
50
|
+
that passes every filter wins, and no shorter candidate's bytes are ever read.
|
|
51
|
+
|
|
52
|
+
## Junction ascent — guards against a budget (`mind/junction.ts`)
|
|
53
|
+
|
|
54
|
+
`junctionContainersFrom` keeps three disciplines apart:
|
|
55
|
+
|
|
56
|
+
- **Phrase-scale reads.** `bytesPrefix(maxContainer + 1)` per visit; a node past
|
|
57
|
+
the cap prunes its branch.
|
|
58
|
+
- **Hub guards, which are true saturations.** A node is not expanded when
|
|
59
|
+
`parentsFirst(bound + 1)` or one `containersSlice` page exceeds `bound`.
|
|
60
|
+
- **The expansion budget, a net.** At most `bound · W` pops, shared across a
|
|
61
|
+
tier's walks. Running out is an abstention (`junctionBudgetExhausted`), and
|
|
62
|
+
the search falls through to the resonance tier.
|
|
63
|
+
|
|
64
|
+
Two tightenings were refuted:
|
|
65
|
+
|
|
66
|
+
- **The lateral-cone limit from `edgeAncestors`** would discard half the
|
|
67
|
+
successful junctions (lateral 1425 of 1426 at `bound` ≈ 570).
|
|
68
|
+
- **Stopping when one side's cone is exhausted** is wrong. A junction can be
|
|
69
|
+
reachable from only one side when that side's seed is a fold sub-node of the
|
|
70
|
+
container. In `test/16`, `cold or hot` is reached from `cold` while the cone
|
|
71
|
+
of `hot` is empty. Exhausting one cone never proves that no junction remains.
|
|
99
72
|
|
|
100
73
|
## Pins
|
|
101
74
|
|
|
102
|
-
- `test/16` — one
|
|
103
|
-
|
|
104
|
-
- `test/27` —
|
|
75
|
+
- `test/16`, `test/34` — refute stopping when one cone is exhausted, in the
|
|
76
|
+
bridge junction and in n-ary cross-region binding.
|
|
77
|
+
- `test/27` — dropping saturated leading and trailing intervals from the climb.
|
|
78
|
+
- `test/49` — refutes treating container fan-out as a hub.
|
|
@@ -1,102 +1,121 @@
|
|
|
1
|
-
# Store —
|
|
2
|
-
|
|
3
|
-
> **Law:**
|
|
4
|
-
>
|
|
5
|
-
>
|
|
6
|
-
> wrappers.
|
|
7
|
-
|
|
8
|
-
##
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
`
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
`
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
`
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
`
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
- `
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
`
|
|
1
|
+
# Store — A Content-Addressed DAG, One Owner of Its Logic
|
|
2
|
+
|
|
3
|
+
> **Law:** a node is named by its content, and equal content is stored once.
|
|
4
|
+
> `AbstractStore` (`src/store.ts`) owns every domain decision: dedup, merge,
|
|
5
|
+
> indexing, containment, batching, caching and maintenance. A backend
|
|
6
|
+
> (`store-sqlite.ts`) implements only thin `_db*`/`_vec*` wrappers.
|
|
7
|
+
|
|
8
|
+
## Nodes
|
|
9
|
+
|
|
10
|
+
| Kind | Id | Stored as |
|
|
11
|
+
| ----------- | ----------------------------------------------- | -------------------------------------------------------- |
|
|
12
|
+
| byte leaf | implicit, `-(byte+1)` in `−256…−1` | nothing — it always exists |
|
|
13
|
+
| branch | dense `0,1,2,…`, minted in order, never deleted | its ordered child ids |
|
|
14
|
+
| flat branch | as a branch | raw bytes, with an empty `kids` marker (`flatKidsBytes`) |
|
|
15
|
+
|
|
16
|
+
An id is an arbitrary mark; content decides which mark a thing gets. A subtree
|
|
17
|
+
shared by a thousand deposits is one node with a thousand parents, so storage
|
|
18
|
+
grows with distinct content, not with volume, and every span the fold produced
|
|
19
|
+
can be addressed.
|
|
20
|
+
|
|
21
|
+
## Interning — dedup, then same bytes, then near, then mint
|
|
22
|
+
|
|
23
|
+
`intern` runs four steps, in order:
|
|
24
|
+
|
|
25
|
+
1. **Exact dedup.** `findLeaf` or `findBranch` looks for equal content: a cache
|
|
26
|
+
first, then a durable probe, so dedup survives a cold cache or a resumed run.
|
|
27
|
+
2. **Same bytes, same node.** A branch whose children name nothing reuses the
|
|
28
|
+
flat node over the same bytes. Its gist is re-captured, because the same
|
|
29
|
+
bytes fold differently standing alone and embedded, and the index must hold
|
|
30
|
+
the gist that a direct query will present (`fold-contract.md`).
|
|
31
|
+
3. **Near dedup.** This applies to branches only, and only against whole
|
|
32
|
+
experiences still in the write buffer. The gist proposes a candidate, which
|
|
33
|
+
must reach the branch's own `identityBar`. The bytes then decide: the two
|
|
34
|
+
must be identical except for one span of at most `W` bytes
|
|
35
|
+
(`differsByOneWindow`). This is the only place two different contents share
|
|
36
|
+
an id.
|
|
37
|
+
4. **Mint** a new id.
|
|
38
|
+
|
|
39
|
+
A refuted alternative: probing the flushed ANN index for near targets. It was
|
|
40
|
+
the dominant training cost, and it was wrong. The 1-bit code ranked a
|
|
41
|
+
byte-distinct branch as nearest and collapsed two different subtrees onto one id
|
|
42
|
+
(`test/02`).
|
|
43
|
+
|
|
44
|
+
## Relations on the nodes
|
|
45
|
+
|
|
46
|
+
- **Continuation edges.** An edge is a unique `(src, dst)` pair, so
|
|
47
|
+
`prevCount(dst)` counts distinct establishing contexts, never repetitions.
|
|
48
|
+
- **Containment.** `addContainer(child, parent)` buffers per child, and the
|
|
49
|
+
flush appends packed pages (`_dbAppendContain`) without rewriting the list.
|
|
50
|
+
`containersSlice` pages through it.
|
|
51
|
+
- **Gist index** (content) and **halo index** (company). Both are RaBitQ-IVF ANN
|
|
52
|
+
over node ids. Gists wait in `_pendingGist` until `indexSubtree` promotes them
|
|
53
|
+
in batches. Halos accumulate exactly in session and persist 2-bit quantized
|
|
54
|
+
(`halo-sketch.md`).
|
|
55
|
+
- **Sketches.** `sketchGet`/`sketchPut` hold durable derived state: a miss costs
|
|
56
|
+
time, never correctness.
|
|
57
|
+
|
|
58
|
+
`leadsSomewhere(id) = hasNext || hasHalo` is the one raw admission predicate: a
|
|
59
|
+
form is knowledge only if it continues somewhere or keeps company.
|
|
60
|
+
|
|
61
|
+
## Probes that do not grow with N
|
|
62
|
+
|
|
63
|
+
Hot paths read through `LIMIT` variants, point probes and capped byte reads, and
|
|
64
|
+
full scans are for maintenance only (`bounded-reads.md`). Flat lookups are
|
|
65
|
+
hashed, then verified:
|
|
66
|
+
|
|
67
|
+
- `flatBranchMayExist` is a negative filter, so its `false` is exact.
|
|
68
|
+
- `findFlatBranch` memoizes hits only, and a miss builds no key.
|
|
69
|
+
- `flatSpans(bytes)` extends one FNV-1a hash per start position, so sweeping a
|
|
70
|
+
span's ends costs O(1) per probe instead of O(span). Recognition's interior
|
|
71
|
+
pass once hashed 251,660,406 bytes for a single response (`test/153.4`).
|
|
72
|
+
|
|
73
|
+
`bytes(id)` and `bytesPrefix(id, cap)` return buffers shared with the caches.
|
|
74
|
+
Callers must never mutate them.
|
|
75
|
+
|
|
76
|
+
## The canonical index — optional, injected, verified
|
|
77
|
+
|
|
78
|
+
`canonAdd`, `canonFind`, `canonCount` and `eachContent` are an optional backend
|
|
79
|
+
capability. Without it, `resolve` has no equivalence fallback.
|
|
80
|
+
|
|
81
|
+
- **The store never learns the equivalence.** The caller injects the
|
|
82
|
+
canonicalizer (`Canon`, for example `textCanon`).
|
|
83
|
+
- **Every candidate is hash-then-verified:** the stored bytes are
|
|
84
|
+
re-canonicalized and compared, so a collision costs a read, never a wrong id.
|
|
85
|
+
- **`buildCanonIndex` runs after training** and is incremental from a watermark.
|
|
86
|
+
A fixture store that skips it under-reports every case variant.
|
|
87
|
+
- **SQLite keeps a negative filter over the canonical hashes** (`canon_bloom`).
|
|
88
|
+
It is persisted in the same transaction as the rows it covers and stamped with
|
|
89
|
+
`canon.upto`. An open whose stamp disagrees rebuilds the filter once
|
|
90
|
+
(`test/36-bloom`).
|
|
91
|
+
|
|
92
|
+
## Caches and maintenance
|
|
93
|
+
|
|
94
|
+
Every in-memory cache is a byte-budgeted `BoundedMap` (`caches.md`). ANN read
|
|
95
|
+
caches are keyed by content (`vecKey`), dropped on any index mutation, and
|
|
96
|
+
cleared at `RESONATE_CACHE_MAX`.
|
|
97
|
+
|
|
98
|
+
Maintenance is incremental:
|
|
99
|
+
|
|
100
|
+
- `compactContentIndex` removes isolated index entries written since its last
|
|
101
|
+
watermark.
|
|
102
|
+
- `repairContentIndex` re-inserts bridge nodes whose gists were evicted before
|
|
103
|
+
they were indexed.
|
|
104
|
+
- `buildCanonIndex` extends the canonical index from a given id.
|
|
97
105
|
|
|
98
106
|
## Adding a backend
|
|
99
107
|
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
108
|
+
1. Subclass `AbstractStore` and implement every
|
|
109
|
+
`protected abstract _db*`/`_vec*` method as a thin wrapper. Never
|
|
110
|
+
re-implement dedup or batching.
|
|
111
|
+
2. Keep the `*First`/`*Slice` methods as real `LIMIT ?` queries, and keep `has*`
|
|
112
|
+
and counts as point probes. Never materialise a list and then slice it.
|
|
113
|
+
3. Run the whole suite with your store substituted.
|
|
114
|
+
|
|
115
|
+
## Pins
|
|
116
|
+
|
|
117
|
+
- `test/08` — storage: the node layout, halo persistence and 2-bit round trip,
|
|
118
|
+
the `haloMass`/`hasHalo` contract, and survival across a reopen.
|
|
119
|
+
- `test/02` — exact round trip: near dedup never merges distinct bytes.
|
|
120
|
+
- `test/36-bloom` — the filters' exactness and persistence.
|
|
121
|
+
- `test/153.4` — the span prober answers exactly what `findFlatBranch` answers.
|
|
@@ -1,79 +1,72 @@
|
|
|
1
|
-
# Thresholds —
|
|
2
|
-
|
|
3
|
-
|
|
4
|
-
perception window
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
`twoEndedSeat(seatCount, size, index)` — the one positional-coordinate algebra
|
|
51
|
-
shared by perception, `fold`, and every synthetic/canonical fold. First half
|
|
52
|
-
uses low seats, second half uses high seats:
|
|
53
|
-
`index < (size+1)/2 ? index : seatCount-size+index`.
|
|
1
|
+
# Thresholds — Derived, Never Tuned
|
|
2
|
+
|
|
3
|
+
> **Law:** every decision cutoff is a formula over the vector dimension `D`, the
|
|
4
|
+
> perception window `W` (`maxGroup`) or the corpus size `N`. `config.ts` holds
|
|
5
|
+
> capacities and budgets only: cache bytes, batch sizes, index parameters, query
|
|
6
|
+
> `k`, ALU precision and the seed. It never holds a threshold.
|
|
7
|
+
|
|
8
|
+
**Why.** A kept memory grows, and a fixed number silently changes meaning as
|
|
9
|
+
`N`, `D` or a span's length changes. A derived bar is measured against the
|
|
10
|
+
medium's own chance and scale, so there is no development set to overfit and no
|
|
11
|
+
calibration that expires with new data. With nothing fitted, a wrong answer
|
|
12
|
+
indicts a law, never a parameter.
|
|
13
|
+
|
|
14
|
+
## Three sources of scale
|
|
15
|
+
|
|
16
|
+
**`D`: chance.** Random unit vectors have cosine about `N(0, 1/D)`, so one noise
|
|
17
|
+
unit is `1/√D`.
|
|
18
|
+
|
|
19
|
+
- `estimatorNoise = 1/√D`
|
|
20
|
+
- `significanceBar = 3/√D`, three noise units above chance
|
|
21
|
+
- `conceptThreshold = 0.5 + 0.5/√D`, for halo concept sharing: the structural
|
|
22
|
+
midpoint plus half a noise unit
|
|
23
|
+
- `mergeThreshold = 1 − 1/√D`
|
|
24
|
+
- `profileCapacity = ⌊√D⌋`, the terms a superposition holds before each falls
|
|
25
|
+
below noise
|
|
26
|
+
|
|
27
|
+
**`W`: the perception quantum.** Below one window, a byte overlap is chance.
|
|
28
|
+
|
|
29
|
+
- `identityBar(D, W, len) = max(mergeThreshold, 1 − W/len)`
|
|
30
|
+
- `reachThreshold = 1 − 1/(2W)`
|
|
31
|
+
- canonical windows of `W−1` and `W` bytes, and `chainReach = W²`
|
|
32
|
+
|
|
33
|
+
**`N`: commonness,** with `corpusN = max(2, edge sources)`.
|
|
34
|
+
|
|
35
|
+
- `hubBound = ⌈√N⌉`, the bound on every `LIMIT` read (`hubCap` on lists)
|
|
36
|
+
- `consensusFloor = ln N + ½`, for pooled climb votes
|
|
37
|
+
- `atomReach = max(1, ⌈N·W/256⌉)`; an atom is a hub when `atomReach > hubBound`
|
|
38
|
+
|
|
39
|
+
`dominates(part, whole) = 2·part > whole` is the one scale-free bar: a strict
|
|
40
|
+
majority.
|
|
41
|
+
|
|
42
|
+
The vector bars are compared against RaBitQ estimates, not exact cosines. That
|
|
43
|
+
is benign for inequality gates over broad regions, and it is why no bar ever
|
|
44
|
+
decides identity (`exact-vs-approximate.md`).
|
|
45
|
+
|
|
46
|
+
The formulas live in `src/geometry.ts`, `src/mind/traverse.ts` (corpus scale)
|
|
47
|
+
and `src/mind/canonical.ts` (windows). A new bar is added there, never to
|
|
48
|
+
`config.ts`.
|
|
54
49
|
|
|
55
50
|
## Two derivations that bite
|
|
56
51
|
|
|
57
|
-
|
|
58
|
-
span tolerates four whole windows of foreign bytes while still
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
back to a noisy concept-hop).
|
|
52
|
+
**`identityBar` must scale with length.** A fixed cosine of `1 − 1/√D` over a
|
|
53
|
+
`4·√D`-byte span tolerates four whole windows of foreign bytes while still
|
|
54
|
+
claiming near-identity. An identity claim may tolerate at most one window (`W`),
|
|
55
|
+
so the bar is `1 − W/len`, floored at `mergeThreshold`. Reusing `mergeThreshold`
|
|
56
|
+
for a whole-span claim silently widens the tolerance as the span grows.
|
|
57
|
+
|
|
58
|
+
**A bar is priced for one quantity only.** `consensusFloor` is calibrated for
|
|
59
|
+
pooled climb votes, where each region contributes at most `ln(N/c) ≤ ln N`, so
|
|
60
|
+
`ln N + ½` asks for corroboration beyond one maximally specific region.
|
|
61
|
+
`chooseNext`'s support count is different: it is bounded by retellings of one
|
|
62
|
+
fact and does not grow with `N`. Gating that count by `consensusFloor` must
|
|
63
|
+
eventually fail, and it did: at N≈325K, a 2-against-1-1-1 corroboration was
|
|
64
|
+
refused, and the answer fell back to a noisy concept hop.
|
|
71
65
|
|
|
72
66
|
## Pins
|
|
73
67
|
|
|
74
|
-
- `test/
|
|
75
|
-
|
|
76
|
-
- `test/
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
abstention at scale.
|
|
68
|
+
- `test/64` — `mergeThreshold`, `identityBar` and `reachThreshold`, and the
|
|
69
|
+
scale-aware floor.
|
|
70
|
+
- `test/40` — `consensusFloor` must not gate `chooseNext`'s support counts.
|
|
71
|
+
- `test/78` — `atomReach`/`atomIsHub`: atoms abstain once the corpus makes them
|
|
72
|
+
hubs.
|